Site Reliability Engineer, Enterprise Technology Services
Summary
Working at Apple offers the chance to do your best work and be part of something truly remarkable. Imagine the possibilities of what you could do here. At Apple, innovative ideas swiftly transform into phenomenal platforms services and customer experiences. Bring passion and dedication to your role and the sky’s the limit.
The Emerging Technology Services team productizes upcoming software and champions it within the Apple community. We’re a dynamic team blending SRE DevOps and Operational Intelligence. Our creations become technology pillars supporting Apple across various domains. From eCommerce to reporting systems, we ensure seamless communication between disparate systems at unprecedented scale and speed. Our support extends to Load Balancers, perimeter security , integrations and API gateways, facilitating platform adoption.
The SRE Engineer position is part of a horizontal Platform Engineering Ops group dedicated to a diverse range of technologies and applications, including highly scalable distributed applications. This role demands strategic engineering and data science skills alongside hands on technical work. The candidate will delve into the security domain, participate in building machine learning pipelines, develop anomaly detection models, explore crypto strategies for privacy and apply data science skills to vast datasets.
Description
We’re seeking a Senior SRE Engineer to design build and scale a modern DevOps and SRE ecosystem from the ground up. This role demands deep hands on experience, strong architectural thinking and the ability to establish GitOps driven cloud native CI/CD platforms using cutting-edge technologies. The ideal candidate will serve as a foundational engineer and technical leader defining standards and reliability practices throughout the organization.
You should have a passion for programming and a solid conceptual understanding of the operating environment including JVM operating systems file systems and network protocols. Technical expertise strong communication skills and teamwork are essential as this role involves collaborating with both technical and non-technical groups within Apple and externally with our supply chain partners.
Responsibilities
- As Site Reliability and Operations Engineer (SRE), you’ll be part of the action—working closely with cross functional teams. You will:
- Engineer scalable, highly available infrastructure aligned with business and reliability objectives.
- Drive KPI, SLA, and SLO performance through measurable reliability and operational targets.
- Optimize system performance, capacity, cost, and resource utilization across growing workloads.
- Automate infrastructure, deployments, remediation, testing, and operational workflows using IaC and AI-driven automation.
- Expand observability through metrics, logs, traces, dashboards, alerting, and actionable service health insights.
- Apply AIOps capabilities to detect anomalies, correlate events, predict failures, and accelerate incident resolution.
- Harden systems through SecOps practices, vulnerability management, access controls, and continuous security improvements.
- Lead incident response, root-cause analysis, reliability improvements, and operational readiness for service expansion.
- Collaborate effectively across team boundaries and mentor junior team members
Minimum Qualifications
- 7+ years of software engineering experience
- Hands on experience with at least one object oriented language, ideally Java or JEE
- Familiarity with automation tools like Ansible and Terraform
- Strong programming and scripting skills in Python, Bash and Lua
- Proficiency in relational and non-relational database fundamentals, including hands-on PL/SQL experience
Preferred Qualifications
- Strong analytical skills are essential for troubleshooting Java and JVM technologies runtime configurations.
- Familiarity with modern web services architectures and cloud platforms like AWS, GCP, Azure and distributed storage systems such as ScaleIO and Amazon S3 is crucial.
- Experience with monitoring and logging tools including Prometheus, Splunk, Grafana and CloudWatch is valuable.
- Understanding of CI/CD, release engineering and DevOps principles is important.
- A good grasp of various machine learning algorithms and patterns is beneficial.
- Knowledge of cryptographic algorithms is also important.
- In-depth experience in writing, understanding and reverse engineering regular expressions to detect patterns is highly desirable.
- A strong understanding of TLS, mTLS and industry-standard secure communications protocols is essential.
- Researching and understanding vulnerabilities and threats from open forums and translating them into system design and implementation to detect and prevent them is a key skill.