Senior SRE Engineer - ASE ADP Compute
Summary
Apple Service Engineering (ASE) teams build and scale the platforms and infrastructure behind many of Apple's services (such as iCloud, iTunes, Siri, and Maps). We are the foundation on which Apple's software developers build the products that our customers love. We are looking for a passionate and dedicated Senior Site Reliability Engineer to provide technical leadership on our team to help ensure our customers have the highest quality Apple Services experience.
Description
The Apple Data Platform (ADP) Compute SRE team is responsible for the core infrastructure, including our legacy bare-metal platforms and modern cloud based infrastructure stack. We partner with both peer SRE teams and several of our world-class software and product engineering teams to support infrastructure reliability, multi-year parallel migrations for Apple properties, as well as the automation, tooling, incident, and process management necessary to ensure smooth 24x7 operations for ADP customers.
Responsibilities
- Serve as Principal/Lead SRE for the ADP SRE team, guiding Individual Contributors and collaborating with senior engineers to ensure service architecture, tooling, design, and implementations reach the highest quality standards, all grounded in empirical data and unique problem-solving expertise
- Identify opportunities to improve technical operations and up skill partner teams by demonstrating best practices, while fostering deep collaboration with software developers across the Data Platform Services to deliver seamless customer experiences
- Provide technical direction for Kubernetes infrastructure, encompassing tooling development (including infrastructure-as-code) and core services, within a cross-functional environment dedicated to applying a consistent incident management process and high availability architecture supported by exhaustive observability metrics and automated deployments
- Utilize Python, Java, or Golang programming skills, enhanced by Generative AI tooling, to rapidly build mission-critical automation and tools, while critically evaluating solutions to balance long-term optimization with immediate business priorities
- Exhibit strong collaboration and presentation abilities to clearly communicate strategic ideas and advocate for the SRE team's deliverables and needs with ASE leadership, ensuring that good ideas are heard and results are rewarded
- Manage production on-call rotations and lead the end-to-end management of all incidents, applying a holistic approach that balances technical excellence with organizational goals
Minimum Qualifications
- 12+ years of experience in Site Reliability Engineering, specifically managing infrastructure and services at scale
- 5+ years of experience in management or technical leadership roles
- 5+ years of proficiency in running applications and managing Kubernetes clusters on Alibaba Cloud, AWS, or GCP
- Proven history of end-to-end project management and delivery
- Demonstrable programming skills for developing software/tools and leading code reviews
- Advanced knowledge of Linux, Networking, and Containers
Preferred Qualifications
- 15+ years of experience in SRE or related work managing infrastructure at scale
- Proficiency with the architecture, deployment, performance tuning, and troubleshooting of open source data analytics or governance technologies such as Spark, Flink, Iceberg, Trino, and/or Druid.
- Experience with scale testing, disaster recovery, and capacity planning
- Ability to define the technical roadmap for infrastructure and drive cross-functional alignment on architectural standards and best practices