Software Engineering Manager - ASE Compute
Summary
People at Apple don't just build products, they craft the kind of experience that has revolutionized entire industries. The diverse collection of our people and their ideas inspire innovation in everything we do. Imagine what you could do here! Join Apple, and help us leave the world better than we found it.
The Apple Service Engineering (ASE) team builds and provides systems and infrastructure that power Apple's services (such as iCloud, Apple Music, Apple Intelligence, and Maps). We are the foundation on which Apple's software developers build the products that our customers love. Our services have to scale globally, stay highly available, and "just work." If you love designing, engineering, and running systems and infrastructure that will help millions of customers, then this is the place for you!
Description
Apple's ASE Compute team builds and operates the private cloud infrastructure that powers Apple services at massive scale. Our platform delivers bare-metal Kubernetes clusters and virtualized environments to thousands of engineers across the company. We are looking for an SRE Manager to lead a team that keeps this infrastructure reliable, performant, and ready for the next order-of-magnitude growth.
This is a hands-on leadership role. You will set the technical direction for reliability and operational excellence while mentoring engineers, driving automation, and partnering closely with software and infrastructure teams to ship improvements that matter.
Responsibilities
- * Lead, grow, and mentor a team of Site Reliability Engineers focused on large-scale Kubernetes and compute infrastructure
- * Own the reliability, availability, and performance of mission-critical cloud platform services
- * Drive incident response, post incident reviews, and systemic improvements that reduce operational toil
- * Partner with software engineering and architecture teams to influence system design for reliability, scalability, and operability
- * Establish and refine SRE practices including SLOs, error budgets, capacity planning, and change management
- * Champion AI-powered tooling and automation to improve incident triage, reduce operational toil, drive capacity efficiency, and accelerate engineering workflows
- * Manage on-call rotations and ensure sustainable, well-supported operational coverage
- * Communicate clearly across teams to build a culture of visibility, transparency, and shared ownership
Minimum Qualifications
- * 3+ years of engineering management experience leading infrastructure or SRE teams
- * Deep experience operating large-scale, multi-tenant Kubernetes environments in production
- * Strong systems background — comfortable troubleshooting across the full stack (network, OS, container runtime, application)
- * Experience with configuration management at scale (Puppet, Ansible, or equivalent)
- * Track record of building high-performing teams through coaching, clear expectations, and psychological safety
- * Demonstrated ability to drive cross-functional initiatives to completion
- * Strong written and verbal communication skills
- Education
- Bachelor's Degree in Computer Science, an engineering-related field, or equivalent related experience
Preferred Qualifications
- * Experience with third-party cloud platforms (AWS, GCP, or Azure)
- * Familiarity with bare-metal provisioning and lifecycle management at datacenter scale
- * Experience with Java, Go, or Python services in production
- * Understanding of cloud-native observability (Prometheus, Thanos, Splunk, or similar)
- * CNCF Certified Kubernetes Administrator (CKA) or equivalent hands-on certification
- * Experience running infrastructure as an internal managed service with defined SLAs