Manager, TechOps / Site Reliability Engineering

AppleApplyPublished 2 days agoFirst seen 3 days ago
Apply

Summary

Do you want to help build some of the largest and most consequential enterprise and customer technology systems in the world? Join Apple’s Information Systems and Technology (IS&T) organization. IS&T is the engine behind everything Apple does for customers and for the people who build for them. It’s Apple’s central nervous system. Supporting 2.5 billion active Apple devices, processing billions of secure transactions, and keeping the technology that defines modern life running flawlessly, IS&T makes the impossible feel effortless.

Do you love building solutions to handle global complexity and immense scale? Imagine what you could do here.

Customer Systems is part of IS&T and drives the technology behind Apple's customer support experience — from contact center operations to the software powering the iconic Genius Bar. The team also builds and operates AppleCare's online support platform, which handles 6 billion visits per year, delivering seamless, high-quality support to Apple customers around the globe.

Description

At Apple, we are dedicated to creating innovative solutions that delight our users through ultra-fast, thoughtfully engineered, and meticulously crafted products. Our team is not just any group — we are a highly motivated, dynamic, and ever evolving collection of individuals driven by a passion for excellence and continuous improvement.

The Customer Systems Operations team is seeking an exceptional Manager to lead a team of TechOps and Site Reliability Engineers responsible for the availability, performance, and resilience of Apple's business-critical operational platforms. The ideal candidate combines deep technical credibility in infrastructure and reliability engineering, strong people leadership, and a proven ability to keep globally distributed, 24x7 operational teams delivering in a fast-paced, highly cross-functional environment.

You will operate at the intersection of engineering, operational excellence, and AI innovation — partnering closely with global engineering managers and business stakeholders to guide technical direction, modernize operational practices, and accelerate reliability outcomes through GenAI-enabled engineering and automation. This is a hands-on leadership role requiring both organizational leadership and direct technical ownership of the team's day-to-day delivery.

Responsibilities

  • Lead and develop a team of TechOps and SRE engineers, owning end-to-end delivery accountability for their operational work, incident response, and reliability initiatives.
  • Provide direct technical guidance on infrastructure design, automation strategy, and incident response, staying closely engaged with the team's day-to-day work.
  • Define and drive the team's technical and operational roadmap, aligned to enterprise reliability and delivery strategy.
  • Establish governance models, engineering standards, and execution frameworks for operational work across the team.
  • Lead the team's response to major incidents, including post-incident reviews and long-term reliability improvements.
  • Champion adoption of GenAI-driven engineering and operational practices to improve automation, productivity, and delivery outcomes.
  • Drive data-driven decision making and AI-enabled operational excellence, including SLO/error-budget tracking and KPI reporting.
  • Manage staffing, workload balancing, and 24x7 on-call/shift coverage plans across a globally distributed team.
  • Hire, coach, and mentor TechOps/SRE engineers while building a culture of ownership, accountability, and continuous improvement.
  • Act as the primary technical liaison and interface between the team and global engineering managers, operations managers, business partners, and cross-functional leaders across regions.
  • Partner with application engineering, product, and business operations leaders to align priorities and execution.
  • Provide clear communication to global managers and leadership on team/program health, risks, and outcomes.

Minimum Qualifications

  • 15+ years of experience in Site Reliability Engineering, TechOps, DevOps, or Infrastructure Engineering, with hands-on background designing and operating highly available, scalable, and secure systems (cloud and on-prem infrastructure, Kubernetes, CI/CD, automation, observability/monitoring platforms).
  • 8+ years directly managing a team of TechOps/SRE/DevOps engineers, including hiring, performance management, coaching, and career development.
  • Experience managing globally distributed teams across time zones, including 24x7 coverage and on-call/shift models.
  • Working knowledge of incident management practices, SLO/error-budget frameworks, and operational excellence disciplines.
  • Familiarity with applying GenAI capabilities to engineering/operational workflows.
  • Ability to operate effectively with senior/global managers and stakeholders while maintaining deep technical engagement with the team's day-to-day work.
  • Exceptional analytical, organizational, and problem-solving skills.

Preferred Qualifications

  • Bachelor's or Master's degree in Computer Science, Engineering, or a related field.
  • Demonstrated ability to build, inspire, and retain high-performing technical operations organizations.
  • Experience in Contact Center, Retail, or other high-availability customer-facing platforms and technologies.