Senior SRE Software Engineer - Software Developer Platform

AppleApplyPublished 1 days agoFirst seen 1 days ago
Apply

Summary

The Apple Services Engineering (ASE) team is one of the most exciting examples of Apple’s long-held passion for combining art and technology. These are the people who power the App Store, Apple TV, Apple Music, Apple Podcasts, and Apple Books. And they do it on a massive scale, meeting Apple’s high expectations with high performance to deliver a huge variety of entertainment in over 35 languages to more than 150 countries. These engineers build secure, end-to-end solutions. They develop the custom software used to process all the creative work, the tools that providers use to deliver that media, all the server-side systems, and the APIs for many Apple services. Thanks to Apple’s unique integration of hardware, software, and services, engineers here partner to get behind a single unified vision. That vision always includes a deep commitment to strengthening Apple’s privacy policy, one of Apple’s core values. Although services are a bigger part of Apple’s business than ever before, these teams remain small, forward-thinking, and cross-functional, offering greater exposure to the array of opportunities here.

Description

Apple Services Engineering infrastructure is BIG. Operating at our scale, across multiple geographically dispersed data centers and servicing hundreds of millions of users presents unique challenges. As an SRE at Apple, you'll need to solve these problems using data, teamwork, and your own expertise. SREs at Apple own the full infrastructure stack; from device driver performance debugging to content delivery network traffic management — our responsibilities are both broad and deep.

ASE runs the majority of its systems on Linux. We run a mix of open source, vendor licensed, and internally developed tools to perform functions such as system configuration management, provisioning, software deployment, logging, and monitoring. You'll learn these tools and have opportunities to improve them. Our team is collaborative; we work closely with the development teams we support to deliver the best results for Apple. We think critically and strive to balance the best solution with the need to get things done for each engineering challenge we face. Good ideas are heard and results are rewarded. Culturally we believe in a close partnership with our development teams and aim to design & build new services together. We're passionate about software and automation in SRE and develop a variety of tooling and infrastructure. Our services run on mixed & hybrid platforms.

Responsibilities

  • * Own the reliability, availability, and performance of mission-critical cloud platform services
  • * Lead, grow, and mentor a team of Site Reliability Engineers focused on developer experience
  • * Drive incident response, post incident reviews, and systemic improvements that reduce operational toil
  • * Partner with software engineering and architecture teams to influence system design for reliability, scalability, and operability
  • * Establish and refine SRE practices including SLOs, error budgets, capacity planning, and change management
  • * Champion AI-powered tooling and automation to improve incident triage, reduce operational toil, drive capacity efficiency, and accelerate engineering workflows
  • * Manage on-call rotations and ensure sustainable, well-supported operational coverage
  • * Communicate clearly across teams to build a culture of visibility, transparency, and shared ownership

Minimum Qualifications

  • * 8+ years in a Site Reliability Engineering, DevOps, or Infrastructure focused role
  • * Advanced knowledge and hands-on experience with source code and artifact management systems, CI/CD infrastructure (GitHub, Artifactory, Jenkins)
  • * Strong systems background — comfortable troubleshooting across the full stack (network, OS, container runtime, application)
  • * Expert-level Go and/or Python, with a track record of shipping software.
  • * Experience with configuration management at scale (Puppet, Ansible, or equivalent)
  • * Experience running infrastructure as an internal managed service with defined SLAs
  • * Demonstrated ability to drive cross-functional initiatives to completion
  • * Strong written and verbal communication skills
  • * Strong sense of ownership and integrity demonstrated through clear communication and collaboration
  • Education
  • * Bachelor's Degree in Computer Science, an engineering-related field, or equivalent related experience

Preferred Qualifications

  • * Experience in managing and scaling distributed systems in a public, private, or hybrid cloud environment
  • * Excellent troubleshooting and problem solving skills, both with and without AI assistance
  • * Experience with scale testing, disaster recovery, and capacity planning
  • * Experience with third-party cloud platforms (AWS, GCP, or Azure)
  • * Troubleshoot complex distributed systems running on both bare metal and hypervisors
  • * Evolve critical, foundational systems to provide next generation features at scale