Site Reliability Engineer

AppleApplyPublished 1 hours agoFirst seen 1 hours ago
Apply

Summary

Joint Mobile Engineering Team (JMET) is a security engineering team that provides critical services for Apple across every product line. From manufacturing to customer-facing operations, the team's services span the entire lifecycle of most Apple hardware. The team designs, implements, and supports services that improve customer safety and privacy through security services tightly coupled with hardware — including server-side solutions that activate Apple devices worldwide and support Apple's efforts in eSIM. The team works closely with cross-functional teams across Apple, as well as carriers and other third parties. Many of the team's services are referenced in the iOS Security Guide or discussed publicly online.

As a Site Reliability Engineer, you will participate in initiatives that are important to the success of upcoming product launches and security initiatives. You will also work with large cross-functional teams to align expectations and validate the work you're doing.

Description

This role is responsible for the reliability, monitoring, and operational health of production data center services. It includes triaging and resolving incidents, building automation to reduce manual work, and partnering with engineering teams to improve system resilience as services scale.

Responsibilities

  • Triage production incidents by severity and business impact, drive mitigation to resolution, and lead root cause analysis through to closure
  • Build and maintain infrastructure supporting data center services
  • Design and maintain monitoring, alerting, and observability coverage across supported services
  • Build automation to reduce manual toil and enable self-healing systems
  • Partner with development and QA teams to prioritize and resolve production defects, and take knowledge transfer on services entering production
  • Participate in capacity planning and architecture discussions to ensure services scale reliably under growth and unexpected demand
  • Apply security practices — access controls, patching, and threat identification — across supported services
  • Share in a 24x7 on-call rotation, responding to and resolving production incidents in a timely manner

Minimum Qualifications

  • Bachelor's degree in Computer Science, a related field, or equivalent practical experience
  • 3+ years of experience in a Site Reliability Engineering, DevOps, or production operations role
  • Experience automating infrastructure tasks using one or more programming or scripting languages
  • Experience with hybrid infrastructure management across on-prem and cloud environments (e.g., AliCloud): compute, networking, storage, and data stores such as Oracle, Cassandra, MongoDB, and Postgres
  • Experience applying security practices: access controls, patching, TLS/SSL, and identifying threats at the application and network level
  • Experience troubleshooting distributed systems failures and driving incidents through to resolution
  • Proficiency in English and Mandarin

Preferred Qualifications

  • Experience with observability and alerting tools (e.g., metrics, tracing, and dashboarding platforms)
  • Experience applying SRE principles — error budgets, SLAs, SLOs, and SLIs — to measure and improve service reliability
  • Experience with container orchestration platforms, including deployments, networking, storage, and secrets management
  • Foundational knowledge of release engineering: CI/CD pipelines, version control, and deployment methodologies
  • Foundational knowledge of distributed systems concepts: scalability, fault tolerance, and trade-offs under load
  • Ability to communicate clearly across cross-functional and global teams, adjusting technical detail for different audiences
  • Ability to work independently and take initiative in ambiguous or fast-changing situations