Site Reliability Engineer — Insight Team

AppleApplyPublished 8 days agoFirst seen 1 days ago
Apply

Summary

The Insight team runs one of Apple's most critical Big Data ecosystems — an exabyte-scale, highly-available infrastructure that underpins manufacturing operations for every Apple product, globally. Every iPhone, iPad, and Mac has touched our systems.
We advance technology by relying on each other's strengths and skills to build something bigger than ourselves. For this reason, team culture is central to our values. We value social skills and integrity as much as technical craft.
We are looking for extraordinary DevOps with experience building large-scale data platforms, analytic tools and solutions which can help take our platform to the next level. Do you excel in a high-demand setting and exceed expectations, in an environment that requires time-management? The right person will prioritize tasks and complete assignments ahead of schedule. While being a great standout colleague, you will also work independently.

Description

The primary mission of this role is to ensure the stability and performance of our production environment, including the secure and reliable execution of all deployment and monitoring processes. We expect this role to go beyond operations by fully embracing a dev-and-ops mindset. Beyond standard support and troubleshooting, you will leverage AI tools and automation to drive efficiency, applying your engineering expertise to continuously upgrade and optimize our Big Data and microservices ecosystem.

Responsibilities

  • You will operate and improve services within a very large-scale, highly available Big Data ecosystem supporting Exabytes level of data with sustained, rapid growth. Your work directly enables the engineering and operations teams that build every Apple product.
  • - Own the reliability, performance, and scalability of services and cloud infrastructure within the Insight ecosystem
  • - Build and advance AIOps capabilities across the ecosystem — developing AI-driven alerting, anomaly detection, LLM-assisted operational tooling, and automated incident triage that measurably improve observability and response
  • - Instrument services for deep observability: dashboards, meaningful alerts, runbooks, and clearly defined SLOs and SLIs that give the team and stakeholders an accurate view of ecosystem health
  • - Partner with incident management to drive effective incident response and conduct thorough post-incident reviews; translate findings into architectural and operational improvements that prevent recurrence
  • - Collaborate with cross-functional engineering teams across Apple's global manufacturing services, communicating effectively across time zones and cultures

Minimum Qualifications

  • BS or MS in Computer Science, Software Engineering, or an equivalent technical discipline
  • Good programming skills in Java, Python or Go
  • Hands-on experience in SRE, DevOps, or platform engineering, with demonstrated ability to work on technical initiatives
  • Experience applying AI/ML techniques to development & operations
  • Proficiency in English for clear technical communication, documentation, and global collaboration
  • Availability to participate in SRE on-call rotations during China morning hours (8:00 AM CST during US Daylight Saving Time, and 9:00 AM CST during US Standard Time)

Preferred Qualifications

  • Cloud Technologies: Hands-on experience with cloud platforms (AWS or GCP), containerization and orchestration (Docker, Kubernetes), and multi-region infrastructure design including data residency considerations and IAM.
  • Distributed Systems & Big Data Platforms: Strong background in microservices, APIs, and messaging/streaming (Kafka, Solace, Pub/Sub), combined with production experience in relational databases (MySQL, PostgreSQL) and big data/storage technologies (Redis, Bigtable, Druid, Elasticsearch, ClickHouse).
  • CI/CD & Observability: Proven ability to build and maintain CI/CD pipelines or GitOps workflows (ArgoCD, Jenkins, GitHub Actions), and operate observability systems (Grafana, Prometheus, Kibana) including service instrumentation, dashboard design, and alert tuning.
  • AI tooling & Problem Solving: Forward-looking experience in developing MCP and AI Agentic flows, paired with a flexible and creative approach to solving complex technical problems.