Site Reliability Engineer, Apple Data Platform / Big Data Platform

AppleApplyPublished 19 hours agoFirst seen 17 hours ago
Apply

Summary

The Apple Services Engineering team (ASE) is one of the most exciting examples of Apple's long-held passion for combining art and technology. These are the people who power the App Store, Apple TV, Apple Music, Apple Podcasts, and Apple Books — at extensive scale, meeting high expectations to deliver a huge variety of entertainment in over 35 languages to more than 150 countries.

Within ASE, the Apple Data Platform SRE team keeps a massive, multi-cloud platform running for thousands of internal engineers building the next generation of data and AI products at Apple. We sit at the intersection of infrastructure, automation, and customer success — running incident response, providing hands-on support to internal teams, and partnering with developers to make cutting-edge services like Spark, Flink, Airflow, Trino, Notebooks, and LLM-based agent platforms reliable at scale.

Description

This is a rare opportunity to build deep expertise across one of the most technically diverse platforms at Apple — while specialising in the big data engines and catalog/governance layers that power analytics and data engineering across the company. As an SRE on Apple Data Platform, you'll operate and support the team's full portfolio, from ML/AI platform services to multi-cloud infrastructure, and grow into the team's go-to expert for big data platform services — including Spark, Flink, Airflow, Trino, Notebooks, REST Catalog services (such as Glue Catalog), and data governance. Just as importantly, you'll be a first point of contact for the internal customers who rely on these services daily — someone who can translate a confusing error or a vague support request into a clear diagnosis and a fast resolution.

We're looking for a self-motivated engineer who thrives on ownership — someone who wants a set of services to call their own, the autonomy to drive their reliability roadmap, and the collaborative instinct to keep that work aligned with the team's broader direction. If you love solving hard operational problems, take genuine satisfaction in helping frustrated customers get unblocked, and want a front-row seat to how Apple's data engineering platform scales, this role offers real room to grow your scope and impact over time.

Responsibilities

  • Operate, monitor, and triage production and non-production environments across the ADP portfolio — data processing, ML/AI, and multi-cloud infrastructure.
  • Participate in a rotating on-call schedule across supported services, including occasional weekday and weekend coverage.
  • Own the operational health of big data platform services as SME — driving reliability, support, and customer guidance for Spark, Flink, Airflow, Trino, Notebooks, REST Catalog, and governance tooling.
  • Serve as a primary point of contact for internal customers via Slack — clearly communicating status, root cause, and next steps during active issues.
  • Screen, triage, and resolve customer-reported service issues and support tickets, prioritizing based on customer impact and urgency.
  • Partner with dev teams across time zones to onboard new services — understanding architecture, then designing monitoring, alerting, and dashboards (Prometheus, Grafana, Splunk).
  • Build automation and self-healing tooling that reduces manual toil and scales the team's operational capacity.
  • Identify, escalate, and resolve production issues to protect platform reliability and customer experience.
  • Champion customer success by helping internal teams understand platform capabilities and adopt tools effectively.
  • Collaborate with SRE and dev partner teams, engineering, and program management to align execution with team and org goals.

Minimum Qualifications

  • Bachelor's Degree in Computer Science, an engineering-related field, or equivalent related experience.
  • 1-4 years in a Site Reliability Engineering, DevOps, or Infrastructure-focused role.
  • Proficient in Python; working knowledge of Golang a plus.
  • Deep understanding of one or more Big Data technologies (Spark, Flink, Airflow, Trino, Notebooks).
  • Experience with Kubernetes and at least one major cloud provider (AWS or GCP).
  • Excellent written and verbal communication skills, with the ability to explain technical issues clearly to non-expert customers.
  • Solid grounding in SRE principles, with prior on-call, production-support, or customer-facing support role experience.

Preferred Qualifications

  • Experience with REST Catalog services (e.g., Glue Catalog) and data governance frameworks.
  • Prior experience in a customer-facing or technical support role, with a demonstrated passion for customer success.
  • Familiarity with observability tooling: Prometheus, Grafana, Splunk, PagerDuty.
  • Working knowledge of CI/CD pipelines and deployment workflows.
  • Experience with S3 and cloud storage/networking fundamentals.
  • Familiarity with data pipeline orchestration and workflow scheduling patterns.
  • A track record of automating manual operations through scripting or tooling.
  • Intellectual curiosity and a drive to keep learning — for yourself, your team, and the org.