Engineering Manager, Data Services, IS&T Ai & Data Platforms

AppleApplyPublished 1 days agoFirst seen 1 days ago
Apply

Summary

Apple's AI & Data Platform (AiDP) Data Services organization is seeking an experienced, versatile engineering leader to head our Data Services SRE organization. This leader will manage a group of engineering managers and technical leads responsible for cross-cutting software and platform tooling that manages a fleet of relational, distributed-SQL, document, search, and analytics engines powering some of Apple's most critical internet services, deployed at massive scale across data centers worldwide. In AiDP, your leadership will directly shape the reliability and scalability of platforms benefiting hundreds of millions of users, and will be critical to the success of some of the most visible current and future Apple features.

Description

The AiDP Data Services SRE organization builds engine-agnostic platform capabilities — provisioning, backup/restore, observability, and self-service — spanning PostgreSQL, CockroachDB, Couchbase, OpenSearch, MongoDB, Oracle, and Apache Doris. As the leader of this organization, you will set technical direction and organizational strategy for engine-agnostic data migration tooling across hybrid-cloud environments, ensuring minimal downtime, guaranteed data integrity, and adherence to local data residency. This role requires exceptional cross-functional leadership, executive-level communication, deep partnership with Core Storage and Analytics leadership, and the ability to build, mentor, and scale a distributed organization of managers and senior engineers.

Responsibilities

  • Lead, mentor, and develop a team of engineering managers and/or senior technical leads across the Data Services SRE organization, driving career growth and organizational health.
  • Set and communicate long-term technical vision and roadmap for engine-agnostic platform capabilities (provisioning, backup/restore, observability, self-service) across a heterogeneous database fleet.
  • Own organizational-level SRE outcomes — reliability targets, incident management maturity, and on-call health — across relational, distributed-SQL, document, search, and OLAP/analytics paradigms.
  • Drive performance engineering strategy and prioritization across heterogeneous engines, partnering with senior ICs and architects on design reviews and technical direction.
  • Oversee service management strategy across bare metal, virtualized (EC2, Ali Cloud ECS), and containerized (K8s) platforms, ensuring consistent operational standards across managers/teams.
  • Sponsor and oversee execution of engine-agnostic data migration programs across hybrid clouds, including cutover strategy, replication/validation, and rollback — ensuring cross-team alignment and risk management at the program level.
  • Guide datacenter architecture and multi-datacenter system design decisions, representing Data Services SRE in broader infrastructure planning and failure-domain strategy discussions.
  • Establish and enforce standards for high-quality documentation — runbooks, architecture diagrams, post-incident reviews, and migration playbooks — across all teams in the organization.
  • Ensure organizational compliance with local data residency, cross-border data transfer, and regulatory requirements applicable to China-based infrastructure, partnering closely with Legal, Security, and regional teams.
  • Manage headcount planning, hiring strategy, budget/resourcing decisions, and cross-org prioritization for the Data Services SRE organization.
  • Represent Data Services SRE in leadership reviews, roadmap planning, and executive stakeholder communication.

Minimum Qualifications

  • BS or MS in Computer Science / related field or equivalent experience, with 12+ years in Site Reliability Engineering / Infrastructure roles, including 4+ years of people management experience.
  • Proven track record building, scaling, and retaining high-performing SRE/infrastructure engineering organizations.
  • Deep hands-on background (prior to or alongside management) with at least two of PostgreSQL, Oracle, MongoDB, CockroachDB, Couchbase, OpenSearch, or Apache Doris, supporting internet-facing production services via On Call and Incident Management.
  • Demonstrated experience overseeing large-scale infrastructure operations with heavy reliance on automation tooling, across Datacenter and Cloud architectures (including Alibaba Cloud/Ali Cloud and/or AWS).
  • Strong executive communication and technical writing skills; ability to translate deep technical context into leadership-level narratives and drive decisions across senior stakeholders.
  • Working knowledge of one or more of the following programming languages: Python, Go — sufficient to engage credibly in technical design discussions and code/architecture reviews.

Preferred Qualifications

  • Experience sponsoring or overseeing the build of database-as-a-service control planes, provisioning APIs, or self-service tooling at an organizational level.
  • Proven track record leading engine-agnostic, cross-cloud / hybrid-cloud data migration programs at scale.
  • Experience defining organization-wide engine-agnostic observability (SLIs/SLOs), backup, and DR standards across a heterogeneous fleet.
  • Experience leading teams operating real-time OLAP/analytics engines (e.g., Apache Doris, ClickHouse, StarRocks) at scale.
  • Proficiency in Mandarin (spoken and written), to support coordination with local Ali Cloud teams and regional vendors/partners.
  • Track record of sponsoring contributions to open-source database projects or internal data-platform tooling.