Software Engineer, Data Services, IS&T Ai & Data Platforms

AppleApplyPublished 1 days agoFirst seen 1 days ago
Apply

Summary

Apple's AI & Data Platform (AiDP) Data Services organization is seeking a motivated database systems engineer to join our Data Services SRE team, focused on our transactional database fleet. Engineers on this team develop and contribute to platform tooling that manages relational and distributed-SQL/document engines powering some of Apple's most critical internet services at massive scale. In AiDP, your work will benefit hundreds of millions of users.

Description

The AiDP Data Services SRE team builds engine-agnostic platform capabilities — provisioning, backup/restore, observability, and self-service — spanning Oracle, PostgreSQL, MongoDB, and CockroachDB. This role involves supporting data migration efforts across hybrid-cloud environments with minimal downtime and guaranteed data integrity, following established runbooks and adhering to local data residency and regulatory requirements. This role requires good communication and collaboration with Core Storage teams and colleagues across a distributed team.

Responsibilities

  • Apply core SRE concepts — monitoring, alerting, and incident management — with growing independence.
  • Build working knowledge of transactional database concepts (consistency models, isolation levels, crash and recovery semantics) across Oracle, PostgreSQL, MongoDB, and CockroachDB.
  • Support performance engineering efforts (profiling, tuning) across transactional engines, under guidance from senior team members.
  • Assist with service management across bare metal, virtualized (EC2, Ali Cloud ECS), and containerized (K8s) style platforms.
  • Execute and support data migration tasks across hybrid clouds, including validation and rollback support.
  • Build understanding of operating systems and datacenter architecture concepts relevant to transactional workloads.
  • Produce and maintain clear documentation — runbooks, architecture diagrams, and migration playbook updates.
  • Support compliance with local data residency and regulatory requirements applicable to China-based infrastructure.

Minimum Qualifications

  • BS or MS in Computer Science / related fields or equivalent work experience, with 3–5 years in a Site Reliability Engineering / Infrastructure focused role.
  • Hands-on production experience with at least one of Oracle, PostgreSQL, MongoDB, or CockroachDB, including exposure to On Call and Incident Management.
  • Exposure to running infrastructure with reliance on automation tooling, across Datacenter and Cloud architectures (including Alibaba Cloud/Ali Cloud and/or AWS).
  • Solid troubleshooting skills, a resourceful first-principles approach to problem solving, and good technical writing habits.
  • Good understanding in one or more of the following programming languages: Python, Go.

Preferred Qualifications

  • Some operational experience with stateful services on Kubernetes (operators, StatefulSets, CSI storage).
  • Exposure to database-as-a-service control planes, provisioning APIs, or self-service tooling.
  • Experience assisting with cross-cloud / hybrid-cloud data migrations for transactional systems.
  • Familiarity with observability concepts (SLIs/SLOs), backup, and DR practices for OLTP/distributed-SQL engines.
  • Proficiency in Mandarin (spoken and written), to support coordination with local Ali Cloud teams and regional vendors/partners.
  • Contributions to open-source database projects or internal data-platform tooling.