Site Reliability Engineer / Devops — Retail Engineering

AppleApplyPublished 1 days agoFirst seen 1 days ago
Apply

Summary

Apple’s IS&T Retail Engineering team is seeking a SDET to own quality across our retail ecosystem — spanning eCommerce backend services, SAP integrations, and cross-functional end-to-end workflows. This is a hands-on technical leadership role: you’ll architect automation frameworks, lead testing strategy across multiple concurrent projects, and drive quality standards organization-wide.

You’ll work at the intersection of backend development, test automation, and program coordination — ensuring that features shipping to apple.com/shop and supporting retail systems meet Apple’s bar for reliability, performance, and customer experience.

Description

You will drive the reliability, deployment, and scalability of compute platforms across on-premises and hybrid cloud environments. Collaborating closely with cross-functional technical and business partners, you will build Infrastructure as Code, optimize container orchestration, and streamline CI/CD delivery pipelines. You will champion automation and operational excellence, ensuring high availability, robust security standards, and proactive observability across large-scale distributed systems.

Responsibilities

  • Deploy, support and monitor compute platforms and application stacks.
  • Ability to understand complex systems and a desire to constantly make things better
  • Explore and evaluate new technologies and solutions.
  • Strong interpersonal skills and ability to work effectively across multiple business and technical teams
  • Develop and fix the application on failures.
  • We promote innovation and use of new technology to further improve our creative output. We’re looking for a talented and passionate person to join this amazing team, if you feel this is you, we’d love to hear from you.

Minimum Qualifications

  • Bachelor’s degree in Computer Science or equivalent field with 7+ years of experience, or Master’s degree with 5+ years of experience.
  • 7+ years of experience in Site Reliability Engineering with a strong focus on building, scaling, and operating large-scale distributed platform services, and Java applications.
  • Strong technical grasp of Open Source technologies designed for large-scale data processing.
  • Proven expertise in designing, analyzing, and troubleshooting complex distributed systems.
  • Proficiency in at least one modern programming or scripting language (Python, Java, Go, Bash, Ansible, or similar).
  • Practical experience designing and deploying end-to-end observability stacks (Prometheus, Grafana, Datadog, OpenTelemetry, ELK, etc.).
  • Demonstrated troubleshooting and problem-solving skills across production software and infrastructure environments.

Preferred Qualifications

  • In-depth understanding of SRE principles, including error budgeting, SLO/SLI/SLA definition, and advanced observability practices (Prometheus, Splunk, Grafana, OpenTelemetry).
  • Advanced programming skills in Java, Python, or Go, with hands-on experience across relational, NoSQL, or OLAP databases and event-driven streaming architectures (Kafka, RabbitMQ).
  • Track record of managing production on-call rotations, critical incident triage, root cause analysis (RCA), and post-incident reviews (PIR).
  • Solid knowledge of enterprise security standards, cryptography, authentication protocols (OAuth, SAML, SSO), and compliance governance.