Site Reliability Engineer / Devops — Retail Engineering
Summary
Apple’s IS&T Retail Engineering team is seeking a SDET to own quality across our retail ecosystem — spanning eCommerce backend services, SAP integrations, and cross-functional end-to-end workflows. This is a hands-on technical leadership role: you’ll architect automation frameworks, lead testing strategy across multiple concurrent projects, and drive quality standards organization-wide.
You’ll work at the intersection of backend development, test automation, and program coordination — ensuring that features shipping to apple.com/shop and supporting retail systems meet Apple’s bar for reliability, performance, and customer experience.
Description
You will drive the reliability, deployment, and scalability of compute platforms across on-premises and hybrid cloud environments. Collaborating closely with cross-functional technical and business partners, you will build Infrastructure as Code, optimize container orchestration, and streamline CI/CD delivery pipelines. You will champion automation and operational excellence, ensuring high availability, robust security standards, and proactive observability across large-scale distributed systems.
Responsibilities
- Deploy, support and monitor compute platforms and application stacks.
- Ability to understand complex systems and a desire to constantly make things better
- Explore and evaluate new technologies and solutions.
- Strong interpersonal skills and ability to work effectively across multiple business and technical teams
- Develop and fix the application on failures.
- We promote innovation and use of new technology to further improve our creative output. We’re looking for a talented and passionate person to join this amazing team, if you feel this is you, we’d love to hear from you.
Minimum Qualifications
- Bachelor’s degree in Computer Science or equivalent field with 7+ years of experience, or Master’s degree with 5+ years of experience.
- 7+ years of experience in Site Reliability Engineering with a strong focus on building, scaling, and operating large-scale distributed platform services, and Java applications.
- Strong technical grasp of Open Source technologies designed for large-scale data processing.
- Proven expertise in designing, analyzing, and troubleshooting complex distributed systems.
- Proficiency in at least one modern programming or scripting language (Python, Java, Go, Bash, Ansible, or similar).
- Practical experience designing and deploying end-to-end observability stacks (Prometheus, Grafana, Datadog, OpenTelemetry, ELK, etc.).
- Demonstrated troubleshooting and problem-solving skills across production software and infrastructure environments.
Preferred Qualifications
- In-depth understanding of SRE principles, including error budgeting, SLO/SLI/SLA definition, and advanced observability practices (Prometheus, Splunk, Grafana, OpenTelemetry).
- Advanced programming skills in Java, Python, or Go, with hands-on experience across relational, NoSQL, or OLAP databases and event-driven streaming architectures (Kafka, RabbitMQ).
- Track record of managing production on-call rotations, critical incident triage, root cause analysis (RCA), and post-incident reviews (PIR).
- Solid knowledge of enterprise security standards, cryptography, authentication protocols (OAuth, SAML, SSO), and compliance governance.