Software Engineer
Summary
Apple is where individual imaginations gather together, committing to the values that lead to great work. Every new product we build, service we create, or Apple Store experience we deliver is the result of us making each other’s ideas stronger. That happens because every one of us shares a belief that we can make something wonderful and share it with the world, changing lives for the better. It’s the diversity of our people and their thinking that inspires the innovation that runs through everything we do. When we bring everybody in, we can do the best work of our lives. Here, you’ll do more than join something — you’ll add something.
Our Operations Team, part of the CS&E organization, plays a critical role in ensuring the stability, reliability, and performance of multiple customer-facing applications that serve millions of users across diverse services. We specialize in sustaining and optimizing our systems, driving operational excellence through automation, cloud infrastructure management, AI-driven insights, and proactive monitoring. Our team thrives on solving challenging problems around Traffic, Deployments, DDoS, WAF, enhancing system resilience, and enabling seamless experiences for our customers. By collaborating closely with development, security, and business teams, we ensure that every application we manage meets Apple's standards of Site Reliability and Security .
Description
We are seeking a skilled DevOps Engineer who is passionate about operational excellence through automation, AI-driven engineering processes, and strong problem-solving abilities. As a DevOps Engineer, you will play a crucial role in ensuring the seamless integration of development and operations processes to deliver best-in-class and highly available systems. Your expertise in cloud platforms, automation, infrastructure management, and applying LLM/AI capabilities to solve real operational challenges will be vital to the success of our projects. We're looking for someone who thrives on breaking down complex, ambiguous problems and engineering practical, scalable solutions — not just executing tasks, but thinking critically about root causes and long-term reliability.
Responsibilities
- Linux, Cloud Infrastructure & Migration: Deep operational expertise across Linux internals (process/memory/CPU diagnostics, networking, kernel-level troubleshooting) combined with ownership of AWS/Cloud infrastructure — architecture, provisioning, disaster recovery , and end-to-end cloud migration (on-prem to AWS), using Terraform/Ansible for infrastructure-as-code at scale.
- Kubernetes/EKS Lifecycle & Container Orchestration: Manage the full lifecycle of Kubernetes/EKS environments — version upgrades, live/rolling patching with zero downtime, node maintenance, cluster optimization, and Helm chart development/troubleshooting.
- Traffic, Network & Security Engineering: Configure and manage NGINX (reverse proxy , load balancing, SSL/TLS, ingress) and GSLB for traffic distribution/failover; implement Shield/WAF-based DDoS protection, threat modeling, and security hardening across the SDLC, including PCI/SOC2/SOX compliance readiness.
- CI/CD, Monitoring & Incident Response: Design and maintain CI/CD pipelines (Jenkins, Spinnaker, GitHub Actions); build actionable monitoring/alerting frameworks (Splunk, Dynatrace, Prometheus, Grafana); participate in production on- call and drive thorough root-cause analysis for P1/P1c incidents.
- AI/LLM-Driven Intelligent Automation: Apply LLM/GenAI-based approaches to automate incident triage, log/anomaly analysis, root cause analysis, runbook generation, and self-healing systems; continuously evaluate and integrate emerging AI/LLM tooling into CI/CD and operations workflows to reduce manual toil.
- Engineering Process Ownership: Design and implement engineering processes that simplify how development teams build, deploy , and operate software, embedding security , compliance, and automation best practices throughout.
Minimum Qualifications
- 5+ years of proven experience as a Senior Ops/SRE/DevOps Engineer or similar role, with demonstrated ownership of production-grade systems.
- Strong Linux systems expertise, hands-on AWS experience, and proficiency managing Kubernetes/EKS (upgrades, patching, Helm).
- Strong programming/scripting skills (Bash, Python, Node, or Java) with proficiency in CI/CD (Jenkins, GitHub Actions, etc.) and observability tooling (Prometheus, Grafana, Splunk).
- Bachelor's Degree in CS or equivalent practical experience.
Preferred Qualifications
- Hands-on experience with LLMs/AI/ML technologies, GenAI tools, agentic frameworks, or AI-driven automation in production environments.
- Experience with NGINX configuration and Shield/WAF/DDoS mitigation strategies at scale.
- Strong problem-solving mindset with demonstrated ability to diagnose complex, ambiguous issues and design scalable, long-term solutions.
- Experience building or integrating LLM-powered agents/copilots for operations, monitoring, or incident response use cases.
- Experience with large-scale on-prem to cloud migrations.
- Familiarity with AEM and Apache upgrades/troubleshooting.
- Experience handling high-priority incidents (P1/P1c) with strong root-cause analysis discipline.