Engineering Project Manager, Edge Services

AppleApplyPublished 3 days agoFirst seen 3 days ago
Apply

Summary

Imagine what you could do here. At Apple, new ideas have a way of becoming extraordinary products, services, and customer experiences very quickly. We are looking for a great service reliability engineer to design, build and operate tools and automated systems to help run the edge infrastructure that delivers Apple's services to customers across China. If you are passionate and curious in how internet and DNS works at a fundamental level and like operating systems at scale, this is the role for you.

Description

You will be in charge of designing and building the infrastructure to support services at Apple scale. Day to day, you will design and tune resolution paths, traffic-steering and failover policies, and capacity; automate operations; harden the platform against load spikes and attacks; and lead the response when incidents occur. You have a passion for automation and making resilient systems that scale. The ideal candidate has a strong knowledge of internet protocols foundation, solid Linux skills, cloud and compute experience, working familiarity with proxies and load balancers, along with strong coding ability in Go and scripting languages .

Minimum Qualifications

  • BS in Computer Science or a related field, or equivalent job-related experience
  • Linux systems administration experience
  • 5+ years operating production infrastructure at scale at a SRE level
  • Hands-on experience operating DNS in production (authoritative and/or recursive)
  • Proficiency in a programming or scripting language such as Go or Python
  • Experience operating systems in cloud platforms
  • Fluent English and Mandarin

Preferred Qualifications

  • Experience with continuous / rapid release engineering
  • Strong tooling and automations development experience
  • Experience working in a 24/7/365 service environment
  • Deep Linux and networking experience: kernel and network-stack performance tuning, packet-level troubleshooting, and operating networks at scale
  • Deep experience running DNS and GSLB at large scale: authoritative and recursive DNS platforms
  • Experience designing for high availability and disaster recovery across regions
  • Depth in cloud platforms and compute, with infrastructure-as-code and CI/CD
  • A track record building SRE practices: SLOs, error budgets, and blameless incident response
  • Experience with AliCloud will be a plus