Software Engineer (Reliability Engineering), IS&T Ai & Data Platforms

AppleApplyPublished 1 days agoFirst seen 1 days ago
Apply

Summary

AI & Data Platforms (AiDP) is IS&T's engine for AI-powered innovation. The team brings together data, application development, and machine learning — including generative AI — along with data services and customer success functions, to help IS&T build solutions more efficiently and streamline the adoption and embedding of generative AI across Apple.

Do you want to help build some of the largest and most consequential enterprise and customer technology systems in the world? Join Apple’s Information Systems and Technology (IS&T) organization. IS&T is the engine behind everything Apple does for customers and for the people who build for them. It’s Apple’s central nervous system. Supporting 2.5 billion active Apple devices, processing billions of secure transactions, and keeping the technology that defines modern life running flawlessly, IS&T makes the impossible feel effortless.

Do you love building solutions to handle global complexity and immense scale? Imagine what you could do here.

AI & Data Platforms (AiDP) is IS&T’s engine for AI-powered innovation. The team brings together data, application development, and machine learning — including generative AI — along with data services and customer success functions, to help IS&T build solutions more efficiently and streamline the adoption and embedding of generative AI across Apple.

Description

We are looking for a talented engineer to join our team and bring passion for building and operating large scale platform and distributed systems leveraging cutting edge open source technologies across hybrid cloud environments.As a software engineer in AiDP reliability engineering you will work on one or many projects related to GenAI, ML, Inference and Big data platform.

Responsibilities

  • Build, enhance, and maintain multi-tenant systems leveraging diverse technologies.
  • Collaborate with cross-functional teams to deliver impactful customer features.
  • Lead projects through full lifecycle, from design discussions to release delivery.
  • Operate, scale, and optimize high-throughput and highly concurrent services.
  • Diagnose, resolve, and prevent production and operational challenges.

Minimum Qualifications

  • BS/MS in computer science or equivalent experience.
  • 5+ years experience programming skills in one of the following areas: Python, Java, or Go.
  • 5+ years experience in Kubernetes, Docker or other container orchestration framework.

Preferred Qualifications

  • Ability to read and explain open source codebase.
  • Experience deploying and managing CI/CD pipelines.
  • Strong expertise in troubleshooting complex production issues.
  • Should be able to understand complex architectures and be comfortable working with multiple teams.
  • Ability to conduct performance analysis and troubleshoot large scale distributed systems.
  • Should be highly proactive with a keen focus on improving uptime/availability of our mission-critical services.
  • Experience with big data technologies - Spark, Flink, Iceberg or emerging GenAI/ML like Ray/MLflow/model serving) technologies.
  • Experience of Linux, database and security concepts.