AIML - Distinguished Engineer, Foundation Model

ApplePublished 1 days agoFirst seen 1 days ago

Summary

Apple is revolutionizing artificial intelligence by developing sophisticated
foundation models that power intelligent features across our product ecosystem.
We are seeking a Distinguished Engineer to set the technical direction for the
systems that power our foundation model training — with an initial focus on the
inference engine that the foundation model team relies on for training, model
evaluation, and other needs in the model development loop.

This is a senior individual-contributor leadership role. You will be one of the
most senior technical voices for foundation model systems at Apple: defining
the vision, driving execution across many teams, and raising the bar for
engineering excellence. The role starts with the inference engine, but we
expect you to move fluidly into adjacent training systems areas as the needs of
the foundation model program evolve.

Description

engineered specifically for Apple silicon and for experiences that are private,
personal, and deeply integrated into the OS. Behind that modeling work sits a
demanding systems layer, and the inference engine is at its center.

Our inference engine is used by the foundation model team throughout the model
development lifecycle: generating and processing data and running rollouts for
training, powering large-scale model evaluation, and serving as an LLM judge
that scores and compares model outputs. These workloads are high throughput,
bursty, and tightly coupled to research iteration — the speed, efficiency, and
reliability of the engine directly set the pace at which the team can train and
improve models.

As a Distinguished Engineer, you will own the technical strategy for this
inference engine and the broader systems that support it. You will partner
closely with modeling and research teams to bring new capabilities into the
development loop, work across many internal teams with very different
requirements, and lead a diverse set of engineers in turning an ambitious
vision into shipped milestones. While inference is the initial focus, you will
also help shape adjacent areas — training infrastructure, data systems, and
evaluation. If you are drawn to hard systems problems where the research and
the infrastructure are inseparable, this is the role.

Responsibilities

  • Set and drive the technical vision and roadmap for the foundation model
  • team's inference engine and the systems around it, used for training,
  • evaluation, and LLM-as-judge workloads.
  • Lead deep work on inference performance, efficiency, and reliability:
  • throughput and latency optimization, batching and scheduling, quantization,
  • speculative decoding, KV-cache management, memory and compute efficiency, and
  • hardware-aware optimization.
  • Architect inference systems that support a wide range of internal use cases —
  • data generation and rollouts for training, offline and large-scale evaluation,
  • and judge/reward scoring — across text, image, speech, and multi-modal models,
  • each with distinct throughput, cost, and quality constraints.
  • Extend your impact into adjacent systems areas — training infrastructure, data
  • pipelines, and evaluation harnesses.
  • Partner with many teams that depend on the engine, translating their diverse
  • needs into a coherent platform, clear interfaces, and a prioritized roadmap.
  • Work closely with ML researchers and modeling teams to co-design models and
  • systems, and to bring state-of-the-art techniques from prototype into the
  • development loop reliably.
  • Lead a diverse set of engineers across teams in setting direction and
  • executing against it; align stakeholders, resolve technical trade-offs, and
  • make the calls that keep large efforts moving.
  • Drive prioritization and milestone delivery across competing demands, balancing
  • near-term research needs against long-term platform investment.
  • Mentor and grow junior and senior engineers; establish engineering
  • standards, review designs, and multiply the impact of the organization.

Minimum Qualifications

  • MS or PhD in Computer Science, Machine Learning, or related technical field,
  • or equivalent industry experience.
  • 15+ years of experience building large-scale ML or distributed systems, with
  • a track record of technical leadership and industry-wide or company-wide
  • impact.
  • Deep, hands-on expertise in foundation model inference engines, with a proven
  • record of improving performance, efficiency, and reliability at scale.
  • Deep experience supporting a diverse set of foundation model inference use
  • cases, each with different throughput, latency, cost, and quality constraints.
  • Breadth beyond inference — the ability to contribute in adjacent systems areas
  • such as training infrastructure, data systems, or evaluation.
  • Deep understanding of GPU/TPU/accelerator architecture, distributed systems, and
  • model optimization (quantization, distillation, compilation, serving).
  • Proficiency with ML frameworks such as JAX, PyTorch, and with
  • inference/serving stacks.
  • Proven experience leading a diverse set of engineers in setting vision and
  • driving execution, including prioritization for milestone deliveries.
  • Demonstrated experience mentoring junior and senior engineers.
  • Demonstrated experience partnering with ML researchers and modeling teams to
  • productionize research.

Preferred Qualifications

  • Experience building or leading inference systems for large language models and
  • multi-modal foundation models at scale.
  • Experience with inference in training, evaluation, or reinforcement-learning
  • loops (e.g., large-scale rollouts, offline eval, or LLM-as-judge / reward
  • scoring).
  • Familiarity with Kubernetes, Docker, and cloud platforms (AWS, GCP, Azure),
  • and with distributed computing frameworks.
  • History of defining technical strategy that shaped an organization's or the
  • industry's direction.