Senior Machine Learning Engineer
Summary
oin the Siri team at Apple! Build and contribute to a product and company that is building products, personal devices, and software designed to enrich people's lives. Work on building and advancing the world's most popular intelligent assistant that helps millions of people get things done — just by asking.
Global Siri works to bring Siri to the next level of intelligence and capability across all languages and markets. We build machine learning models, systems, and software that understand the intents of hundreds of millions of users and their billions of requests to Siri on Apple devices such as iPhone, iPad, Apple Watch, Mac, AirPods, HomePod, Vision Pro, and Apple TV. On the Global Siri team, we develop ML models and algorithms for both on-device and server-side applications, ultimately delivering product-critical models that surprise and delight customers around the world in the languages they speak. We build high-efficiency on-device Large Language Models and advanced parameter-efficient fine-tuning technologies that deliver personal, contextual, and highly private user experiences at scale. We are seeking an experienced and technically deep Machine Learning Engineer to lead the modeling, adaptation, and engineering optimization of on-device models that power features used by hundreds of millions of people worldwide.
Description
As a Machine Learning Engineer focusing on on-device LLMs and adapter technologies, you will tackle some of the most challenging problems in modern applied machine learning: delivering state-of-the-art agentic capabilities, reasoning, and task completion under strict on-device compute, memory, latency, and context constraints.
You will lead the end-to-end lifecycle of on-device LLM adapters and sub-models—ranging from training recipe formulation, data curation, and architecture tuning to rigorous metric evaluation, error analysis, and on-device runtime optimization. You will work closely with cross-functional teams including Foundation Model teams and Product teams to scale Apple Intelligence across diverse languages and domains.
Responsibilities
- Model Training & Adaptation: Design, train, and fine-tune high-performance on-device LLMs and modular adapter architectures (e.g., LoRA, rank scaling, routing) for specific Siri and Apple Intelligence tasks.
- Context & Efficiency Optimization: Innovate and implement techniques for prompt/context compression, KV cache optimization, tool-schema compaction, and efficient token representations to fit strict context-window limits.
- Evaluation & Metric Hillclimbing: Build and maintain rigorous offline and online evaluation pipelines; design benchmarks for agentic tool use, goal completion, and multi-turn conversational accuracy; conduct deep failure analysis to drive continuous model hillclimbing.
- Production Deployment & Stability: Partner with platform engineers to land models and adapters into production builds, validate performance across OTA updates, and monitor runtime behavior at scale.
Minimum Qualifications
- 5+ years of hands-on industry experience building, training, and shipping production-grade deep learning and NLP systems at scale.
- Deep Hands-on LLM Experience: Proven track record in training and fine-tuning Large Language Models (e.g., SFT, DPO/RLHF, PEFT/LoRA). Strong intuitive understanding of model behavior, loss dynamics, and scaling laws.
- Advanced Engineering Skills: Exceptional coding skills in Python and proficiency with modern deep learning frameworks (PyTorch, DeepSpeed, Megatron-LM, Hugging Face, vLLM). Familiarity with high-performance C++ or Swift runtime integration is a plus.
- Data & Evaluation Rigor: Solid experience designing domain-specific datasets, synthetic data generation pipelines, automated evaluation harnesses, and conducting systematic root-cause failure analysis.
- Collaboration & Execution: Strong ownership, cross-functional communication, and project management skills; ability to navigate complex engineering constraints and deliver against tight product deadlines.
Preferred Qualifications
- On-Device / Edge Deployment Experience: Deep understanding of edge-AI constraints and deployment pipelines (e.g., CoreML, ONNX, TensorRT-LLM, Apple Neural Engine / Metal optimization).
- Agentic Modeling & Tool-Calling: Direct experience building agentic workflows, multi-turn reasoning, and function/tool-calling capabilities in compact/dense models (e.g., 2B–9B parameter class).
- Proven research track record via publications in top-tier conferences (NeurIPS, ICML, ICLR, ACL, EMNLP) or demonstrated history of shipping industry-leading GenAI products.
Apple is an equal opportunity employer that is committed to inclusion and diversity, and thus we treat all applicants fairly and equally. Apple is committed to working with and providing reasonable accommodation to applicants with physical and mental disabilities.