Sr. Machine Learning Engineer, ML Systems Evaluation Engineering
Summary
Join the team redefining what a deeply personal and integrated assistant can be.
As part of the Siri organization, you will help shape one of the world's most widely used AI assistants, powered by our next-generation of Apple Intelligence, with capabilities like personal context understanding and on-screen awareness, built with privacy from the ground up. Your work will have direct, meaningful impact for users across iOS, iPadOS, macOS, watchOS, and visionOS.
This is a rare opportunity to build at the intersection of cutting-edge AI and human-centered design, shipping technology that is centered around users and their needs.
Join the SCALE (Statistical Coverage and Large-scale Evaluation) team at Apple and contribute to a highly accomplished group that evaluates Siri and AI/ML models at scale to delight and inspire our users globally. If you are excited by large-scale systems, rigorous statistical evaluation, natural language understanding, and shaping the future of AI, we want you here! You’ll do more than join something, you’ll add something!
Description
We are seeking a highly skilled Senior Machine Learning Engineer specializing in Conversational AI. Our goal is to deliver offline evaluation insights that drive model development and improve the end-user experience, all while upholding Apple's strict privacy standards.
In this pivotal role, you will collaborate with cross-functional teams to curate and evolve high-quality evaluation datasets for state-of-the-art models. With the advent of Apple Intelligence, you will tackle novel challenges in evaluating highly personalized user experiences.
Responsibilities
- Your key focus will be leveraging Large Language Models (LLMs) to automatically evaluate model changes (LLM-as-a-Judge) and assess the naturalness of digital assistant conversations.
- You will also harness generative AI to create adversarial scenarios that anticipate edge cases, enabling robust, forward-looking evaluation.
- Your work, and the datasets you generate, will directly shape the next generation of Siri and Apple products.
Minimum Qualifications
- 7+ years of professional experience applying machine learning to real-world problems and crafting scalable data solutions, specifically in natural language products.
- Proven experience managing large-scale datasets for ML training and/or evaluation.
- Excellent programming skills in Python
- MS/PhD in Machine Learning, Computer Science, or equivalent experience in a related field
Preferred Qualifications
- The qualifications that will benefit a candidate to be successful in your role. Please list no more than 6-8.
- Deep domain knowledge in Conversational AI and a strong understanding of the end-to-end ML product lifecycle.
- Expertise in defining and measuring evaluation coverage for Large Language Models (LLMs) and agentic systems.
- Track record of delivering large-scale, cross-functional ML product or platform outcomes.
- Excellent problem-solving, critical thinking, and communication skills to drive alignment across teams.
- Experience with systems engineering; in-depth understanding of interdependencies of ML and SW components