Benchmarking Project Lead, Siri Evaluation

ApplePublished 1 days agoFirst seen 1 days ago

Summary

Join the team redefining what a deeply personal and integrated assistant can be. As part of the Siri organization, you will help shape one of the world's most widely used AI assistants, powered by our next-generation of Apple Intelligence, with capabilities like personal context understanding and on-screen awareness, built with privacy from the ground up. Your work will have direct, meaningful impact for users across iOS, iPadOS, macOS, watchOS, and visionOS. This is a rare opportunity to build at the intersection of cutting-edge AI and human-centered design, shipping technology that is centered around users and their needs.

Evaluation is at the heart of how we build our product. As Siri AI becomes more and more powerful and offers ever richer experiences to our users, our evaluations have to keep pace. We are seeking a senior manager to help drive our evaluation efforts. The role will involve managing teams working on evaluation development and data science, and leading high-impact initiatives to bring state-of-the-art agentic evaluation to the whole Siri team.

Description

As part of the work on next generation Siri, we are developing novel measurements of its quality. To ensure that the evaluation systems we are building are reliable, we plan benchmarking their accuracy on a wide range of features, locales, and platforms using humans in the loop.

Responsibilities

  • Design and execute efficient data collection processes using humans in the loop
  • Lead annotation efforts for various languages and devices
  • Plan and manage the budget of annotator resourcing with Finance, Annotation Ops, and International team partners.
  • Track and improve the quality of the human judgements through revision of annotator training materials, clear annotation questions, annotator training, and efficient reviewing mechanisms
  • Collaborate with other engineering teams to design and build a tooling ecosystem for managing and browsing rich datasets

Minimum Qualifications

  • Agentic Coding proficiency to achieve data-science, data collection and visualisation tasks
  • Good understanding of metrics, crowd science, data collection, annotation analysis, statistics
  • Ability to work independently and cross-functionally to integrate in partner team reporting systems and pipelines
  • Excellent communication skills and the ability to thrive in a highly collaborative work environment

Preferred Qualifications

  • Attunement to computational linguistics, language quality, human in the loop evaluation
  • Good engineering practices to create sustainable and easy to use data management pipelines
  • Python experience and other tools for data collection and visualisation