Staff Machine Learning Engineer, Siri Attention and Invocation
Summary
As part of Siri Attention and Invocation, we collaborate to deliver the next revolution in human-computer interaction, to inspire and create groundbreaking technology for large-scale systems spanning speech, vision, and generative AI to overcome real-world challenges through innovation and user-centered design that improves the daily experience of millions of our customers.
Description
We are seeking an exceptional Staff Machine Learning Engineer to lead the development of audio and video generation capabilities that bring conversational agents to life. In this role, you will drive the technical vision for generating realistic, expressive synthetic speech and visual representations, mentor senior and junior engineers, and shape the roadmap for multimodal generative experiences across our products.
Responsibilities
- Design and implement end-to-end systems for generating synthetic audio (speech, acoustics, sound) and/or video for agents interaction
- Lead complex generative AI projects from research prototype through large-scale production deployment
- Establish best practices for model evaluation, perceptual quality assessment, and monitoring of generative outputs
- Drive technical decisions and architecture for multimodal (audio/video) generation systems, balancing quality, latency, and scalability
- Identify high-impact opportunities where advances in generative modeling can meaningfully improve agents experiences
- Collaborate closely with research, product, design, and infrastructure teams to translate cutting-edge techniques into shipped features
- Mentor and provide technical guidance to engineers across the team, raising the bar for ML engineering practices
- Stay current with emerging techniques in generative modeling (e.g., diffusion, autoregressive, end-to-end architectures) and evaluate their applicability to our systems
Minimum Qualifications
- Master’s or PhD in Computer Science, Electrical Engineering, Machine Learning, or a related field, or equivalent practical experience
- Deep hands-on experience with generative audio and/or video architectures (e.g., diffusion models, autoregressive models, GANs, VAEs, end-to-end neural synthesis)
- Demonstrated ability to lead complex, ambiguous projects from research through production, and to make sound technical tradeoffs under real-world constraints (quality, latency, compute)
- Experience evaluating generative model outputs, including both objective metrics and perceptual/subjective quality assessment
Preferred Qualifications
- Proven experience building and shipping machine learning systems in production, with significant focus on generative modeling
- Excellent collaboration and communication skills, with a track record of working across research, engineering, and product teams
- Strong software engineering skills, with experience designing scalable ML systems and pipelines (e.g., Python, PyTorch/TensorFlow, distributed training infrastructure)
- Experience with speech synthesis (TTS), voice conversion, audio acoustics/background modeling, or conversational AI systems
- Experience with generative video/animation techniques (e.g., facial animation, lip-sync, avatar rendering, video diffusion)
- Publications in generative modeling, speech, audio, or computer vision at top-tier venues (e.g., NeurIPS, ICML, ICASSP, CVPR, Interspeech)
- Experience deploying real-time or low-latency generative models at scale
- Familiarity with multimodal modeling (joint audio-visual generation, cross-modal conditioning)
- Prior experience mentoring engineers or leading technical direction for a team