Software Development Engineer in Test, Siri AI QE
Summary
Join the team redefining what a deeply personal and integrated assistant can be.
As part of the Siri organization, you will help shape one of the world's most widely used AI assistants, powered by our next-generation of Apple Intelligence. Your work spans personal context understanding, on-screen awareness, and system-level actions built with privacy at the core. In this role, you will hold direct responsibility for shipping features with uncompromised quality across iOS, iPadOS, macOS, watchOS, and visionOS.
We are looking for a top-tier Software Development Engineer in Test who pairs strong software engineering and test architecture fundamentals with cutting-edge expertise in agentic AI. If you are an organized, self-directed engineer who thrives as an end-to-end Feature Quality DRI—knowing exactly when to write robust deterministic code and when to deploy agentic harnesses—this is the opportunity to make an outsized impact.
Description
In Siri AI Quality Engineering, you will ensure Siri delivers the right experience in the right context—across all Apple devices, system configurations, and real-world conditions.
As a Software Development Engineer in Test in Siri AI Quality Engineering, you will serve as a Feature Quality Directly Responsible Individual (DRI), shaping end-to-end quality strategy from early feature design stages through public release.
In this role, you will architect comprehensive validation strategies that balance robust, deterministic test automation with novel agentic test harnesses designed to evaluate non-deterministic, contextual AI interactions. You will exercise sound engineering judgment—knowing when to write clean, maintainable programmatic automation and when to deploy agentic workflows—while coordinating targeted manual validation to capture what automation cannot yet reach, systematically shrinking manual overhead every release cycle.
Responsibilities
- Feature Quality Ownership (DRI): Serve as the end-to-end Quality DRI for key Siri capabilities, shaping quality strategy, test architecture, risk assessment, and go/no-go shipping criteria across platforms from design to release.
- Pragmatic Automation & Harness Architecture: Design, build, and maintain robust test automation in Python and Swift/XCTest. Architect shared test harnesses and dynamic execution pipelines integrated into CI/CD.
- Agentic Test Solutions: Architect and deploy agentic testing solutions (tool-calling loops, multi-agent evaluation, dynamic state exploration) to validate complex, multi-turn, contextual Siri interactions that traditional scripts cannot adequately test.
- Engineering Judgment: Demonstrate clear discernment in tooling—knowing when to write deterministic, maintainable code versus when to leverage LLM-driven agents and evals.
- Real-Device & Ecosystem Validation: Oversee comprehensive testing across the physical Apple ecosystem (iPhone, iPad, Mac, Apple Watch, Apple Vision Pro), validating system states, sensor inputs, connectivity, and on-device intelligence.
- Triage & Cross-Functional Alignment: Systematically triage complex failure modes across ML models, OS frameworks, and client apps. Drive root-cause isolation, partner with engineering teams to land fixes, and report high-signal quality metrics to leadership.
Minimum Qualifications
- Bachelor's or Master's degree in Computer Science or a related field.
- 8+ years of experience in software development or test engineering, with a proven track record of technical leadership and high personal organization as a Feature DRI or Quality Lead.
- Strong programming fundamentals in Python and/or Swift (or Objective-C/C++), including experience developing modular test harnesses, CLI tools, and automated pipelines.
- Hands-on experience building or operationalizing GenAI / Agentic systems: practical experience with LLM orchestration, structured tool/function calling, prompt/context engineering, and agent execution harnesses.
- Deep understanding of test engineering: test planning, risk-based testing, parameterized test frameworks, dynamic test generation, and CI/CD automation.
- Experience with physical device testing and hardware/software integration: testing real consumer hardware across varied OS configurations, environments, and ecosystem interactions.
- Strong organizational and communication skills: structured in thought, highly self-directed in tracking dependencies, proactively driving cross-regional handoffs, and presenting clear quality signals to cross-functional stakeholders.
Preferred Qualifications
- Experience designing evaluation frameworks for non-deterministic AI systems: golden datasets, LLM-as-a-judge pipelines, rubric design, semantic similarity scoring, and regression/drift detection over time.
- Familiarity with Apple platforms and developer ecosystems (Xcode, XCTest, macOS/iOS system internals).
- Experience developing autonomous exploratory testing agents or autoresearch/eval loops where agents discover edge cases and surface actionable bugs.
- Strong background in system-level or OS-level feature validation (inter-process communication, background daemons, privacy boundaries, power/performance constraints).
Apple is an equal opportunity employer that is committed to inclusion and diversity, and thus we treat all applicants fairly and equally. Apple is committed to working with and providing reasonable accommodation to applicants with physical and mental disabilities.