Research Engineer, Advancing Agent Quality, DeepMind

GooglePublished 1 days agoFirst seen 1 days ago

At Google, research-focused Software Engineers are embedded throughout the company, allowing them to setup large-scale tests and deploy promising ideas quickly and broadly. Ideas may come from internal projects as well as from collaborations with research programs at partner universities and technical institutes all over the world.

From creating experiments and prototyping implementations to designing new architectures, engineers work on real-world problems including artificial intelligence, data mining, natural language processing, hardware and software performance analysis, improving compilers for mobile platforms, as well as core search and much more. But you stay connected to your research roots as an active contributor to the wider research community by partnering with universities and publishing papers.

The Agent Quality team defines, measures, and advances long-horizon, tool-using capabilities for Google’s frontier agents. We build foundational evaluation systems and data flywheels for Gemini Spark—Google’s flagship agent for deep research, multi-step problem solving, and workspace integration. Our work encompasses pioneering environment creation (including contexts, autoraters, human evaluations), developing advanced diagnostic tools for agent improvement, and harvesting trajectories and reward signals into SFT and RL post-training to recursively improve core Gemini models.

Artificial intelligence will be one of humanity’s most transformative inventions. At Google DeepMind, we are a pioneering AI lab with exceptional interdisciplinary teams focused on advancing AI development to solve complex global challenges and accelerate high-quality product innovation for billions of users. We use our technologies for widespread public benefit and scientific discovery, ensuring safety and ethics are always our highest priority.

We are pushing the boundaries across multiple domains. Our global teams offer diverse learning opportunities and varied career pathways for those driven to achieve exceptional results through collective effort.

Individual pay is determined by factors including job-related skills, experience, and relevant education or training.

US: $174000 - $252000 (USD) + 15% bonus target + equity + benefits

Learn more about benefits at Google.

Responsibilities

  • Design, build, and scale realistic agent environments and task suites. Research and develop Agentic AutoRaters (AR), calibrate them against human evaluation, and benchmark agent capabilities against industry-leading frontier models.
  • Enable agents to automatically generate and iterate on verifiers (such as automated test cases), facilitating effective exploration and iterative problem-solving.
  • Develop trajectory analysis frameworks and diagnostic tooling to identify root-cause agent failure modes (e.g., passivity, hallucination, or brittle tool execution), and automatically optimize agent harnesses to drive continuous self-improvement.
  • Harvest complex multi-turn interaction trajectories into high-quality datasets and reward signals to power SFT and RL flywheels for frontier Gemini models.

Minimum qualifications:

  • 2 years of experience with agentic AI workflows, frameworks, and approaches.
  • 1 year of experience with machine learning and deep learning.
  • Experienece in software engineering and cloud-based development.
  • Experience with Python, TensorFlow, PyTorch, or similar ML frameworks.

Preferred qualifications:

  • In-depth knowledge of machine learning algorithms, including supervised learning and reinforcement learning.