Software Engineer, Gemini Eval Infra, DeepMind

GoogleApplyPublished 2 hours agoFirst seen 1 hours ago
Apply

At Google DeepMind our mission is to build the world's first general-purpose learning agent. Central to this mission is the complex task of measuring the intelligence of our prototypes. As a Software Engineer, you will be working with the cutting edge AI agents developed by our exceptional team of Machine Learning and Neuroscience research scientists. Your responsibilities will include everything from creating systems for agent testing using 2D and 3D games to developing test problems within physics simulators. You will create graphical visualization of results, build competitive agent leaderboards and test new algorithms on robots. To succeed in this role you will need to have a strong foundation in software engineering and enjoy working on a wide range of challenging problems within a mission-driven team.

The mission of the GenAI Eval Platform team is to provide the high-velocity, and scalable evaluation backbone that powers the development and deployment of Google’s frontier AI models.

The team builds the evaluation platform used across Google DeepMind (GDM), in close collaboration with research teams. Evaluations are the steering wheel of AI progress: they determine model capability, safety, reasoning depth, and readiness for worldwide launch.

Artificial intelligence will be one of humanity’s most transformative inventions. At Google DeepMind, we are a pioneering AI lab with exceptional interdisciplinary teams focused on advancing AI development to solve complex global challenges and accelerate high-quality product innovation for billions of users. We use our technologies for widespread public benefit and scientific discovery, ensuring safety and ethics are always our highest priority.

We are pushing the boundaries across multiple domains. Our global teams offer diverse learning opportunities and varied career pathways for those driven to achieve exceptional results through collective effort.

Individual pay is determined by factors including job-related skills, experience, and relevant education or training.

US: $207000 - $300000 (USD) + 20% bonus target + equity + benefits

Learn more about benefits at Google.

Responsibilities

  • Design and optimize distributed evaluation execution engines capable of orchestrating large volumes of inference steps across TPU and GCU pools with high throughput and low latency.
  • Build foundational abstractions to evaluate complex Large Language Model (LLM) agent loops, tool use, and automated LLM-as-a-judge rating systems.
  • Design robust error classification, automated retry policies, and observability dashboards to maintain strict Service Level Objectives (SLOs) for evaluation pipeline success rates.
  • Partner closely with GDM research scientists and Data Science teams to anticipate frontier model evaluation requirements and translate them into elegant infrastructure solutions.
  • Mentor fellow engineers, set high standards for code quality (Python in Google3), and advocate testing and system design practices.

Minimum qualifications:

  • Bachelor's degree in Computer Science, Electrical Engineering, a related technical field, or equivalent practical experience.
  • 8 years of experience with software development or 5 years of experience with an advanced degree.
  • 5 years of experience with data structures and algorithms.
  • Experience with distributed systems, software architecture, and system design.

Preferred qualifications:

  • MBA or Master's degree in Computer Science, Software Engineering, or a related field.

  • 8 years of experience in software development, with a focus on large-scale distributed systems and data pipelines.