Research Evaluation Lead

Rhoda AIApplyPublished 1 days agoFirst seen 1 days ago
Apply

At Rhoda AI, we’re building the next generation of generalist intelligent robots. We own the full robotics stack from high-performance hardware and robot systems to the infrastructure and state-of-the-art foundation world models that control our robots. Our robots are designed to be generalists capable of operating in complex, real-world environments and handling long-tail edge cases, made possible by our cutting edge research and end-to-end system design. We've raised over $450M and are investing aggressively in model research, infrastructure, hardware development, and manufacturing scale-up to make generalist robotics a reality.

Mission

Own end-to-end robot model evaluation for Research. Turn research questions into consistent, high-quality, repeatable evals that enable fast iteration and trusted results.

Responsibilities

  • Translate research intent into clear eval protocols, trial plans, and success criteria.
  • Own execution end-to-end: model handoff → station readiness → pilot execution → QA → results.
  • Train and manage eval pilots; ensure consistency across people, shifts, and stations.
  • Distinguish model failures from hardware, setup, operator, or data-quality issues.
  • Maintain eval setups, resets, randomization, metadata, and experiment traceability.
  • Track quality, throughput, and bottlenecks; continuously improve the eval process.
  • Partner closely with Research, Robot Data, and Eval Platform teams.

What we’re looking for

  • Some understanding of robotics / ML experimentation.
  • Computer science background or hands-on experience with coding
  • Rigorous, detail-oriented, and able to understand the intent behind an experiment, not just execute instructions.
  • Strong hands-on execution and ownership.
  • Experience in robotics testing, data collection, lab operations, or QA preferred.

Success looks like

A researcher can hand off a model and research question and receive a trusted, standardized eval result with sufficient trials and QA.