Machine Learning Evaluation Engineer

AppleApplyPublished 1 days agoFirst seen 1 days ago
Apply

Summary

We are looking for a highly motivated Machine Learning Evaluation Engineer to join our team and help define and drive the evaluation of advanced machine learning and computer vision technologies. You will work closely with algorithm, data, and engineering teams to build scalable evaluation methodologies, uncover model weaknesses, and turn complex data into actionable insights.

You will play a key role in ensuring that our ML systems deliver high-quality, robust experiences across diverse real-world scenarios. The ideal candidate combines strong fundamentals in classical machine learning and computer vision with hands-on experience in metrics design, data analysis, failure analysis, visualization, and evaluation infrastructure. You should also be comfortable leveraging modern AI-assisted tools and workflows to improve engineering efficiency and accelerate analysis.

Description

In this role, you will define and drive the evaluation strategy for computer vision and machine learning algorithms used in complex product experiences. You will work closely with algorithm engineers to understand system behavior, identify the most meaningful quality signals, and develop evaluation frameworks that reflect real-world performance.

You will design metrics and evaluation methodologies that go beyond aggregate accuracy and help the team understand performance across important data slices and scenarios. You will analyze large-scale datasets to identify gaps in data quality and coverage and develop strategies to improve the representativeness of training and evaluation data.

A significant part of this role will involve failure analysis. You will investigate model failures, identify recurring patterns, develop failure taxonomies, and determine whether issues are driven by data, labeling, algorithm limitations, environmental conditions, or other system-level factors. You will build tools and visualizations that enable engineers to efficiently explore failures and understand the underlying root causes.

You will also develop scalable and automated workflows for evaluation, analysis, and reporting. You will be expected to leverage modern AI-assisted tools where appropriate to accelerate data analysis, visualization development, coding, and workflow automation while maintaining technical rigor and reproducibility.

As a member of the team, you will help shape evaluation best practices, influence algorithm and data decisions, and drive improvements across the ML development lifecycle. You should be comfortable navigating ambiguity, independently identifying opportunities for improvement, and partnering with cross-functional teams to deliver high-quality ML systems.

Minimum Qualifications

  • BS and a minimum of 3 years relevant industry experience
  • 3+ years of relevant industry experience in machine learning, computer vision, algorithm evaluation, data science, or a related field.
  • Strong background in machine learning and computer vision, including classical ML and statistical modeling techniques.
  • Proven experience designing and implementing evaluation methodologies and quality metrics for ML algorithms.
  • Strong expertise in failure analysis and root-cause analysis, with the ability to identify systematic model weaknesses and translate findings into actionable recommendations.
  • Experience analyzing data quality, diversity, representativeness, and coverage gaps across large and complex datasets.
  • Strong programming skills in Python with experience in data processing, analysis, and visualization.
  • Experience building automated and scalable evaluation pipelines and tooling.
  • Ability to develop clear and insightful visualizations and dashboards that communicate model performance, regressions, and failure patterns.
  • Strong understanding of statistical analysis, experimentation, and performance measurement.
  • Ability to independently drive ambiguous technical problems from problem definition through analysis and recommendations.
  • Experience using AI-powered development and analysis tools to improve productivity, accelerate data exploration, automate repetitive workflows, and improve engineering efficiency.

Preferred Qualifications

  • MS in Computer Science, Computer Engineering, Electrical Engineering, Statistics, Applied Mathematics, or a related technical field (Advanced degree is a plus)
  • Experience evaluating computer vision algorithms such as object detection, classification, tracking, segmentation, pose estimation, or hand tracking.
  • Experience with large-scale ML datasets and data pipelines.
  • Experience developing internal tools or platforms for ML evaluation and analysis.
  • Experience with deep learning frameworks and modern ML systems.
  • Experience with synthetic data, data augmentation, or automated data generation.
  • Familiarity with model monitoring, regression detection, and production ML quality systems.
  • Experience applying generative AI or AI-assisted workflows to engineering and analytical tasks.
  • Experience mentoring engineers or providing technical leadership for complex evaluation initiatives.
  • Excellent communication skills and the ability to collaborate effectively across algorithm, engineering, data, and product teams.