Software Engineer, Systems & ML Infrastructure - MSL FAIR Foundations
Description
Meta is seeking Software Engineers to join the Frontier Evals Research team within Meta Superintelligence Labs. Evaluations are a critical part of AI progress at Meta Superintelligence Labs, determining what capabilities get built, which features get prioritized, and how quickly our models improve. As a Systems and ML Infrastructure Engineer on this team, you will build the platforms and services that enable reliable evaluation of our most advanced AI models across text, vision, audio, and beyond. You'll work alongside researchers and engineers to turn rapidly evolving research workflows into scalable, dependable infrastructure. This is a highly technical software engineering role focused on distributed systems, developer infrastructure, and production-grade ML platforms. You will design and own systems for scheduling and executing evaluation workloads, managing datasets and model artifacts, monitoring correctness and performance, and making results reproducible and easy to consume. The infrastructure you build will directly support research decisions and major model lines within MSL, making reliability, scalability, operational excellence, and engineering rigor paramount. You will succeed by moving quickly in an open-ended research environment while building durable systems, reducing operational toil, and creating abstractions that help researchers iterate faster. If you are passionate about building the technical foundation for frontier AI development and thrive in fast-paced, high-impact environments, we encourage you to apply.
Responsibilities
Design, build, and operate scalable infrastructure for running evaluations across large model fleets, datasets, modalities, and compute environments Develop orchestration, scheduling, data, and artifact-management systems that make the evaluation workflows reliable and reproducible Build APIs, abstractions, and developer tools that allow researchers to launch, debug, compare, and interpret evaluations efficiently Improve system reliability through testing, observability, capacity planning, performance optimization, and automated failure recovery Partner with research and engineering teams to translate new evaluation requirements into reusable platform capabilities
Qualifications
Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience 3+ years of software engineering experience in building backend, distributed, data, or machine learning infrastructure Proficiency in Python, C++, or another systems programming language Experience designing, implementing, and operating reliable services, platforms, or data-processing systems Experience independently delivering medium- to large-scale technical projects from design through production operation Demonstrated knowledge of software engineering practices, including testing, code review, observability, incident response, and performance analysis Ability to work effectively with researchers and engineers and to adapt to rapidly changing requirements Experience building infrastructure for large-scale machine learning training, inference, evaluation, or data processing Experience with distributed compute systems, workflow orchestration, containers, cluster schedulers, or cloud infrastructure Experience with performance profiling, resource efficiency, reliability engineering, and production observability Familiarity with language model post-training workflows, including supervised fine-tuning, reinforcement learning, evaluation, and inference, and the infrastructure needed to support them at scale Experience building internal platforms or developer tools used by multiple teams in fast-moving technical environments
Compensation: $154,003/year to $217,000/year + bonus + equity + benefits