Compute Orchestration & Scheduling
Microsoft AI is looking for engineers to build the compute infrastructure powering frontier-model development. The Compute Orchestration & Scheduling team owns cluster orchestration, workload scheduling, resource allocation, quota management, and the systems that enable fast initialization and reliable fault recovery on next-generation GPU supercomputers.
Our team prepares infrastructure for the next generation of distributed training and inference across multiple locations and a rapidly growing fleet of accelerators, including NVIDIA Grace Blackwell, Vera Rubin, and AMD GPU platforms, alongside the CPU, network, and storage systems they depend on.
Key challenges include topology-aware placement, large-scale distributed training, fast and reliable initialization of training runtimes, rapid recovery for long-running AI jobs, observability into cluster behavior, and improved utilization across increasingly large and diverse accelerator fleet. Better scheduling efficiency and platform reliability translate directly into more effective compute for AI research and products.
You will work closely with researchers, model engineers, hardware architects, and infrastructure teams to turn frontier-model requirements into scalable platform capabilities. We value engineers who navigate ambiguity, remove roadblocks, and deliver improvements to users quickly and iteratively.
Responsibilities
- Develop and tune the compute infrastructure stack allocating NVIDIA Grace Blackwell (GB) and Vera Rubin (VR) resources to AI workloads.
- Scale GPU clusters across hardware generations to thousands of accelerators and beyond.
- Use operational data and workload insights to inform the compute (GPU and CPU) roadmap for large-scale AI research.
- Partner with model-development teams to improve the infrastructure used to train and serve AI models.
- Find practical ways around roadblocks and deliver improvements rapidly, iterating with users in a fast-paced, design-driven environment.
- Embody Microsoft’s culture and values.
Qualifications
Required Qualifications
- A bachelor’s degree in computer science or a related technical field and at least six years of engineering experience writing code in languages such as C, C++, Python, Go, or JavaScript; or equivalent practical experience.
Preferred Qualifications
- Significant additional engineering experience, with a master’s degree or equivalent practical experience, building production software and distributed systems.
- Experience with Ray, Kubernetes, Kueue, Volcano, or another AI-focused system for scheduling, scaling, or fault tolerance is especially relevant.
This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.
Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.