Cloud Hardware Development Manager, AWS Gen AI & ML Servers

Amazon Web ServicesApplyPublished 21 hours agoFirst seen 1 hours ago
Apply
AWS operates the world's largest fleet of GPU-accelerated servers powering AI/ML training and inference at cloud scale. Our team defines the server architectures, drives the hardware designs, and owns the fleet quality for these platforms — from component selection through datacenter operations. If you want to shape the physical hardware that frontier models train on, this is the role.

We are seeking a Cloud Hardware Development Manager to lead a team of hardware, systems development engineers, and technical program managers responsible for the design, validation, manufacturing, and fleet operations of GPU-accelerated server platforms. You will set technical direction for your team, manage ODM and silicon supplier partnerships, drive hardware programs from concept through production ramp, and own fleet reliability metrics. This role requires deep hardware expertise to guide technical decisions, combined with people leadership to hire, develop, and retain a high-performing engineering team.

What You Will Do
You will lead hardware engineers who define the platforms running the world's largest AI workloads. You will set technical strategy for your team's hardware portfolio working alongside product and leadership teams, making trade-offs between schedule, cost, reliability, and performance. You will guide your engineers through ambiguous design problems while removing blockers and ensuring delivery. You will own the end-to-end lifecycle of your team's hardware, from architecture definition through fleet operations, and be accountable for fleet quality metrics including annualized failure rates and system availability.

Why You Will Love It
The world's most advanced frontier models train on the hardware your team designs. You will build and lead the engineers who define GPU server platforms at unprecedented scale. Your decisions on architecture, components, and quality directly impact fleet reliability for customers within weeks of deployment. The team is deeply technical and high-trust, with direct access to senior leadership and the freedom to set technical direction.

The Ideal Candidate
You combine deep hardware expertise with strong people leadership. You have designed or led teams that delivered server hardware at scale, and you can still dive into a schematic review or signal integrity issue when needed. You set high standards for your team and your partners, make data-driven decisions, and communicate clearly to both engineers and executives. You build teams that are stronger because of your presence but don't require it to be successful.

Key job responsibilities
Technical Leadership & Strategy
  • Set technical direction for your team's hardware portfolio across GPU-accelerated server platforms, making architecture and component trade-offs aligned with customer requirements and business goals
  • Guide system-level design decisions across thermal, mechanical, power delivery, signal integrity, and accelerator subsystems — including trade-offs on cooling architecture (liquid vs. air), power budget allocation, and PCIe/interconnect topology
  • Drive design reviews, qualification gates, and go/no-go decisions with deep technical judgment
  • Own fleet reliability and availability for your team's hardware; drive continuous improvement to reduce failure rates
People Leadership & Development
  • Hire, develop, and retain a team of hardware and system software engineers; build a high-performing team culture focused on engineering excellence and customer obsession
  • Set clear goals, provide regular feedback, and actively coach engineers through career growth as part of 1:1s
  • Perform calibrations and promotion assessments; ensure your team's technical bar remains high
  • Create an inclusive team environment where engineers can do their best work
Program Delivery & ODM Management
  • Drive hardware programs from concept through manufacturing and fleet deployment, ensuring milestones are met across NPI phases (Program Initiation, Design, Qualification, Pilot, Post-GA)
  • Lead ODM and silicon supplier partnerships: establish technical standards, drive design reviews, manage EVT/DVT/PVT builds, and hold partners accountable for quality and schedule
  • Identify program risks early, escalate with data and proposed mitigations, and unblock cross-functional dependencies
Fleet Operations & Quality
  • Own operational metrics for your team's server platforms: annualized failure rates, system availability, manufacturing escape rates
  • Drive root cause analysis of fleet-wide failures and ensure corrective actions flow back into design requirements and qualification criteria for future platforms
  • Establish closed-loop feedback systems connecting field failure data to upstream design and manufacturing improvements
Cross-Functional Alignment
  • Align with EC2 architecture teams on instance requirements, workload characterization, and platform roadmaps
  • Partner with firmware, software, test automation, and datacenter operations teams to ensure hardware is debuggable, serviceable, and automation-ready
  • Communicate technical strategy, program status, and risk posture to senior leadership
May require occasional (

A day in the life
You start the day in a 1:1 with one of your engineers, coaching them through a design trade-off on power delivery for a next-gen accelerator platform. Mid-morning, you lead a design review with your ODM partner, driving closure on open signal integrity items from the DVT build. In the afternoon, you review fleet reliability data with your team, identifying a component failure trend and aligning on corrective actions. You end the day in a planning session with senior leadership, presenting your team's hardware roadmap and resource needs for the next platform generation.

About the team
The Hardware Engineering AI/ML Ultraserver platform team is a group of engineers and technical program managers directly responsible for launching and maintaining GPU-accelerated servers in the AWS fleet. Located in Seattle, Cupertino, and Austin, we work with internal engineering teams, ODMs, and design partners to deliver next-generation AI/ML infrastructure deployed in datacenters worldwide. We move fast with small, empowered teams.

Basic Qualifications

  • Bachelor's degree or above in Electrical or Mechanical Engineering, or Bachelor's degree in computer science, engineering, mathematics or equivalent
  • 7+ years of hardware development experience for server, compute, networking, or storage platforms
  • 1+ years of experience managing hardware engineering teams
  • Experience driving hardware development programs (servers, racks, networking, or storage) through full product lifecycle including design, validation, manufacturing, and deployment
  • Experience working with ODMs or manufacturing partners through product development and production
  • Experience in one or more server technologies: thermal/mechanical design, power delivery, high-speed signal integrity, reliability, or accelerator subsystems

Preferred Qualifications

  • Master's degree or above in Electrical Engineering, Mechanical Engineering, or a related field
  • 5+ years of experience managing multi-discipline hardware development teams delivering products at datacenter scale
  • Experience leading ODM and silicon supplier partnerships across multiple geographies, establishing technical standards and driving quality frameworks
  • Track record of owning fleet quality metrics (annualized failure rates, system availability) and driving design improvements based on operational data
  • Knowledge of datacenter infrastructure constraints including networking, power, space, and cooling
  • Experience with GPU/accelerator server platforms, NPI processes (EVT, DVT, PVT), or hardware qualification programs
  • Demonstrated ability to communicate technical strategy and influence senior executives through written narratives
  • Track record of hiring, developing, and retaining high-performing hardware engineering talent
Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company’s reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.


USA, CA, Cupertino - 201,300.00 - 272,400.00 USD annually
USA, TX, Austin - 175,100.00 - 236,900.00 USD annually
USA, WA, Seattle - 175,100.00 - 236,900.00 USD annually