Staff Product Manager - Compute

Lambda Labs•Published 3 hours ago•First seen 3 hours ago

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.

If you'd like to build the world's best AI cloud, join us.

Note: This position requires presence in our Bellevue or San Francisco office location 4 days per week; Lambda's designated work from home day is currently Tuesday.

About the Role

Compute is the product at Lambda. Customers come to us for graphics processing units (GPUs) they can get, hold, and run hard, and nearly everything else we sell depends on the compute layer working well.

As a product manager on the Compute team, you will own a meaningful part of defining the future of how we provide compute to our customers. The domain covers the shape of what customers rent, meaning instance families, sizes, and tenancy, across both GPUs and central processing units (CPUs). It covers how customers get capacity and hold onto it. It covers what they build production automation against, and how fast and predictably that behaves. It covers what happens once the workload is running, which means placement, performance isolation, hardware failure, maintenance, and the signals customers need to run their own operations. You will work across these areas depending on which customer needs are burning hottest.

Most of this is the feature set a mature compute cloud already has and Lambda does not have yet. A neocloud is not a hyperscaler, though, and copying that list wholesale is the wrong instinct. Part of the job is judgment about which of those capabilities matter here and which carry cost we should not pay. The other part is deciding where serving AI workloads well means building something the hyperscalers never needed.

Great product managers at Lambda are defined by three things: insight, influence, and execution. Insight means you look at the data, determine what it means for customers and business, and then figure out what to do about it. But, a great idea doesn't mean anything in a vacuum. That is where influence comes in. Influence means you take that idea and get others to want to buy into it; you win over engineers, designers, executives, and partners without relying on authority. But a great idea that everyone is excited about doesn't matter unless it is delivered to customers. Execution means you work with the right people to get the idea launched, then measure and iterate. We hire product managers who learn new domains fast and reason rigorously from evidence. Deep compute platform experience at a hyperscaler or neocloud, across both GPUs and CPUs, is highly desired.

If you have built compute primitives at a cloud provider and want to do it again somewhere the answers are not settled, we'd love to hear from you.

We value diverse backgrounds, experiences, and skills, and we are excited to hear from candidates who can bring unique perspectives to our team. If you do not exactly meet this description but believe you may be a good fit, please still apply and help us understand your readiness for this role.

What You'll Do

  • Find the Real Problem: Get close to customers, deals, escalations, and support on On-Demand GPU Instances and 1-Click Clusters, and work out what is actually blocking them rather than what they asked for.
  • Decide What Matters: Rank the work and sequence it, and be honest about what Lambda is not doing this year and why.
  • Write the Definition: Produce requirements, user stories, and acceptance criteria precise enough that engineering builds from them without a translation layer.
  • Price and Package It: Decide how your part of compute is priced, packaged, and committed to across NVIDIA GPU generations, and how customers compare the options.
  • Deliver With Engineering: Work through the build with the compute and control plane teams, make the tradeoff calls that come up mid-flight, and keep scope honest against the date.
  • Land the Launch: Set launch criteria that cover the operational readiness an infrastructure product needs, and get documentation, pricing, sales, and support in place before it goes live rather than after.
  • Measure and Iterate: Define what success looks like before launch, then go find out whether it happened, using adoption and utilization rather than opinion.
  • Make the Call Stick: Take a position on contested tradeoffs, write it down well enough that people can disagree with it precisely, and keep owning the decision after it is made.
  • Work as a Group: Partner with the other product managers on compute and with the teams that own storage, networking, orchestration, and commerce, so the pieces land as one product.

You

  • Have 7+ years of product management experience, including time on a compute platform at a hyperscaler, neocloud, or comparable cloud provider.
  • Have shipped compute primitives that external customers used at scale, like instance types, capacity products, provisioning interfaces, or placement and isolation controls.
  • Understand GPU and CPU infrastructure well enough to reason with engineers about launch paths, hardware failure modes, placement, and performance.
  • Have owned a product where day two behavior mattered as much as launch, including failure handling, maintenance, and what customers are told when something breaks.
  • Have owned pricing, packaging, or commitment terms for an infrastructure product.
  • Can tell which parts of a hyperscaler playbook transfer to an AI cloud and which are cost without benefit.
  • Can reason about utilization and capacity economics, and make the call when customer flexibility and hardware utilization pull against each other.
  • Able to define iterative plans that move an organization from the current state towards the desired outcome.
  • Can turn ambiguous customer, technical, and commercial inputs into product definition that multiple teams execute against.
  • Communicate plainly, write well, and make decisions easier for people who do not all share the same context.

Nice to Have

  • Experience with interruptible or preemptible capacity, or with reservations and committed use products.
  • Experience with the surfaces customers automate against, like public APIs, infrastructure as code providers, machine images, or snapshots.
  • Experience with hardware failure handling across a large hardware footprint, like spare pools, node replacement, or fault containment.
  • Experience with performance isolation and placement on shared hardware, including what a provider can honestly guarantee.
  • Experience with bare metal or dedicated host products, including isolation between tenants.
  • Experience with distributed training or large-scale inference workloads and what they demand of the compute layer.

Salary Range Information

The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.

About Lambda

  • Founded in 2012, with 500+ employees, and growing fast
  • Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove
  • We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG
  • Our values are publicly available: https://lambda.ai/careers
  • We offer generous cash & equity compensation
  • Health, dental, and vision coverage for you and your dependents
  • Wellness and commuter stipends for select roles
  • 401k Plan with 2% company match (USA employees)
  • Flexible paid time off plan that we all actually use

Equal Opportunity Employer

Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.

Compensation: USD 323,000 - 430,000 per year