Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.
If you'd like to build the world's best AI cloud, join us.
*Note: This position requires presence in our Bellevue or San Francisco office location 4 days per week; Lambda's designated work from home day is currently Tuesday.
A customer runs distributed training across 64 to 1,024+ GPUs (Graphics Processing Units). When a job slows down, customers need one question answered instantly: is it my code, NCCL (NVIDIA Collective Communications Library) configuration issue, or a bad InfiniBand link? Observability is how Lambda answers that question. It lets customers distinguish their code, their configuration, and our platform in minutes, and that transparency is a core reason teams trust Lambda with their largest training runs. This role makes that transparency real.
You will be the product manager who owns observability as a product across Lambda's cloud. That means everything customers see: GPU and cluster health, utilization, job-level telemetry, and InfiniBand fabric metrics. It also means the platform telemetry layer underneath, which powers Lambda's internal fleet operations and reliability work. You will partner daily with SRE (Site Reliability Engineering), fleet engineering, and the console and API (Application Programming Interface) teams to bring one coherent observability experience to customers and operators alike. This role reports to the Head of Platform Product Management.
Great product managers at Lambda are defined by three things: insight, influence, and execution. Insight means you look at the data, determine what it means for customers and business, and then figure out what to do about it. But, a great idea doesn’t mean anything in a vacuum. That is where influence comes in. Influence means you take that idea and get others to want to buy into it; you win over engineers, designers, executives, and partners without relying on authority. But a great idea that everyone is excited about doesn’t matter unless it is delivered to customers. Execution means you work with the right people to get the idea launched, then measure and iterate. We hire product managers who learn new domains fast and reason rigorously from evidence. Deep observability domain experience is a strong plus, but insight, influence, and execution are the bar.
If you love turning raw telemetry into products customers rely on, and you want your work to be the reason a research team trusts their 1,024-GPU training run, we'd love to hear from you. We value diverse backgrounds, experiences, and skills, and we are excited to hear from candidates who can bring unique perspectives to our team. If you do not exactly meet this description but believe you may be a good fit, please still apply and help us understand your readiness for this role.