
Technical Product Manager - AI Compute Platform
NebiusAbout Nebius:
Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.
Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.
Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.
Our customers build the frontier of AI on top of Nebius — training state-of-the-art models, running production inference at scale, shipping the research and products that define where the field is going next.
We are building the AI cloud that the people building the frontier of AI choose deliberately — not on price, not on raw capacity, but on how it works to use it day to day. To do that, we are growing the AI Compute Platform product team and hiring multiple Technical Product Managers across the full surface of the platform.
Your scope will be defined by what you bring. We will match your technical strengths, customer experience, and product instincts to the area of the platform where you can have the most impact. The platform is broad — and at our scale, every slice is mission-critical.
If you want to help build the best AI cloud in the world — and you have the technical depth to engage engineering leaders as a peer (not as a translator) and the comfort to talk to customers directly — this team is for you.
The platform you'll help build:
- Hardware platforms & launch — bringing new GPU and CPU platforms (GB300, Vera Rubin, ARM/Grace, future generations) to production with full launch readiness across the stack.
- Cluster lifecycle & fleet operations — new region launches, 100,000+ GPU cluster bring-up, platform sharding and allocation architecture, release engineering, host-lifecycle automation, operational efficiency.
- Reliability & Mission Control — autohealing, health checks, SLA, fault-tolerant training, MTTR reduction, customer trust at scale, observability as a product.
- Customer experience & developer surface — Compute APIs, console, CLI, IMDS and in-VM signals, self-service workflows, notifications, customer-facing observability, unified UX across the product line.
- GPU & InfiniBand foundational services — drivers, firmware, NCCL, IB/RoCE, NVLink topology, the foundational layer everything else builds on.
- Managed runtime platforms — Soperator (Slurm-on-Kubernetes) and MK8S (Managed
Opens the company's application page
Listed via
Jobicy
jobicy.com
Similar roles
Design & Tech
Related reads from TCHNX

How Klarna Engineered the Psychology of Painless Spending
Buy Now, Pay Later apps have mastered the art of invisible payment infrastructure. We dissect Klarna's interface design choices and reveal the behavioural psychology making instalment debt feel effortless.

Haptic Feedback Is Reshaping Digital Interfaces
After years of pursuing flat, buttonless designs, tech companies are rediscovering the value of tactile interaction. A new wave of products proves that touching isn't just feeling it's understanding.

Figma Config 2026: AI Takes Over, But Are Designers On Board?
Figma’s Config 2026 shift to an "AI-first" model prioritises automation and dev tools over core design mechanics. This sparked backlash from UX/UI designers, who feel Figma is chasing flashy trends while ignoring basic quality-of-life updates like better grids and page organisation.
