Skip to main content
Jensen Huang, NVIDIA CEO, and Baseten CEO discuss AI inference and cloud scaling, highlighting NVIDIA's investment.

Editorial illustration for NVIDIA puts USD 150 M into Baseten, backing Jensen Huang’s inference‑first pivot

NVIDIA Backs Baseten with $150M AI Inference Bet

NVIDIA puts USD 150 M into Baseten, backing Jensen Huang’s inference‑first pivot

Updated: 3 min read

Jensen Huang has a new mantra: inference is the future. Not training, not the brute-force calculus of building ever-larger models, but the quiet, relentless work of running them. NVIDIA is now putting its money where its CEO’s mouth is, pouring $150 million into Baseten.

The startup sits at the intersection of that shift, optimizing GPU infrastructure for the H100 and the next-gen B200 chips. It’s a bet that the real money lies in deployment, not development. And at a $5 billion valuation, the market agrees.

CapitalG’s presence in the round adds a wrinkle, Alphabet’s venture arm backing a NVIDIA-adjacent platform, even as Google builds its own AI stack. Rivals can still find common ground. Inference is that big.

For NVIDIA, the investment reinforces a strategic pivot championed by chief executive Jensen Huang, who has repeatedly argued that inference will ultimately become a much larger market than model training. As enterprises move from experimentation to full-scale deployment, demand for reliable and cost-efficient inference infrastructure is accelerating, placing companies like Baseten at the centre of this transition. Baseten's platform is optimised for NVIDIA's latest GPU architectures, including the H100 and next-generation B200 chips.

By enabling high-performance inference workloads on these GPUs, Baseten effectively extends NVIDIA's ecosystem, helping ensure its hardware remains the default choice as AI adoption spreads across enterprises. CapitalG's participation adds a competitive dimension, given Alphabet's own investments in AI infrastructure and model deployment. Nevertheless, the collaboration underlines the strategic importance of inference, even among industry rivals.

At a $5 billion valuation, Baseten now joins a small group of AI infrastructure startups commanding premium multiples.

This is not a passive bet. It is a declaration. Jensen Huang has spent years insisting that the future of AI lies not in building bigger models, but in running them at scale.

With $150 million into Baseten, NVIDIA is putting hard capital behind that conviction. The deal does more than inflate a valuation to $5 billion; it anchors a new axis of competition. CapitalG’s involvement, despite Alphabet’s own inference ambitions, proves that even rivals recognise the inevitability of this shift.

Baseten’s job is clear: make NVIDIA’s hardware indispensable for the workloads that will define the next decade. If inference truly dwarfs training, then this investment is not a hedge. It is a throne.

Common Questions Answered

How does Baseten's inference stack improve AI model performance?

[baseten.co](https://baseten.co/resources/guide/the-baseten-inference-stack) reveals that their inference stack optimizes every layer of AI model deployment, from hardware to software. The stack combines open-source techniques with proprietary enhancements to deliver low latency, high throughput, and cost-efficient model serving across different AI modalities.

What performance improvements did Baseten achieve with NVIDIA Blackwell GPUs?

[nvidia.com](https://www.nvidia.com/en-us/customer-stories/baseten-cloud-scaling-ai-inference/) reports that Baseten achieved 5× higher throughput for high-traffic endpoints and up to 38% faster LLM serving with NVIDIA Blackwell GPUs. The improvements enable Baseten to serve more user requests with the same GPU infrastructure and reduce latency for large language models.

Why did Baseten raise $150 million in its Series D funding round?

[baseten.co](https://baseten.co/blog/announcing-baseten-150m-series-d) indicates the funding was raised to push the boundaries of performant, reliable, and cost-efficient AI inference. The company aims to build infrastructure that supports production-grade AI systems with fast models, interchangeable compute, and flexible, Pythonic runtimes.

LIVE14:31MCP's new authorization protocols make it "enterprise ready