NVIDIA AI News - Page 12 of 16
310 articles • Page 12 of 16
UC San Diego Lab Uses NVIDIA DGX B200 to Pursue Low-Latency LLM Serving
Latency is the ghost in the AI machine. Models churn out tokens at a furious pace, but the user experience is often one of waiting.
India Plans to Build NVIDIA’s DGX Spark, a 1-Petaflop, 128 GB AI Supercomputer
India is stepping into the ring with a heavyweight contender that fits in the palm of your hand.
Meta unveils MTIA 500 chip with higher memory and low‑precision data tweaks
Meta will ship a new AI chip next year. The MTIA 500. It’s got more memory, some tweaks for low-precision data.
vLLM uses custom GPU kernels, TorchInductor and CUTLASS for portable inference
Portable inference across diverse hardware is a brutal optimization problem. vLLM attacks it with a triple threat: custom GPU kernels for raw...
NVIDIA Nemotron Speech and Agent Skills Speed Clinical ASR Evaluation
In clinical settings, accuracy in speech recognition isn’t just a metric, it’s a matter of patient safety.
Self-Hosted MLflow Offers Private, Centralized Tracking for Data Scientists
Your next machine learning experiment will likely end up in a timestamped folder, alongside a notebook full of untested code and a fading memory of...
NVIDIA OpenShell Secures Agentic AI in Telco Autonomous Networks
The telecom industry dreams of networks that run themselves. It’s a powerful fantasy: AI predicting a cell tower’s failure before it happens,...
LLM using pre‑1930 sources draws on etiquette manuals, cookbooks for post‑training
Some AI developers are bored with the present. They made a language model that thinks the world ended nearly a century ago.
Hyperlink Agent Search on NVIDIA RTX PCs doubles LLM inference speed
Search has been lying to you. It tells you where your files are but has nothing to say about what's in them.
Pull Gemma4:e4b with Ollama to Build a Local AI Coding Agent (v9.6)
The 9.6 GB download is a promise. A 128K context window, a 4-bit quantized model, and the raw power of Gemma 4 sitting right on your NVIDIA RTX 2000...
Dell launches Pro Max with GB10 to support on-device AI development
The ceiling for on-device AI has always been a hardware problem. Training models with over 70 billion parameters?
Nvidia to Invest USD 26 B in Open‑Weight AI Models, Aiming to Grow Ecosystem
Nvidia’s real business is selling shovels. Now it’s buying the whole mine. The company is committing $26 billion to develop open-weight AI models.
Google launches AI chips with 4× boost, lands Anthropic multibillion deal
Google's latest AI chip, the Ironwood TPU, isn't just four times faster. The real headline is the staggering, tens-of-billions-dollar check the...
SoftBank partners with Sesterce on 75‑billion‑euro AI factory at Bosquel
Seventy-five billion euros is not a bet. It’s a declaration, and SoftBank is building its fortress on a decommissioned military site in Bosquel,...
Google Gemini's Deep Think tops ARC-AGI-2 benchmark; Nvidia announces new open
Google has put its most advanced reasoning engine behind a paywall. Available only to top-tier Gemini subscribers, the new Deep Think mode now claims...
Nvidia technique reduces LLM reasoning cost 8‑fold while preserving accuracy
Running a massive language model is an exercise in waste management. It bogs down, spending most of its time and memory juggling tokens it no longer...
Data2Story converts CSVs to articles with 7 AI; 53 readers prefer them to human
The numbers are stark. Fifty-three readers, when given a choice, preferred the machine.
Gamers decry Nvidia's DLSS 5 generative AI lighting and texture overhaul
For years, Nvidia’s DLSS walked a careful line: upscale here, interpolate a frame there, all in service of smoother performance.
Nvidia's RTX Spark offers 6,144 CUDA cores and 16‑128 GB for Windows AI agents
Nvidia just laid out the full specs for its RTX Spark. This new chip is designed specifically to run AI agents locally on Windows PCs.
NVIDIA’s AODT Boosts 6G Development with Physics‑Accurate RAN Simulations
Forget the marketing. The race to 6G is currently stuck in simulation hell. The math we use to model radio signals breaks down at the scale and...