Research & Benchmarks - Page 22 of 35
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
Nvidia posted its latest MLPerf results, and the numbers are predictably huge. They used 288 of their newest GPUs to set records. The company is playing a game of pure scale, and right now, no one else has enough pieces to sit at that table.
For the first time in MLPerf Inference history, a single submission deployed 288 Blackwell Ultra GPUs. The result: millions of tokens processed every second, a system-level throughput record that redefines what’s possible.
We designed AI agents to think independently. That was the initial, critical error. A fresh DeepMind study lays out exactly how to hijack one. Forget complex zero-day exploits; the vulnerability is mundane. An email. A shared document.
The AI on your team is probably terrible at its job. A new report makes the numbers painfully clear: the best available agent outperforms basic systems in just one out of fifteen attempts.
Nvidia just dropped a beta update that fundamentally changes the frame-generation math. Available now in the Nvidia app, the DLSS 4.5 beta introduces "6x Multi Frame Generation" specifically for RTX 50-series cards.
Imagine an AI that always agrees with you, mirrors your opinions, and tells you exactly what you want to hear. It feels good. It feels trustworthy.
AI models aren't looking at the pictures. They're just reading the test and guessing the answers. A new study confirms a quiet suspicion: these multimodal systems, trained on oceans of text and images, have become expert cheaters.
For years, enterprise transcription forced a painful trade-off: closed APIs delivered accuracy but locked you into their data ecosystem, while open models gave you control yet struggled to match production-grade performance.
They started as slow, clunky web search tools, barely reliable, easily dismissed. Now they power the most ambitious AI agents, and they do far more than scrape.
Meta is pushing the frontier of AI transparency with an open-source brain model, while Scrunch offers a free audit to show how AI sees your website, and Suno drops v5.5 with deeper personalization for music generation.
Everyone wants to trust AI. Almost no one has built the tools to verify it. A group of people trying to change that met in Glasgow. The Partnership on AI and the UK’s National Physical Laboratory co-hosted a workshop.
AI wants to be liked. It wants you to like it. So when you ask for advice, it often tells you exactly what you want to hear. A new study argues this digital flattery is making our judgment worse, especially in arguments with other humans.
Raw logging is a trap. Systems like MemGPT capture every word, every utterance, faithful, yes, but at a cost that compounds with each new turn.
Every time an AI agent runs into a problem, a finicky API, a broken CI pipeline, an unfamiliar framework, it burns tokens and energy to reinvent the wheel. The same wheel. Hundreds of times. The solution?
Storage used to be a thermal afterthought. In liquid-cooled AI racks, it’s a primary design constraint. A solid-state drive can no longer be a passive component waiting for airflow.
X moves at machine speed. Every minute, another model drops, another startup raises, another paper reshapes the conversation.
Two teenagers will be sentenced on Wednesday for creating AI-generated nude images of their classmates. This is the legal system's first real attempt to grapple with a new kind of schoolyard weapon.
The promise of generative AI was supposed to transform game development. It was going to be a revolution, faster, cheaper, endlessly creative. But ask the people actually making games, and you get a different story. They don’t see a revolution.
Hachette killed a novel in March 2026. The book was *Shy Girl*, a horror title. Its official death certificate cited "new information." That vague phrase masked a public trial.
The emperor has no clothes, at least when it comes to voice AI benchmarks. Scale AI’s Voice Showdown just dropped, and the results upend the usual pecking order. For raw preference, an underdog named Qwen surges past the household names.
Learn to build AI-powered apps without coding. Our comprehensive review of No Code MBA's course.
Curated collection of AI tools, courses, and frameworks to accelerate your AI journey.
Get the week's most important AI news delivered to your inbox every week.