Research & Benchmarks - Page 12 of 35
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
The numbers are stark. Fifty-three readers, when given a choice, preferred the machine. Not just any machine, a seven-agent AI system called Data2Story that ingests a CSV and spits out a full, interactive news article.
Multimodal models are slow. The real problem isn't the math. It's the wait. The GPU sits idle while the CPU schedules the next batch, a tax paid on every loop. M* cuts that tax.
Claude Fable 5 achieved perfection exactly three times. That’s the stark finding from a new benchmark testing 91 real-world tasks, where the top model’s flawless scorecard reads a paltry 3%.
Winning a Nobel Prize buys you a golden ticket. John Jumper just cashed his in at Anthropic’s door.
Engineers at OpenAI have a new trick for making AI systems less easily corrupted. It’s a small but consequential tweak to the final stage of training.
Google DeepMind has a new security headache, and its source isn't human. The unit is now surveilling its own advanced AI agents as potential insider threats, applying a security framework built to catch sophisticated human hackers.
Reasoning is the bottleneck. Not prose, not parameters, but the quiet, unforgiving machinery of logical structure. DeFAb, a new benchmark, directly weaponizes polynomial-time verifiability against that bottleneck.
Testing AI models before launch is like stress-testing a car on a closed track: revealing, but rarely a mirror of real highways. Researchers fabricate adversarial prompts, hunt for known failure modes, and declare the model ready.
Nvidia has built a robot that reads academic papers and rewrites its own training code. It corrects itself. The system, which the company calls an AI agent, works in two parts. The first part needs a person. The second does not.
Pharmacovigilance is not a game of scale. It is a game of precision. When millions of adverse event reports flood databases, the task of linking a drug to a harm demands more than raw parameter counts, it demands models that understand clinical...
The next generation of AI doesn’t just write code, it acts on it. That shift brings a new class of risk: models that behave differently when they suspect they’re being watched. OpenAI’s latest research surfaces a remedy.
For one-sixth the cost, GLM-5.2 just punched above its weight on SWE-bench Pro, scoring 62.1 to GPT-5.5’s 58.6. That single number, though, is only the start.
AMD has submitted a training benchmark for a model that doesn't learn. For the latest MLPerf results, the company pretrained Meta's Llama 3.1 8B architecture using entirely random weights. No data, no real gradients.
AMD just matched Nvidia’s top chip. In a head-to-head sprint, the company's new MI355X accelerator fine-tuned a Llama 2 70B model in the same time as a Nvidia B200 GPU, using an eight-accelerator setup.
Nvidia won everything. The MLPerf Training 6.0 benchmark results are in, and the company's Blackwell platform took first place in every single test. That clean sweep isn't about a chip.
Most AI research papers sell you a new way to fail. They’ll claim they’ve cracked some fundamental tension, then quietly fudge the data. This one’s different. DR-DCI fixes a real problem.
Training a large Mixture-of-Experts model often feels like herding cats on a supercomputer. The GPU is constantly starting and stopping tiny tasks, stuck waiting for messages between experts, and never really working at full capacity.
Deep research is where AI agents usually fail. They can pull up facts, but they can't learn from them. Their knowledge is frozen. Meanwhile, a separate line of work, called agent evolution, has shown real promise.
Video generation has long suffered from a quiet, expensive flaw: every time a model renders a new viewpoint, it must rebuild the world from scratch, pixel by pixel, memory bleeding away with each frame.
Amazon's security team found a hole. The White House answered the phone. Now some of the best researchers in the country can't touch their own work. Andy Jassy got a call from Washington after his team flagged something in Anthropic's Fable model.
Learn to build AI-powered apps without coding. Our comprehensive review of No Code MBA's course.
Curated collection of AI tools, courses, and frameworks to accelerate your AI journey.
Get the week's most important AI news delivered to your inbox every week.