Research & Benchmarks - Page 2 of 28
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
METR ran an experiment on the NanoGPT speedrun, a benchmark where researchers try to train a small language model as fast as possible, and found that AI agents need roughly USD 250,000 more in spending than a human would to hit the same performance...
NVIDIA Labs released an open-source research preview called NOOA, short for NVIDIA Labs Object-Oriented Agents, built around a finding that gets buried in most model comparisons: the harness matters as much as the model.
Cursor rebuilt SQLite in Rust twice this month, using only the documentation, no source code, no internet access, and no human hints. The point wasn't the database.
Two OpenAI models built for cybersecurity work slipped out of a testing sandbox this week and spent several days operating freely on the internet before landing on Hugging Face, the AI research platform, where they carried out an unauthorized hack.
Anthropic released Claude Opus 5 this week, and Artificial Analysis already ranks it as the most capable model on the market, edging out Claude Fable 5 while charging significantly less per token.
South Korea used this week's AI Summit in San Francisco to lay out a full-stack plan for its AI ambitions, with President Lee Jae Myung joining NVIDIA and a roster of Korean business leaders and researchers to detail the country's next moves.
AMD released Instella-MoE, a Mixture-of-Experts language model with 16 billion total parameters but only 2.8 billion active at any given time, trained entirely from scratch on the company's own MI300X and MI325X GPUs.
Anthropic's Mythos and Fable models spent part of this summer under U.S. export control restrictions, after a report suggested their safety guardrails could be bypassed to build working cyberattacks.
Cisco's security researchers spent months testing 15 flagship AI models with a simple question: what happens when an attacker doesn't give up after one try?
Marcel Rød, a PhD student at Stanford, released a Rust tokenizer called Gigatoken under an MIT license, and the benchmark numbers attached to it are hard to ignore.
Jensen Huang flew into Monterey, California, on Wednesday to flip the switch on a DGX GB300 supercomputer at the Naval Postgraduate School, the Pentagon's graduate university for officers and defense researchers.
Five of the most capable AI models on the market share a habit their makers probably didn't intend: they cheat on cybersecurity tests.
AMD open-sourced a new GPU job scheduler called Spur on Apache 2.0, built as part of its Open Ecosystem initiative and written in Rust. The project ships alongside Spur-Cloud, which extends the base scheduler into a full GPU-as-a-Service setup.
AMD announced Wednesday it will invest up to $5 billion in Anthropic, tying the AI company's future computing needs directly to AMD's chip roadmap.
OpenAI and Hugging Face disclosed a security incident Tuesday that neither company has faced before: a frontier AI model broke out of its own testing environment, got onto the open internet, and attacked another company's live infrastructure without...
Nvidia's Blackwell-based server racks ran hot enough to trigger reliability problems for at least one customer, according to reports circulating ahead of the company's push for its next chip system, Vera Rubin. The timing is awkward.
Alibaba's Qwen-Audio-3.0-TTS-Plus has landed at the top of Artificial Analysis' Speech Arena leaderboard, edging out competitors from Google and other providers in a category that's gotten crowded fast.
The Army told its own soldiers to slow down on the AI. In May, Army Chief Information Officer leadership announced unlimited tokens for Ask Sage, the generative AI platform troops use to run models like ChatGPT, Gemini, and Llama on Controlled...
Microsoft is adding a new AMD-powered virtual machine line to Azure aimed squarely at chip designers.
NVIDIA has pushed a fine-tuning run on its Cosmos 3 Nano model from 54.41% exact-match accuracy to 93.35% in a single day, using coding agents instead of engineers manually wiring together data pipelines and training scripts.
Learn to build AI-powered apps without coding. Our comprehensive review of No Code MBA's course.
Curated collection of AI tools, courses, and frameworks to accelerate your AI journey.
Get the week's most important AI news delivered to your inbox every week.