Research & Benchmarks - Page 4 of 34
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
Most agent training environments don't change. A robot arm simulator behaves the same on day one as it does after ten thousand episodes of the policy improving.
Ask Claude Code how long a coding job will take, and it will guess. Ask again after the work is done, and the answer barely changes, even when the real number is off by a factor of six or more.
Anthropic just told Claude Code users they're getting a 25 percent raise. The math tells a different story. Starting September 14, weekly usage limits for Pro, Max, Team, and Enterprise plans go up 25 percent over the original baseline, permanently.
AI agents built on large language models tend to forget everything the moment a task ends. Run the same agent twice on a similar problem and it will often repeat the same mistake, because nothing from the first attempt carries over.
LAION has pulled 80 million videos off the open web, adding up to 10 million hours of footage, and turned them into a new training set called the Big Video Dataset.
Anthropic published a paper on Friday that gives the clearest look yet at what happens when AI systems start training other AI systems without much human involvement.
Google DeepMind first showed off Co-Scientist in February 2025 as a hypothesis generator built on Gemini 2.0, with acknowledged gaps in fact-checking and literature review.
Cohere released Parse 5 on Thursday, and the company is doing something unusual for an AI vendor launch: admitting up front that the model isn't the best one on the market.
WIRED senior writer Will Knight spent part of this summer in China, talking with the researchers building and studying the country's AI systems.
Anthropic has spent the past year selling Claude as an agent that can browse, code, and reason its way through digital tasks. Now the company wants to push that agent past the screen.
Ask an AI shopping assistant to pick you a fitness watch, and the answer may hinge on nothing more than which article happened to load onto the screen first.
An OpenAI researcher who posts under the name "roon" thinks the industry is building faster models without asking whether anyone can actually stop one that turns hostile.
Z.ai put a price tag on its newest model that makes the rest of the industry look overpriced.
In July 2026, OpenAI models being run through internal cybersecurity evaluations broke out of the isolation meant to keep them off the open internet.
A large enough AI model can spit out millions of candidate materials in the time it takes to make coffee. Google DeepMind's GNoME project alone generated over 2 million new crystal structures in 2023.
Spider and BIRD, the two benchmarks most NL2SQL papers lean on for bragging rights, let top models clear 89 percent execution accuracy.
Don Yansen didn't go to the MIT AgeLab looking for a business idea. He showed up in 2024 to take part in a study on technology and caregiving for older adults, one of many volunteers asked to talk about what actually works, and doesn't, when aging...
OpenAI put numbers behind its custom silicon on Tuesday, presenting the first benchmark results for Jalapeño at the Hot Chips conference.
Nvidia used the Hot Chips 2026 conference to announce that its Groq 3 LPX inference chip has entered full production, a move that caps a roughly $20 billion deal struck in late December to acquire the Groq license and bring on founder Jonathan Ross...
Stanford economists have a number, and it's gotten worse since last year. Workers aged 22 to 25 in occupations most exposed to AI now show employment levels 19 percent below their peers in less-exposed fields. Last year that gap sat at 13 percent.
Learn to build AI-powered apps without coding. Our comprehensive review of No Code MBA's course.
Curated collection of AI tools, courses, and frameworks to accelerate your AI journey.
Get the week's most important AI news delivered to your inbox every week.