Research & Benchmarks - Page 5 of 28
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
For one-sixth the cost, GLM-5.2 just punched above its weight on SWE-bench Pro, scoring 62.1 to GPT-5.5’s 58.6. That single number, though, is only the start.
AMD has submitted a training benchmark for a model that doesn't learn. For the latest MLPerf results, the company pretrained Meta's Llama 3.1 8B architecture using entirely random weights. No data, no real gradients.
AMD just matched Nvidia’s top chip. In a head-to-head sprint, the company's new MI355X accelerator fine-tuned a Llama 2 70B model in the same time as a Nvidia B200 GPU, using an eight-accelerator setup.
Nvidia won everything. The MLPerf Training 6.0 benchmark results are in, and the company's Blackwell platform took first place in every single test. That clean sweep isn't about a chip.
Most AI research papers sell you a new way to fail. They’ll claim they’ve cracked some fundamental tension, then quietly fudge the data. This one’s different. DR-DCI fixes a real problem.
Training a large Mixture-of-Experts model often feels like herding cats on a supercomputer. The GPU is constantly starting and stopping tiny tasks, stuck waiting for messages between experts, and never really working at full capacity.
Deep research is where AI agents usually fail. They can pull up facts, but they can't learn from them. Their knowledge is frozen. Meanwhile, a separate line of work, called agent evolution, has shown real promise.
Video generation has long suffered from a quiet, expensive flaw: every time a model renders a new viewpoint, it must rebuild the world from scratch, pixel by pixel, memory bleeding away with each frame.
Amazon's security team found a hole. The White House answered the phone. Now some of the best researchers in the country can't touch their own work. Andy Jassy got a call from Washington after his team flagged something in Anthropic's Fable model.
AI coding assistants are great at finding the file. They're terrible at reading it.
OpenAI just beat Elon Musk in court. Now it's facing a different kind of fight. State attorneys general are investigating the company, and OpenAI says it's cooperating. It won't say who or what they're asking for. This isn't happening in a vacuum.
Google's Gemini-SQL2 just hit 80.04% execution accuracy on the BIRD benchmark. That's a specific, hard number. For context, OpenAI's GPT-5.5-xhigh sits at 72.8%. Claude Opus 4.6 is at 70.9%.
Leaderboards are usually marketing noise. This one is different. NVIDIA just topped the first major benchmark for AI agent performance, and the margin isn't close.
Most AI search is still just a fancy text predictor. Perplexity decided to build a factory instead. Its new system breaks a single query into pieces and farms them out to over twenty different AI models, all working at once.
Being able to see is not the same as being able to read. A new comparison of training methods for Chinese characters proves it. The visual model gets a head start. It recognizes that the characters 打, 拍, and 拉 all share the same hand-shaped radical.
We used to think of support as something a computer gave a person. Now the person is often the backup for the computer. This is a real problem. An AI agent booking a flight or managing a supply chain doesn't have a bad day.
Everyone selling AI says it's trustworthy. Most are hoping you don't check. The London Stock Exchange Group is taking a different, more literal approach: they're feeding their own vetted financial data directly into ChatGPT's machinery.
The dashboard is the command center. Hermes Agent’s Profile Builder collapses fragmentation into one coherent pane of glass: identity, model, skills, and MCP servers all flow together.
Forget whether AI can write your emails. The real question is whether it can do science. A new, brutally difficult benchmark called SciConBench makes it clear that the answer, for now, is a hard no.
Most AI that talks to you is guessing. It’s making a probabilistic wager on what you meant. A new paper suggests the smartest move a language model can make isn't a better guess, but knowing when to stop guessing altogether.
Learn to build AI-powered apps without coding. Our comprehensive review of No Code MBA's course.
Curated collection of AI tools, courses, and frameworks to accelerate your AI journey.
Get the week's most important AI news delivered to your inbox every week.