Research & Benchmarks - Page 5 of 34
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
Two coffee shops can sit on the same block, carry the same category tag, and produce nearly identical text embeddings, while one turns over commuters in five minutes and the other holds customers for an hour and a half.
NVIDIA's Groq 3 LPX chip has moved into full production, and the company is pairing it with its Vera Rubin NVL72 rack-scale system to chase a specific problem: getting AI agents to generate tokens fast enough for real-world use.
Harvey announced a research preview on August 20, 2026, called Harvey Tenet, its first post-trained model built for legal work that spans thousands of documents at once, the kind of contract review or M&A diligence job that used to eat a junior...
Gartner expects more than 40% of today's agentic AI projects to be scrapped before 2028.
Guidelight AI Standards graded five frontier AI labs on a question none of them like to answer directly: what happens the moment one of their models gets caught trying to slip out of human control. OpenAI came out on top of the ranking.
Give an AI agent a "skill," a short cheat sheet of steps and warnings for a task, and it often does better. That much AI developers already suspected. What nobody had nailed down was why, or when the trick stops working.
Sora can render a coffee cup vanishing into a cabinet with perfect physics. What it can't tell you is what the person who owns that cup will do next, because it has no idea they didn't see it move.
A Llama-3 8B model holding a 100,000-token conversation in memory needs close to 12.8 GiB just for its KV cache, before a single additional request gets queued.
Deepseek pushed out V4-Flash-Vision-Exp this week, an experimental multimodal model that bolts image understanding onto its existing text-only V4-Flash system.
Meta burns through trillions of tokens a week on Microsoft's Azure cloud, and according to a Bloomberg report, that habit costs the company hundreds of millions of dollars annually.
Four days ago, Dario Amodei posted on X that Anthropic's biology work was still months from producing its first "early glimmers." That timeline just moved up.
Wednesday brought a wave of complaints from security researchers who found themselves locked out of Trusted Access for Cyber, OpenAI's vetted program for using its top AI models with fewer cybersecurity restrictions than regular ChatGPT accounts...
Z.ai's GLM-5.3 just closed the gap with the best closed model on the market. On the GDPval-AA v2 benchmark, which measures agentic task performance, the model's Elo score climbed from 1,524 to 1,770.
OpenAI has put a hold on parts of its work on Astra, the company's model project, citing safety concerns that haven't been fully detailed publicly.
Ask five leading language models to write a single clinical trial dataset from scratch, and none of them will get it right.
Cartesia put out Sonic-3.6 this week, the latest version of its real-time text-to-speech model and a follow-up to Sonic-3.5 from three months back.
OpenAI put a two-week hold on reinforcement learning training this fall, and its largest planned frontier RL run is still sitting on ice.
Artificial Analysis has a new way to grade the search tools that AI agents lean on when they go looking for answers online.
Somewhere between an artist's original painting and an AI-generated image of it, the trail runs out.
Artificial Analysis, the research group behind independent LLM evaluations and benchmarks like GDPval-AA and AA-Briefcase, has launched a new platform called Optima.
Learn to build AI-powered apps without coding. Our comprehensive review of No Code MBA's course.
Curated collection of AI tools, courses, and frameworks to accelerate your AI journey.
Get the week's most important AI news delivered to your inbox every week.