Research & Benchmarks - Latest AI News & Updates
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
Gartner expects more than 40% of today's agentic AI projects to be scrapped before 2028.
Guidelight AI Standards graded five frontier AI labs on a question none of them like to answer directly: what happens the moment one of their models gets caught trying to slip out of human control. OpenAI came out on top of the ranking.
Give an AI agent a "skill," a short cheat sheet of steps and warnings for a task, and it often does better. That much AI developers already suspected. What nobody had nailed down was why, or when the trick stops working.
Sora can render a coffee cup vanishing into a cabinet with perfect physics. What it can't tell you is what the person who owns that cup will do next, because it has no idea they didn't see it move.
A Llama-3 8B model holding a 100,000-token conversation in memory needs close to 12.8 GiB just for its KV cache, before a single additional request gets queued.
Deepseek pushed out V4-Flash-Vision-Exp this week, an experimental multimodal model that bolts image understanding onto its existing text-only V4-Flash system.
Meta burns through trillions of tokens a week on Microsoft's Azure cloud, and according to a Bloomberg report, that habit costs the company hundreds of millions of dollars annually.
Four days ago, Dario Amodei posted on X that Anthropic's biology work was still months from producing its first "early glimmers." That timeline just moved up.
Wednesday brought a wave of complaints from security researchers who found themselves locked out of Trusted Access for Cyber, OpenAI's vetted program for using its top AI models with fewer cybersecurity restrictions than regular ChatGPT accounts...
Z.ai's GLM-5.3 just closed the gap with the best closed model on the market. On the GDPval-AA v2 benchmark, which measures agentic task performance, the model's Elo score climbed from 1,524 to 1,770.
OpenAI has put a hold on parts of its work on Astra, the company's model project, citing safety concerns that haven't been fully detailed publicly.
Ask five leading language models to write a single clinical trial dataset from scratch, and none of them will get it right.
Cartesia put out Sonic-3.6 this week, the latest version of its real-time text-to-speech model and a follow-up to Sonic-3.5 from three months back.
OpenAI put a two-week hold on reinforcement learning training this fall, and its largest planned frontier RL run is still sitting on ice.
Artificial Analysis has a new way to grade the search tools that AI agents lean on when they go looking for answers online.
Somewhere between an artist's original painting and an AI-generated image of it, the trail runs out.
Artificial Analysis, the research group behind independent LLM evaluations and benchmarks like GDPval-AA and AA-Briefcase, has launched a new platform called Optima.
Yogyakarta got a new building block in Indonesia's AI ambitions this week. Universitas Gadjah Mada, Indosat Ooredoo Hutchison and NVIDIA opened the UGM Indosat NVIDIA AI Technology Center, the first AI research hub housed inside an Indonesian...
Anthropic and OpenAI have spent much of 2024 and 2025 arguing that their models are closing in on doing AI research without much human help. A new paper from Princeton and the UK AI Security Institute says the evidence doesn't hold up.
A coding agent doesn't spend most of its time thinking. It spends its time running commands, reading files, checking outputs and formatting results, hundreds of small steps that follow from a single plan.