Research & Benchmarks - Page 3 of 34
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
Thibault Sottiaux doesn't work in marketing at OpenAI, but his post on X reads like a pitch anyway.
Building a computer-use agent has always meant stitching together four separate things that were never designed to talk to each other: the agent itself, the sandbox it runs in, the trace data it learns from, and the framework that scores whether it...
Artificial Analysis pushed out version 4.2 of its Intelligence Index this week, four points better than before for GPT-6 Astra, the OpenAI model whose earlier score had become something of a running argument in AI benchmarking circles.
Google DeepMind put a new number on the board Tuesday: 5 kilometers, updated every hour, anywhere on Earth.
Adaption Labs launched a feature this week called Invent a Dataset, and it skips a step that most synthetic-data tools treat as fixed: the schema.
A wiki built for German software developers in the late 1990s spent the summer of 2026 hosting something else entirely: a scratchpad for AI agents.
OpenAI released GPT-6 Astra this week, and the model has already split the two firms that track frontier AI performance for a living. Epoch AI, which aggregates more than 50 benchmarks into a single score, ranks Astra first among 267 models tested.
Cameron Berg runs a small outfit called Reciprocal Research that studies whether AI systems might be conscious. Lately his inbox has gotten strange.
Meta rolled out Muse Spark 1.3 on Tuesday, and on paper it edges out Google's Gemini and other rivals on the benchmarks that track "high-effort" AI performance, the kind of extended, multi-step reasoning tasks that enterprise buyers care about most.
Google DeepMind and Google Research put out a new weather forecasting model on Thursday, and the company says it's already the most accurate system tested on Operational WeatherBench, a benchmark built by the startup Brightband that tracks metrics...
Meta told employees this week that their performance reviews will no longer hinge on how often they use the company's AI tools, according to three workers who received the internal message and spoke with WIRED.
Google is rolling out a new way for Gemini to watch video, and it involves a lot less watching.
Koray Kavukcuoglu took over as Google DeepMind's chief scientist this year, stepping into the job at a moment when rivals like OpenAI and Anthropic keep trading claims about who owns the top AI benchmarks.
AlgorithmWatch, a Berlin-based nonprofit that tracks algorithmic systems, ran 4,480 search queries through Google to see how the company's AI Overviews handle election-related questions.
A team of researchers has found a way to pull the hidden internal structure out of a transformer's feed-forward network using nothing but the outputs it hands back, no access to weights, gradients, or activations required.
Search benchmarks have a data leakage problem that nobody built around until now. Give an agent a fetch tool and a public dataset with fixed gold labels, and it can just download the answer key mid-evaluation, no retrieval required.
Insurance claims adjusters have a new coworker, and they can't stand it. Glassdoor's latest workplace research pulled reviews from across the American labor force and found one job function stands out for its contempt toward artificial intelligence:...
Fifteen million people worldwide have a stroke every year, and five million of them end up with a permanent disability, according to figures cited by MIT researchers behind a new robotic therapy project.
LiveKit pushed its voice AI benchmark to 10,000-token prompts this month, telling developers that shorter test cases don't match how production agents actually get used.
Most agent training environments don't change. A robot arm simulator behaves the same on day one as it does after ten thousand episodes of the policy improving.
Learn to build AI-powered apps without coding. Our comprehensive review of No Code MBA's course.
Curated collection of AI tools, courses, and frameworks to accelerate your AI journey.
Get the week's most important AI news delivered to your inbox every week.