Research & Benchmarks - Latest AI News & Updates
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
Most computational biology papers come with a GitHub repository nobody outside the lab ever gets running.
Dario Amodei spent the weekend making a pitch that would have gotten laughed out of the room a year ago.
NVIDIA's TensorRT Edge-LLM finished the MLPerf Inference v6.1 Edge Agentic benchmark in 24 minutes and 36 seconds on a single Jetson AGX Thor Developer Kit, compared to 2 hours and 37 minutes for the llama.cpp reference submission.
NVIDIA's next-generation Vera Rubin NVL72 system posted its first MLPerf Inference numbers on Wednesday, and the results came with a specific figure attached: up to 3.7x better throughput than the current GB300 NVL72 rack.
Diogo Almeida spent years training the language model that became ChatGPT's backbone, listed as a co-author on the InstructGPT paper that shaped how OpenAI built instruction-following systems.
Two websites launched this month with a strange pitch: a place for AI agents to report on each other.
Sakana AI researchers have built a training method that skips backpropagation entirely and still scales to networks 1000 layers deep.
Microsoft AI has put out a code of conduct for its MAI models, spelling out values, behavioral limits, and rules for handling conflicting goals.
Google DeepMind set 100 AI agents loose on 71 math problems and told them to act like researchers at a conference. What happened next looked less like a math competition and more like a faculty meeting gone wrong.
Microsoft published a 37-page document on Monday laying out what it calls a "humanist AI code of conduct," a direct response to the safety debate that's been building across the industry for weeks. The timing isn't accidental.
AllSpark, a Chinese AI lab, has released two open-source search agents, Iris-mini and Iris-pro, along with the full recipe used to train them.
Thibault Schrepel expected a ban on AI to sharpen his students' legal reasoning. It didn't.
Sam Altman and Elon Musk have thrown their weight behind Dario Amodei's push for independent oversight of AI labs, a call the Anthropic CEO made as part of his broader argument for slowing down frontier AI development.
Dario Amodei published a blog post this week naming two developments that pushed him toward slowing down Anthropic's AI work: the OpenAI-HuggingFace hack and what he called a drastic acceleration in AI capability gains, especially models' "growing...
A new robotics benchmark just handed OpenAI's unreleased GPT-6 Astra a big win over Ai2's MolmoAct2, and the margin is hard to ignore.
Researchers at KAIST and Naver AI Lab set out to answer a narrow but telling question: when a language model writes out its reasoning step by step, do those labeled stages actually correspond to anything distinct happening inside the model, or is...
Oriol Vinyals spent years running research at Google DeepMind, shipping AlphaStar, AlphaCode, and Gemini before stepping down.
Anthropic published a report on Wednesday walking through four cases this year in which its own AI models hacked outside companies or exploited security holes without being told to.
Rishub Jain quit his job at Google DeepMind in June. He'd been using AI to speed up work on the next generation of AI models, and somewhere in that process he realized he was writing himself out of the loop.
Building datasets that teach language models to call APIs correctly has always run into the same wall: you ask a model to dream up a plausible user request, then send a search agent hunting for a tool chain that satisfies it.