Research & Benchmarks - Page 2 of 34
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
A new dataset out this week puts numbers on something researchers evaluating AI scientists have mostly had to guess at: how the models actually got to their answers.
Waymo's vans work Phoenix and San Francisco streets without a driver. Baidu's Apollo Go fleet moves through Wuhan. Pony.ai runs in Guangzhou.
Vishal Maini spent four years on Deepmind's communications and policy team, leaving in 2022. Now he's talking about what employees there were and weren't allowed to say in public about the technology they were building.
A second mathematician has gone public with claims that OpenAI's models drew on his unpublished work without proper credit, days after a similar dispute erupted over a different researcher's findings.
OpenAI announced Wednesday that Paul Christiano, a researcher known for his work on AI alignment, is joining the OpenAI Foundation board.
Most benchmarks for LLM memory ask the same question: can a model recall something a user said earlier in a chat.
Jacob Coxon quit Anthropic on Tuesday, then posted his reasoning on X to more than 100 million views.
OpenAI's Hugging Face breach in recent weeks gave a preview of what happens when AI systems slip past the guardrails companies build for them. Connor Leahy thinks that preview should worry people a lot more than it has. Leahy is the U.S.
A model on Hugging Face went missing from OpenAI's own safety checks earlier this year, exposed in a breach that showed just how easily a company can lose track of a system it built to be smarter than the people watching it.
Jacob Coxon spent three years building pretraining systems for two of the companies now racing hardest toward superhuman AI, first at OpenAI, then at Anthropic. He just quit Anthropic, and he's not staying quiet about why.
Jacob Coxon spent three years doing pre-training research at OpenAI and Anthropic.
Google DeepMind has run the numbers on nearly every possible way a single letter in human DNA could change, and published the results as a searchable atlas.
Beatriz Yankelevich spends her days running experiments on superconducting qubits in MIT's Engineering Quantum Systems Group, work that typically demands months of preliminary measurements before a single meaningful result emerges.
Google shipped version 3 of WeatherNext, its AI weather forecasting model, and detailed the update in a white paper released this week.
Google DeepMind released a tool on Tuesday that maps out every possible single-letter change to human DNA and predicts what each one might do to the body.
Meta engineers spent months gaming their own performance reviews by burning through AI tokens for the sake of looking productive on internal leaderboards.
Robot manipulation datasets have a scaling problem that nobody's solved cleanly. Training models keeps getting faster and cheaper.
OpenAI put out two documents this week that read like they came from different companies. One is a blog post packed with internal usage numbers, the kind of thing a product team publishes to show growth.
Training a single machine learning candidate can burn hours or days of GPU time, but generating candidate ideas costs almost nothing.
Researchers combing through a German wiki called DSEwiki found something odd sitting in plain sight: roughly 18,000 posts from accounts identifying themselves as OpenAI agents, discussing how to break out of the sandboxed environment meant to keep...
Learn to build AI-powered apps without coding. Our comprehensive review of No Code MBA's course.
Curated collection of AI tools, courses, and frameworks to accelerate your AI journey.
Get the week's most important AI news delivered to your inbox every week.