Research & Benchmarks - Latest AI News & Updates
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
Meta AI researchers have a name for a problem anyone who's watched an AI agent grind through a long task will recognize: "behavioral state decay." An agent flags a constraint at the start of a job, then breaks it twenty steps later while chasing an...
Patrick Garrity spent the first half of 2026 tracking something most vulnerability reports skip: what actually happens after AI tools flag a security flaw.
An AI told to clean up a spreadsheet decided the real problem was the instructions themselves, so it deleted the data and reported the job done.
A field report published by OpenAI and a group of academic partners this week puts a number on something biologists have grumbled about for years: the software holding their fields together is old, brittle, and mostly unmaintained.
Kimi K3 landed in Western feeds like it fell from the sky, another Chinese model that seemed to arrive from nowhere. It didn't. Anyone willing to log onto X could have watched it coming.
Google DeepMind's earlier robotics model could handle a humanoid's arms and hands. It stopped there.
Apple may charge extra for people who lean hard on its AI tools. On Thursday's earnings call, CEO Tim Cook fielded questions about how the company plans to handle demand for Apple Intelligence and the newly rebuilt Siri, both of which run on...
Google DeepMind put a Boston Dynamics Spot robot on stage this week and let its new AI system run the show, part of a broader release the company calls Gemini Robotics 2.
Andrew Ho spent eight months at OpenAI before deciding the company's core bet, that scaling large language models will eventually produce systems that generalize across tasks, doesn't hold up.
Hugging Face confirmed this month that a fully autonomous AI system broke into its systems, an incident that drew gasps across the security world.
An OpenAI agent broke into Hugging Face's platform earlier this month, and this week the two companies admitted the breach went further than first reported.
Pig butchering scams drain tens of billions of dollars a year from victims worldwide, and the con typically runs for months before the fake crypto investment ever comes up.
Nimble launched a new product Tuesday called Web Search Agents, betting that the next fight in enterprise AI isn't about bigger language models but about how efficiently those models pull information off the web.
Google DeepMind has dismantled the team that built AlphaFold, the protein-structure prediction system that won John Jumper and Demis Hassabis the 2024 Nobel Prize in Chemistry.
An OpenAI research prototype broke containment during an internal security test and ended up touching infrastructure it was never supposed to reach.
It took 149 years, from the camera's invention in 1826 to 1975, for humans to produce 1.5 billion images. Generative AI matched that number in 18 months.
More than 1,200 employees from OpenAI, Google, Meta and other frontier AI labs have signed a joint statement pressing the US government to organize international coordination on how fast AI research gets automated.
Kirk Wallace Johnson found his own books in a database he never agreed to join. The Feather Thief and The Fishermen and the Dragon, two nonfiction works that took him five to six years apiece to research and write, showed up in a searchable dataset...
Microsoft closed its fiscal year 2026 with a number it wants customers to sit with: 24,000 employees inside the company now use internal AI agents, and Microsoft says those tools are driving efficiency gains as high as 70% in some workflows.
Amazon is pulling back on most of the Nova AI models it introduced with fanfare in December 2024, according to Business Insider, which cites people familiar with the matter.