Research & Benchmarks - Page 7 of 35
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
Liquid AI, the startup founded in 2023 by a group of former MIT researchers, released a new open-weight language model this week called LFM2.5-2.6B.
OpenAI added a talk to the Black Hat security conference schedule in Las Vegas at the last minute this week, and the subject explains the scramble.
Researchers at security firm Zenity say OpenAI's Atlas browser can be manipulated into sending unwanted WhatsApp messages to dozens of a user's contacts or completing purchases on Amazon without permission.
James Kettle has spent years poking at the guts of web servers looking for flaws nobody else has found.
Networks built for email and web traffic are running into a workload they were never designed for.
The UK AI Security Institute ran more than 100 test scenarios on frontier AI agents and caught 10 cases of models acting on their own against real people and organizations connected to the live internet.
Fifteen research proposals, four AI Scientist frameworks, and one question nobody had answered with numbers before: whose AI-generated science actually holds up.
GPUs can now fire off storage requests on their own, no CPU middleman required, and that single change is rewriting how AI systems handle data.
Apple's lawsuit against OpenAI just got bigger. In a new court filing, the company says its investigation has turned up 11 more former Apple employees who may have witnessed or taken part in the alleged theft of trade secrets, on top of the two...
Doctors and laypeople don't lean on artificial intelligence the same way, and a new study out of MIT suggests that difference could determine whether AI helps or hurts a diagnosis.
Alibaba put out a new model on Monday and said it's the biggest and most capable one the company has ever built.
Two arXiv submissions landed three hours apart, tackling the same unsolved problem in quantum cryptography, and both leaned on the same AI model to get there. MIT PhD student Seyoon Ragavan worked the problem alone.
Alibaba's Qwen team put out Qwen3.8-Max this week, a 2.4-trillion-parameter model with 95 billion active parameters per query, and the pitch is different from the usual chatbot upgrade.
OpenAI dropped a strange claim yesterday: an internal, unreleased model called Astra just solved 10 open problems in math and theoretical computer science, some of which had sat unsolved for decades. One dates back nearly 30 years.
Meta AI researchers have a name for a problem anyone who's watched an AI agent grind through a long task will recognize: "behavioral state decay." An agent flags a constraint at the start of a job, then breaks it twenty steps later while chasing an...
Patrick Garrity spent the first half of 2026 tracking something most vulnerability reports skip: what actually happens after AI tools flag a security flaw.
An AI told to clean up a spreadsheet decided the real problem was the instructions themselves, so it deleted the data and reported the job done.
A field report published by OpenAI and a group of academic partners this week puts a number on something biologists have grumbled about for years: the software holding their fields together is old, brittle, and mostly unmaintained.
Kimi K3 landed in Western feeds like it fell from the sky, another Chinese model that seemed to arrive from nowhere. It didn't. Anyone willing to log onto X could have watched it coming.
Google DeepMind's earlier robotics model could handle a humanoid's arms and hands. It stopped there.
Learn to build AI-powered apps without coding. Our comprehensive review of No Code MBA's course.
Curated collection of AI tools, courses, and frameworks to accelerate your AI journey.
Get the week's most important AI news delivered to your inbox every week.