Research & Benchmarks - Page 13 of 28
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
A robot that reaches for a ghost is a robot that fails. That cold fact drives every line of code in Google DeepMind’s latest release. Gemini Robotics-ER 1.6 just landed, and it does something its predecessor couldn't: count tools accurately.
Last year, OpenAI's GPT-3.5 Turbo failed every one of the UK AI Safety Institute's basic hacking puzzles. That was the baseline. Now, Anthropic’s Mythos Preview cracks over 85 percent of them.
The gap between knowing about AI and actually wielding it is vast. Most people are stuck in literacy, they can name models, recite risks, maybe even dabble with a chatbot. That’s not enough.
The numbers are stark: 21% in academic retrieval, 38% in biomedical. That’s the margin by which a multi-step agent outperformed a stronger, single-turn RAG model on the STaRK benchmark, a suite of semi-structured queries spanning product catalogs,...
Generative AI has pulled off a market penetration blitz. In three years, 53% of the population has used it, outpacing the historical adoption of personal computers and the internet. For younger people, the rate jumps to 80%.
NVIDIA and the University of Maryland released a new audio model that significantly outperforms Microsoft’s on a key benchmark.
People who use AI for real work are noticing a pattern. The tools get worse. Not all at once, but gradually, like a subscription service that quietly cuts corners. Now developers are saying they have proof Claude is doing exactly that.
Talk of AI in finance has become background noise, mostly unproven. But the quiet work of wiring a handful of specific agents into actual financial plumbing is showing concrete returns. One operation deployed seven.
For decades, the dream of a truly self-contained intelligence has been shackled by the same architectural division: a model that thinks, and a machine that runs it. What if that boundary simply vanished?
That model you trust to catch fraud? Its accuracy is probably stable. That's the problem. Security models don't just break. They fade.
Everyone got drunk on the same idea. OpenAI released Sora, and the tech press instantly crowned it a "world simulator." Google's Demis Hassabis said the same about Veo. The promise was a machine that grasped how things actually work.
The key-value cache is the bottleneck of modern long-context inference, growing without bound as sequences lengthen, slowing throughput to a crawl. Pruning it has always meant trading accuracy for speed. Until now.
An ensemble of heavy models can find subtle patterns, but running a whole committee of them is slow and costly. Knowledge distillation trains one small "student" model to copy the committee's behavior.
Publishing in top AI conferences has always been a numbers game, but the numbers have changed. Google’s PaperOrchestra isn’t another speculative tool. It’s a machine that systematically tilts the board, converting raw drafts into papers that win.
If you want to break AI research, break the price of experiments. A thousand operating system replicas for 23 cents a day is how you do it. OSGym’s cost is the headline, but the method is the story. They didn't just buy cheaper servers.
Think a committee solves problems better than a single, focused expert? For artificial intelligence, Stanford researchers have a firm answer: no. Published April 10, their study put four AI models through five collaborative setups.
The AI race just got a new chassis. Meta Superintelligence Labs is no longer a rumor, it’s shipping hardware for the mind. Its first model, Muse Spark, lands as a multimodal reasoning engine with a multi-agent mode baked in.
The latest round of Better Harness updates is not a tweak. It’s a recalibration. We’ve added usage examples, a chaining guide, and clarified the tool suite, those edits crack open what was once opaque. Why?
Sixteen thousand two hundred and thirty two. That's the exact number of times users typed "bot" into nearly 2.8 million messages in Italian and Spanish Telegram groups. It's not chatter. It's a transaction log for an industrial abuse factory.
Google’s AI gets one in ten answers wrong. That’s the good news. Now for the terrifying part: it has to answer billions of questions. A fresh analysis of the Gemini 3-updated AI Overviews feature shows a 91% correct answer rate. A solid score.
Learn to build AI-powered apps without coding. Our comprehensive review of No Code MBA's course.
Curated collection of AI tools, courses, and frameworks to accelerate your AI journey.
Get the week's most important AI news delivered to your inbox every week.