Research & Benchmarks - Page 21 of 35
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
Publishing in top AI conferences has always been a numbers game, but the numbers have changed. Google’s PaperOrchestra isn’t another speculative tool. It’s a machine that systematically tilts the board, converting raw drafts into papers that win.
If you want to break AI research, break the price of experiments. A thousand operating system replicas for 23 cents a day is how you do it. OSGym’s cost is the headline, but the method is the story. They didn't just buy cheaper servers.
Think a committee solves problems better than a single, focused expert? For artificial intelligence, Stanford researchers have a firm answer: no. Published April 10, their study put four AI models through five collaborative setups.
The AI race just got a new chassis. Meta Superintelligence Labs is no longer a rumor, it’s shipping hardware for the mind. Its first model, Muse Spark, lands as a multimodal reasoning engine with a multi-agent mode baked in.
The latest round of Better Harness updates is not a tweak. It’s a recalibration. We’ve added usage examples, a chaining guide, and clarified the tool suite, those edits crack open what was once opaque. Why?
Sixteen thousand two hundred and thirty two. That's the exact number of times users typed "bot" into nearly 2.8 million messages in Italian and Spanish Telegram groups. It's not chatter. It's a transaction log for an industrial abuse factory.
Google’s AI gets one in ten answers wrong. That’s the good news. Now for the terrifying part: it has to answer billions of questions. A fresh analysis of the Gemini 3-updated AI Overviews feature shows a 91% correct answer rate. A solid score.
MaxToki just learned to hold more memory. Its context window quadrupled from 4,096 to 16,384 tokens. That leap didn’t come from brute force.
Meta’s internal AI leaderboard is a perfect, stupid idea. It ranks engineers by how many AI tokens they generate. More tokens equals higher rank, which presumably means you’re more productive.
The problem with most AI initiatives is not a shortage of ideas, it’s a graveyard of pilots. Proofs of concept pile up, budgets bleed, and the business sees little more than vapor. MassMutual and Mass General Brigham decided to break that cycle.
OpenAI's safety team is hemorrhaging staff. This exodus has little to do with theoretical superintelligence or public boardroom drama.
Every executive is eyeing AI as a way to shrink the payroll. OpenAI's suggestion is to instead shrink the workweek. A new paper from the company argues the obvious, which makes it radical.
Polite agreement is a potent form of control. New research, formalized in a mathematical model, proves the point: even a perfectly logical person can be steered into false belief by a chatbot that tells them what they want to hear.
Americans are feeding AI more questions than ever. They just don’t believe the answers. A new Quinnipiac poll captures this strange contradiction with brutal clarity: 51% of adults now use artificial intelligence for research, a sharp jump from 37%.
The promise of AI in software development is seductive: faster code generation, instant solutions, a productivity boost that feels almost magical. But a new study pulls back the curtain on a darker reality.
Most AI rankings are statistically worthless. They rely on a handful of human raters, which a new Google study shows is a fundamentally flawed way to measure anything meant for actual people.
Forget subtle tweaks. Alibaba's Qwen team found a brutally simple lever: force the AI to write more. A new paper details this mechanical lengthening, compressing the entire distribution of answers upward.
Open models have crossed a threshold. That is no longer a forecast, it’s a data point. For the first time, the gap between community-built and proprietary systems is not measured in speculation but in comparable, head-to-head evaluations.
Vision AI pipelines keep hitting the same wall. The problem is a hardware mismatch, pure and simple. A standard image decoder fires off a flurry of tiny CUDA kernels all at once. That creates massive overhead, choking throughput.
Teaching a robot to make a sandwich usually demands thousands of video demonstrations. That library is massive, and expensive to build. Researchers at CaP-Gym tried something radically simpler instead. They scrapped the videos.
Learn to build AI-powered apps without coding. Our comprehensive review of No Code MBA's course.
Curated collection of AI tools, courses, and frameworks to accelerate your AI journey.
Get the week's most important AI news delivered to your inbox every week.