Research & Benchmarks - Page 21 of 28
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
A new benchmark lands, and GPT-5.2 sits at the top. OpenAI's FrontierScience test is designed to be grueling: each problem demands three to five hours of work, scored on a ten-point rubric, with the model grading itself.
The road to safe, autonomous mobility has never been more data-hungry. Robotaxis must navigate infinite edge cases, icy crosswalks, jaywalking pedestrians, the chaotic dance of a busy intersection.
The silence of a text dataset is deceptive. It never mumbles, never trips over a word, never battles the hum of an air conditioner in the background.
Most customer service AI is a scripted dead end. Fastweb and Vodafone built a bossy one that thinks. For 9.5 million of their Italian customers, a request doesn't trigger a canned reply.
GPT-5.2 Thinking doesn’t answer your questions. It builds your project. For web developers tired of babysitting AI through incomplete, half-baked logic, this model arrives as something rarer: a partner that reasons end-to-end.
The AI learning landscape is crowded with hours-long lectures and dense textbooks. But one YouTube channel cuts through the noise, serving complex concepts in under a minute, stripped of fluff yet never dumbed down.
The consultants missed the point. Budgets get spent on real, expensive problems. Sharp companies knew this. They'd mock up a fix with duct tape, test it against actual numbers, and learn. AI didn't start that fire. It poured gasoline on it.
Every developer building AI workflows is chasing multi-agent systems right now. The pitch is compelling: divide the labor, conquer the complexity. New research from Google and MIT throws cold water on that plan. It often backfires.
Companies keep trying to build AI that sounds less robotic. According to a new study, their best efforts mostly make it worse.
AI2 just sharpened its scalpel. With the launch of Olmo 3.1 32B Think, the institute has hacked significant chunks off math and reasoning benchmarks: a five-plus-point surge on AIME, four-plus on ZebraLogic.
The machine never lies, but it does edit, tweak, and rewrite in shades of gray. Pangram’s latest detector cuts through that ambiguity. It claims 99.98% accuracy. And it finally acknowledges what we all know: AI writing isn’t binary.
Every ChatGPT prompt doesn't drain a bottle of water. That stat is basically an urban legend, and it distracts from a less dramatic truth. Experts who track this stuff say data centers often use less water than people assume.
Power depends on sand. Specifically, the kind you refine into silicon for computer chips. That simple, fragile fact has become the latest arena for geopolitical maneuvering.
46.4% on Humanity’s Last Exam. 66.1% on DeepSearchQA. And 59.2% on BrowseComp, our best ever. The new Gemini Deep Research agent doesn’t just inch ahead; it clears the bar. These aren’t incremental gains.
The number 70% is a haunting figure for anyone betting on AI in the enterprise. Google’s new FACTS benchmark doesn’t just measure how smart a model is, it measures how often it tells the truth.
Your coding agents shouldn’t be blind. Yet, every time they triage a failed execution, they’re flying without a cockpit. LangSmith Fetch changes that.
The consultant of tomorrow won’t be buried in technical documentation. They’ll be freed by an AI that knows the answers, before they even ask.
Machine learning pipelines are notoriously finicky. Choose the wrong scaler, a suboptimal model, or an ill-suited set of hyperparameters, and your results suffer.
The physics of intelligence has a shortcut. A large, expensive model, bloated with billions of parameters, can be coaxed into teaching a far smaller one almost everything it knows.
The most important prompt isn’t the one you feed the AI. It’s the one you feed the prompt itself. Inside Google, a researcher named Anna has cracked a quiet art form: meta-prompting. She doesn’t just ask Gemini to make a video.
Learn to build AI-powered apps without coding. Our comprehensive review of No Code MBA's course.
Curated collection of AI tools, courses, and frameworks to accelerate your AI journey.
Get the week's most important AI news delivered to your inbox every week.