Research & Benchmarks - Page 13 of 35
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
AI coding assistants are great at finding the file. They're terrible at reading it.
OpenAI just beat Elon Musk in court. Now it's facing a different kind of fight. State attorneys general are investigating the company, and OpenAI says it's cooperating. It won't say who or what they're asking for. This isn't happening in a vacuum.
Google's Gemini-SQL2 just hit 80.04% execution accuracy on the BIRD benchmark. That's a specific, hard number. For context, OpenAI's GPT-5.5-xhigh sits at 72.8%. Claude Opus 4.6 is at 70.9%.
Leaderboards are usually marketing noise. This one is different. NVIDIA just topped the first major benchmark for AI agent performance, and the margin isn't close.
Most AI search is still just a fancy text predictor. Perplexity decided to build a factory instead. Its new system breaks a single query into pieces and farms them out to over twenty different AI models, all working at once.
Being able to see is not the same as being able to read. A new comparison of training methods for Chinese characters proves it. The visual model gets a head start. It recognizes that the characters 打, 拍, and 拉 all share the same hand-shaped radical.
We used to think of support as something a computer gave a person. Now the person is often the backup for the computer. This is a real problem. An AI agent booking a flight or managing a supply chain doesn't have a bad day.
Everyone selling AI says it's trustworthy. Most are hoping you don't check. The London Stock Exchange Group is taking a different, more literal approach: they're feeding their own vetted financial data directly into ChatGPT's machinery.
The dashboard is the command center. Hermes Agent’s Profile Builder collapses fragmentation into one coherent pane of glass: identity, model, skills, and MCP servers all flow together.
Forget whether AI can write your emails. The real question is whether it can do science. A new, brutally difficult benchmark called SciConBench makes it clear that the answer, for now, is a hard no.
Most AI that talks to you is guessing. It’s making a probabilistic wager on what you meant. A new paper suggests the smartest move a language model can make isn't a better guess, but knowing when to stop guessing altogether.
Every time a foundation-model agent remembers, it also exposes. That tension, between personalization and privacy, defines a new frontier in agent memory research.
When you train a credit scoring model, you need a tool that ranks risk across a portfolio, not just flags defaults at a single point. Model 5 posted the best numbers for penalized PR-AUC, recall, and F1-score.
The ONNX graph is a labyrinth of nodes, each one a decision point. But when you’re chasing FP8 performance, the path isn’t just about layout, it’s about fusion.
Forget automation. The real sales pitch has pivoted to strategy. Software vendors now hawk systems that don't just complete tasks—they decide which tasks are worth doing. Hand a model an objective like "cut costs," and let it figure out the how.
The most tedious part of federated learning research isn’t the thinking, it’s the iterating. You define a hypothesis, code a variant, run the experiment, log the result, and do it again. And again. The loop is essential but exhausting.
Most AI cost analysis misses the point. The expensive part isn't starting a conversation with the model. It's letting it finish. Inference cost splits in two. Prefill is cheap and fast: you throw your prompt in, the model builds its initial state.
Forget the tidy benchmarks. AI is being thrown against actual scientific work now, with messy data and no clear finish line.
Machine learning models crave a simple fight. Give them a World Cup match, and they'll happily pick the stronger side. But a draw? They despise the very idea.
Reddit conducted a quiet, unsettling experiment. For a period, the platform allowed moderators to secretly flood specific forums with comments generated entirely by artificial intelligence. These comments were crafted to argue like humans.
Learn to build AI-powered apps without coding. Our comprehensive review of No Code MBA's course.
Curated collection of AI tools, courses, and frameworks to accelerate your AI journey.
Get the week's most important AI news delivered to your inbox every week.