Research & Benchmarks - Page 14 of 35
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
The laptop you bought last year is obsolete. Nvidia, the trillion-dollar chip architect for data centers, is now redesigning the personal computer. Its core feature will be an intelligence you never requested.
The numbers are deceptively tight. Three independent leaderboards, Open LLM v2, a twelve-benchmark suite, LiveBench, all converge on a narrow band of effective dimension between 2.86 and 4.80. That is the competitive frontier.
The National Science Foundation just wrote a $50 million check to MIT. For the physicists at the Institute for Artificial Intelligence and Fundamental Interactions, it’s vindication.
We’ve built a labyrinth. Each corridor holds a single, brilliant AI tool. The real problem isn't their intelligence. It's the toll of navigating between them. Your actual work slides to the background.
Every map tells a lie. The good ones show you how. Geospatial machine learning patches the gaps in our data, stitching a picture from scraps.
Every new data science tool promises to free you from drudgery. This one might actually do it. The job is being rewired from the inside, not by a single clever algorithm but by a new kind of software worker. Call it an agent.
Diagnosing Alzheimer's disease hinges on spotting the subtle shift from normal aging to mild impairment, and finally to dementia. A study just posted to arXiv harnesses standard clinical tests to do exactly that, with striking accuracy.
MIT's new AI can finally read a chart. That's not a small thing. Business intelligence is drowning in bar graphs and line plots—data that's instantly clear to any analyst but has always been gibberish to a machine.
A brain-computer interface translates a thought into a click. It’s fragile. Researchers can sabotage the signal with a whisper of digital noise, turning a command for “yes” into “no.” The usual fix is to build a bigger, heavier AI model to withstand...
Mathematics is a discipline of deliberate, grinding thought. There are no shortcuts, only better ideas. Now, a blunt instrument called the Leiden Declaration makes the case that this foundational act is under threat.
Transformer models keep collecting benchmarks. Their latest trophy comes from biomechanics, for predicting the hidden forces inside a walking person's hip. A new paper in *Gait2Hip-60* confirms one now leads the ranking.
Large language models generate Python. They produce C++. Some even output basic quantum assembly. Yet they cannot converse with a quantum computer. Not truly.
The attention mechanism has a secret life, one that depends not just on architecture but on the optimizer that trains it.
A new fault line is cracking through social science research. It’s not about theory. It has nothing to do with methodology. This split is about access, and a recent study puts stark numbers to it.
Recipe AIs are boring. Ask one what goes with chicken and it will list garlic, lemon, thyme. This is because it has read a million recipes and is averaging them out. It knows what humans say goes together, not why.
AI search agents are supposed to be explorers. Instead, they’re more like detectives who only follow the evidence they already expect to find.
OpenAI has decided to weaponize one of its most advanced AI models against the next pandemic. It’s giving the thing away for free. Governments, academic labs, and small teams can now apply for access to the company’s life-sciences model.
Everyone knows AI agents write code. We’ve missed what that code actually is. It's not their final product. It's their working memory, their plan, their entire method of reasoning. A new review paper makes this blunt argument.
A human glances at a banana and a photograph, and the task is instantly clear. A robot, staring at the same scene, drowns in noise. It processes every pixel, every shadow, every irrelevant corner, and gets lost.
Friday at CVPR 2026 isn’t just another afternoon in Exhibit Hall A & F, it’s a microcosm of the field’s most urgent tensions.
Learn to build AI-powered apps without coding. Our comprehensive review of No Code MBA's course.
Curated collection of AI tools, courses, and frameworks to accelerate your AI journey.
Get the week's most important AI news delivered to your inbox every week.