Research & Benchmarks - Page 3 of 28
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
Meta pulled the plug on an internal leaderboard that tracked how much AI tokens its employees were burning through, after the company's AI spending put it on pace to hit billions of dollars in costs by 2026.
Demis Hassabis wants a new referee for AI, and he's borrowing the playbook from Wall Street.
A German research consortium coordinated by the KI Bundesverband has released Soofi S 30B-A3B, an open-source language model trained entirely on Deutsche Telekom's Industrial AI Cloud in Munich.
Virtual screening runs against millions to billions of compounds, and co-folding models like OpenFold3 often produce the most accurate structures in the batch. The catch is cost.
Motors scream. Pans sizzle. A child on a swing arcs higher with every push. Hanns Christoph Nägerl's lab at the University of Innsbruck just broke that universal rule.
Sysdig researchers said last week they'd found the first documented case of "agentic ransomware," an extortion campaign called JadePuffer where an AI agent handled the technical work of a real cyberattack on its own.
A team testing agent configurations ran into a familiar problem: change the model, rewrite the prompt, swap a retrieval tool, and the average score barely moves. One version edges out another by two points.
Building a better federated learning algorithm is a grind. Researchers must make a cascade of small, interlocking decisions—how clients update models, how the server combines them—and comparing approaches is notoriously messy.
Most AI benchmarks are useless now. The good ones are too easy. The latest large language models score over 90% on tests that were considered hard just a few years ago, rendering them incapable of showing which system is actually better.
Searching for a specific simulation model today is a brute-force nightmare. It's like hunting for one uniquely stamped brick in a vast, unmarked warehouse. A new arXiv study, "How Can AI Find My Model?", maps a way out.
Most AI agents follow instructions. A new one rewrites its own instructions, then runs them to see if they work. Researchers call it RSEA, a Recursive Self-Evolving Agent. Its mind is a three-part stack written in plain language.
The weekly grind of market research is brutal. You open a dozen tabs, skim endless articles, and try to stitch together a coherent brief from fragmented signals. Hours vanish. The result? Often a shallow summary, not a strategic insight.
Meta AI’s latest brain-to-text model doesn’t read minds, it reads MEG signals, and it reads them disturbingly well. Brain2Qwerty v2 hits 61% word accuracy, a jump that leaves prior non-invasive methods, stuck at 8%, in the dust.
The MIT Keller Gallery will host “Beyond Data‑Driven Aesthetics” through June 30, a show that pulls together philosophy, mathematics, computer science and design computation into tangible installations and interactive visualizations.
Twenty dollars a month. That’s the price of a streaming subscription, a couple of coffees, or, if you’re a developer, unrestricted access to MiniMax’s coding models across an entire ecosystem of tools.
Sina built a model with three billion parameters that beats giants at logic puzzles. It also stumbles over basic facts. This isn't an accident. It's the point. The VibeThinker-3B is a deliberately small model.
Memory is the current quagmire for AI agents. Too much slows them down, too little makes them forgetful, and every week a new framework claims it's solved the puzzle. MRAgent is the latest to throw its hat in, promising to do more with less.
For years, if you wanted serious AI, you bought from Nvidia. That tidy, lucrative arrangement is now fracturing.
The path to deploying a production-ready NVIDIA AI‑Q Blueprint on Oracle Cloud Infrastructure begins not with code, but with capacity.
Your RAG evaluation scores are probably going up. This is not necessarily good news. Everyone chasing better benchmarks is now wrestling with a familiar ghost: overfitting. You see a number climb, you declare progress.
Learn to build AI-powered apps without coding. Our comprehensive review of No Code MBA's course.
Curated collection of AI tools, courses, and frameworks to accelerate your AI journey.
Get the week's most important AI news delivered to your inbox every week.