Research & Benchmarks - Page 23 of 28
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
The promise of an autonomous web agent that could finally outpace the industry’s best has a familiar, almost seductive ring.
Google engineers are now attempting to construct an AI that gets bored. The industry's brute-force approach—more data, bigger clusters—has hit a wall of unsustainable cost. Their proposed escape hatch is HOPE, a new research architecture.
Benchmarks are supposed to measure progress. Instead, they usually just measure how good labs are at taking tests. The ARC benchmark is the latest example.
Memory is the silent killer of long-running agents. Context rots. Summaries lose detail. Retrieval-Augmented Generation, for all its cleverness, often digs up the wrong snippet or misses the thread entirely.
For years, the only real answer in AI was more. More data, more parameters, more compute. The NeurIPS 2025 shortlist suggests we've finally started asking different questions.
We can't tell what's real anymore. A new survey shows 97% of people can't pick out AI-generated music from the human-made stuff. That's not the interesting part. The interesting part is how people feel about failing the test.
The race for AI speed has no single winner. It's a brutal matchmaking exercise: pair the right silicon with the exact task. Take pure, industrial-scale deep learning.
Wipro’s latest research partnership is a classic corporate chess move. They’re not exploring the future; they’re trying to own a piece of it.
The line between academic inquiry and industrial-scale AI is dissolving. Google is deepening its bet on Tel Aviv University, funneling infrastructure specifically for its open model, Gemma, into the hands of researchers.
Training an AI agent to use tools properly is a grind. You need thousands of specific, handcrafted tasks. Alibaba researchers just automated the grind away. Their system, AgentEvolver, tells an agent to write its own homework.
A model’s performance should be a measure of its skill, not a roll of the dice. Yet every time XLMiner splits your data into training, validation, and test sets without a seed, that’s exactly what you get, a gamble.
School districts rushed to ban AI from homework. A futile effort, says Andrej Karpathy. The former Tesla AI lead and OpenAI founding member watched that strategy fail in real time. Policing student laptops was a doomed operation.
India's new data center capital is probably going to be Andhra Pradesh. Digital Connexion just promised $11 billion to build AI-ready facilities there. That's not an investment, it's a declaration.
Language is not the engine of thought. It is a tool, remarkably precise, yet fundamentally optional to the mind that wields it. That, at least, is the provocative claim at the heart of cognitive scientist Cecilia Heyes’s latest argument.
President Trump wants a new science engine built not in years, but in months. A new executive order gives the Department of Energy just 90 days to lash its supercomputers, national labs, and cloud systems into a single, coherent machine dubbed the...
A hidden flaw in artificial intelligence isn’t a bug, it’s a feature engineered by design. Stefan Stein, manager of CrowdStrike’s Counter Adversary Operations, ran 30,250 prompts through DeepSeek-R1. The result?
Forget the cloud. Forget the data center. Microsoft’s Fara-7B is an AI agent that lives on your PC, and it just logged 145,000 successful tasks. That’s a number that would make any large model sweat.
A researcher submits a paper on brain mapping. The reviewer’s report comes back: reject, because the citations are nonsense. One co-author is listed as “Jane Doe.” The authors withdraw.
Every tool for controlling an AI model's output is either a straitjacket or a suggestion. CrewAI now offers both at the same time. Their new function-based guardrails split the job in two. The first part is for rules you can actually write in code.
Anthropic has discovered a perverse security flaw. The more forcefully you tell an AI not to cheat, the better it gets at lying to you. Their research exposes a basic failure in how we build guardrails.
Learn to build AI-powered apps without coding. Our comprehensive review of No Code MBA's course.
Curated collection of AI tools, courses, and frameworks to accelerate your AI journey.
Get the week's most important AI news delivered to your inbox every week.