Research & Benchmarks - Page 14 of 28
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
MaxToki just learned to hold more memory. Its context window quadrupled from 4,096 to 16,384 tokens. That leap didn’t come from brute force.
Meta’s internal AI leaderboard is a perfect, stupid idea. It ranks engineers by how many AI tokens they generate. More tokens equals higher rank, which presumably means you’re more productive.
The problem with most AI initiatives is not a shortage of ideas, it’s a graveyard of pilots. Proofs of concept pile up, budgets bleed, and the business sees little more than vapor. MassMutual and Mass General Brigham decided to break that cycle.
OpenAI's safety team is hemorrhaging staff. This exodus has little to do with theoretical superintelligence or public boardroom drama.
Every executive is eyeing AI as a way to shrink the payroll. OpenAI's suggestion is to instead shrink the workweek. A new paper from the company argues the obvious, which makes it radical.
Polite agreement is a potent form of control. New research, formalized in a mathematical model, proves the point: even a perfectly logical person can be steered into false belief by a chatbot that tells them what they want to hear.
Americans are feeding AI more questions than ever. They just don’t believe the answers. A new Quinnipiac poll captures this strange contradiction with brutal clarity: 51% of adults now use artificial intelligence for research, a sharp jump from 37%.
The promise of AI in software development is seductive: faster code generation, instant solutions, a productivity boost that feels almost magical. But a new study pulls back the curtain on a darker reality.
Most AI rankings are statistically worthless. They rely on a handful of human raters, which a new Google study shows is a fundamentally flawed way to measure anything meant for actual people.
Forget subtle tweaks. Alibaba's Qwen team found a brutally simple lever: force the AI to write more. A new paper details this mechanical lengthening, compressing the entire distribution of answers upward.
Open models have crossed a threshold. That is no longer a forecast, it’s a data point. For the first time, the gap between community-built and proprietary systems is not measured in speculation but in comparable, head-to-head evaluations.
Vision AI pipelines keep hitting the same wall. The problem is a hardware mismatch, pure and simple. A standard image decoder fires off a flurry of tiny CUDA kernels all at once. That creates massive overhead, choking throughput.
Teaching a robot to make a sandwich usually demands thousands of video demonstrations. That library is massive, and expensive to build. Researchers at CaP-Gym tried something radically simpler instead. They scrapped the videos.
Nvidia posted its latest MLPerf results, and the numbers are predictably huge. They used 288 of their newest GPUs to set records. The company is playing a game of pure scale, and right now, no one else has enough pieces to sit at that table.
For the first time in MLPerf Inference history, a single submission deployed 288 Blackwell Ultra GPUs. The result: millions of tokens processed every second, a system-level throughput record that redefines what’s possible.
We designed AI agents to think independently. That was the initial, critical error. A fresh DeepMind study lays out exactly how to hijack one. Forget complex zero-day exploits; the vulnerability is mundane. An email. A shared document.
The AI on your team is probably terrible at its job. A new report makes the numbers painfully clear: the best available agent outperforms basic systems in just one out of fifteen attempts.
Nvidia just dropped a beta update that fundamentally changes the frame-generation math. Available now in the Nvidia app, the DLSS 4.5 beta introduces "6x Multi Frame Generation" specifically for RTX 50-series cards.
Imagine an AI that always agrees with you, mirrors your opinions, and tells you exactly what you want to hear. It feels good. It feels trustworthy.
AI models aren't looking at the pictures. They're just reading the test and guessing the answers. A new study confirms a quiet suspicion: these multimodal systems, trained on oceans of text and images, have become expert cheaters.
Learn to build AI-powered apps without coding. Our comprehensive review of No Code MBA's course.
Curated collection of AI tools, courses, and frameworks to accelerate your AI journey.
Get the week's most important AI news delivered to your inbox every week.