Research & Benchmarks - Page 18 of 35
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
Compressing the key-value cache without destroying signal fidelity has long been a tug-of-war between quantization error and the noise that quantization itself introduces.
Seventy-six patients walked into a Boston emergency room. Two attending physicians examined them, made their calls. So did two AI models from OpenAI, o1 and 4o.
A 2021 design is outperforming its 2026 successor , across every tested configuration, and by a wide margin.
The US government has declared China's best AI model is falling behind. It's a political verdict. DeepSeek V4 is the current Chinese champion.
Medical AI has one job: get the answer right. A new system from Google DeepMind mostly does, which makes its single, critical failure that much more important.
Anthropic asked a new version of its Claude model to solve 99 problems in bioinformatics. On 76 tasks that at least one human expert could handle, Claude matched their performance. On the 23 that stumped every expert, it got about a third right.
The barrier between imagination and execution just collapsed. You don’t need a single line of code to breathe life into a voice agent, console.x.ai’s playground gives you a blank slate and two paths to build.
Artificial intelligence vision systems are plagued by a stubborn paradox: scrub one bias, and another pops up elsewhere. This relentless game of algorithmic Whac-a-Mole has frustrated researchers for years.
The pitch was simple: keep your life locked on your phone. For on-device AI, that's the whole sell. Privacy isn't a bonus; it's the box. But making that promise work? A nightmare. Weak processors choke. Spotty signals drop.
Elon Musk built an empire selling the future. Now he’s in court selling a story about being its victim.
The bottleneck has always been memory. In structural biology, protein complexes sprawl across thousands of residues, and single‑GPU limits have kept those models caged.
DeepSeek's V4 Pro model just hit the scene, boasting performance that punches at the level of titans like OpenAI's GPT-5.4. It's a major technical statement.
Benchmarks are a game of inches, until they’re not. Poolside AI just took a yard. Its new Laguna XS.2 and M.1 models scored 68.2% and 72.5% on SWE-bench Verified. The bigger number is flashy. The smaller one is the trapdoor.
Tech's oldest trick is following the crowd. Oracle just lit the crowd on fire, torching its legacy enterprise software business to chase artificial intelligence. Not the kind that writes sonnets. This is a pivot into the plumbing.
When an AI agent calls a tool, who watches the watcher? The answer, it turns out, is another agent, but that fixer can break things too. This new architecture splits the work: one agent executes, a second reviews.
Conventional wisdom says you can't keep packing more knowledge into a language model's fixed architecture without a crash. At some point, concepts should start to interfere and performance should plateau.
The Pentagon has a new name on its wish list: Google’s Gemini AI, for secret work. A group of Google’s employees intend to keep it off that list.
Building a functional AI is a grind, a marathon of specialties from wrangling clean data to tuning arcane algorithms.
Drug discovery is a digital quagmire. A single promising molecule might get shoved through fifty different software tools for screening, optimization, and testing. Getting an AI to reliably run that gauntlet? Nearly impossible.
The data warehouse was built for reports. The data lake was built for storage. Neither was built for the mess of thousands of dashboards, each a custom silo, each demanding its own maintenance, each frustrating the users who just need answers.
Learn to build AI-powered apps without coding. Our comprehensive review of No Code MBA's course.
Curated collection of AI tools, courses, and frameworks to accelerate your AI journey.
Get the week's most important AI news delivered to your inbox every week.