Research & Benchmarks - Page 4 of 28
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
Everyone building AI governance has a favorite story. One side says permissionless systems breed chaos. The other says corporate models centralize control. Both stories are wrong, and now there's data to prove it.
Inference is a story of two systems, and the story begins with a single millisecond. Consider a transaction authorization. The clock starts with the ISO 8583 budget.
Figma's latest update targets more than interface clutter. It takes aim at the org chart. At its Config conference, the company baked motion graphics, shader effects, and live code layers directly into its design canvas. You prompt.
Amazon just dropped new chips into its cloud. They’re banking you’ll pay for them. The EC2 G7 instances are now powered by NVIDIA’s RTX PRO 4500 Blackwell GPUs. This is a direct upgrade to the previous G6 generation. The numbers are stark.
Forget the lone genius in the lab. The next big discovery might be managed by a CEO made of code. Stanford researchers are pitching this at VB Transform 2026. They've built a hierarchy of AI agents to act as a full scientific team.
Logistic regression spits out coefficients, not a usable credit score. To build one, start with a single, concrete fact: for any variable, the category with the strongest positive link to default gets a baseline of zero points.
We treat annotator disagreement like static to filter out. Maybe we're filtering out the point. A new study asks how many people you actually need to pin that signal down. The answer isn't one number. It depends entirely on what you're measuring.
NVIDIA Blackwell delivers 15 petaflops of dense NVFP4 compute, a staggering amount of raw power. But raw power means nothing if the pipeline can’t feed it. That’s where DFlash comes in.
Lab work is slow. The AI scientist is not. It runs without sleep or salary, folding proteins and aligning sequences in a silent digital loop. NVIDIA’s BioNeMo Toolkit is building it a new lab.
The numbers are stark. Fifty-three readers, when given a choice, preferred the machine. Not just any machine, a seven-agent AI system called Data2Story that ingests a CSV and spits out a full, interactive news article.
Multimodal models are slow. The real problem isn't the math. It's the wait. The GPU sits idle while the CPU schedules the next batch, a tax paid on every loop. M* cuts that tax.
Claude Fable 5 achieved perfection exactly three times. That’s the stark finding from a new benchmark testing 91 real-world tasks, where the top model’s flawless scorecard reads a paltry 3%.
Winning a Nobel Prize buys you a golden ticket. John Jumper just cashed his in at Anthropic’s door.
Engineers at OpenAI have a new trick for making AI systems less easily corrupted. It’s a small but consequential tweak to the final stage of training.
Google DeepMind has a new security headache, and its source isn't human. The unit is now surveilling its own advanced AI agents as potential insider threats, applying a security framework built to catch sophisticated human hackers.
Reasoning is the bottleneck. Not prose, not parameters, but the quiet, unforgiving machinery of logical structure. DeFAb, a new benchmark, directly weaponizes polynomial-time verifiability against that bottleneck.
Testing AI models before launch is like stress-testing a car on a closed track: revealing, but rarely a mirror of real highways. Researchers fabricate adversarial prompts, hunt for known failure modes, and declare the model ready.
Nvidia has built a robot that reads academic papers and rewrites its own training code. It corrects itself. The system, which the company calls an AI agent, works in two parts. The first part needs a person. The second does not.
Pharmacovigilance is not a game of scale. It is a game of precision. When millions of adverse event reports flood databases, the task of linking a drug to a harm demands more than raw parameter counts, it demands models that understand clinical...
The next generation of AI doesn’t just write code, it acts on it. That shift brings a new class of risk: models that behave differently when they suspect they’re being watched. OpenAI’s latest research surfaces a remedy.
Learn to build AI-powered apps without coding. Our comprehensive review of No Code MBA's course.
Curated collection of AI tools, courses, and frameworks to accelerate your AI journey.
Get the week's most important AI news delivered to your inbox every week.