Research & Benchmarks - Page 11 of 35
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
Searching for a specific simulation model today is a brute-force nightmare. It's like hunting for one uniquely stamped brick in a vast, unmarked warehouse. A new arXiv study, "How Can AI Find My Model?", maps a way out.
Most AI agents follow instructions. A new one rewrites its own instructions, then runs them to see if they work. Researchers call it RSEA, a Recursive Self-Evolving Agent. Its mind is a three-part stack written in plain language.
The weekly grind of market research is brutal. You open a dozen tabs, skim endless articles, and try to stitch together a coherent brief from fragmented signals. Hours vanish. The result? Often a shallow summary, not a strategic insight.
Meta AI’s latest brain-to-text model doesn’t read minds, it reads MEG signals, and it reads them disturbingly well. Brain2Qwerty v2 hits 61% word accuracy, a jump that leaves prior non-invasive methods, stuck at 8%, in the dust.
The MIT Keller Gallery will host “Beyond Data‑Driven Aesthetics” through June 30, a show that pulls together philosophy, mathematics, computer science and design computation into tangible installations and interactive visualizations.
Twenty dollars a month. That’s the price of a streaming subscription, a couple of coffees, or, if you’re a developer, unrestricted access to MiniMax’s coding models across an entire ecosystem of tools.
Sina built a model with three billion parameters that beats giants at logic puzzles. It also stumbles over basic facts. This isn't an accident. It's the point. The VibeThinker-3B is a deliberately small model.
Memory is the current quagmire for AI agents. Too much slows them down, too little makes them forgetful, and every week a new framework claims it's solved the puzzle. MRAgent is the latest to throw its hat in, promising to do more with less.
For years, if you wanted serious AI, you bought from Nvidia. That tidy, lucrative arrangement is now fracturing.
The path to deploying a production-ready NVIDIA AI‑Q Blueprint on Oracle Cloud Infrastructure begins not with code, but with capacity.
Your RAG evaluation scores are probably going up. This is not necessarily good news. Everyone chasing better benchmarks is now wrestling with a familiar ghost: overfitting. You see a number climb, you declare progress.
Everyone building AI governance has a favorite story. One side says permissionless systems breed chaos. The other says corporate models centralize control. Both stories are wrong, and now there's data to prove it.
Inference is a story of two systems, and the story begins with a single millisecond. Consider a transaction authorization. The clock starts with the ISO 8583 budget.
Figma's latest update targets more than interface clutter. It takes aim at the org chart. At its Config conference, the company baked motion graphics, shader effects, and live code layers directly into its design canvas. You prompt.
Amazon just dropped new chips into its cloud. They’re banking you’ll pay for them. The EC2 G7 instances are now powered by NVIDIA’s RTX PRO 4500 Blackwell GPUs. This is a direct upgrade to the previous G6 generation. The numbers are stark.
Forget the lone genius in the lab. The next big discovery might be managed by a CEO made of code. Stanford researchers are pitching this at VB Transform 2026. They've built a hierarchy of AI agents to act as a full scientific team.
Logistic regression spits out coefficients, not a usable credit score. To build one, start with a single, concrete fact: for any variable, the category with the strongest positive link to default gets a baseline of zero points.
We treat annotator disagreement like static to filter out. Maybe we're filtering out the point. A new study asks how many people you actually need to pin that signal down. The answer isn't one number. It depends entirely on what you're measuring.
NVIDIA Blackwell delivers 15 petaflops of dense NVFP4 compute, a staggering amount of raw power. But raw power means nothing if the pipeline can’t feed it. That’s where DFlash comes in.
Lab work is slow. The AI scientist is not. It runs without sleep or salary, folding proteins and aligning sequences in a silent digital loop. NVIDIA’s BioNeMo Toolkit is building it a new lab.
Learn to build AI-powered apps without coding. Our comprehensive review of No Code MBA's course.
Curated collection of AI tools, courses, and frameworks to accelerate your AI journey.
Get the week's most important AI news delivered to your inbox every week.