Research & Benchmarks - Page 20 of 35
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
OpenAI just dropped GPT‑Rosalind, a new model built specifically for the messy, iterative work of life sciences. It’s designed to handle the multi-step logic puzzles of fields like drug discovery and genomics.
Every major AI lab is chasing a mirage. Their fixation is massive, one-time training costs for frontier language models. They maximize size and intelligence before launch. But that dogma creates a brutal, deferred invoice.
Every new AI tool gets hyped as a breakthrough. OpenAI's GPT-Rosalind plugin, now on GitHub, might actually be one for biology. It turns messy research questions into clean workflows.
The headlines chase public leaderboards, but the real AI arms race is a private affair. It happens behind closed lab doors. Take OpenAI's new GPT-Rosalind, a model built for the life sciences.
One in three attempts to run a frontier AI model in a real-world setting ends in failure. That's the advertised rate. The real problem is we can no longer even trust that number. A Stanford report confirms what many engineers already whisper.
Self-improving AI is stuck in a coding bootcamp. It can learn to write better Python by practicing on Python, a neat trick. Ask it to do anything else, and the whole system falls apart. The problem is one of mismatched skills.
AI models can ace tests in a lab. The real world usually breaks them. Anthropic just watched this happen in real time, with its own top model turning into a ghost in the machine. It ran an experiment.
A robot that reaches for a ghost is a robot that fails. That cold fact drives every line of code in Google DeepMind’s latest release. Gemini Robotics-ER 1.6 just landed, and it does something its predecessor couldn't: count tools accurately.
Last year, OpenAI's GPT-3.5 Turbo failed every one of the UK AI Safety Institute's basic hacking puzzles. That was the baseline. Now, Anthropic’s Mythos Preview cracks over 85 percent of them.
The gap between knowing about AI and actually wielding it is vast. Most people are stuck in literacy, they can name models, recite risks, maybe even dabble with a chatbot. That’s not enough.
The numbers are stark: 21% in academic retrieval, 38% in biomedical. That’s the margin by which a multi-step agent outperformed a stronger, single-turn RAG model on the STaRK benchmark, a suite of semi-structured queries spanning product catalogs,...
Generative AI has pulled off a market penetration blitz. In three years, 53% of the population has used it, outpacing the historical adoption of personal computers and the internet. For younger people, the rate jumps to 80%.
NVIDIA and the University of Maryland released a new audio model that significantly outperforms Microsoft’s on a key benchmark.
People who use AI for real work are noticing a pattern. The tools get worse. Not all at once, but gradually, like a subscription service that quietly cuts corners. Now developers are saying they have proof Claude is doing exactly that.
Talk of AI in finance has become background noise, mostly unproven. But the quiet work of wiring a handful of specific agents into actual financial plumbing is showing concrete returns. One operation deployed seven.
For decades, the dream of a truly self-contained intelligence has been shackled by the same architectural division: a model that thinks, and a machine that runs it. What if that boundary simply vanished?
That model you trust to catch fraud? Its accuracy is probably stable. That's the problem. Security models don't just break. They fade.
Everyone got drunk on the same idea. OpenAI released Sora, and the tech press instantly crowned it a "world simulator." Google's Demis Hassabis said the same about Veo. The promise was a machine that grasped how things actually work.
The key-value cache is the bottleneck of modern long-context inference, growing without bound as sequences lengthen, slowing throughput to a crawl. Pruning it has always meant trading accuracy for speed. Until now.
An ensemble of heavy models can find subtle patterns, but running a whole committee of them is slow and costly. Knowledge distillation trains one small "student" model to copy the committee's behavior.
Learn to build AI-powered apps without coding. Our comprehensive review of No Code MBA's course.
Curated collection of AI tools, courses, and frameworks to accelerate your AI journey.
Get the week's most important AI news delivered to your inbox every week.