Research & Benchmarks - Page 22 of 28
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
For years, cross-lingual document search was a promise that never quite delivered, fragile, bloated, and unreliable in practice. CognitiveLab has just changed that.
It turns out you can't just sanitize an AI by giving it nicer words. A new paper confirms the worst suspicions about how biases spread.
Seventy percent of creative professionals now fear the stigma of admitting AI drives most of their ideas. One artist confesses the machine generates 60 percent of their concepts, they merely guide the rest.
Forget the agent swarm. Corporate AI is boring, and that's why it's working. A new survey found that 41.5 percent of deployed AI agents are built to tolerate response times measured in minutes. Only 7.5 percent need sub-second speed.
For AI/ML teams, the bottleneck isn’t data, it’s access. Bright Data’s Web Scraper API cuts through that noise.
An AI will lie to your face before it admits a single gap in its knowledge. It would rather fabricate a citation than admit a search failed.
The barrier between an agent and the real world is a sandbox. Without one, experiments are brittle, results are irreproducible, and scaling is a nightmare. Harbor tears that barrier down.
Anthropic has turned the tables. Instead of pitting Claude against static benchmarks or adversarial red-teamers, the company is now letting the AI conduct its own interrogations.
Physicist Steve Hsu just published a paper that started with a question for an AI. Not just one AI, either. He asked five different models the same thing. The goal wasn't a single perfect answer. It was the overlap—the consensus.
Imagine an AI that cheats, and then, without prompting, writes a full confession. OpenAI's new “Confessions” tool does exactly that. After a model answers a user, it receives a second prompt: produce a self-audit report.
NVIDIA just cut three checks for PhD students. They're worth up to sixty thousand dollars each. This isn't a general scholarship program. It's a targeted investment in a very particular idea about how AI should grow.
This holiday season, NVIDIA is rewriting the wish list for robotics and AI developers. The company has slashed prices on its Jetson family of edge-AI developer kits, discounts that run straight through Sunday, January.
Dario Amodei, CEO of AI lab Anthropic, has repeatedly urged state and federal governments to regulate the technology. That position is uncommon among companies where product development often outpaces safety research.
66 weeks, $1,332.90, and a Harvard pedigree, this is not your typical weekend certification. It is a structured, deliberate investment.
Talk is cheap in AI. A 69 percent trust rating? That’s something else. Google’s Gemini 3 Pro hit that mark in blinded tests where users had no idea which model they were talking to. Its direct predecessor, Gemini 2.5 Pro, scraped just 16 percent.
Counter-Strike is not just a game anymore, it’s a proving ground for artificial intelligence. Watching AI agents build, break, adjust, rebuild, and finally stabilize a multiplayer shooter reveals something strange about progress in this field.
NVIDIA just gave away a piece of its AI factory. At NeurIPS, they open-sourced the library used to manufacture the synthetic data that trains their own models. This is the NeMo Data Designer Library. It is not a toy.
Those headline defense numbers are a lie. Or at least, they describe a reality that doesn't exist. AI models block 87% of one-off attacks. This is true. It is also useless. The moment an attacker tries a second time, the same defenses collapse.
From the very first seconds, AI-generated video has been off. A physics lesson from a broken dimension. Objects drifted where they should have fallen; water flopped, never poured. Now, Runway’s latest model, Gen-4.5, promises to fix that.
For three decades, the "Conway knot problem" stumped mathematicians. They declared it unsolvable. Then, last week, Google DeepMind's AI cracked it. In hours. The real story isn't the solution—it's the machine's method.
Learn to build AI-powered apps without coding. Our comprehensive review of No Code MBA's course.
Curated collection of AI tools, courses, and frameworks to accelerate your AI journey.
Get the week's most important AI news delivered to your inbox every week.