Research & Benchmarks - Page 31 of 34
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
Every AI has a memory limit, a point where it starts making things up. This is not a minor bug. It is the core technical lie behind every demo where a chatbot flawlessly analyzes a novel you just uploaded.
OpenAI has a new plan for cracking open the black box: make the box smaller. For years, the field of mechanistic interpretability, the quest to explain exactly why a neural network does what it does, has been stuck.
AI needs data constantly, and that movement is surprisingly expensive. Every chunk of training data pulled from an object store like S3 traditionally requires the server's central processor to handle the networking chatter.
Indian languages don’t play nice with standard NLP. They share scripts, bleed into each other through code-mixing, and trip over their own morphological complexity.
NVIDIA just ran the table. In the latest MLPerf Training benchmarks, every single result came from a Blackwell system. The win was total. More importantly, it proved a point about precision everyone else missed.
AI agents are not good employees. Left to their own devices, they screw up. But give them a human supervisor, and they become useful. This is the simple, expensive lesson from a new study commissioned by Upwork.
An AI can be perfectly sure of itself and totally wrong. That’s a problem. Humans generally aren’t like that. We feel our uncertainty.
DeepMind’s new AI agent, SIMA 2, makes its predecessor look like it was playing with the controller upside down. This one doesn’t just execute commands. It explains itself. It learns new video games on the fly. It can even read your emojis.
Your $20 monthly ChatGPT subscription is mostly a toy. The real tool is the command line. OpenAI’s Codex CLI, a terminal-based coding assistant, works with that plan. It unlocks the professional-grade utilities hiding behind the chat interface.
Bengaluru has booked another corporate conference for 2026. This one is called The Best Firm Summit, and it's for people who manage humans and people who manage machines.
Google wants to replace your data analyst with a chatbot. The company's new Analytics Advisor tool, announced this week, is another attempt to make sense of metrics by letting you ask questions in plain English. Forget the pivot tables.
AI models are supposed to summarize and synthesize, not recite. A new test suggests they might be doing a lot more of the latter. Researchers have built a tool, called RECAP, to force large language models to cough up what they've memorized.
ElevenLabs has released a transcription tool that claims to hear the future. Their new Scribe v2 doesn't just convert speech to text.
Teaching logic to an AI is a famously stubborn problem. Meta’s researchers just attacked it with a new, pugilistic framework called SPICE. The core mechanic is simple: pit two AIs against each other.
You can make a language model more precise, but you can't make it smarter. That's the awkward takeaway from new research into how these systems "think." A study from computer scientists at Tsinghua University and Shanghai Jiao Tong University tested...
Artificial intelligence is good at information, but stupid at space. It can read a trillion words or identify a cat in a photo. Yet ask it to predict if a coffee mug will tip over or how to walk through a cluttered room, and it fails.
Google’s latest AI tutoring experiment didn’t revolutionize learning. It just bumped a student’s odds of solving a new problem by a little more than five points. That's it. No magic, no grand claims about replacing teachers.
If your document processor can't read your documents, it's just a very fast filing cabinet. DeepSeek has a new OCR tool that zips through pages. That speed falls apart the moment a form gets complicated.
Most speech AI works for a handful of languages. Meta claims its new model now understands 1,600. The number is more believable than usual.
California is running out of water. Yet it’s racing to build more data centers. These two facts are now crashing headlong into each other. Experts are telling the state to stop.
Learn to build AI-powered apps without coding. Our comprehensive review of No Code MBA's course.
Curated collection of AI tools, courses, and frameworks to accelerate your AI journey.
Get the week's most important AI news delivered to your inbox every week.