Research & Benchmarks - Page 30 of 34
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
Academic AI research, performance benchmarks, scientific breakthroughs, and peer-reviewed studies advancing artificial intelligence frontiers.
The enterprise is finally getting serious about agents, and ServiceNow just raised the bar.
Forget the specialized tools. OpenAI’s newest model doesn’t use a custom math engine or a separate code interpreter. It just uses reinforcement learning, the kind you’d train a game-playing bot with, and a lot of raw compute.
Weather forecasting just got a major upgrade. WeatherNext 2’s forecast data is now live in Earth Engine and BigQuery, and early access on Vertex AI is open for custom model inference. This isn’t a simple tweak.
The music blog Stereogum has been a digital mainstay for nearly two decades, a place where the conversation about new albums, obscure bands, and the culture of listening felt like a genuine hang. But the internet’s economic ground has shifted.
In the arms race of AI, size has long been the ultimate advantage, until now. DeepEyesV2, a smaller open-source model, is punching well above its weight class. How? Not by memorizing more data, but by knowing when to reach for a tool.
Most AI conversations feel like talking to someone with short-term memory loss. You give it your name. It asks for it again two lines later. The promise of a continuous, intelligent dialogue keeps breaking against simple amnesia.
Every AI has a memory limit, a point where it starts making things up. This is not a minor bug. It is the core technical lie behind every demo where a chatbot flawlessly analyzes a novel you just uploaded.
OpenAI has a new plan for cracking open the black box: make the box smaller. For years, the field of mechanistic interpretability, the quest to explain exactly why a neural network does what it does, has been stuck.
AI needs data constantly, and that movement is surprisingly expensive. Every chunk of training data pulled from an object store like S3 traditionally requires the server's central processor to handle the networking chatter.
Indian languages don’t play nice with standard NLP. They share scripts, bleed into each other through code-mixing, and trip over their own morphological complexity.
NVIDIA just ran the table. In the latest MLPerf Training benchmarks, every single result came from a Blackwell system. The win was total. More importantly, it proved a point about precision everyone else missed.
AI agents are not good employees. Left to their own devices, they screw up. But give them a human supervisor, and they become useful. This is the simple, expensive lesson from a new study commissioned by Upwork.
An AI can be perfectly sure of itself and totally wrong. That’s a problem. Humans generally aren’t like that. We feel our uncertainty.
DeepMind’s new AI agent, SIMA 2, makes its predecessor look like it was playing with the controller upside down. This one doesn’t just execute commands. It explains itself. It learns new video games on the fly. It can even read your emojis.
Your $20 monthly ChatGPT subscription is mostly a toy. The real tool is the command line. OpenAI’s Codex CLI, a terminal-based coding assistant, works with that plan. It unlocks the professional-grade utilities hiding behind the chat interface.
Bengaluru has booked another corporate conference for 2026. This one is called The Best Firm Summit, and it's for people who manage humans and people who manage machines.
Google wants to replace your data analyst with a chatbot. The company's new Analytics Advisor tool, announced this week, is another attempt to make sense of metrics by letting you ask questions in plain English. Forget the pivot tables.
AI models are supposed to summarize and synthesize, not recite. A new test suggests they might be doing a lot more of the latter. Researchers have built a tool, called RECAP, to force large language models to cough up what they've memorized.
ElevenLabs has released a transcription tool that claims to hear the future. Their new Scribe v2 doesn't just convert speech to text.
Teaching logic to an AI is a famously stubborn problem. Meta’s researchers just attacked it with a new, pugilistic framework called SPICE. The core mechanic is simple: pit two AIs against each other.
Learn to build AI-powered apps without coding. Our comprehensive review of No Code MBA's course.
Curated collection of AI tools, courses, and frameworks to accelerate your AI journey.
Get the week's most important AI news delivered to your inbox every week.