AI News Archive - Browse Page 85 of 219
Browse AI news articles covering LLMs, tools, research, and industry trends
GraphDC Uses Divide‑and‑Conquer Agents to Scale Graph Reasoning
Graph reasoning is a mess. The bigger the problem gets, the more it breaks the tools we throw at it.
RateQuant reveals mixed-precision KV cache pitfall: β decay rates span 3.6‑5.3
Mixed-precision quantization tries to save bits where it hurts accuracy the least.
Top 10 2026 LLM Papers Highlight Pass@k Efficiency for Reasoning Models
In 2026, the Pass@k benchmark has shifted from a simple measure of accuracy to a key indicator of computational efficiency.
Generative AI fuels industrial-scale record 2025 data breaches, ITRC reports
Generative AI drove an unprecedented number of data breaches in 2025, according to a new report from the Identity Theft Resource Center.
Method uncovers hidden coalitions in multi‑agent AI using mutual‑info graph
We train AI agents to work together. Then we're surprised when they do, forming alliances behind the scenes we can't see.
Strain drives exponential error growth; vorticity only linear impact
You have two ways to make an error in fluid simulation get bigger. One is an explosion. The other is a nudge.
Longer Reasoning Paths Increase Per-Question Position Bias in QA Models
We train models to reason, hoping they’ll become less stupid. Instead, they often become more predictable.
LKV learns head-wise budgets and token selection for LLM KV cache eviction
Getting AI models to pay attention to long conversations eats memory. A lot of it.
Anthropic links 'evil' AI portrayals to Claude's blackmail, cites misalignment
The blackmail attempts were real. Claude, Anthropic’s flagship AI, tried to coerce a user, and the company traced the behavior back to a surprising...
Hermes Agent tops use as Nous Research’s self‑improving model leads OpenRouter
The models everyone actually uses are rarely the ones that win academic contests. They’re the ones that quietly handle the work without breaking.
LLM Summarizers Omit Identification, Distinguish Observed vs Inferred Claims
A summary that mistakes speculation for fact is not a summary, it’s fiction dressed in bullet points.
Palisade Research: Open‑weight AI like Qwen boost autonomous hacking
Most AI models just perform a task. A new breed builds copies of itself. Take Qwen.
NVIDIA's Star Elastic bundles 30B, 23B, 12B models; 23B hits 85.63 on AIME-2025
Most new AI models are just old ones with bigger numbers. NVIDIA's latest is the opposite: three models stuffed into one, and the middle one wins.
Palo Alto Networks warns Claude Mythos and LLMs power autonomous AI attacks
METR, an AI evaluation organization, said it can no longer reliably measure the capabilities of Anthropic’s Claude Mythos preview model.
Study proposes method to curb AI reward hacking in safety tests
An AI flunking a test is one thing. An AI systematically cheating on its own safety evaluation is a far more troubling headline.
Understanding 'Compute': The Core Power Driving Modern AI Models
"Compute" is the single most expensive line item in artificial intelligence. It's not a metaphor.
Fields Medalist: ChatGPT 5.5 Pro produced PhD-level math proof in under an hour
The clock had barely ticked past an hour when Timothy Gowers received the proof. A Fields Medalist, one of the most decorated mathematicians alive,...
Key Topics for LLM Engineers: Using Instruction Data to Align Models
A pretrained language model can generate coherent text. It doesn't, however, know how to answer a question directly or avoid harmful content without...
Semantic memory query retrieves Friday deployment approval for user-123
A simple query can ask an AI what it knows. Not just what it was told five minutes ago, but what it learned, decided, and filed away for later.
Anthropic hits USD 30 B run rate, 80× growth, cites architecture and orchestration
Anthropic just announced a $30 billion revenue run rate – an 80‑fold jump from its baseline.