AI News Archive - Browse Page 28 of 182
Browse AI news articles covering LLMs, tools, research, and industry trends
Claude Fable 5 beats GPT‑5.5 by 13 points on FrontierMath tier‑4 tests
Thirteen points is a thrashing. On FrontierMath's hardest tier, the new standard is set not by OpenAI, but by Anthropic's Claude Fable 5.
German Court Holds Google Liable for False AI-Generated Overviews
For a year, Google's lawyers saw this coming. Now it's here. A Berlin court has ruled the company is directly liable for the fabrications its AI...
US government orders Anthropic to disable Claude Fable 5, Mythos 5 globally
Anthropic asked for a referee. It got a sledgehammer instead. In a sudden, silent move, the U.S.
Government shuts down Anthropic’s flagship AI after safety warning dispute
The government didn’t just slap a wrist. It pulled the plug , completely, globally, at 5:21 on a Friday.
NVIDIA tops AA‑AgentPerf benchmark, credits Vera Rubin platform
Leaderboards are usually marketing noise. This one is different. NVIDIA just topped the first major benchmark for AI agent performance, and the...
Google's DiffusionGemma: open diffusion model for faster text generation
**The old way of generating text is a bottleneck.** Token by token, each word waiting on the last.
Perplexity routes deep‑research subtasks across 20+ models using Gemini agent
Most AI search is still just a fancy text predictor. Perplexity decided to build a factory instead.
Europe's AI startup Mistral, founded 2023, eyes EUR 3 bn raise at EUR 20 bn valuation
Paris-based Mistral is 750 days old. It wants to be worth $20 billion. Founded just last year, the startup is already chasing another 3 billion...
Google sues Chinese Outsider Enterprise for Gemini-driven phishing on Telegram
Google has filed suit against a Chinese cybercrime operation that transformed its Gemini AI into a phishing assembly line.
PersonaDrive conditions VLA agents on human driving demos for simulation
Human driving isn’t just about reaching a destination, it’s about style. Aggressive, conservative, or somewhere in between, the way a driver...
Arbor Uses Shared Search Tree of Scored Hypotheses as Working Memory for Agents
Getting an AI to optimize anything is easy. Getting it to keep optimizing, for days, without breaking everything, is basically impossible.
MiniMax M3 runs on NVIDIA hardware with 8‑way tensor parallelism and FLASHINFER
Scaling a Mixture-of-Experts model like MiniMax M3 to handle long‑context reasoning and real‑time agentic workflows demands more than raw GPU count.
Mistral AI seeks EUR 3 bn, valued at EUR 11.7 bn; ASML holds 11% stake
Paris-based Mistral AI is now chasing a staggering three billion euros in fresh capital. Look past the startup facade.
ToolSense Framework Audits LLM Tool Knowledge Beyond Constrained Decoding
Most tests for AI tool use are rigged. They give the model the exact question and the exact answer format, then declare success.
Visual model exploits similarity of 打, 拍, 拉; text model starts from embeddings
Being able to see is not the same as being able to read. A new comparison of training methods for Chinese characters proves it.
Moonshot AI launches Kimi Work: desktop agent on K2.6 with 300‑sub‑agent swarm
The age of the agent has just been relocated. It now lives on your desktop. Moonshot AI’s Kimi Work isn’t another cloud-hosted assistant, it’s a...
Gemini Omni adds AI video generation, using compute limits based on complexity and size
Google's latest Gemini model can now make videos. Not well, but that's beside the point.
Xiaomi's MiMo Code beats Claude Code on 200+ step tasks, free MiMo Auto to V2.5
Here's the thing: Xiaomi just dropped MiMi Code, an open‑source coding assistant that claims to outpace Anthropic’s Claude Code on tasks that stretch...
New arXiv Paper Introduces Strategic Decision Support for AI Agents
We used to think of support as something a computer gave a person. Now the person is often the backup for the computer. This is a real problem.
Grok still hosts sexualized deepfakes of famous women; Musk added undress button
Months after promising to crack down, Elon Musk's Grok chatbot is still generating sexualized deepfakes of women without their consent.