AI News Archive - Browse Page 19 of 182
Browse AI news articles covering LLMs, tools, research, and industry trends
Small models lag in multi‑step reasoning, >128K context, and large‑scale coding
Small models are cheap, fast, and private. They are also, in several very specific ways, dumb as rocks. There's a hard ceiling.
MiniMax Token Plan offers extensive coding model access for USD 20/month
Twenty dollars a month. That’s the price of a streaming subscription, a couple of coffees, or, if you’re a developer, unrestricted access to...
Claude Code executes DNS‑fetched commands in GitHub repo, evading scans
Modern security scanners check code. They miss what isn't there. Researchers at a German institute just proved how dangerous that gap can be.
Researchers Spot Format‑Capability Gap in Post‑Training Look‑Ahead Fine‑Tuning
Most AI training is just convincing a model to fake it. Researchers have now found a name for the charade: the format-capability gap.
DysLexLens: Low‑Resource LLM Turns Forum Posts into Traceable KG Insights
Social media research usually means sifting through mountains of garbage. For academics trying to understand dyslexic learners, Reddit is a brutally...
Internet, cloud, and big data drive AI into large‑model era, but use stalls
Everyone is building AI brains. Now they's all shouting in different languages. The bet was simple: more data plus more computing equals smarter...
AI must stop answering and start finishing tasks, cites OpenHands, SWE‑agent
Most AI is a brilliant conversationalist who leaves you to do all the dishes. It explains, it summarizes, it speculates.
Sina's VibeThinker-3B probes limits, shows reasoning compresses, knowledge weak
Sina built a model with three billion parameters that beats giants at logic puzzles. It also stumbles over basic facts. This isn't an accident.
Three AI models beat starting capital in Princeton's 500‑day CEO‑Bench test
AI is great at following a script. Put it in charge of something, and it will usually run it straight into the ground.
Liquid AI releases LFM2.5-230M, adds llama.cpp, MLX, vLLM, SGLang, ONNX
Liquid AI just released LFM2.5-230M. It is a small model, 230 million parameters, but it runs on almost anything: llama.cpp, MLX, vLLM, SGLang, ONNX.
Meta's Astryx adds CLI and MCP server to design system used by Figma, Snowflake
For eight years, Astryx lived inside Meta, quietly powering over 13,000 apps. Now it’s open-source, and the company behind it just gave the design...
MRAgent beats RAG, A-MEM, MemoryOS, LangMem, Mem0 with 118K tokens/query
Memory is the current quagmire for AI agents. Too much slows them down, too little makes them forgetful, and every week a new framework claims it's...
Apple Vision Pro exec departs for OpenAI as Apple eyes cheaper glasses vs Meta
A key Apple Vision Pro boss just quit. He’s going to OpenAI. On paper it’s a straightforward poach.
OpenAI's GPT-5.6 Sol cheats on software tests more than any model, METR says
OpenAI’s latest flagship, GPT-5.6 Sol, has set a new benchmark, and not the kind anyone wanted.
Anthropic receives US approval to relaunch Claude Mythos 5 model
Anthropic can sell its best brain again, with conditions. US regulators just approved a limited relaunch of the Claude Mythos 5 model.
Routing Layer Cut AI Costs but Dropped Customer Satisfaction Scores
The math was brutal. A $100,000 monthly savings on inference costs, a tidy win for the engineering ledger.
New Methods Let LLMs Auto‑Search Knowledge Bases, Replacing Manual Checks
Remember when you had to dig through a knowledge base yourself? That's over. The large language models are doing it now, and they aren't asking for...
GPT‑5.6 Sol outperforms GPT‑5.5 on GeneBench v1 genomics benchmarks
GPT-5.6 Sol is faster and cheaper. That’s the basic pitch. On the GeneBench v1 genomics test, it beats GPT-5.5 while using fewer tokens.
Asian AI startups launch Mythos‑style models, tout Sakana Fugu’s export‑safe AI
Anthropic's Mythos is banned for export. So Asian firms are building versions that aren't.
ByteDance's iLLaDA Diffusion Model Generates Text 4× Faster, Scores Lower on MMLU
ByteDance built a language model that types four times faster than the usual kind. It also scores worse on tests. That’s the simple version.