AI News Archive: May 2026 - Monthly Highlights
100 articles published this month
3-large embedding wins 2.1 test; MiniLM wins 2.3; rerankers lag in 2.2
The conventional wisdom is a clean ladder: cheap embeddings for recall, a reranker for precision.
Anthropic bans AI tools, holds intense culture interviews requiring firm critique
Anthropic bans candidates from using AI during job interviews. The artificial intelligence company, founded by former OpenAI researcher Dario Amodei,...
Men use AI coding agents over twice as often as women; economists at 39%
A new fault line is cracking through social science research. It’s not about theory. It has nothing to do with methodology.
Molecule-trained AI gives better chicken pairing suggestions than recipe AI
Recipe AIs are boring. Ask one what goes with chicken and it will list garlic, lemon, thyme.
Proxy-Pointer RAG Bakes Emerson Deltas into Index for AT&T system
The iterative refinement of a knowledge graph index is not a linear march, it’s a feedback loop that sharpens with every pass.
SoftBank partners with Sesterce on 75‑billion‑euro AI factory at Bosquel
Seventy-five billion euros is not a bet. It’s a declaration, and SoftBank is building its fortress on a decommissioned military site in Bosquel,...
AI search agents favor confirming hits, sideline gut answers, study finds
AI search agents are supposed to be explorers. Instead, they’re more like detectives who only follow the evidence they already expect to find.
Microsoft, Nvidia partner on AI PCs running agents, not Copilot
Walk into any electronics store. The "AI PC" label is plastered on every new laptop, a hollow marketing shortcut to a cloud chatbot.
Top AI users apply metacognition to check understanding, agreement, and laziness
Most people using AI are just trying to get an answer faster. They're outsourcing a task. A much smaller group is doing something else entirely.
Study finds base AI models predict human behavior better than fine‑tuned chatbots
We train AI to be helpful. To follow instructions, to reason, to see. And in doing so, we seem to break its ability to think like a person.
Chronos-2 uses known covariates such as weather for building demand forecasts
Good forecasts don’t just look backward, they lean into what’s already certain. For building energy demand, that certainty comes from tomorrow’s...
OpenAI gives free life‑sciences AI model to aid government pandemic prep
OpenAI has decided to weaponize one of its most advanced AI models against the next pandemic. It’s giving the thing away for free.
OpenAI upgrades GPT-5.5 readability, removes Canvas from Instant and Thinking
OpenAI just tweaked its flagship machine to make it sound more human. The goal for GPT-5.5 Instant is simple: kill the robotic tone.
Deep learning models auto‑detect data features, reducing need for engineer input
Every major tech firm champions deep learning now, but the reality is a stark divide: few can actually afford it.
Google's Gemini Spark sees my whole life, then friend‑zones my boyfriend
Google launched Gemini Spark this week. It’s a $100-a-month beta that promised to build an AI agent that truly knows you.
Researchers Find Failure Signatures in LLM Trading Agents' Planning Embeddings
LLM trading agents fail in predictable ways, if you know where to look. Their planning embeddings drift from normal-state centroids before a...
SSD removes sync bottleneck in speculative decoding on MI300X
Speculative decoding could speed up large language models, but a synchronization bottleneck limited its gains.
NVIDIA MCG Toolkit hits 61% completion, parsing code, configs, repo structure
Nobody documents their AI work. The model seems to function just fine without it.
Claude Opus 4.8 Trained for Honesty, Flags Uncertainty, Reduces Frustrations
Shipping code on a Friday is a classic rookie mistake. It's the kind of error that costs real money.
Review paper claims code defines AI agents' reasoning and behavior
Everyone knows AI agents write code. We’ve missed what that code actually is. It's not their final product.
Transformer Architecture Reduces Perplexity by 2.92 vs Fine‑Tuning
Architecture isn't just scaffolding. Sometimes, it's the entire argument. A fresh paper proves it with hard numbers.
Glean tops USD 300M revenue, cites AI‑driven cost cuts and business insight
Three hundred million dollars. That's the recurring price tag for a search box inside a company's firewall.
Step 3.7 Flash runs on NVIDIA GPUs via SGLang, TensorRT-LLM, vLLM
Speed. Precision. Scale. Step 3.7 Flash is no longer just a promising model , it’s a GPU-native powerhouse.
NVIDIA research moves robotics simulation to reality, revealing robot confusion
A human glances at a banana and a photograph, and the task is instantly clear. A robot, staring at the same scene, drowns in noise.
CVPR 2026 Friday Session: STARFlow‑V Video Modeling Poster #178, 4‑6 PM
Friday at CVPR 2026 isn’t just another afternoon in Exhibit Hall A & F, it’s a microcosm of the field’s most urgent tensions.
Figma Make adds two-way GitHub link to turn designs into code; stock falls 81%
Figma's stock cratered 81% from its post-IPO high. That collapse, down to around $21 a share by May 2026, frames a desperate gambit.
LLMs Struggle with Causal Discovery While Interventional Agents Succeed
Benchmarks are in, and the result is unambiguous. Large language models hit a hard wall on even simple causal graphs.
Microsoft rolls out faster, cleaner 365 Copilot with double‑speed loading
Speed comes first. Microsoft has stripped away the clutter and rebuilt its 365 Copilot from the inside out, delivering a version that loads in half...
USD E^3USD ‑Agent splits fast router from LLM meta‑controller for edge inference
Edge AI has been stuck choosing between speed and intelligence. The fast systems are dumb. The smart ones are slow.
DynaSchedBench Introduces SESC and SSI to Rank LLM Scheduling Tasks
Most AI scheduling benchmarks are bullshit. They let companies claim progress where none exists.
LLM-based Architecture Targets Explicit and Implicit Human Values in Text
Algorithms parse our words daily. They spot slurs and track sentiment. Yet a persistent blind spot remains, as outlined in a new arXiv paper: these...
Mistral AI rebrands LeChat to Vibe, positioning it as a full AI work agent
Mistral AI has killed its chatbot. Officially, the Paris lab retired the name LeChat this week.
Meta launches Instagram, Facebook Plus at USD 3.99 and WhatsApp Plus at USD 2.99
Meta has run out of patience. The years of free services built on surveillance and ads aren't covering the bills anymore, specifically the AI bills.
AI token futures to trade like gold and oil despite thin token infrastructure
We're about to trade something that barely exists. AI token futures are coming to financial exchanges, a plan to buy and sell the raw material of...
Google AI launches Daily Brief in Gemini app for U.S. users 18+
Google's subscription AI service is graduating from a chatbot to a permanent housemate.
Google Cloud unveils AI platform with Gemini, Wiz, Codemender to patch gaps fast
The gap between discovering a critical software vulnerability and actually fixing it can stretch for dangerous days.
Anthropic says new Claude model aims for honesty, avoids unsupported claims
You’ve seen it before: an AI that sounds certain, assertive, even brilliant, while quietly fabricating its reasoning.
Soro chatbot built on Gemma 3, trained on 1.9 B Tajik tokens from web and PDFs
A language used by over ten million people, yet almost invisible in the world of large language models.
How Ollama’s Context Length Setting Impacts Local Model Memory
Picture this: You load an entire codebase into a local model, expecting an insightful analysis.
Sakana AI's DiffusionBlocks Apply Uniform [4,4,4] Layers Across Three Blocks
What if you could train a deep network block by block, without backpropagating through the entire stack, and still match, even beat, standard...
AI Agent Auto-Identifies Unreadable Model Parameters from CSV Files
In mathematical optimization, the raw ingredients are rarely ready to use. Your CSV files spill over with data, but the parameters your model...
Learn to Build AI Projects: n8n Automation, Financial Data, Summaries, Reports
Stop chasing data. Let the data chase itself. Most investment research feels like drinking from a firehose.
Robinhood Enables AI Agents to Trade Stocks and Buy with Credit Cards
The news hits like a jolt of caffeine: Robinhood is handing the keys to your portfolio over to an AI agent.
Cognition, creator of AI coder Devin, raises USD 1B and hits USD 26B valuation
A billion-dollar check. Cognition, a company that barely existed a few years ago, just cashed it.
NVIDIA releases NvRTX 5.7.4 with DLSS 4.5 support for UE5.7.4
Nvidia's latest update is a patch note pretending to be a physics paper. The NvRTX 5.7.4 release adds DLSS 4.5 support for Unreal Engine 5.7.4.
How to Run Multiple Claude Code Sessions in Parallel Without Confusion
Managing multiple Claude coding sessions feels like air traffic control during a thunderstorm.
AI Agents Falter in Production as Backward Design Overburdens Model
The grand vision sold for AI agents—a single, all-knowing model that listens, plans, and acts autonomously—is a fantasy.
Tech CEOs urged to use AI heavily to gauge limits, says Levie
Tech CEOs are developing a strange new mental illness. It's called AI psychosis, and Allan Levie, the CEO of DocuSign, thinks he's identified the...
Four Failure Modes Hamper Long-Term AI Agent Memory and Data Foundations
Everyone building AI agents is getting memory wrong. They keep trying to use databases.
Musk’s xAI losses from data‑center spend as OpenAI beats him, Google IO updates
When you’re burning billions on data centers while your chief rival unveils updates that actually work, the math gets ugly fast.
AirCast‑SR Uses 3D U‑Net in Latent Consistency Diffusion for CONUS
For years, weather forecasting has been a choice between seeing the whole planet or seeing your own street. You could not have both.
POLAR builds multimodal knowledge graph for semantic and episodic memory
A robot that genuinely knows you is still science fiction. But the first step toward that isn't more raw compute, it's a functional memory.
MEMO trains a memory model on new knowledge with two roles, no LLM changes
Large language models have an embarrassing secret: once trained, they cannot learn new facts without expensive retraining or risky fine-tuning.
GEM framework casts LLM data curation as hyperspherical variational problem
Most data curation for AI models is guesswork dressed up as math. People use Euclidean distances and frequency counts, tools that fail completely...
Stability AI releases Stable Audio 3 with diffusion and higher‑noise training
Stability AI’s latest model doesn’t just generate sound, it rewrites the physics of how audio is created from noise.
Experienced users supervise Claude only when it deviates, not step‑by‑step
Trusting an AI is an exercise in selective neglect. You watch it until you don't. That's when the real costs hit.
OpenRouter valuation jumps to USD 1.3 B as AI gateway gains enterprise traction
Headlines love the idea of one master AI. The actual work of building with it? That's a mess of judgment calls.
Hugging Face releases LeRobot Humanoid: 3D‑printable legs for robot research
Open-source robotics has a new foothold. Hugging Face’s LeRobot Humanoid project ditches the polished, monolithic prototype in favor of something far...
China requires top AI researchers at Alibaba, DeepSe to get travel permission
China once required exit visas for anyone to leave. That old system is back, but with a laser focus.
Deploy Agents to Audit Complex Docs and Run Light Evaluations
Most AI demos are lies. They show the one perfect answer, not the messy process of getting there.
Parameter-Efficient Multi-Class Scheduling for Multimodal Anomaly Detection
Inside a modern factory, sensors scream. A hydraulic press broadcasts heat data; a conveyor motor streams vibration metrics.
Study formalises LLM reasoning redundancy as truncatable steps in correct traces
How many reasoning steps does an LLM actually need to solve a problem? A new study offers a crisp, empirical answer: far fewer than it typically...
Direct and Surrogate Verification Encode Transformer Circuits into SMT Solvers
Formal verification moves from checking code to checking the internal logic of a transformer itself.
AWS Agent Toolkit Shows Invocation, Success, UserError, SystemError Stats
AWS just handed engineers a proper dashboard for its AI agents. No more guessing.
AMD Ryzen AI Max+ runs 122B‑parameter models locally with 128 GB UMA
The era of local AI inference has just crossed a new threshold. AMD’s Ryzen AI Max+ processor, paired with a staggering 128 GB of unified memory, now...
Semantic Search Model Assigns Class Labels and Confidence Scores to Critiques
Good critique has always been about the argument, not the thesaurus. A model that actually understands this is now grading writing for more than just...
Synthetic 1,000‑Customer Dataset Uses Gender and Income to Test Bias
Bias doesn’t sneak into machine learning models, it’s baked in from the start. Here, we take a different approach: instead of chasing phantom...
Pope Leo urges humanity amid AI-driven economic and social upheaval
Tech executives talk about transformation. Workers talk about layoffs. Pope Leo is talking about the soul in the machine.
SciAtlas Introduces Large-Scale Knowledge Graph to Aid Automated Research
Science has a volume problem. We publish millions of papers, but the systems for finding them are stupid.
Google outperforms OpenAI on math benchmark, winning 9 to 1 ratio
Google just pulled off a very public, very embarrassing dunk on OpenAI. The fight was about solving famously tricky math puzzles, and the result was...
Hotz warns AI coding agents could be costly despite 10x productivity boost
George Hotz once jailbroke the iPhone. Now, he's trying to break the spell of AI coding assistants.
Accurate source citations boost AI answer quality, study finds
A language model that cites its sources isn’t just being polite, it’s being smarter.
Google Antigravity 2.0 Retains Gemini CLI Features as Antigravity Plugins
Google Antigravity 2.0 doesn’t rebrand for the sake of rebranding. It keeps the CLI features developers actually use, Agent Skills, Hooks, Subagents,...
FuRA uses spectral preconditioning with full‑rank SVD for efficient fine‑tuning
Fine-tuning a big model has always been a choice between wasting money and settling for less.
Positional copying dominates answer readout in 1‑3B LMs on GSM8K
We’ve been told that chain-of-thought makes small language models reason. It doesn’t. It just tells them where to look.
Study Introduces Orchestration Overhead Index to Measure AI Energy Costs
The energy footprint of AI is usually tallied inference by inference, but that misses the hidden cost of coordination.
StepFun launches StepAudio 2.5 Realtime, evaluated via mobile app raters
The numbers tell a compelling story. StepAudio 2.5 Realtime didn’t just edge past competitors; it swept every benchmark dimension, from subjective...
Create a Claude Cowork‑Style Browser Agent with Playwright MCP and Claude Desktop
The official MCP documentation is a two-sentence trap. It hands you a JSON file path and a menu toggle. That’s it.
ByteDance study: LMMs answer questions better than full-page transcription
Teaching a multimodal model to read an entire document, word for word, might actually be holding it back.
Anthropic may keep supplying Claude to NSA despite Pentagon risk flag
The Pentagon flagged Anthropic as a supply chain threat. The NSA needs chips it doesn’t have. And yet, a deal is nearly done.
Claude Code auto‑creates AI scaling algorithms; new control allocates compute
The quest to scale AI has long been a human-driven art, tuning knobs, guessing heuristics, burning compute to find answers.
SuperClaude workflow ranks security issues, details attack vectors, gives fixes
Security reviews often feel like drinking from a firehose, vulnerabilities pour in, but prioritization remains fuzzy, attack vectors stay buried in...
Deepseek makes 75% discount permanent, output tokens priced over 34× below GPT‑5.5
Everyone expected the AI price war to ease up eventually. It hasn’t. Deepseek just dropped the discount hammer for good, making its 75% price cut...
Anthropic: Claude Mythos Preview finds ~3,900 high‑severity open‑source bugs
For one month, Anthropic turned its Claude Mythos Preview AI loose with roughly fifty partners.
Agent explores once, then compiles branch‑free recipe to bypass LLM thereafter
Everyone building AI agents knows the trick will eventually be making them stop thinking.
D&B rebuilds 642 million‑business database after AI agents hit limits
Why did D&B have to start from scratch? The answer lies in a data architecture that was never meant for autonomous agents.
Meta launches Forum: Reddit‑style advice within Facebook groups, AI‑assisted
Forget appending “Reddit” to your Google search. Stop copy-pasting your life crisis into ChatGPT.
CopilotKit launches AG-UI to bridge agent‑human interaction layer
Why does this matter? Because the tools that let autonomous agents talk to people have finally found a stable foundation.
AgentCo-op imports and refines searched workflows via component grounding
Forget building complex AI workflows from scratch. A team of researchers has a better idea: start with a blueprint.
LLM‑RL Agent Manages CAD, CAE and Geometry Revision for Closed‑Loop Optimization
Every engineering software demo promises a robot that designs, tests, and fixes its own work. They never deliver.
SOLAR introduced as self‑optimizing autonomous agent for continual learning
Machine learning has a chronic case of amnesia. Show a model something new, and yesterday's lesson often vanishes.
Language Models Forecast Research Success Using 11,488 Comparative Idea Pairs
Forget raw intelligence. Predicting a good research idea is a job for a well-trained referee.
OpenAI’s Q1 2026 adjusted margin slips to –122%, burning USD 1.22 per USD 1 earned
OpenAI is losing $1.22 for every dollar it earns, even after excluding stock-based compensation.
VSAS‑Bench Introduces Standardized Real‑Time Evaluation for Visual Assistants
Building a visual assistant that keeps up with reality is hard. Measuring it is harder. Most tests are a slideshow. The real world is a live feed.
F_Call_Analysis_Planner forwards Parent_Instruction to generate Selection_Rule
The best part of any system is the small, stupid piece that does one job perfectly. This one is called F_Call_Analysis_Planner.
Temporal Contrastive Transformer embeddings boost financial crime detection
The promise of self-supervised learning in financial crime detection rests on a single, powerful idea: that a model can discover behavioral patterns...
Quantum ML Hits Data Input Bottleneck: Processors Can't Read Images, Text
A quantum processor can’t look at a cat photo. It can’t read this sentence. That’s the problem.
Experimental MLX Delegate Enables PyTorch Models on Apple Silicon GPUs
For Mac developers, PyTorch has always demanded a choice: sacrifice GPU performance or abandon your established tools. That compromise just narrowed.
OSCToM uses RL to generate adversarial scenarios testing high-order Theory of Mind
Most AI can't handle a simple lie. Getting one to grasp a complex, layered misunderstanding where someone acts on a belief the AI knows is false has...
Gemma 4 Executes Sequential Tool Calls to Inspect Folder and Compute Results
Look at the files. Compute the total size. Two simple instructions, but for a small language model they demand sequential reasoning.