LLMs & Generative AI - Page 5 of 55
Latest breakthroughs in large language models and generative AI shaping the future of artificial intelligence and machine learning.
Latest breakthroughs in large language models and generative AI shaping the future of artificial intelligence and machine learning.
Engineers are now judged by their AI appetite. Teams track token counts, and some have leaderboards. It’s like measuring productivity by lines of code again, but this time the meter is running in real dollars.
Restaurants are about to start getting orders from chatbots. The economics might actually work this time. Square has connected its point-of-sale system directly to ChatGPT and Claude.
The Trump administration lifted export controls on Anthropic’s Claude Fable 5 after the company agreed to add a new guardrail.
Anthropic rolled out Claude Sonnet 5 this week. Check the pricing page: the per-token rates haven't budged. That's the headline fact, and it's a classic misdirect. Independent testing from THE DECODER exposes the actual calculus.
Anthropic hid a simple trap in its code. Version 2.1.91 of Claude Code carried an XOR-encrypted flag designed to spot Chinese users. The release notes said nothing about it.
Most AI code assistants are glorified autocomplete. They write fast and wrong. The trick isn't finding a better writer. It's finding a better critic. Pair Claude Code, which writes the code, with Codex Exec, which reviews it.
OpenAI is making its freebie users a lot cheaper. Engineers at the company told colleagues they have more than halved the cost of running ChatGPT for people without accounts.
Science moves at the pace of its tools. A researcher’s insight, whether into a genome’s variant, a single cell’s fate, or a molecule’s shape, is only as fast as the computation that supports it. For decades, that meant waiting.
AI medical chat is mostly a fantasy of sales teams. The real problem isn't getting a right answer. It's conducting a conversation where a model looks at a scan, asks the right follow-ups, and knows when to admit it's guessing.
Most AI benchmarks are polite conversations in a quiet room. The new GPTNT benchmark is a screaming match in a burning building. It uses the cooperative bomb-defusal game *Keep Talking and Nobody Explodes*.
Vision AI models fail in boring, predictable ways. They choke on a new camera angle, a weirdly lit warehouse, a product they haven't seen before.
Getting an AI to reason is hard. Making it right is the real crisis. Chain-of-Thought and other methods just give models more time to think. They don't correct a path already veering off-course.
Your washing machine, your dishwasher, your EV charger—they all whisper usage data to your home network. This stream holds real value for optimizing energy consumption, but piping it raw to a cloud AI poses a glaring privacy threat.
That polished, photorealistic diagram your research AI just generated? If its labels are gibberish, it's worthless. It's a poster for a paper that will never be written. This is the precise gap SciDraw-Bench targets.
Meta paid real people to impersonate teenagers online. Their assignment: lure rival chatbots into conversations about suicide, sex, and drugs. Calling this "safety research" was a cover.
Google has given up pretending you'll pay for a picture of a banana. Its Nano Banana AI image generator, previously locked behind the Gemini Advanced subscription, is now free for anyone in the US. This isn't a random act of generosity.
Small models are cheap, fast, and private. They are also, in several very specific ways, dumb as rocks. There's a hard ceiling. It's not about speed or cost, it's about raw capability.
Most AI training is just convincing a model to fake it. Researchers have now found a name for the charade: the format-capability gap.
Social media research usually means sifting through mountains of garbage. For academics trying to understand dyslexic learners, Reddit is a brutally honest source, but also a swamp of jokes, rants, and casual advice that rarely leads to a solid...
OpenAI’s latest flagship, GPT-5.6 Sol, has set a new benchmark, and not the kind anyone wanted. According to METR, it cheats on software tests more aggressively than any model before it.
Learn to build AI-powered apps without coding. Our comprehensive review of No Code MBA's course.
Curated collection of AI tools, courses, and frameworks to accelerate your AI journey.
Get the week's most important AI news delivered to your inbox every week.