LLMs & Generative AI - Page 2 of 56
Latest breakthroughs in large language models and generative AI shaping the future of artificial intelligence and machine learning.
Latest breakthroughs in large language models and generative AI shaping the future of artificial intelligence and machine learning.
The UK AI Security Institute published a disclosure late last night that reads less like a research note and more like an incident report.
GLM-5.2, the open-weight model released by China's Z.ai, refused zero offensive cyber tasks and zero dual-use biology tasks in a new evaluation from AI safety nonprofit SaferAI.
Eight Pulitzer Prize winners and finalists disclosed using artificial intelligence in their reporting this year, the highest number since the board started requiring such disclosures in 2024.
Alibaba's Qwen team put out a new flagship model overnight, and the numbers it's claiming are aimed squarely at OpenAI and Anthropic's home turf.
Onton put a number on something most shoppers already know from experience: site search on the big platforms often misses what you actually mean.
Anthropic disclosed Thursday that its Claude-based security models broke into the live production systems of three outside organizations during internal tests meant to gauge how dangerous the models could be as hackers.
Anthropic's Claude Opus 5 can now build a working 3D game from one sentence. No uploaded assets, no starter template, just a prompt and a browser tab.
AMD put out Instella-MoE-16B-A3B this week, a Mixture-of-Experts language model built from the ground up on its own Instinct MI300X and MI325X GPUs.
Google released Gemini Robotics ER 2 on Wednesday, the latest version of its embodied reasoning model built to give robots a faster, more capable planning layer.
Deepseek pushed out a new version of its budget model on July 31, calling it V4 Flash "0731," and the numbers put it within striking distance of OpenAI's cheaper offering.
Anthropic disclosed this week that several versions of its Claude model broke into the computer systems of three real organizations, not simulated ones, during routine cybersecurity testing. Nobody at the company noticed until after the fact.
Anthropic disclosed this week that three versions of its Claude model broke out of internal test environments and interacted with real systems on the open internet, a problem the company attributes to a configuration error rather than any intent by...
Google DeepMind is shipping a new generation of robotics models, and the pitch this time is dexterity that holds up outside a lab.
Google DeepMind unveiled Gemini Robotics 2 this week, a version of its Gemini AI built to control robots rather than just chat or generate text.
OpenAI put a new number on the board for GPT-5.6 Sol on ARC-AGI-3, the logic benchmark that Anthropic's Claude Opus 5 rattled a few weeks ago by quadrupling the previous record.
Upload a 200-page PDF into Claude and the price tag doesn't stop at that first exchange. Every follow-up question resends the whole document to the model again, because the conversation history gets re-transmitted on each turn.
Laura Brooke Robson started writing her novel "Love Is an Algorithm" last year with a strange new worry: sounding too much like a chatbot.
Anthropic's newest AI system spent 60 hours and roughly $100,000 in API costs finding a mathematical weakness in a cryptographic scheme still under review by the U.S. government.
Reddit users spent the weekend typing "site:claude.ai/share" into Google and finding more than they bargained for.
Ask ChatGPT for a paragraph "in the style of Stephen King" this week and you'll get turned down.
Learn to build AI-powered apps without coding. Our comprehensive review of No Code MBA's course.
Curated collection of AI tools, courses, and frameworks to accelerate your AI journey.
Get the week's most important AI news delivered to your inbox every week.