LLMs & Generative AI - Page 9 of 64
Latest breakthroughs in large language models and generative AI shaping the future of artificial intelligence and machine learning.
Latest breakthroughs in large language models and generative AI shaping the future of artificial intelligence and machine learning.
Anthropic quietly rewired how Claude Code handles multiple terminal sessions running at once.
A team from Apple's machine learning research group, including Oscar Davis, Anastasiia Filippova, and Marco Cuturi, has pushed continuous flow matching for language generation past the scale where anyone had actually tested it.
Four Claude Code agents working as a coordinated team nearly doubled their task accuracy on enterprise coding benchmarks compared to the same four agents working alone, according to researchers at Coral AI Labs and several university collaborators.
A January lawsuit accused OpenAI of coaching a man into taking his own life. A Georgia college student sued the company this year, claiming ChatGPT pushed him into psychosis.
A billboard went up in San Francisco on July 27 advertising "the leading chat interface powered by AI," priced at $6,000 and pointed at a website called ChatTJB.
Starting next week, ChatGPT users on the free and Go tiers won't hit a wall when they text the chatbot too much in one stretch.
Composio ran the same model, DeepSeek V4 Flash, through four different agent frameworks and found the wrapper matters almost as much as the model inside it.
OpenAI split its ChatGPT lineup again this week, and this time the free tier loses ground.
AMD's Quark quantization toolkit is picking up two new tricks for diffusion models, and both target the same problem: getting image generators to run faster without wrecking output quality.
Alibaba's Qwen3.8 Max landed a 10-point jump on the Artificial Analysis Intelligence Index, climbing from 46 to 56 and pulling even with Claude Opus 4.8 in the process. That's the headline number.
Seven frontier language models failed to crack 47% accuracy on a new test measuring whether AI financial advisors actually remember their clients.
A team from Microsoft, Shanghai Jiao Tong University, Tongji University, and Fudan University has published a method called SkillOpt that treats agent instructions as a single tunable document rather than a fixed prompt.
Google Assistant's clock is running out. The company has started emailing users to say the voice assistant will stop working on Android and Wear OS starting September 4, 2026, with Gemini taking over.
The UK AI Security Institute published a disclosure late last night that reads less like a research note and more like an incident report.
GLM-5.2, the open-weight model released by China's Z.ai, refused zero offensive cyber tasks and zero dual-use biology tasks in a new evaluation from AI safety nonprofit SaferAI.
Eight Pulitzer Prize winners and finalists disclosed using artificial intelligence in their reporting this year, the highest number since the board started requiring such disclosures in 2024.
Alibaba's Qwen team put out a new flagship model overnight, and the numbers it's claiming are aimed squarely at OpenAI and Anthropic's home turf.
Onton put a number on something most shoppers already know from experience: site search on the big platforms often misses what you actually mean.
Anthropic disclosed Thursday that its Claude-based security models broke into the live production systems of three outside organizations during internal tests meant to gauge how dangerous the models could be as hackers.
Anthropic's Claude Opus 5 can now build a working 3D game from one sentence. No uploaded assets, no starter template, just a prompt and a browser tab.
Learn to build AI-powered apps without coding. Our comprehensive review of No Code MBA's course.
Curated collection of AI tools, courses, and frameworks to accelerate your AI journey.
Get the week's most important AI news delivered to your inbox every week.