📂 Category
LLMs & Generative AI News Archive - Page 4 of 13
1,263 articles in this category • Page 4 of 13
- 301. Correlated errors cut panel accuracy 8‑22 points; top judge matches panel
- 302. OpenAI's GPT-5.5-Cyber Beats Anthropic Mythos, Starts Patching Initiative
- 303. Pull Gemma4:e4b with Ollama to Build a Local AI Coding Agent (v9.6)
- 304. Anthropic, Micron to Design AI Memory Architecture for Performance, Efficiency
- 305. Sakana's Fugu multi-model hits frontier performance, cites geopolitical edge
- 306. Guide to Using Claude Code for Browser Navigation and Its Simple Mechanics
- 307. Combining hidden neurons still yields a line, highlighting activation's role
- 308. Three NLTK tricks, including MWETokenizer, preserve domain terms in NLP
- 309. Sakana AI Fugu Ultra aims to match models; base Fugu low‑latency coding and chat
- 310. Samsung Deploys ChatGPT and Codex in Software, Marketing, Product, Manufacturing
- 311. Retrieval quality quickly becomes bottleneck for parametric memory’s long‑term weights
- 312. AI agents pick tools using function and parameter descriptions, study shows
- 313. Tip: Ask Clarifying Questions First to Refine ChatGPT Prompts
- 314. Altman says researchers underestimated scaling, calls LeCun's LLM view a dead end
- 315. Convert FP16 LLM to 4‑bit Q4_K_M on Windows AMD Radeon GPUs via llama.cpp
- 316. IEEE launches five‑course online program on large language models
- 317. CUDA Kernel Keeps Corpus on GPU, Cutting Retrieval Latency in RAG
- 318. Anthropic adds live dashboards, Cloudflare‑compatible code to Claude
- 319. Gemma-2-2B-Instruct with Llama-3.1-8B-Instruct cuts 99.9 tokens on 248‑prompt test
- 320. Modeling multi-agent deliberation as closed-loop system with hidden anchors
- 321. Lightweight model cuts RMSE in meteorology, carbon flux, soil moisture, grids
- 322. Understanding JSON Mode, Function Calling, and Structured Output in LLMs
- 323. Audit AI tools: inventory, VPN/zero-trust, continuous fingerprinting
- 324. Adobe adds AI Assistant to Photoshop, Premiere, Illustrator, InDesign, Frame.io
- 325. Claude Fable (Mythos) 5 shows limited bug‑finding and refactoring aid
- 326. Helion adopts LFBO with on‑the‑fly Random Forest for autotuning
- 327. NAVI‑Orbital performs first in‑orbit autonomous vision‑language inference
- 328. TurboQuant and OSCAR vie in KV cache compression race at ICLR 2026
- 329. Study probes if language models can hypothesize new math structures
- 330. NVIDIA XR AI Enables Real‑Time Multimodal Agents for AR Glasses
- 331. PrologMCP Launches as Task-Agnostic Open-Source Server for LLM Agents
- 332. Reconfigure OpenClaw on Mac Mini to Deploy a Local LLM Model
- 333. Roadmap to LLM Engineer in 2026: Foundations, Prompting, Fine‑Tuning, Alignment
- 334. Attention output GEMM reduces blended Fprop speedup to 1.47× in NVFP4 training
- 335. Estonian institute benchmarks AI models' vulnerability to Russian propaganda
- 336. Study quantifies AI agent trust formation, breakage, recovery in survival game
- 337. UP‑NRPA Allows Dynamic Customization of Dialogue Strategies Without Offline RL
- 338. Mobile NPU powers on‑device diffusion LLM with Multi‑Block Speculative Decoding
- 339. Orchestra‑o1 Enables Efficient Omnimodal Agent Collaboration
- 340. Vision LLMs Expand PDF Parsing to Charts, Diagrams, and Tables
- 341. Claude Fable 5 beats GPT‑5.5 by 13 points on FrontierMath tier‑4 tests
- 342. German Court Holds Google Liable for False AI-Generated Overviews
- 343. Google's DiffusionGemma: open diffusion model for faster text generation
- 344. Google sues Chinese Outsider Enterprise for Gemini-driven phishing on Telegram
- 345. PersonaDrive conditions VLA agents on human driving demos for simulation
- 346. ToolSense Framework Audits LLM Tool Knowledge Beyond Constrained Decoding
- 347. Gemini Omni adds AI video generation, using compute limits based on complexity and size
- 348. Xiaomi's MiMo Code beats Claude Code on 200+ step tasks, free MiMo Auto to V2.5
- 349. OpenAI hires Sottiaux in 2024, shifts from internal tools to ChatGPT overhaul
- 350. Low Kruskal-Rank Adaptation Shows Matrix Rank Stays r, Kruskal Rank Falls to 1
- 351. Anthropic apologizes for invisible guardrails on Claude Fable, first Mythos model
- 352. AI pre‑mediation matched professional mediators in multi‑issue negotiation test
- 353. AVLLMs Mirror VLM and VideoLLM Sequential Flow in Audio‑Visual Tasks
- 354. vLLM uses custom GPU kernels, TorchInductor and CUTLASS for portable inference
- 355. Claude Fable declines basic biology queries; Opus 4.8 responds
- 356. Run DiffusionGemma on NVIDIA GPUs for high‑throughput text generation
- 357. SynIB Introduces Information Bottleneck to Boost Multimodal Synergy
- 358. Understanding AgentOps: Discipline and the agentops.ai Platform Explained
- 359. Grab, CJ ENM, LiveKit praise Gemini 3.5 Live Translate for quality and accuracy
- 360. Apple's top AI concept mirrors vibe coding, using Shortcuts as a model
- 361. CoCoNuT paradigm expands residual stream for latent‑space, multi‑path reasoning
- 362. OmniMem adds modality-aware memory allocation for audio‑visual LLMs
- 363. PathoSage Introduces Three‑Stage Framework for Patch‑Level Pathology Reasoning
- 364. Apple unveils third‑gen foundation model, AFM 3 Cloud shows 36% boost
- 365. NVFP4 recipe speeds JAX/MaxText training on NVIDIA Blackwell and Rubin
- 366. Weaker LLMs Accidentally Delete Content, Shrinking Documents Over Time
- 367. Four New Specific Techniques to Boost Productivity with Claude Code
- 368. Jensen Huang sees token market segmenting into distinct value tiers
- 369. OpenAI to revamp ChatGPT, shift to business customers, rival Anth
- 370. MLP Networks Fit High-Frequency Functions One Oscillation at a Time
- 371. SafeGene Introduces Reusable Safety-Adapter for Cross-Task Model Families
- 372. FAIR-Calib Introduces Two-Stage PTQ Framework for Diffusion LLM Quantization
- 373. Elmes* Automates Fine-Grained Rubric Building for LLMs in Niche Education
- 374. Lean4Agent launches FormalAgentLib to model and verify workflow consistency
- 375. Study Finds No One-Size-Fits-All Strategy for Multi-Agent Communication
- 376. xAI used Anthropic’s Claude via personal accounts after access revoked for months
- 377. Study examines temporal preference concepts in large language models
- 378. Errorquake-10k Benchmark Scores 10,000 LLM Responses on 0-4 Severity Scale
- 379. Three SpaCy Tricks Speed Up Production-Grade Text Processing
- 380. Zhipu AI employs Muon Optimizer and Muon Split in GLM-4.5 and GLM-5 pretraining
- 381. Anthropic says Claude writes >90% of its code; AI pause button urged
- 382. Choosing AI Models: Prioritize Real‑World Needs Over Benchmark Rankings
- 383. ELI releases LLM benchmark showing top models resist Russian propaganda
- 384. AI trust certification trial in Fintech, Banking, Insurance, Health, US, Vietnam
- 385. SMAC-Talk Adds Natural Language to StarCraft Multi-Agent Challenge for LLMs
- 386. Spectral transfer identity s=αγ ties curvature exponent to Hessian decay
- 387. ChatHealthAI Aligns Structured EHR Data with Frozen LLM for Clinical Reasoning
- 388. Study Explores Graph Scaffolds as Reasoning Aid for Large Language Models
- 389. NVIDIA releases Cosmos 3 with Super‑Text2Image and Nano‑Policy‑DROID
- 390. Guide: Run a Claude Managed Agent Task End‑to‑End via Session Stream
- 391. Microsoft unveils Surface NVIDIA RTX Spark Dev Box for AI agent development
- 392. Nvidia builds RTX Spark supercomputer chips with Microsoft for AI agents
- 393. LLM-derived valence direction aligns with EEG signals in 123 subjects
- 394. gSMILE Framework Tackles LLM Transparency by Mapping Prompt Responses
- 395. DAStatFormer extracts 24 ANOVA-selected features per channel, slashing data size
- 396. Test-Time Prompt Optimization Turns Demonstrations into Rewards for VLM Models
- 397. BitsMoE uses SVD to keep basis unquantized, allocating bits to expert spectral factors
- 398. Claude Code Leads Feature Set as Codex Adopts Similar Tools for Coding
- 399. MiniMax-M3 launches, beats GPT-5.5 and Gemini 3.1 Pro on benchmarks, costs 5‑10%
- 400. Turing Award winner Richard Sutton: Pure generative AI cannot do real science