📂 Category
LLMs & Generative AI News Archive - Page 3 of 11
1,086 articles in this category • Page 3 of 11
- 201. Errorquake-10k Benchmark Scores 10,000 LLM Responses on 0-4 Severity Scale
- 202. Three SpaCy Tricks Speed Up Production-Grade Text Processing
- 203. Zhipu AI employs Muon Optimizer and Muon Split in GLM-4.5 and GLM-5 pretraining
- 204. Anthropic says Claude writes >90% of its code; AI pause button urged
- 205. Choosing AI Models: Prioritize Real‑World Needs Over Benchmark Rankings
- 206. ELI releases LLM benchmark showing top models resist Russian propaganda
- 207. AI trust certification trial in Fintech, Banking, Insurance, Health, US, Vietnam
- 208. SMAC-Talk Adds Natural Language to StarCraft Multi-Agent Challenge for LLMs
- 209. Spectral transfer identity s=αγ ties curvature exponent to Hessian decay
- 210. ChatHealthAI Aligns Structured EHR Data with Frozen LLM for Clinical Reasoning
- 211. Study Explores Graph Scaffolds as Reasoning Aid for Large Language Models
- 212. NVIDIA releases Cosmos 3 with Super‑Text2Image and Nano‑Policy‑DROID
- 213. Guide: Run a Claude Managed Agent Task End‑to‑End via Session Stream
- 214. Microsoft unveils Surface NVIDIA RTX Spark Dev Box for AI agent development
- 215. Nvidia builds RTX Spark supercomputer chips with Microsoft for AI agents
- 216. LLM-derived valence direction aligns with EEG signals in 123 subjects
- 217. gSMILE Framework Tackles LLM Transparency by Mapping Prompt Responses
- 218. DAStatFormer extracts 24 ANOVA-selected features per channel, slashing data size
- 219. Test-Time Prompt Optimization Turns Demonstrations into Rewards for VLM Models
- 220. BitsMoE uses SVD to keep basis unquantized, allocating bits to expert spectral factors
- 221. Claude Code Leads Feature Set as Codex Adopts Similar Tools for Coding
- 222. MiniMax-M3 launches, beats GPT-5.5 and Gemini 3.1 Pro on benchmarks, costs 5‑10%
- 223. Turing Award winner Richard Sutton: Pure generative AI cannot do real science
- 224. Google I/O 2026 Showcases Gemini‑Powered Infinite Scaler and Code Countdown
- 225. Gemini App Targets General Users—Students, Writers, Marketers, and More
- 226. Study fine-tunes honest and deceptive variants of five transformers with LoRA
- 227. 3-large embedding wins 2.1 test; MiniLM wins 2.3; rerankers lag in 2.2
- 228. Proxy-Pointer RAG Bakes Emerson Deltas into Index for AT&T system
- 229. Study finds base AI models predict human behavior better than fine‑tuned chatbots
- 230. Chronos-2 uses known covariates such as weather for building demand forecasts
- 231. OpenAI upgrades GPT-5.5 readability, removes Canvas from Instant and Thinking
- 232. Deep learning models auto‑detect data features, reducing need for engineer input
- 233. Google's Gemini Spark sees my whole life, then friend‑zones my boyfriend
- 234. Researchers Find Failure Signatures in LLM Trading Agents' Planning Embeddings
- 235. SSD removes sync bottleneck in speculative decoding on MI300X
- 236. Claude Opus 4.8 Trained for Honesty, Flags Uncertainty, Reduces Frustrations
- 237. Transformer Architecture Reduces Perplexity by 2.92 vs Fine‑Tuning
- 238. Step 3.7 Flash runs on NVIDIA GPUs via SGLang, TensorRT-LLM, vLLM
- 239. LLMs Struggle with Causal Discovery While Interventional Agents Succeed
- 240. DynaSchedBench Introduces SESC and SSI to Rank LLM Scheduling Tasks
- 241. LLM-based Architecture Targets Explicit and Implicit Human Values in Text
- 242. Google AI launches Daily Brief in Gemini app for U.S. users 18+
- 243. Google Cloud unveils AI platform with Gemini, Wiz, Codemender to patch gaps fast
- 244. Anthropic says new Claude model aims for honesty, avoids unsupported claims
- 245. Soro chatbot built on Gemma 3, trained on 1.9 B Tajik tokens from web and PDFs
- 246. How Ollama’s Context Length Setting Impacts Local Model Memory
- 247. NVIDIA releases NvRTX 5.7.4 with DLSS 4.5 support for UE5.7.4
- 248. How to Run Multiple Claude Code Sessions in Parallel Without Confusion
- 249. POLAR builds multimodal knowledge graph for semantic and episodic memory
- 250. MEMO trains a memory model on new knowledge with two roles, no LLM changes
- 251. GEM framework casts LLM data curation as hyperspherical variational problem
- 252. Experienced users supervise Claude only when it deviates, not step‑by‑step
- 253. Deploy Agents to Audit Complex Docs and Run Light Evaluations
- 254. Parameter-Efficient Multi-Class Scheduling for Multimodal Anomaly Detection
- 255. Study formalises LLM reasoning redundancy as truncatable steps in correct traces
- 256. Direct and Surrogate Verification Encode Transformer Circuits into SMT Solvers
- 257. AWS Agent Toolkit Shows Invocation, Success, UserError, SystemError Stats
- 258. AMD Ryzen AI Max+ runs 122B‑parameter models locally with 128 GB UMA
- 259. Semantic Search Model Assigns Class Labels and Confidence Scores to Critiques
- 260. Hotz warns AI coding agents could be costly despite 10x productivity boost
- 261. Accurate source citations boost AI answer quality, study finds
- 262. FuRA uses spectral preconditioning with full‑rank SVD for efficient fine‑tuning
- 263. Positional copying dominates answer readout in 1‑3B LMs on GSM8K
- 264. StepFun launches StepAudio 2.5 Realtime, evaluated via mobile app raters
- 265. Anthropic may keep supplying Claude to NSA despite Pentagon risk flag
- 266. Claude Code auto‑creates AI scaling algorithms; new control allocates compute
- 267. SuperClaude workflow ranks security issues, details attack vectors, gives fixes
- 268. Anthropic: Claude Mythos Preview finds ~3,900 high‑severity open‑source bugs
- 269. Meta launches Forum: Reddit‑style advice within Facebook groups, AI‑assisted
- 270. SOLAR introduced as self‑optimizing autonomous agent for continual learning
- 271. VSAS‑Bench Introduces Standardized Real‑Time Evaluation for Visual Assistants
- 272. F_Call_Analysis_Planner forwards Parent_Instruction to generate Selection_Rule
- 273. OSCToM uses RL to generate adversarial scenarios testing high-order Theory of Mind
- 274. Alibaba's Qwen3.7-Max runs 35 hrs, self‑monitors reward‑hacking, supports Claude Code
- 275. Gemini 3.5 Flash Shows Fast Responses in Free Account Tests
- 276. Pricing Change Alters Complaint Language, Skews Classifier Accuracy
- 277. Claude skill helps data scientists spot 5‑6 PM weekday usage spikes in 2026
- 278. Deepseek launches Deepseek Code to compete with Claude Code and OpenAI's Codex
- 279. Robotics may get a ChatGPT moment with massive human‑generated training data
- 280. Microservice Architecture Unites OCR, Classification, and LLM Pipelines
- 281. Proposal Calls for Data Probes to Study Impact of Training Data on LLMs
- 282. Isotonic calibration gets O(n⁻¹/³) sample complexity, cost‑optimal LLM routing
- 283. Basis Spline Decoupling Enables Compression of Transformer Models
- 284. LLM Retrieves Median 2020 Inflation Expectation, Drowning Prompt Guidance
- 285. QuickReduce FP4 delivers ~4.1× speedup over RCCL at TP=4 for large messages
- 286. Study Uses SHARP and New Error Framework to Assess PHRs in Health AI
- 287. Google's Gemini 3.5 Flash, pricier, adds 11 Omniscience points hallucinations 61%
- 288. Alibaba launches Qwen3.5‑LiveTranslate‑Flash: 60‑language translation in 2.8 s
- 289. I/O 2026 unveils Gemini Omni for universal creation, Gemini 3.5 Flash debut
- 290. EKS Hosts Multistage Multimodal Recommender; DLRM Personalizes Rankings
- 291. Gemini 3.5 Flash Enhances Web UI, Graphics and AI Studio Animations
- 292. Fresh Web Data Grounds LLMs, Highlighting RAG's Production Limits
- 293. ANNEAL lets neuro‑symbolic agents patch knowledge graphs without weight changes
- 294. Claude Cowork: Guide to Turning Q1 Sales Data into a Structured Word Report
- 295. Activation steering reveals latent bias in LLMs, reinjection restores decisions
- 296. 95% of task‑specific generative AI pilots never reach production
- 297. 5 Practical Uses of Local Language Models Highlight Code‑First Approach
- 298. SkillSmith extracts fine-grained boundaries so agents run only needed components
- 299. Quantized LLMs Show Emerging Bias, Masking Gradual Degradation
- 300. AgentStop cuts GPU power, heat and battery drain by ending AI agents early