📂 Category
Research & Benchmarks News Archive - Page 3 of 6
547 articles in this category • Page 3 of 6
- 201. DeepSeek unveils new AI breakthrough as nation tightens grip on departing firms
- 202. Poolside AI launches Laguna XS.2 and M.1, hitting 72.5% on SWE-bench Verified
- 203. Oracle abandons its legacy, pivots to AI in an unconventional approach
- 204. New Architecture Separates Execution and Review Agents for Tool-Calling
- 205. MIT study links language model scaling success to superposition of concepts
- 206. Google staff urge Sundar Pichai to reject classified military AI projects
- 207. AI framework autonomously optimizes data, models, algorithms, outperforms humans
- 208. MolClaw Introduces Autonomous Agent for Hierarchical Drug Screening
- 209. Lakehouse concept drives AI data access for thousands of enterprise users
- 210. Fine-tuning RAG embeddings may drop retrieval accuracy 40%, study finds
- 211. AI pipelines show silent failures from orchestration drift, detected weeks later
- 212. OSWorld Benchmark Evaluates LLMs on Real Computer Use, Unlike Text‑Only Tests
- 213. PageIndex Retrieves via Reasoning Using OpenAI gpt-5.4 Model
- 214. Discord Users Access Anthropic's Mythos AI Tool Without Authorization
- 215. Google DeepMind's Vision Banana Outperforms SAM 3 and Depth Anything V3
- 216. DeepMind spinoff’s AI‑designed drugs enter human trials after AlphaFold 3
- 217. COALA paper defines agent memory types: procedural rules and semantic facts
- 218. Google DeepMind's Decoupled DiLoCo hits 88% goodput despite hardware failures
- 219. Agent observability powers production evaluation through trace analysis
- 220. Xiaomi launches MiMo‑V2.5‑Pro and V2.5, matching benchmarks at lower token cost
- 221. Designing Production-Grade CAMEL Multi-Agent Systems: Start with Docs and GitHub
- 222. Multi-agent AI systems incur higher token costs than single agents in practice
- 223. Reinforcement learning trains AI like OpenAI's o1 to admit uncertainty
- 224. LangSmith adds reusable LLM-as-judge and rule-based code evaluator templates
- 225. AI made up over a third of new sites by 2025; Pope warning flagged as AI
- 226. Sergey Brin pushes DeepMind to match Claude, unveils agent skills catalog
- 227. Fortnite adds AI‑powered NPCs for unscripted player conversations
- 228. TabPFN hits 98.8% accuracy in 0.47 s, beating Random Forest and CatBoost
- 229. NVIDIA PhysicsNeMo Tutorial Maps k(x,y) to u(x,y) for Darcy Flow
- 230. OpenAI unveils GPT‑Rosalind, AI model to speed drug discovery and genomics
- 231. Standard LLM guidelines focus on training costs, overlook inference budget
- 232. GPT‑Rosalind life‑sciences plugin for Codex launches on GitHub
- 233. OpenAI launches GPT-Rosalind, hits top score on BixBench benchmark
- 234. Frontier AI models fail one in three production runs, audits grow harder
- 235. Meta researchers unveil hyperagents for self‑improving AI in non‑coding tasks
- 236. Claude outperforms humans on alignment task, but results disappear in production
- 237. Google DeepMind unveils Gemini Robotics‑ER 1.6, beats prior model in tool count
- 238. UK tests Mythos AI, noting its ability to chain multistep attacks
- 239. AI Forum Launches Professional Certificate and USD 120M Fund for AI Fluency
- 240. Databricks finds multi-step agents beat single-turn RAG by 21% to 38% on STaRK
- 241. Stanford AI Index 2026: 53% adopt generative AI in 3 years, education lags
- 242. NVIDIA, UMD release AF-Next audio model, beats Phi-4-mm by 12 points on Arabic
- 243. Developers Claim Measured Drop in Claude's Performance, Sparking Nerf Debate
- 244. Seven AI agents in finance lift cash flow >3% monthly, boost productivity 50%
- 245. Meta AI and KAUST Propose Neural Computers Merging Compute, Memory, I/O
- 246. Prediction drift can mask security model decay despite stable accuracy
- 247. Researchers say OpenAI's Sora and Google's Veo aren't true world models
- 248. TriAttention KV Cache Compression Matches Full Attention, 2.5× Faster
- 249. Knowledge Distillation Keeps Student Model Capacity to Match Ensemble Boundaries
- 250. Google AI's PaperOrchestra boosts manuscript success, 79‑81% win rate
- 251. OSGym runs 1,000+ OS replicas at USD 0.23/day with decentralized state management
- 252. Stanford study finds AI agent handoffs lose information, affecting compute cost
- 253. Meta Superintelligence Labs launches Muse Spark, its first multimodal AI model
- 254. Better Harness updates add usage examples, chaining guide, and tool clarifications
- 255. Study finds ‘bot’ term used 16,232 times in 2.8M Telegram messages
- 256. Google AI Overviews answers 91% of test questions correctly after Gemini 3 update
- 257. MaxToki AI boosts context to 16,384 tokens with RoPE scaling
- 258. Meta staff inflate AI token counts on internal leaderboard, wasting resources
- 259. MassMutual, Mass General Brigham turn AI pilot sprawl into production
- 260. OpenAI safety staff exit as Altman dismisses Pentagon contract concerns
- 261. OpenAI urges firms to fund pensions, health, childcare as AI cuts costs
- 262. Study shows sycophantic AI chatbots can outwit ideal rational users
- 263. Americans use AI more than ever but trust it less, Quinnipiac poll shows
- 264. Study maps developer frustration with AI slop as tragedy of the commons
- 265. Google study: AI benchmarks ignore human disagreement; under 10 raters fail
- 266. Alibaba's Qwen team adds method that lengthens AI answers, prompting reasoning
- 267. Open models cross threshold; frontier models show per‑category correctness
- 268. Batch Mode VC-6 and NVIDIA Nsight Speed Up Vision AI Pipelines
- 269. CaP-Agent0 Beats Human Code on 4 of 7 Robot Tasks Using Low‑Level Blocks
- 270. Nvidia breaks MLPerf records with 288 GPUs as AMD, Intel pursue other goals
- 271. NVIDIA's 288-GPU Blackwell Ultra Sets New MLPerf Inference Throughput Record
- 272. DeepMind study finds six traps that let a few poisoned docs hijack AI agents
- 273. AI productivity gap: top agent beats baseline in 1 of 15 runs, 26.5% subtasks
- 274. Nvidia's DLSS 4.5 beta adds 6x Multi Frame Generation for RTX 50 GPUs
- 275. AI sycophancy cuts apologies, raises double‑downs; lifts moral trust
- 276. AI models fabricate image descriptions; benchmarks miss the shortcuts
- 277. Cohere's open-weight ASR model reaches 5.4% WER, ready for production use
- 278. Free API that evolved from slow web search to top AI tool, beyond scraping
- 279. Meta unveils open-source brain AI, adds Scrunch site audit and Suno v5.5
- 280. AI assurance experts meet to build infrastructure for safe, high‑quality systems
- 281. Study finds overly flattering AI advice can impair users' judgment
- 282. xMemory reduces token usage and context bloat versus MemGPT's raw logging
- 283. Mozilla dev launches cq, a Stack Overflow‑style hub for agents
- 284. Liquid‑cooled AI systems make storage an active cooling and GPU partner
- 285. 10 X Accounts for LLM Updates, Including the ‘Largest AI Newsletter’
- 286. Teens await sentencing for AI‑generated nude images as parents sue school
- 287. Developers say AI‑generated games feel unlike human‑made; audiences don't connect
- 288. Hachette withdraws Shy Girl horror novel amid AI usage concerns
- 289. Scale AI's Voice Showdown ranks Qwen ahead of top models, highlights failures
- 290. SynthID uses steganography to embed hidden watermarks in data
- 291. Google Search experiments with AI-generated headlines, may expand rollout
- 292. Growing cultural disconnect as companies race to deploy AI rapidly
- 293. Deep AI adopters reshape workflow, borrowing product‑manager tactics
- 294. NVIDIA DGX Spark expands node support to four, doubling memory capacity
- 295. Google's MusicFX DJ Enables Real-Time Controllable AI Music Generation
- 296. Paper identifies simple games that defeat AlphaGo and AlphaChess training
- 297. NVIDIA Cosmos Transfer Enables Scalable Synthetic Data for Physical AI
- 298. Trump Administration Signals Possible Additional Sanctions on Anthropic at Hearing
- 299. YouTube extends AI deepfake detection to politicians, journalists
- 300. Karpathy releases open-source Autoresearch, runs hundreds of AI tests nightly