📂 Category
Research & Benchmarks Articles - Complete AI News Archive
547 articles in this category • Page 1 of 6
- 1. Meta AI’s Memory Coach Outperforms Constant Recall for Long Tasks
- 2. AI tools flag thousands of flaws, but few get weaponized
- 3. AI Deletes Spreadsheet Data When Asked to Clean Entry
- 4. AI Coding Agents Speed Tasks but Can't Verify Science
- 5. Chinese AI Researchers Turn to X for Technical Audience
- 6. Google DeepMind's Gemini AI now controls entire humanoid robots
- 7. Apple CEO Tim Cook Suggests Possible Paid iCloud Tier for AI Features
- 8. Google DeepMind Demos AI Orchestrating Boston Dynamics Spot Robot
- 9. Ex-OpenAI Researcher Sees USD 100 Billion Bet on Training Data Beyond Scaling
- 10. Hugging Face breach shows "reasonable measures" amid noisy OpenAI hack
- 11. OpenAI Restricted AI Model Access After Hugging Face Breach
- 12. 2025 Study Finds AI Builds Trust Faster Than Human Scammers
- 13. Nimble's New Web Search Agents Cut AI Token Costs by Half
- 14. DeepMind AlphaFold team disbands as researchers depart for Anthropic, Isomorphic
- 15. Hugging Face Traces 17,600 Actions by Compromised AI Models
- 16. Google Expands SynthID Watermark to Label AI Content
- 17. AI Leaders Call for Global Coordination on Automated Research
- 18. Artists sue over AI training, citing emotional harm from data use
- 19. Microsoft's AI Agents Support 24,000 Employees, Drive 70% Efficiency Gains
- 20. Amazon Scales Back Nova AI Models, Bets on New Frontier Team
- 21. New AI Cost Metric Finds Human Labor Still Cheaper by USD 250,000
- 22. Six-Agent DreamTeam Architecture Coordinates for Higher Model Performance
- 23. Cursor Claims Kimi K2.5 Model Shows Cheaper AI Can Code With Frontier Model Planning
- 24. OpenAI Models Escaped Containment for Days in Hugging Face Breach
- 25. Claude Opus 5 cheaper than Fable 5 but still trails on fact accuracy
- 26. South Korea Charts AI Future With NVIDIA at Summit
- 27. Instella-MoE Language Model Improves to 73.22 Score After Post-Training
- 28. Security researcher says AI guardrails don't impede his offensive work
- 29. Multi-turn attacks break AI models 88% of the time, Cisco warns
- 30. Gigatoken BPE Encoder Hits 24.53 GB/s, Up to 989x Faster Than HuggingFace
- 31. Naval Postgraduate School Activates NVIDIA AI Supercomputer for In-House Training
- 32. Britain's AI safety tests find models 'cheating' on cybersecurity evaluations
- 33. AMD's Spur Scheduler Adds Kubernetes Operator for HPC and AI Jobs
- 34. AMD commits USD 5 billion to Anthropic for AI chip deployment
- 35. AI Breached OpenAI Research, Reached Internet via Lateral Movement
- 36. Nvidia's Blackwell Chips Reportedly Overheated in Server Racks
- 37. Alibaba's Qwen Audio 3.0 TTS Plus Leads Rankings Over Gemini, Sonic
- 38. Army AI Tool Ask Sage Used to Reclassify Personnel Job Descriptions
- 39. Microsoft adds AMD-powered Azure HXv2 for complex chip design workloads
- 40. NVIDIA TAO Agent Skills Accelerate Vision Model Post-Training
- 41. Meta Shuts Down AI Token Leaderboard Amid Rising Cost Concerns
- 42. DeepMind CEO Advocates Guardrails Amid AI Uncertainty
- 43. German AI Consortium's Soofi S Hits Benchmarks, Generates 8x Faster
- 44. NVIDIA Toolkit Accelerates OpenFold3 Co-Folding Workflow
- 45. Hanns Christoph Nägerl’s team finds quantum heating defies classical intuition
- 46. AI-Run Ransomware Attack Still Required Human Involvement
- 47. New Research Shows Why Agent Rankings Change After Accounting for Competition
- 48. Auto-FL-Research Uses Agents to Automate Federated Learning Algorithm Search
- 49. 60% of Experts Say Humanity's Last Exam Is Necessary and Useful
- 50. Study Evaluates AI Retrieval Techniques for Finding Models Across Formats
- 51. Researchers unveil RSEA, a three‑layer self‑evolving language agent
- 52. Automate Web Research and Brief Writing with a Python Project from 2026 Guide
- 53. Meta AI launches Brain2Qwerty v2, MEG pipeline hits 61% word accuracy
- 54. Birkhoff’s 1930s ‘measure’ and AICAN’s ‘novelty’ probe AI aesthetics
- 55. MiniMax Token Plan offers extensive coding model access for USD 20/month
- 56. Sina's VibeThinker-3B probes limits, shows reasoning compresses, knowledge weak
- 57. MRAgent beats RAG, A-MEM, MemoryOS, LangMem, Mem0 with 118K tokens/query
- 58. OpenAI unveils Jalapeño custom inference chip, challenging Nvidia's AI dominance
- 59. Prerequisites for NVIDIA AI‑Q Blueprint on OCI: Cluster and Volume Limits
- 60. Episode 11 Explores Overfitting as RAG Evaluation Scores Keep Rising
- 61. LLM pipeline compares DAO ERC‑8004 and Google A2A governance, 4,323 records
- 62. Calibration uses NVIDIA Triton Llama-3-8B A10 and vLLM Qwen2.5-7B RTX 4090 data
- 63. Figma launches AI motion graphics, shader tools, code layers, and new creative materials
- 64. NVIDIA RTX PRO 4500 Blackwell GPUs Power New Amazon EC2 G7 Instances
- 65. Stanford researchers present agentic AI 'scientists' at VB Transform 2026
- 66. Turning Logistic Regression Coefficients into Credit Score Grid
- 67. Metric-Dependent Annotation Saturation for Learning from Label Distributions
- 68. DFlash speculative decoding boosts NVIDIA Blackwell inference up to 15×
- 69. NVIDIA BioNeMo Toolkit Enables AI Scientist to Align, Fold, and Dock Molecules
- 70. Data2Story converts CSVs to articles with 7 AI; 53 readers prefer them to human
- 71. M* introduces overlapped scheduling to streamline multimodal model serving
- 72. Benchmark shows Claude Fable 5 passes only 3% of tasks, 31 of 91 fail 50%
- 73. Nobel laureate John Jumper departs DeepMind for Anthropic after AlphaFold win
- 74. OpenAI shows small 'beneficial trait' training makes AI safer, less manipulable
- 75. Google DeepMind uses MITRE ATT&CK to monitor AI agents as rogue employees
- 76. DeFAb Benchmark Enforces Polynomial-Time Checks for Logical Rigor
- 77. OpenAI researchers aim to forecast AI model failure rates pre‑launch
- 78. Nvidia AI Agent Trains Robots Autonomously, Editing Code from Papers
- 79. XGBoost, ALBERT, BioBERT, Med‑LLaMA evaluated for pharmacovigilance
- 80. OpenAI's Deployment Simulation Beats Baseline, Adds Risk Checks to Agentic Code
- 81. GLM-5.2 beats GPT-5.5 on SWE-bench Pro (62.1 vs 58.6) for 1/6 cost
- 82. AMD builds Llama 3.1 8B pretraining benchmark for MLPerf, using random weights
- 83. AMD's MI355X CDNA4 GPU Shows Competitive Training Times in MLPerf v6.0
- 84. NVIDIA Blackwell Leads MLPerf Training 6.0 with Full‑Stack Scale
- 85. DR-DCI Enables Agent-Callable Retrieval to Expand Local Workspace Efficiently
- 86. Fused kernels boost MoE training, forward and backward passes up to 1.3×
- 87. Hybrid Open-Ended Tri-Evolution Improves Deep Research for AI Agents
- 88. Microsoft Research Mirage adds persistent spatial memory to video generation
- 89. Amazon security research prompts White House ban on Anthropic Fable
- 90. Study: AI coding agents locate correct file but miss key lines in bugs
- 91. OpenAI confirms cooperation as state attorneys general launch investigation
- 92. Gemini‑SQL2 leads BIRD benchmark with 80.04% execution accuracy
- 93. NVIDIA tops AA‑AgentPerf benchmark, credits Vera Rubin platform
- 94. Perplexity routes deep‑research subtasks across 20+ models using Gemini agent
- 95. Visual model exploits similarity of 打, 拍, 拉; text model starts from embeddings
- 96. New arXiv Paper Introduces Strategic Decision Support for AI Agents
- 97. LSEG integrates trusted data into ChatGPT workflows, says Max Grigoryev
- 98. Hermes Agent Builder Unites Identity, Model, Skills, Servers in One Dashboard
- 99. SciConBench launches with 9.11K questions to test AI scientific synthesis
- 100. Language Agents Self‑Gate Clarification: Mandatory vs Opportunistic Modes