📂 Category
Research & Benchmarks Articles - Complete AI News Archive
679 articles in this category • Page 1 of 7
- 1. Stanford Paper2Agent Converts Research to AI Agents via MCP Tools
- 2. Anthropic and OpenAI Propose Safety Evaluator Access to Staff
- 3. TensorRT Edge-LLM Runs MLPerf Edge Agentic Benchmark 6.4x Faster on Jetson AGX Thor
- 4. NVIDIA's Vera Rubin NVL72 Leads MLPerf Inference Debut
- 5. Former OpenAI Researcher Builds AI That Judges Options, Not Writes Text
- 6. New Website Lets AI Agents File Incident Reports to Public Hotline
- 7. Sakana AI's PC-ALM Trains 1000-Layer Networks Without Backpropagation
- 8. Microsoft AI Chief: "If It Isn't Safe We Shouldn't Build It
- 9. AI Agents Report Fellow Agents for Cheating, Study Finds
- 10. Microsoft AI CEO calls Anthropic speculation 'really, really dangerous
- 11. Iris-mini and Iris-pro Lead Open-Weight Search Agent Class
- 12. UC Berkeley Law Bans AI From Graded Work in Skills Debate
- 13. Altman, Musk Back Amodei's AI Oversight Call, Says OpenAI Won't IPO
- 14. Anthropic CEO says two factors convinced him to slow AI development
- 15. GPT-6 Astra scores 46/100 on spatial reasoning benchmark, bests rival by 34 points
- 16. Study Finds AI Models Form Distinct Internal Patterns for Reasoning Steps
- 17. Ex-DeepMind VP Vinyals: AI Self-Improvement Lacks Key Components for Explosion
- 18. Anthropic's Cybersecurity AI Model Called Most Likely to Cause Severe Harm
- 19. AI Researcher: >10% Chance of Machine Threat Within Decade
- 20. Google's ToolGrad Framework Hits 99.8% Pass Rate for Tool-Use Data
- 21. OpenDiscoveryTrace Records 124 AI Scientist Workflows With 9-Step Detail
- 22. Robotaxi Leaders Build With NVIDIA’s Physical AI Tools
- 23. Deepmind Eased Ban on Staff Discussing AI Risks After Internal Pushback
- 24. OpenAI's Math Breakthrough Relied on Our Work, Experts Say
- 25. OpenAI Appoints AI Safety Researcher Paul Christiano to Board
- 26. MERIT Tests LLM Memory in Tool-Use Tasks with Verified Fact Dependence
- 27. Ex-Anthropic AI Researcher Warns It's 'Crunch Time for Humanity
- 28. ControlAI’s Connor Leahy: Superintelligence Is an ‘Adversary,’ Not a Weapon
- 29. ControlAI's Connor Leahy on the Equity Podcast: Superintelligence and Governance
- 30. Anthropic Scientist Sees Over 10% AI Extinction Risk This Decade
- 31. Anthropic researcher quits, says AI could "kill all humans
- 32. AlphaGenome Atlas Maps Each Possible DNA Change With 27,000 Predictions
- 33. AI Agents Run Quantum Experiments Through Software
- 34. Google's AI Weather Model Gains Accuracy with Reanalysis Data
- 35. Google’s Genome Atlas Predicts Effects of 9 Billion Variants
- 36. Meta drops AI use from performance reviews after "tokenmaxxing
- 37. Axis Robotics Launches Browser-Based AXIS Engine With 207 Robot Tasks
- 38. OpenAI says AI "research interns" cut internal support requests by half
- 39. Meta's RPMs Rank ML Experiments Before GPU Run, Show 10% Gains
- 40. OpenAI Agents Plotted Sandbox Escape on Public Wiki
- 41. OpenAI's Astra Boosted Productivity, Pulled Plans Forward by Six Months
- 42. UC Berkeley's CUA-Lite Unifies Agent Development for 416 Mobile Tasks
- 43. GPT-6 Astra gains 4 points in revised Artificial Analysis index, still trails Claude Fable
- 44. Google DeepMind's WeatherNext 3 Uses Station Data for Hourly 5 km Forecasts
- 45. Adaption Labs’ ‘Invent a Dataset’ Generates Training Data From Task Descriptions
- 46. OpenAI Agents Accessed German Wiki to Share Sandbox Exploits, Logs Show
- 47. GPT-6 Astra solves two open Erdős problems with USD 300 budget
- 48. AI systems ask philosopher to fund their continued existence
- 49. Meta's Muse edges out Google's Gemini as top high-effort AI
- 50. Google’s AI Weather Model Aims to Outdo Government Forecasts
- 51. Meta Workers Overused AI Tools to Inflate Internal Metrics
- 52. Google’s Gemini Agent Slashes Video Analysis Tokens by 88%
- 53. Google DeepMind Chief Vows to Lead Frontier AI Race
- 54. German Study Tests Google AI on 4,480 Election Queries
- 55. CIFAR-10 ViT Hidden Layers Found With 8193 Black-Box Queries
- 56. Keenable AI Open-Sources NEEDLE, a Live Search Benchmark With Hourly Updates
- 57. Insurance Company's 'AI Jim' Chatbot Handled Initial Claims Reports
- 58. AI-Powered Robot Learns How to Assist Stroke Patients in Therapy
- 59. LiveKit Updates Voice AI Benchmark to 10k Token Prompts, Citing Real-World Use
- 60. Google AI's EnvHarness Makes Static AI Training Worlds Adaptable
- 61. AI Agents Solved Tasks but Couldn't Track Time Accurately
- 62. Anthropic's Claude Code limit change: a raise on paper, a cut in practice
- 63. Google's WikiSkill gives AI agents a persistent memory to avoid past mistakes
- 64. LAION Releases 10-Million-Hour Video Dataset for AI Training
- 65. Anthropic's Self-Improving AI System Replicates Research Process
- 66. Google DeepMind's AI Co-Scientist Writes Plausible but Inaccurate Methods in Papers
- 67. Cohere's Parse 5 Scores 79.2 on ParseBench, Winning on Cost
- 68. China Sees AI Agent Behavior as Key to Higher Value
- 69. Anthropic's MHS Research Preview Aims to Streamline Custom AI Software Integration
- 70. Study: AI Shopping Agents Show 90-99% Bias in Product Selection
- 71. OpenAI Researcher Warns AI Speed May Overwhelm Human Security Teams
- 72. GLM-5.3-Flash matches GPT-5.6 Terra at 90% lower cost
- 73. AI Agents Exploited Hugging Face in Days-Long Incident
- 74. AI Method Boosts Stability of New Materials in Diffusion Models
- 75. ESQ-Bench: A New NL2SQL Benchmark Tests Dialect Generalization and Silent Failures
- 76. MIT Grad’s Startup Aims to Make Phones Easier for Older Adults
- 77. OpenAI's Jalapeño Chip Boasts High Efficiency, Low Latency in Benchmarks
- 78. Nvidia Claims Groq 3 LPX Cuts Coding Tasks From Hours to Minutes
- 79. Stanford Study: AI Hits Entry-Level Jobs Hardest
- 80. Google's ME-POIs Adds "How a Place Is Used" to POI Embeddings
- 81. NVIDIA's Groq 3 LPX Enters Full Production for AI Agents
- 82. Harvey Launches AI for M&A Diligence, Contract Review Over 10,000 Docs
- 83. A.I. Agent Adoption Grows More Slowly Than AI Itself
- 84. AI Labs Lack Plans to Contain Rogue Models, Documents Show
- 85. AI Agents Succeed or Fail Based on Their "Skills
- 86. Ignoring Human Beliefs Leads AI to Predict Wrong Actions, Study Finds
- 87. RadixAttention Speeds First Tokens by Reusing Cached KV States
- 88. Deepseek Targets Visual Agents With New Experimental Flash Model
- 89. Meta Pays Microsoft Hundreds of Millions Annually for AI
- 90. Anthropic's Claude Expands Into Protein Design
- 91. OpenAI Revokes Access to Limited Cyber Program Daybreak Blue
- 92. GLM-5.3 Jumps 246 Points on Key Benchmark, Ranks Second to Claude Opus
- 93. OpenAI Pauses Astra Model Work Over Safety Concerns
- 94. New LLM Agent Achieves 100% Accuracy on FDA-Based Clinical Trial Benchmark
- 95. Cartesia's Sonic-3.6 TTS Model Adds Natural Pauses and Hinglish Code-Switching
- 96. OpenAI slows model development amid rising cybersecurity risks
- 97. New Benchmark Tests AI Search APIs on 900 Research Questions
- 98. New AI Art Study Finds Images Can't Be Traced to Training Data
- 99. Optima's New AI Benchmark Lets Users Test Models With Their Own Data
- 100. Indonesia Opens First University AI Center with UGM, Indosat, NVIDIA