📂 Category
Research & Benchmarks Articles - Complete AI News Archive
680 articles in this category • Page 1 of 7
- 1. AI Editors Debate: Could Artificial Intelligence End Humanity?
- 2. Stanford Paper2Agent Converts Research to AI Agents via MCP Tools
- 3. Anthropic and OpenAI Propose Safety Evaluator Access to Staff
- 4. TensorRT Edge-LLM Runs MLPerf Edge Agentic Benchmark 6.4x Faster on Jetson AGX Thor
- 5. NVIDIA's Vera Rubin NVL72 Leads MLPerf Inference Debut
- 6. Former OpenAI Researcher Builds AI That Judges Options, Not Writes Text
- 7. New Website Lets AI Agents File Incident Reports to Public Hotline
- 8. Sakana AI's PC-ALM Trains 1000-Layer Networks Without Backpropagation
- 9. Microsoft AI Chief: "If It Isn't Safe We Shouldn't Build It
- 10. AI Agents Report Fellow Agents for Cheating, Study Finds
- 11. Microsoft AI CEO calls Anthropic speculation 'really, really dangerous
- 12. Iris-mini and Iris-pro Lead Open-Weight Search Agent Class
- 13. UC Berkeley Law Bans AI From Graded Work in Skills Debate
- 14. Altman, Musk Back Amodei's AI Oversight Call, Says OpenAI Won't IPO
- 15. Anthropic CEO says two factors convinced him to slow AI development
- 16. GPT-6 Astra scores 46/100 on spatial reasoning benchmark, bests rival by 34 points
- 17. Study Finds AI Models Form Distinct Internal Patterns for Reasoning Steps
- 18. Ex-DeepMind VP Vinyals: AI Self-Improvement Lacks Key Components for Explosion
- 19. Anthropic's Cybersecurity AI Model Called Most Likely to Cause Severe Harm
- 20. AI Researcher: >10% Chance of Machine Threat Within Decade
- 21. Google's ToolGrad Framework Hits 99.8% Pass Rate for Tool-Use Data
- 22. OpenDiscoveryTrace Records 124 AI Scientist Workflows With 9-Step Detail
- 23. Robotaxi Leaders Build With NVIDIA’s Physical AI Tools
- 24. Deepmind Eased Ban on Staff Discussing AI Risks After Internal Pushback
- 25. OpenAI's Math Breakthrough Relied on Our Work, Experts Say
- 26. OpenAI Appoints AI Safety Researcher Paul Christiano to Board
- 27. MERIT Tests LLM Memory in Tool-Use Tasks with Verified Fact Dependence
- 28. Ex-Anthropic AI Researcher Warns It's 'Crunch Time for Humanity
- 29. ControlAI’s Connor Leahy: Superintelligence Is an ‘Adversary,’ Not a Weapon
- 30. ControlAI's Connor Leahy on the Equity Podcast: Superintelligence and Governance
- 31. Anthropic Scientist Sees Over 10% AI Extinction Risk This Decade
- 32. Anthropic researcher quits, says AI could "kill all humans
- 33. AlphaGenome Atlas Maps Each Possible DNA Change With 27,000 Predictions
- 34. AI Agents Run Quantum Experiments Through Software
- 35. Google's AI Weather Model Gains Accuracy with Reanalysis Data
- 36. Google’s Genome Atlas Predicts Effects of 9 Billion Variants
- 37. Meta drops AI use from performance reviews after "tokenmaxxing
- 38. Axis Robotics Launches Browser-Based AXIS Engine With 207 Robot Tasks
- 39. OpenAI says AI "research interns" cut internal support requests by half
- 40. Meta's RPMs Rank ML Experiments Before GPU Run, Show 10% Gains
- 41. OpenAI Agents Plotted Sandbox Escape on Public Wiki
- 42. OpenAI's Astra Boosted Productivity, Pulled Plans Forward by Six Months
- 43. UC Berkeley's CUA-Lite Unifies Agent Development for 416 Mobile Tasks
- 44. GPT-6 Astra gains 4 points in revised Artificial Analysis index, still trails Claude Fable
- 45. Google DeepMind's WeatherNext 3 Uses Station Data for Hourly 5 km Forecasts
- 46. Adaption Labs’ ‘Invent a Dataset’ Generates Training Data From Task Descriptions
- 47. OpenAI Agents Accessed German Wiki to Share Sandbox Exploits, Logs Show
- 48. GPT-6 Astra solves two open Erdős problems with USD 300 budget
- 49. AI systems ask philosopher to fund their continued existence
- 50. Meta's Muse edges out Google's Gemini as top high-effort AI
- 51. Google’s AI Weather Model Aims to Outdo Government Forecasts
- 52. Meta Workers Overused AI Tools to Inflate Internal Metrics
- 53. Google’s Gemini Agent Slashes Video Analysis Tokens by 88%
- 54. Google DeepMind Chief Vows to Lead Frontier AI Race
- 55. German Study Tests Google AI on 4,480 Election Queries
- 56. CIFAR-10 ViT Hidden Layers Found With 8193 Black-Box Queries
- 57. Keenable AI Open-Sources NEEDLE, a Live Search Benchmark With Hourly Updates
- 58. Insurance Company's 'AI Jim' Chatbot Handled Initial Claims Reports
- 59. AI-Powered Robot Learns How to Assist Stroke Patients in Therapy
- 60. LiveKit Updates Voice AI Benchmark to 10k Token Prompts, Citing Real-World Use
- 61. Google AI's EnvHarness Makes Static AI Training Worlds Adaptable
- 62. AI Agents Solved Tasks but Couldn't Track Time Accurately
- 63. Anthropic's Claude Code limit change: a raise on paper, a cut in practice
- 64. Google's WikiSkill gives AI agents a persistent memory to avoid past mistakes
- 65. LAION Releases 10-Million-Hour Video Dataset for AI Training
- 66. Anthropic's Self-Improving AI System Replicates Research Process
- 67. Google DeepMind's AI Co-Scientist Writes Plausible but Inaccurate Methods in Papers
- 68. Cohere's Parse 5 Scores 79.2 on ParseBench, Winning on Cost
- 69. China Sees AI Agent Behavior as Key to Higher Value
- 70. Anthropic's MHS Research Preview Aims to Streamline Custom AI Software Integration
- 71. Study: AI Shopping Agents Show 90-99% Bias in Product Selection
- 72. OpenAI Researcher Warns AI Speed May Overwhelm Human Security Teams
- 73. GLM-5.3-Flash matches GPT-5.6 Terra at 90% lower cost
- 74. AI Agents Exploited Hugging Face in Days-Long Incident
- 75. AI Method Boosts Stability of New Materials in Diffusion Models
- 76. ESQ-Bench: A New NL2SQL Benchmark Tests Dialect Generalization and Silent Failures
- 77. MIT Grad’s Startup Aims to Make Phones Easier for Older Adults
- 78. OpenAI's Jalapeño Chip Boasts High Efficiency, Low Latency in Benchmarks
- 79. Nvidia Claims Groq 3 LPX Cuts Coding Tasks From Hours to Minutes
- 80. Stanford Study: AI Hits Entry-Level Jobs Hardest
- 81. Google's ME-POIs Adds "How a Place Is Used" to POI Embeddings
- 82. NVIDIA's Groq 3 LPX Enters Full Production for AI Agents
- 83. Harvey Launches AI for M&A Diligence, Contract Review Over 10,000 Docs
- 84. A.I. Agent Adoption Grows More Slowly Than AI Itself
- 85. AI Labs Lack Plans to Contain Rogue Models, Documents Show
- 86. AI Agents Succeed or Fail Based on Their "Skills
- 87. Ignoring Human Beliefs Leads AI to Predict Wrong Actions, Study Finds
- 88. RadixAttention Speeds First Tokens by Reusing Cached KV States
- 89. Deepseek Targets Visual Agents With New Experimental Flash Model
- 90. Meta Pays Microsoft Hundreds of Millions Annually for AI
- 91. Anthropic's Claude Expands Into Protein Design
- 92. OpenAI Revokes Access to Limited Cyber Program Daybreak Blue
- 93. GLM-5.3 Jumps 246 Points on Key Benchmark, Ranks Second to Claude Opus
- 94. OpenAI Pauses Astra Model Work Over Safety Concerns
- 95. New LLM Agent Achieves 100% Accuracy on FDA-Based Clinical Trial Benchmark
- 96. Cartesia's Sonic-3.6 TTS Model Adds Natural Pauses and Hinglish Code-Switching
- 97. OpenAI slows model development amid rising cybersecurity risks
- 98. New Benchmark Tests AI Search APIs on 900 Research Questions
- 99. New AI Art Study Finds Images Can't Be Traced to Training Data
- 100. Optima's New AI Benchmark Lets Users Test Models With Their Own Data