📂 Category
Research & Benchmarks News Archive - Page 3 of 7
680 articles in this category • Page 3 of 7
- 201. DFlash speculative decoding boosts NVIDIA Blackwell inference up to 15×
- 202. NVIDIA BioNeMo Toolkit Enables AI Scientist to Align, Fold, and Dock Molecules
- 203. Data2Story converts CSVs to articles with 7 AI; 53 readers prefer them to human
- 204. M* introduces overlapped scheduling to streamline multimodal model serving
- 205. Benchmark shows Claude Fable 5 passes only 3% of tasks, 31 of 91 fail 50%
- 206. Nobel laureate John Jumper departs DeepMind for Anthropic after AlphaFold win
- 207. OpenAI shows small 'beneficial trait' training makes AI safer, less manipulable
- 208. Google DeepMind uses MITRE ATT&CK to monitor AI agents as rogue employees
- 209. DeFAb Benchmark Enforces Polynomial-Time Checks for Logical Rigor
- 210. OpenAI researchers aim to forecast AI model failure rates pre‑launch
- 211. Nvidia AI Agent Trains Robots Autonomously, Editing Code from Papers
- 212. XGBoost, ALBERT, BioBERT, Med‑LLaMA evaluated for pharmacovigilance
- 213. OpenAI's Deployment Simulation Beats Baseline, Adds Risk Checks to Agentic Code
- 214. GLM-5.2 beats GPT-5.5 on SWE-bench Pro (62.1 vs 58.6) for 1/6 cost
- 215. AMD builds Llama 3.1 8B pretraining benchmark for MLPerf, using random weights
- 216. AMD's MI355X CDNA4 GPU Shows Competitive Training Times in MLPerf v6.0
- 217. NVIDIA Blackwell Leads MLPerf Training 6.0 with Full‑Stack Scale
- 218. DR-DCI Enables Agent-Callable Retrieval to Expand Local Workspace Efficiently
- 219. Fused kernels boost MoE training, forward and backward passes up to 1.3×
- 220. Hybrid Open-Ended Tri-Evolution Improves Deep Research for AI Agents
- 221. Microsoft Research Mirage adds persistent spatial memory to video generation
- 222. Amazon security research prompts White House ban on Anthropic Fable
- 223. Study: AI coding agents locate correct file but miss key lines in bugs
- 224. OpenAI confirms cooperation as state attorneys general launch investigation
- 225. Gemini‑SQL2 leads BIRD benchmark with 80.04% execution accuracy
- 226. NVIDIA tops AA‑AgentPerf benchmark, credits Vera Rubin platform
- 227. Perplexity routes deep‑research subtasks across 20+ models using Gemini agent
- 228. Visual model exploits similarity of 打, 拍, 拉; text model starts from embeddings
- 229. New arXiv Paper Introduces Strategic Decision Support for AI Agents
- 230. LSEG integrates trusted data into ChatGPT workflows, says Max Grigoryev
- 231. Hermes Agent Builder Unites Identity, Model, Skills, Servers in One Dashboard
- 232. SciConBench launches with 9.11K questions to test AI scientific synthesis
- 233. Language Agents Self‑Gate Clarification: Mandatory vs Opportunistic Modes
- 234. Study Defines Privacy-Utility Frontier for Agent Memory via PR and AER
- 235. Model 5 tops penalized PR-AUC, recall and F1-score in scoring model training
- 236. NVIDIA Nsight Designer Streams ONNX Editing and TensorRT Engine Build
- 237. AI moves beyond automation to plan, optimize and execute business initiatives
- 238. NVIDIA FLARE Auto-FL Enables Agent-Led Coding in Controlled Experiments
- 239. Multiverse reduces inference cost by favoring low‑cost prefill over decoding
- 240. AI agents solve neuroscience pipeline tasks on datasets larger than benchmarks
- 241. ML models predict World Cup outcomes, but miss draws, capture team strength
- 242. Reddit releases AI comment archive to study LLM persuasion tactics
- 243. Nvidia plans PC reboot, Apple unveils smart glasses on Vergecast
- 244. Open LLM v2, 12‑benchmark suite, LiveBench show d_eff 2.86‑4.80
- 245. NSF renews MIT AI‑physics institute, adds museum and hackathon outreach
- 246. From Prompt Tools to Workflow‑Driven AI: Managing Learning Curves
- 247. Geospatial ML Models Show Uneven Reliability Across Sparse Strata
- 248. Agents automate data retrieval, cleaning, analysis, modeling and reporting
- 249. Explainable ML Classifies Alzheimer's Early in 1,641 ADNI Subjects
- 250. MIT researchers train AI to read charts, streamlining downstream workflows
- 251. Lightweight CNN Boosts Adversarial Robustness in EEG‑Based Brain‑Computer Interfaces
- 252. Hundreds sign Leiden Declaration as AI threatens mathematicians' profession
- 253. Transformer tops Gait2Hip-60 benchmark with 0.819 R² in hip force prediction
- 254. QASM-Eval Introduces First Dataset for Training LLMs on OpenQASM-3
- 255. Parallax adds learned covariance correction to linear attention, retains softmax
- 256. Men use AI coding agents over twice as often as women; economists at 39%
- 257. Molecule-trained AI gives better chicken pairing suggestions than recipe AI
- 258. AI search agents favor confirming hits, sideline gut answers, study finds
- 259. OpenAI gives free life‑sciences AI model to aid government pandemic prep
- 260. Review paper claims code defines AI agents' reasoning and behavior
- 261. NVIDIA research moves robotics simulation to reality, revealing robot confusion
- 262. CVPR 2026 Friday Session: STARFlow‑V Video Modeling Poster #178, 4‑6 PM
- 263. USD E^3USD ‑Agent splits fast router from LLM meta‑controller for edge inference
- 264. Sakana AI's DiffusionBlocks Apply Uniform [4,4,4] Layers Across Three Blocks
- 265. AI Agent Auto-Identifies Unreadable Model Parameters from CSV Files
- 266. Learn to Build AI Projects: n8n Automation, Financial Data, Summaries, Reports
- 267. AI Agents Falter in Production as Backward Design Overburdens Model
- 268. Hugging Face releases LeRobot Humanoid: 3D‑printable legs for robot research
- 269. Synthetic 1,000‑Customer Dataset Uses Gender and Income to Test Bias
- 270. SciAtlas Introduces Large-Scale Knowledge Graph to Aid Automated Research
- 271. Google outperforms OpenAI on math benchmark, winning 9 to 1 ratio
- 272. ByteDance study: LMMs answer questions better than full-page transcription
- 273. Language Models Forecast Research Success Using 11,488 Comparative Idea Pairs
- 274. Researchers use triplet loss to train high-quality Horn logic embeddings
- 275. Positive-IC 46.4% indicates negative bias; |IC| just under 0.02 after two runs
- 276. AgentNLQ released as a general‑purpose NL2SQL agent; accuracy lags human writers
- 277. SuperAI Conference Highlights Growing AI Startup and Infrastructure Scene in Asia
- 278. CODEX Agent Adds AI‑Q Deep Research Skill from GitHub Repository
- 279. AI models learn chemistry; talent and collaborations offset location concerns
- 280. Real-Time Diffusion on Apple M3 Ultra: CoreML, Quantization, Neural Engine
- 281. Evaluating AI Agents: Does the Engine Grasp Instructions and Reason Facts?
- 282. AgentWall adds runtime safety layer for local AI agents' actions
- 283. LangSmith Engine automates agent debugging; OpenAI's Frontier offers platform
- 284. VideoWorld paper links prediction, simulation, reasoning in robotics
- 285. Channel-independent tolerate modalities but falter on within-modality gaps
- 286. Study Finds Current ToM Benchmarks Overlook First‑Person, Dynamic Interaction
- 287. Graph‑Enhanced RAG Architecture Cuts Latency in Meta‑Scale Production
- 288. OpenClaw founder runs 100 AI agents for USD 1.3M/month code, review PRs, find bugs
- 289. New benchmark shows AI video generators look realistic but lack reasoning
- 290. Researchers train AI model achieving near-full performance using 12.5% of experts
- 291. RecursiveMAS cuts multi-agent inference time 2.4×, slashes token use 75%
- 292. ArXiv to ban authors of papers with unchecked LLM‑generated content
- 293. BenchJack proposes secure-by-design AI benchmark audit with eight flaw taxonomy
- 294. 12‑Metric AI Agent Eval Harness Built in 9‑14 Days Across 100+ Deployments
- 295. Google DeepMind adds Gemini-powered cursor to Chrome for visual queries
- 296. BaLoRA adds Bayesian uncertainty to low‑rank adaptation, but lags fine‑tuning
- 297. Community review tools guide novices in AI research, study finds
- 298. Tilde Research's Aurora optimizer beats Muon and NorMuon at 340M scale
- 299. OpenAI unveils Daybreak to secure Codex, with industry and government rollout
- 300. New embeddings prioritize preferential similarity over semantics for clustering