📂 Category
Research & Benchmarks News Archive - Page 2 of 6
547 articles in this category • Page 2 of 6
- 101. Study Defines Privacy-Utility Frontier for Agent Memory via PR and AER
- 102. Model 5 tops penalized PR-AUC, recall and F1-score in scoring model training
- 103. NVIDIA Nsight Designer Streams ONNX Editing and TensorRT Engine Build
- 104. AI moves beyond automation to plan, optimize and execute business initiatives
- 105. NVIDIA FLARE Auto-FL Enables Agent-Led Coding in Controlled Experiments
- 106. Multiverse reduces inference cost by favoring low‑cost prefill over decoding
- 107. AI agents solve neuroscience pipeline tasks on datasets larger than benchmarks
- 108. ML models predict World Cup outcomes, but miss draws, capture team strength
- 109. Reddit releases AI comment archive to study LLM persuasion tactics
- 110. Nvidia plans PC reboot, Apple unveils smart glasses on Vergecast
- 111. Open LLM v2, 12‑benchmark suite, LiveBench show d_eff 2.86‑4.80
- 112. NSF renews MIT AI‑physics institute, adds museum and hackathon outreach
- 113. From Prompt Tools to Workflow‑Driven AI: Managing Learning Curves
- 114. Geospatial ML Models Show Uneven Reliability Across Sparse Strata
- 115. Agents automate data retrieval, cleaning, analysis, modeling and reporting
- 116. Explainable ML Classifies Alzheimer's Early in 1,641 ADNI Subjects
- 117. MIT researchers train AI to read charts, streamlining downstream workflows
- 118. Lightweight CNN Boosts Adversarial Robustness in EEG‑Based Brain‑Computer Interfaces
- 119. Hundreds sign Leiden Declaration as AI threatens mathematicians' profession
- 120. Transformer tops Gait2Hip-60 benchmark with 0.819 R² in hip force prediction
- 121. QASM-Eval Introduces First Dataset for Training LLMs on OpenQASM-3
- 122. Parallax adds learned covariance correction to linear attention, retains softmax
- 123. Men use AI coding agents over twice as often as women; economists at 39%
- 124. Molecule-trained AI gives better chicken pairing suggestions than recipe AI
- 125. AI search agents favor confirming hits, sideline gut answers, study finds
- 126. OpenAI gives free life‑sciences AI model to aid government pandemic prep
- 127. Review paper claims code defines AI agents' reasoning and behavior
- 128. NVIDIA research moves robotics simulation to reality, revealing robot confusion
- 129. CVPR 2026 Friday Session: STARFlow‑V Video Modeling Poster #178, 4‑6 PM
- 130. USD E^3USD ‑Agent splits fast router from LLM meta‑controller for edge inference
- 131. Sakana AI's DiffusionBlocks Apply Uniform [4,4,4] Layers Across Three Blocks
- 132. AI Agent Auto-Identifies Unreadable Model Parameters from CSV Files
- 133. Learn to Build AI Projects: n8n Automation, Financial Data, Summaries, Reports
- 134. AI Agents Falter in Production as Backward Design Overburdens Model
- 135. Hugging Face releases LeRobot Humanoid: 3D‑printable legs for robot research
- 136. Synthetic 1,000‑Customer Dataset Uses Gender and Income to Test Bias
- 137. SciAtlas Introduces Large-Scale Knowledge Graph to Aid Automated Research
- 138. Google outperforms OpenAI on math benchmark, winning 9 to 1 ratio
- 139. ByteDance study: LMMs answer questions better than full-page transcription
- 140. Language Models Forecast Research Success Using 11,488 Comparative Idea Pairs
- 141. Researchers use triplet loss to train high-quality Horn logic embeddings
- 142. Positive-IC 46.4% indicates negative bias; |IC| just under 0.02 after two runs
- 143. AgentNLQ released as a general‑purpose NL2SQL agent; accuracy lags human writers
- 144. SuperAI Conference Highlights Growing AI Startup and Infrastructure Scene in Asia
- 145. CODEX Agent Adds AI‑Q Deep Research Skill from GitHub Repository
- 146. AI models learn chemistry; talent and collaborations offset location concerns
- 147. Real-Time Diffusion on Apple M3 Ultra: CoreML, Quantization, Neural Engine
- 148. Evaluating AI Agents: Does the Engine Grasp Instructions and Reason Facts?
- 149. AgentWall adds runtime safety layer for local AI agents' actions
- 150. LangSmith Engine automates agent debugging; OpenAI's Frontier offers platform
- 151. VideoWorld paper links prediction, simulation, reasoning in robotics
- 152. Channel-independent tolerate modalities but falter on within-modality gaps
- 153. Study Finds Current ToM Benchmarks Overlook First‑Person, Dynamic Interaction
- 154. Graph‑Enhanced RAG Architecture Cuts Latency in Meta‑Scale Production
- 155. OpenClaw founder runs 100 AI agents for USD 1.3M/month code, review PRs, find bugs
- 156. New benchmark shows AI video generators look realistic but lack reasoning
- 157. Researchers train AI model achieving near-full performance using 12.5% of experts
- 158. RecursiveMAS cuts multi-agent inference time 2.4×, slashes token use 75%
- 159. ArXiv to ban authors of papers with unchecked LLM‑generated content
- 160. BenchJack proposes secure-by-design AI benchmark audit with eight flaw taxonomy
- 161. 12‑Metric AI Agent Eval Harness Built in 9‑14 Days Across 100+ Deployments
- 162. Google DeepMind adds Gemini-powered cursor to Chrome for visual queries
- 163. BaLoRA adds Bayesian uncertainty to low‑rank adaptation, but lags fine‑tuning
- 164. Community review tools guide novices in AI research, study finds
- 165. Tilde Research's Aurora optimizer beats Muon and NorMuon at 340M scale
- 166. OpenAI unveils Daybreak to secure Codex, with industry and government rollout
- 167. New embeddings prioritize preferential similarity over semantics for clustering
- 168. Baidu's Ernie 5.1 Cuts 94% Pre‑Training Costs Using Once‑For‑All Framework
- 169. Hermes Agent tops use as Nous Research’s self‑improving model leads OpenRouter
- 170. Palisade Research: Open‑weight AI like Qwen boost autonomous hacking
- 171. Study proposes method to curb AI reward hacking in safety tests
- 172. Build Python Vector Search with Cosine Similarity for Scale‑Invariant Matching
- 173. AI success shifts from 95% accuracy to latency, cost, and reliability
- 174. Apple Workshop Shows ML with Homomorphic Encryption, Georgia Institute, CISPA
- 175. OpenAI opens GPT-5.5-Cyber to vetted security researchers, adds three tiers
- 176. LightSeek launches TokenSpeed, cutting LLM latency by half vs TensorRT-LLM
- 177. CLIP-FP8 Model Matches CLIP-FP16 Quality; Patch Embedding Quantizers Matter
- 178. Automation updates AI context morning with active threads, key dates, note
- 179. Google DeepMind buys minority stake in EVE Online studio for AI testing
- 180. Meta AI releases NeuralBench, benchmark for 36 EEG tasks, 94 datasets
- 181. iTARFlow Shows Competitive Performance on ImageNet 64‑256px Resolutions
- 182. CreativityBench benchmark introduces 4K‑entity affordance KB to test LLM creativity
- 183. Self-Attentive Meta-Optimizer Adds Gradient Alignment and Group-Adaptive Rates
- 184. Local edits in LLM-driven NAS can trigger broader performance shifts
- 185. Groq‑Powered Agentic Assistant Uses Sub‑Agent to Catalog 2024‑25 SLMs
- 186. AI autoencoders and joint communications‑sensing rank among 6G enablers
- 187. Anthropic adds 'dreaming' feature to Claude Managed Agents for memory recall
- 188. Anthropic's USD 200 B, five‑year Google Cloud deal makes up >40% of backlog
- 189. MRC retires paths, then probes to confirm failures and recovery
- 190. eOptShrinkQ enables near‑lossless KV cache compression with spectral denoising
- 191. Harvard study finds OpenAI's o1 and 4o outdiagnose ER doctors in 76‑patient test
- 192. 2021 EDEN-unbiased quantizer beats 2026 successor in average accuracy
- 193. US benchmark shows China lagging; Deepseek model underperforms private tests
- 194. Google DeepMind AI co‑clinician beats GPT‑5.4 in blind tests, lags docs
- 195. Anthropic benchmark says Claude matches experts, 23 tasks remain ambiguous
- 196. Grok Voice Think Fast 1.0 lets non‑programmers design agents via console.x.ai
- 197. WPI professor Gerych offers solution to AI vision ‘Whac‑a‑mole’ bias dilemma
- 198. New method advances privacy‑preserving AI training on consumer devices
- 199. Musk says he was duped, warns AI could kill us, xAI to IPO via SpaceX in June
- 200. NVIDIA BioNeMo wraps CPU layer with DistributedTriangleMultiplication