📂 Category
Research & Benchmarks News Archive - Page 4 of 7
680 articles in this category • Page 4 of 7
- 301. New embeddings prioritize preferential similarity over semantics for clustering
- 302. Baidu's Ernie 5.1 Cuts 94% Pre‑Training Costs Using Once‑For‑All Framework
- 303. Hermes Agent tops use as Nous Research’s self‑improving model leads OpenRouter
- 304. Palisade Research: Open‑weight AI like Qwen boost autonomous hacking
- 305. Study proposes method to curb AI reward hacking in safety tests
- 306. Build Python Vector Search with Cosine Similarity for Scale‑Invariant Matching
- 307. AI success shifts from 95% accuracy to latency, cost, and reliability
- 308. Apple Workshop Shows ML with Homomorphic Encryption, Georgia Institute, CISPA
- 309. OpenAI opens GPT-5.5-Cyber to vetted security researchers, adds three tiers
- 310. LightSeek launches TokenSpeed, cutting LLM latency by half vs TensorRT-LLM
- 311. CLIP-FP8 Model Matches CLIP-FP16 Quality; Patch Embedding Quantizers Matter
- 312. Automation updates AI context morning with active threads, key dates, note
- 313. Google DeepMind buys minority stake in EVE Online studio for AI testing
- 314. Meta AI releases NeuralBench, benchmark for 36 EEG tasks, 94 datasets
- 315. iTARFlow Shows Competitive Performance on ImageNet 64‑256px Resolutions
- 316. CreativityBench benchmark introduces 4K‑entity affordance KB to test LLM creativity
- 317. Self-Attentive Meta-Optimizer Adds Gradient Alignment and Group-Adaptive Rates
- 318. Local edits in LLM-driven NAS can trigger broader performance shifts
- 319. Groq‑Powered Agentic Assistant Uses Sub‑Agent to Catalog 2024‑25 SLMs
- 320. AI autoencoders and joint communications‑sensing rank among 6G enablers
- 321. Anthropic adds 'dreaming' feature to Claude Managed Agents for memory recall
- 322. Anthropic's USD 200 B, five‑year Google Cloud deal makes up >40% of backlog
- 323. MRC retires paths, then probes to confirm failures and recovery
- 324. eOptShrinkQ enables near‑lossless KV cache compression with spectral denoising
- 325. Harvard study finds OpenAI's o1 and 4o outdiagnose ER doctors in 76‑patient test
- 326. 2021 EDEN-unbiased quantizer beats 2026 successor in average accuracy
- 327. US benchmark shows China lagging; Deepseek model underperforms private tests
- 328. Google DeepMind AI co‑clinician beats GPT‑5.4 in blind tests, lags docs
- 329. Anthropic benchmark says Claude matches experts, 23 tasks remain ambiguous
- 330. Grok Voice Think Fast 1.0 lets non‑programmers design agents via console.x.ai
- 331. WPI professor Gerych offers solution to AI vision ‘Whac‑a‑mole’ bias dilemma
- 332. New method advances privacy‑preserving AI training on consumer devices
- 333. Musk says he was duped, warns AI could kill us, xAI to IPO via SpaceX in June
- 334. NVIDIA BioNeMo wraps CPU layer with DistributedTriangleMultiplication
- 335. DeepSeek unveils new AI breakthrough as nation tightens grip on departing firms
- 336. Poolside AI launches Laguna XS.2 and M.1, hitting 72.5% on SWE-bench Verified
- 337. Oracle abandons its legacy, pivots to AI in an unconventional approach
- 338. New Architecture Separates Execution and Review Agents for Tool-Calling
- 339. MIT study links language model scaling success to superposition of concepts
- 340. Google staff urge Sundar Pichai to reject classified military AI projects
- 341. AI framework autonomously optimizes data, models, algorithms, outperforms humans
- 342. MolClaw Introduces Autonomous Agent for Hierarchical Drug Screening
- 343. Lakehouse concept drives AI data access for thousands of enterprise users
- 344. Fine-tuning RAG embeddings may drop retrieval accuracy 40%, study finds
- 345. AI pipelines show silent failures from orchestration drift, detected weeks later
- 346. OSWorld Benchmark Evaluates LLMs on Real Computer Use, Unlike Text‑Only Tests
- 347. PageIndex Retrieves via Reasoning Using OpenAI gpt-5.4 Model
- 348. Discord Users Access Anthropic's Mythos AI Tool Without Authorization
- 349. Google DeepMind's Vision Banana Outperforms SAM 3 and Depth Anything V3
- 350. DeepMind spinoff’s AI‑designed drugs enter human trials after AlphaFold 3
- 351. COALA paper defines agent memory types: procedural rules and semantic facts
- 352. Google DeepMind's Decoupled DiLoCo hits 88% goodput despite hardware failures
- 353. Agent observability powers production evaluation through trace analysis
- 354. Xiaomi launches MiMo‑V2.5‑Pro and V2.5, matching benchmarks at lower token cost
- 355. Designing Production-Grade CAMEL Multi-Agent Systems: Start with Docs and GitHub
- 356. Multi-agent AI systems incur higher token costs than single agents in practice
- 357. Reinforcement learning trains AI like OpenAI's o1 to admit uncertainty
- 358. LangSmith adds reusable LLM-as-judge and rule-based code evaluator templates
- 359. AI made up over a third of new sites by 2025; Pope warning flagged as AI
- 360. Sergey Brin pushes DeepMind to match Claude, unveils agent skills catalog
- 361. Fortnite adds AI‑powered NPCs for unscripted player conversations
- 362. TabPFN hits 98.8% accuracy in 0.47 s, beating Random Forest and CatBoost
- 363. NVIDIA PhysicsNeMo Tutorial Maps k(x,y) to u(x,y) for Darcy Flow
- 364. OpenAI unveils GPT‑Rosalind, AI model to speed drug discovery and genomics
- 365. Standard LLM guidelines focus on training costs, overlook inference budget
- 366. GPT‑Rosalind life‑sciences plugin for Codex launches on GitHub
- 367. OpenAI launches GPT-Rosalind, hits top score on BixBench benchmark
- 368. Frontier AI models fail one in three production runs, audits grow harder
- 369. Meta researchers unveil hyperagents for self‑improving AI in non‑coding tasks
- 370. Claude outperforms humans on alignment task, but results disappear in production
- 371. Google DeepMind unveils Gemini Robotics‑ER 1.6, beats prior model in tool count
- 372. UK tests Mythos AI, noting its ability to chain multistep attacks
- 373. AI Forum Launches Professional Certificate and USD 120M Fund for AI Fluency
- 374. Databricks finds multi-step agents beat single-turn RAG by 21% to 38% on STaRK
- 375. Stanford AI Index 2026: 53% adopt generative AI in 3 years, education lags
- 376. NVIDIA, UMD release AF-Next audio model, beats Phi-4-mm by 12 points on Arabic
- 377. Developers Claim Measured Drop in Claude's Performance, Sparking Nerf Debate
- 378. Seven AI agents in finance lift cash flow >3% monthly, boost productivity 50%
- 379. Meta AI and KAUST Propose Neural Computers Merging Compute, Memory, I/O
- 380. Prediction drift can mask security model decay despite stable accuracy
- 381. Researchers say OpenAI's Sora and Google's Veo aren't true world models
- 382. TriAttention KV Cache Compression Matches Full Attention, 2.5× Faster
- 383. Knowledge Distillation Keeps Student Model Capacity to Match Ensemble Boundaries
- 384. Google AI's PaperOrchestra boosts manuscript success, 79‑81% win rate
- 385. OSGym runs 1,000+ OS replicas at USD 0.23/day with decentralized state management
- 386. Stanford study finds AI agent handoffs lose information, affecting compute cost
- 387. Meta Superintelligence Labs launches Muse Spark, its first multimodal AI model
- 388. Better Harness updates add usage examples, chaining guide, and tool clarifications
- 389. Study finds ‘bot’ term used 16,232 times in 2.8M Telegram messages
- 390. Google AI Overviews answers 91% of test questions correctly after Gemini 3 update
- 391. MaxToki AI boosts context to 16,384 tokens with RoPE scaling
- 392. Meta staff inflate AI token counts on internal leaderboard, wasting resources
- 393. MassMutual, Mass General Brigham turn AI pilot sprawl into production
- 394. OpenAI safety staff exit as Altman dismisses Pentagon contract concerns
- 395. OpenAI urges firms to fund pensions, health, childcare as AI cuts costs
- 396. Study shows sycophantic AI chatbots can outwit ideal rational users
- 397. Americans use AI more than ever but trust it less, Quinnipiac poll shows
- 398. Study maps developer frustration with AI slop as tragedy of the commons
- 399. Google study: AI benchmarks ignore human disagreement; under 10 raters fail
- 400. Alibaba's Qwen team adds method that lengthens AI answers, prompting reasoning