📂 Category
Research & Benchmarks News Archive - Page 5 of 6
547 articles in this category • Page 5 of 6
- 401. GPT-5.2 Thinking emerges as collaborative AI for end-to-end web builds
- 402. YouTube channel serves AI concepts in under-minute clips for fast learning
- 403. AI Ends Build-vs-Buy Debate, Focus Shifts to Real Business Impact
- 404. Google, MIT study finds multi-agent AI often loses context in sequential tasks
- 405. Researchers find complex AI persona tactics hurt meaning in development
- 406. AI2 releases Olmo 3.1 32B Think, up 5+ points on AIME and 4+ on ZebraLogic
- 407. Pangram 3.0 AI detector reports 99.98% accuracy, adds four usage tiers
- 408. Experts say data centers' water use is less risky than public perceives
- 409. U.S. Leads Pax Silica Initiative Launched at Summit to Secure Silicon Supply
- 410. Gemini Deep Research agent posts top results on HLE, DeepSearchQA, leads BrowseComp
- 411. Google's FACTS benchmark shows 70% factuality ceiling across four tests
- 412. LangSmith Fetch lets Claude Code, Cursor agents debug from terminal
- 413. SAP deploys 95%-accurate AI to redefine consultant role by 2030
- 414. TPOT evolves ML pipelines via genetic algorithms in four steps
- 415. Model distillation cuts latency 2-3× and lowers costs by double-digit percentages
- 416. Googler details meta-prompt technique that guides Gemini to craft Veo videos
- 417. CognitiveLab unveils NetraEmbed, 150% accuracy gain, adds ColNetraEmbed
- 418. Student AI models can inherit bias and harmful traits from teacher models
- 419. 70% of Creatives Fear Stigma as AI Drives Majority of Their Ideas – Anthropic
- 420. Corporate AI agents favor simple workflows; 41.5% accept minute-range latency
- 421. Bright Data API Delivers Seamless AI/ML Integration and Anti-Bot Protection
- 422. AI agents claim sources verified despite dead links; 14 error types logged
- 423. Harbor Framework Enables Sandbox Agent Execution on Docker, Modal, Daytona
- 424. Anthropic puts Claude in the interviewer's chair for AI testing
- 425. Physicist Steve Hsu releases paper on AI-assisted physics using GPT-5 idea
- 426. OpenAI trials “Confessions” tool that makes models generate self-audit reports
- 427. NVIDIA offers up to USD 60,000 fellowships to PhD students for model collaboration
- 428. NVIDIA cuts prices on Jetson edge-AI developer kits for holiday shoppers
- 429. Anthropic faces pressure as CEO Dario Amodei backs AI regulation
- 430. Harvard Data Course Runs 66 Weeks, Costs USD 1,332.90 (~Rs 1.18 Lakh)
- 431. Gemini 3 Pro tops trust, ethics, safety at 69% vs 16% for Gemini 2.5
- 432. Counter-Strike Sets New Benchmark for Vibe Coding, Says Ex-Mixpanel CEO
- 433. NVIDIA open-sources NeMo Data Designer for synthetic AI datasets at NeurIPS
- 434. AI models stop 87% of attacks but only 8% of attempts; Qwen3-32B hits 86.18%
- 435. Runway's Gen-4.5 text-to-video AI claims unprecedented physical accuracy
- 436. AI Solves 30-Year-Old Math Problem, Showcasing Perplexity's Patent Search Tool
- 437. OpenAGI agent says it beats OpenAI and Anthropic; study deems over-optimistic
- 438. Google's self-modifying model needs extra engineering, smarter compute for complex training
- 439. ARC benchmark declines as labs tune AI to optimize its specific logic
- 440. General Agentic Memory uses dual-agent design, beats RAG on benchmarks
- 441. NeurIPS 2025: Top 4 Papers Highlight Shift From Bigger Models to Limits
- 442. 97% Can't Distinguish AI Music; 71% Surprised, 51% Uncomfortable
- 443. TPUs Designed for Deep Learning Can Outperform GPUs in Many Workloads
- 444. Wipro partners with IISc and FSID for AI and quantum research collaboration
- 445. Google expands AI partnership with Tel Aviv University, infrastructure for Gemma
- 446. Alibaba's AgentEvolver lifts tool-use accuracy ~30% via auto-generated tasks
- 447. Set Seed in XLMiner: Use Integer 12345, 42, 2024 for Consistent Partitions
- 448. Karpathy says AI-homework crackdown failed, urges in-class grading shift
- 449. Digital Connexion to Invest USD 11 Billion in Andhra Pradesh AI Data Centres
- 450. Cecilia Heyes labels language a 'cognitive gadget' for precise social learning
- 451. DOE orders cloud, labs, and network integration for AI Genesis mission in 90 days
- 452. CrowdStrike's Stein finds DeepSeek-R1 adds 50% more bugs on Chinese prompts
- 453. Microsoft's Fara-7B AI agent, rival to GPT-4o, runs on PC, logs 145k tasks
- 454. Authors retract brain-mapping paper after reviewers flag fabricated citations
- 455. CrewAI Introduces Function-Based Guardrails for Rule-Based Output Constraints
- 456. Anthropic finds strict anti-hacking prompts increase AI sabotage and lying
- 457. M-GRPO Boosts Coordination in Multi-Agent Training Over Single-Agent GRPO
- 458. Google DeepMind hires ex-Boston Dynamics CTO to create Gemini AI for any robot
- 459. NotebookLM Turns Complex Spreadsheets into Presentation Insights
- 460. Use Temporal Patterns: Plot Timestamps to Spot Seasonality, Trends, Shifts
- 461. MIT Energy Initiative conference highlights storage research priorities
- 462. ServiceNow uses LangSmith, knowledge graph and MCP to orchestrate agents
- 463. OpenAI researcher details new AI model using general RL, no code interpreters
- 464. WeatherNext 2 data now in Earth Engine, BigQuery; Vertex AI early access opens
- 465. Stereogum persists amid streaming, AI and dwindling ad revenue as ads dry up
- 466. DeepEyesV2 Beats Larger Open-Source Models by Leveraging Search Tools
- 467. Google AI agents: consistency, context, short-term session history, long-term memory
- 468. Researchers push Context Engineering 2.0 as AI moves from Era 2.0 to 3.0
- 469. OpenAI finds sparse models aid debugging, may boost mechanistic interpretability
- 470. RDMA Cuts CPU Use in S3-Compatible Storage, Boosting AI Performance
- 471. Indian language ID proves tough; authors release baseline ML models
- 472. NVIDIA Blackwell Wins All MLPerf Training v5.1 Benchmarks with FP4 Accuracy
- 473. Upwork study finds AI agents outperform alone when paired with humans
- 474. Human-aligned AI models show greater robustness and reliability, study finds
- 475. DeepMind AI agent explores new games, explains its actions better than SIMA 1
- 476. OpenAI Codex CLI Works with ChatGPT Plan and VS Code Extension
- 477. Bengaluru Hosts The Best Firm Summit 2026 for HR and AI Leaders
- 478. Google AI Advisors Let Users Probe Performance with Conversational “Why” Queries
- 479. RECAP tool shows Claude 3.7 reproduces ~3,000 words from The Hobbit and Harry Potter
- 480. ElevenLabs' Scribe v2 delivers real-time, negative-latency transcription
- 481. Meta's SPICE framework beats baselines, boosts math and general reasoning
- 482. Study finds reasoning LLMs are more efficient but not more capable
- 483. AI vision pioneer aims to extend models from data to space understanding
- 484. LearnLM tutoring boosts student problem-solving by 5.5 percentage points
- 485. DeepSeek OCR Fast but Fails Complex Forms; Choose Proven Architecture
- 486. Meta's Omnilingual ASR hits sub-10% error on 78% of 1,600 languages
- 487. Experts advise locating new US data centers outside water-stressed California
- 488. Space Data Centers: Companies Harness Sunlight, Cooling, No Permits for AI
- 489. Pharma Cautious as AI Promises Faster Drug Discovery and Smarter Trials
- 490. Google's Veo-3 fakes surgical videos; 1.78 handling, 1.64 tissue, lowest logic
- 491. Amazon launches beta AI translation for self-published Kindle books
- 492. Interactive AI Agent Uses OpenAI Function Schemas for Rapid ML Tasks
- 493. ComputeEval 2025.2 expands to 232 CUDA challenges, upping LLM test difficulty
- 494. Google's Ironwood TPU to be generally available on Cloud in weeks
- 495. OpenAI Says It Won’t Seek Government Backstop for Infrastructure, CFO Friar Says
- 496. German Commons opens pipeline to free AI datasets from copyright limbo
- 497. SerpApi Converts Live Search Results into Structured API Data for ML Pipelines
- 498. Forest Listeners lets users explore Amazon and Atlantic forests to find species
- 499. Ex-Microsoft Chair warns AI will cut entry-level jobs at Bengaluru summit
- 500. Databricks study: AI judges need people focus, not just tech development