AI Europe
  • Europe
  • Europa
  • Britain
  • France
  • Germany
  • Italy
  • Spain
  • Japan
  • Canada
  • Africa
  • People
  • AI
  • …
    • Afrique
    • Netherlands
    • Poland
  • Agentic AI
  • AGI
  • AI
  • Anthropic
  • Google
  • Microsoft
  • OpenAI
  • xAI
AI Europe
  • Europe
  • Europa
  • Britain
  • France
  • Germany
  • Italy
  • Spain
  • Japan
  • Canada
  • Africa
  • People
  • AI
  • …
    • Afrique
    • Netherlands
    • Poland

Browsing Tag

Benchmark

10 posts
AAgentic AI
AWS Releases Aws-Bench to Evaluate Agents on Cloud Tasks
Read More

AWS Releases Aws-Bench to Evaluate Agents on Cloud Tasks

  • August 22, 2026
AWS has recently released aws-bench, an open-source benchmark to evaluate how accurately AI agents complete real AWS tasks,…
AAnthropic
Anthropic Reportedly in Talks to Acquire Rival Decart for $6 Billion, Bolstering Compute Ahead of IPO — BigGo Finance
Read More

Anthropic Reportedly in Talks to Acquire Rival Decart for $6 Billion, Bolstering Compute Ahead of IPO — BigGo Finance

  • August 13, 2026
AI startup Anthropic is reportedly in talks to acquire rival Decart AI in a deal valued at approximately…
GGoogle
Google Doubles Down on Faster, Cheaper AI — but No Sign of Gemini 3.5 Pro
Read More

Google Doubles Down on Faster, Cheaper AI — but No Sign of Gemini 3.5 Pro

  • July 21, 2026
Google is rolling out some new AI models — but not the one that many were expecting. The…
GGoogle
Eros GenAI Researchers Win Grand Prize in Google DeepMind × Kaggle AGI Benchmark Hackathon
Read More

Eros GenAI Researchers Win Grand Prize in Google DeepMind × Kaggle AGI Benchmark Hackathon

  • July 17, 2026
GAUGE benchmark evaluates whether frontier AI models recognise uncertainty and act on it responsibly LONDON , UNITED KINGDOM,…
AAgentic AI
Image description
Read More

AI search agents don’t fail at searching, they fail at asking the right questions when queries get ambiguous

  • July 5, 2026
AI search agents rarely fail at multi-step research tasks because of the search itself. Their real problem is…
AAgentic AI
문혜원 기자
Read More

KAIST quantifies AI agents using up to 136 times more power per query

  • July 5, 2026
A Korean research team has quantitatively analyzed, for the first time in the world, the power consumption and…
OOpenAI
ChatGPT
Read More

ChatGPT Pro Is Splitting Into Three: GPT-5.6 Benchmark Reveals Luna, Terra, Sol Pro

  • July 2, 2026
An OpenAI research paper published on June 30 named three configurations of GPT-5.6 that the company has never…
AAgentic AI
Image description
Read More

Only three AI models finished above starting capital in a 500-day startup survival test

  • June 28, 2026
To test exactly these skills, the researchers developed CEO-Bench. The benchmark simulates a realistic example of this kind of…
OOpenAI
Getty Images Stock Soars After OpenAI Licensing Agreement | June 2026 - News and Statistics
Read More

Getty Images Stock Soars After OpenAI Licensing Agreement | June 2026 – News and Statistics

  • June 22, 2026
Jun 22, 2026 Getty Images Holdings Inc. experienced a stock surge of up to 145% on Monday following…
OOpenAI
Life Science
Read More

OpenAI Life Science Benchmark Reveals AI Passes Only 1 in 3 Scientific Research Tasks

  • June 18, 2026
OpenAI published LifeSciBench on June 17, 2026, a 750-task evaluation built with 173 PhD-level scientists to test whether…
AI Europe
www.europesays.com