AI Europe
  • Europe
  • Europa
  • Britain
  • France
  • Germany
  • Italy
  • Spain
  • Japan
  • Canada
  • Africa
  • People
  • AI
  • …
    • Afrique
    • Netherlands
    • Poland
  • Agentic AI
  • AGI
  • AI
  • Anthropic
  • Google
  • Microsoft
  • OpenAI
  • xAI
AI Europe
  • Europe
  • Europa
  • Britain
  • France
  • Germany
  • Italy
  • Spain
  • Japan
  • Canada
  • Africa
  • People
  • AI
  • …
    • Afrique
    • Netherlands
    • Poland

Browsing Tag

AI benchmarks

29 posts
AAGI
NIMI Mathematical Landscape
Read More

ARC-AGI-3 Gets Open-Source Agent That Writes Python World Models Instead of Neural Weights

  • July 31, 2026
A German academic research group today posted an open-source AI agent to arXiv that approaches ARC-AGI-3 — the…
AAgentic AI
file photo taken 15 2019 AI robot
Read More

Princeton Gives AI Agents Unpublished Questions: Original Scientists Grade Results

  • July 30, 2026
A Princeton-led team posted a paper today describing the first AI evaluation designed around a deceptively simple idea:…
GGoogle
The letter emerges amid ongoing debate around autonomous systems and safety, dispute over Chinese AI models, etc. (Image for representation: Magnific)
Read More

OpenAI, Google, Meta staff sign letter calling for mechanisms to pace AI progress | Technology News

  • July 29, 2026
4 min readNew DelhiUpdated: Jul 29, 2026 01:10 PM IST As US-China tensions over AI escalate, a group…
MMicrosoft
Markus Kasanmascheff
Read More

Microsoft Releases MAI-Cyber-1-Flash Cybersecurity AI Model With Model Routing

  • July 28, 2026
TL;DR Model Launch: Microsoft has introduced MAI-Cyber-1-Flash and is placing its configuration inside MDASH, the company’s model-routing security…
AAnthropic
Markus Kasanmascheff
Read More

Claude Opus 5 Targets Fable Performance at Lower Cost

  • July 25, 2026
TL;DR Opus Launch: Anthropic has announced Claude Opus 5 as a lower-cost near-Fable option. Token Pricing: The company says…
AAnthropic
Claude Opus 5 outscores Fable 5 on 8 of 13 benchmarks at half the token price
Read More

Claude Opus 5 outscores Fable 5 on 8 of 13 benchmarks at half the token price

  • July 25, 2026
Anthropic’s Claude Fable 5 is arguably the most influential LLM to launch in recent memory, marking the biggest…
AAnthropic
The Blueprint
Read More

Claude Opus 5 delivers top coding scores at half the per-task cost

  • July 24, 2026
Anthropic has launched Claude Opus 5, a new flagship model that the company says delivers state-of-the-art performance on…
MMicrosoft
Markus Kasanmascheff
Read More

Microsoft Weighs Chinese Kimi K3 Model for Selected Copilot Tasks

  • July 24, 2026
TL;DR Copilot Evaluation: Microsoft is weighing Kimi K3 for selected Copilot requests, but no deployment decision is confirmed.…
MMicrosoft
Markus Kasanmascheff
Read More

Microsoft Reportedly Trains Sales Team to Target AI Labs

  • July 17, 2026
TL;DR Internal Briefing: Microsoft executives have outlined a sales plan for employees to challenge OpenAI, Anthropic, and Google.…
AAgentic AI
Qualcomm Says Agents Will Decide Where AI Runs
Read More

Qualcomm Says Agents Will Decide Where AI Runs

  • July 10, 2026
Editor’s note: This is AI Impact, Newsweek’s weekly newsletter where each week, we will explore how business leaders…
xxAI
Markus Kasanmascheff
Read More

Grok 4.5 Launches for Coding Agents as SpaceXAI Tests Lower Prices

  • July 9, 2026
TL;DR Launch: SpaceXAI and Cursor have released Grok 4.5, positioning it for coding agents, long-running tool use, and…
AAgentic AI
Patronus team
Read More

Patronus AI lands $50M to build ‘digital worlds’ that stress-test AI agents

  • June 25, 2026
AI agents are becoming more sophisticated. They are evolving from answering questions to autonomously executing multi-step complex tasks.…
AI Europe
www.europesays.com