Compare to

Discover how Anthropic's Claude Opus 5 and Google's Gemini 3.8 Flash stack up against each other in this comprehensive comparison of two leading AI language models. Released in July 2026 and September 2026 respectively, these models represent significant advancements in artificial intelligence, with Claude Opus 5 offering a 1,000,000-token context window and Gemini 3.8 Flash offering a 1,048,576-token context window.

Explore their capabilities, pricing, and performance metrics, with Claude Opus 5 achieving 1,493 on LMArena Elo and Gemini 3.8 Flash scoring 1,493, making this comparison essential for developers and organizations seeking the right AI solution for their specific needs.

Models Overview

Anthropic Claude Opus 5
Google Gemini 3.8 Flash

Provider

The company that provides the model.
AnthropicGoogle

Context Length

Maximum number of tokens the model can process
1M1.05M

Maximum Output

Maximum number of tokens the model can generate in one response
128K65.54K

Release Date

When the model was first released.
24-07-202602-09-2026

Knowledge Cutoff

When the model's training data ends.
2026-05Unknown

Open Source

Whether the model weights are openly available.
FALSEFALSE

Pricing Comparison

Compare the pricing of Anthropic's Claude Opus 5 and Google's Gemini 3.8 Flash to determine the most cost-effective solution for your AI needs. Prices are the standard API tier per million tokens, as published by each provider as of September 2026.

Anthropic Claude Opus 5
Google Gemini 3.8 Flash

Input Cost

Cost per million input tokens
$5 / 1M tokens$0.75 / 1M tokens

Output Cost

Cost per million tokens generated
$25 / 1M tokens$3.75 / 1M tokens

Comparing Benchmarks and Performance

Compare the performances of Anthropic's Claude Opus 5 and Google's Gemini 3.8 Flash on industry benchmarks. Scores are the ones the providers and public leaderboards report; a benchmark neither reports is left out.

Anthropic Claude Opus 5
Google Gemini 3.8 Flash

LMArena Elo

Crowd-sourced blind preference rating on the LMArena text leaderboard.
1,4931,493

Terminal-Bench 2.1

Agentic tasks completed in a real terminal.
Benchmark not available89.4%

Humanity's Last Exam

Expert-written questions at the frontier of human knowledge.
56.3%Benchmark not available

Sources — Claude Opus 5: platform.claude.com, arena.ai, anthropic.com; Gemini 3.8 Flash: ai.google.dev, blog.google, arena.ai, deepmind.google.

Compare More Models