Wordle
Play 20 recent Wordle answers turn by turn with standard gray/yellow/green feedback. Invalid guesses still cost a turn, scores are capped at 10 turns per puzzle, and lower is better.
Leaderboard
Intelligence Index- 13.6000
openai/gpt-5.3-codex
- 23.6500
anthropic/claude-opus-5
- 33.7000
openai/gpt-6-astra
- 43.7500
google/gemini-3.8-flash
- 53.7500
openai/gpt-5.6-sol
- 63.8000
x-ai/grok-4.6
- 73.9500
anthropic/claude-fable-5.1
- 84.0000
google/gemini-3.7-flash
- 94.0000
openai/gpt-5.5
- 104.0000
openai/gpt-5.6-luna
- 114.0000
x-ai/grok-4.5
- 124.0500
anthropic/claude-opus-4.6
- 134.0500
anthropic/claude-sonnet-5
- 144.0500
thinkingmachines/inkling-small
- 154.0500
z-ai/glm-5.3-flash
- 164.1000
openai/gpt-5.6-terra
- 174.2000
openai/gpt-oss-120b
- 184.2500
qwen/qwen3.8-flash
- 194.3000
anthropic/claude-opus-4.7
- 204.5500
anthropic/claude-fable-5
- 214.5500
inception/mercury-2
- 224.6500
anthropic/claude-sonnet-4.5
- 234.7500
z-ai/glm-5.3
- 244.9500
inclusionai/ling-3.0-flash
- 255.4500
anthropic/claude-opus-4.8
- 265.5500
anthropic/claude-opus-4.5
- 276.8000
google/gemini-3-flash-preview
- 286.8500
anthropic/claude-opus-4.1
- 297.2500
moonshotai/kimi-k3
- 307.7000
qwen/qwen3.8-27b
- 318.2000
google/gemini-3.1-flash-lite-preview
- 328.7000
moonshotai/kimi-k2.6
- 339.2000
anthropic/claude-haiku-4.5
- 349.3500
openai/gpt-5.4
- 359.6000
google/gemma-4-31b-it
- 369.6000
x-ai/grok-4.20-beta
- 379.6500
google/gemini-3.5-flash-lite
- 389.7000
deepseek/deepseek-v3.2
- 399.7000
openai/gpt-5.4-mini
- 409.7500
mistralai/mistral-medium-3.1
- 419.7500
openai/gpt-5.1
- 429.8500
mistralai/mistral-large-2512
- 439.9000
upstage/solar-pro4
- 449.9500
google/gemma-4-26b-a4b-it
- 459.9500
meta-llama/llama-4-maverick
- 4610.0000
amazon/nova-2-lite-v1
- 4710.0000
amazon/nova-lite-v1
- 4810.0000
amazon/nova-micro-v1
- 4910.0000
amazon/nova-pro-v1
- 5010.0000
google/gemini-3.1-pro-preview
- 5110.0000
meta-llama/llama-4-scout
- 5210.0000
minimax/minimax-m2.5
- 5310.0000
minimax/minimax-m2.7
- 5410.0000
mistralai/mistral-small-2603
- 5510.0000
moonshotai/kimi-k2.5
- 5610.0000
openai/gpt-oss-20b
- 5710.0000
qwen/qwen3.5-27b
- 5810.0000
upstage/solar-pro-3
- 5910.0000
z-ai/glm-5
Prompt
The model is told to reply with exactly one 5-letter word per turn, that any extra text is penalized, and that duplicate letters are allowed.
Score
Lower is better. Each puzzle score is the turn the word is solved on, or 10 if the model never solves it within 10 turns. Invalid guesses still count as turns.
Execution
Benchmark runners execute locally, simulate the Wordle judge deterministically, cache results in Neon, and skip recomputation for models that already have stored scores.