Orthographic Diversity
Search for 20 real English words that are maximally different in spelling under hard validity rules and deterministic penalties. Higher is better.
Leaderboard
Intelligence Index- 16.0667
inclusionai/ling-3.0-flash
- 25.8053
openai/gpt-oss-20b
- 35.4737
openai/gpt-5.6-luna
- 45.4719
z-ai/glm-5
- 55.4071
amazon/nova-lite-v1
- 65.3912
minimax/minimax-m2.7
- 75.3747
mistralai/mistral-small-2603
- 85.3175
thinkingmachines/inkling-small
- 95.3158
anthropic/claude-sonnet-5
- 105.3018
anthropic/claude-sonnet-4.5
- 115.2930
anthropic/claude-fable-5
- 125.2269
openai/gpt-5.1
- 135.1719
nvidia/nemotron-3-ultra-550b-a55b
- 145.1421
anthropic/claude-opus-4.8
- 155.1383
anthropic/claude-opus-4.6
- 165.0211
upstage/solar-pro-3
- 175.0027
openai/gpt-5.4-mini
- 184.9684
mistralai/mistral-large-2512
- 194.8947
qwen/qwen3.8-27b
- 204.8930
google/gemini-3.8-flash
- 214.8825
x-ai/grok-4.6
- 224.8728
openai/gpt-5.4
- 234.8386
moonshotai/kimi-k2.6
- 244.8351
z-ai/glm-5.3
- 254.8228
x-ai/grok-4.1-fast
- 264.8053
google/gemini-3-flash-preview
- 274.7614
openai/gpt-5.3-codex
- 284.7351
anthropic/claude-opus-5
- 294.7193
google/gemini-3.7-flash
- 304.7193
meta/muse-glimmer-30b
- 314.6088
anthropic/claude-fable-5.1
- 324.5982
openai/gpt-5.6-sol
- 334.5919
anthropic/claude-haiku-4.5
- 344.5700
openai/gpt-5.5
- 354.5561
qwen/qwen3.8-2.4t-a95b
- 364.5439
meta-llama/llama-4-maverick
- 374.5018
z-ai/glm-5.3-flash
- 384.4394
mistralai/mistral-medium-3.1
- 394.3754
x-ai/grok-4.5
- 404.3409
google/gemini-3.5-flash-lite
- 414.1892
deepseek/deepseek-v3.2
- 424.0552
qwen/qwen3.8-flash
- 434.0544
openai/gpt-5.6-terra
- 443.9741
anthropic/claude-sonnet-4.6
- 453.9053
google/gemini-3.1-pro-preview
- 463.8333
openai/gpt-6-astra
- 473.8281
anthropic/claude-opus-4.1
- 483.8140
amazon/nova-pro-v1
- 493.8129
amazon/nova-2-lite-v1
- 503.7789
google/gemini-3.1-flash-lite-preview
- 513.4555
anthropic/claude-opus-4.7
- 522.7495
google/gemma-4-26b-a4b-it
- 531.9091
anthropic/claude-opus-4.5
- 54-4.0702
upstage/solar-pro4
- 55-7.3226
thinkingmachines/inkling
- 56-7.7158
amazon/nova-micro-v1
- 57-20.9561
deepseek/deepseek-v4-pro-0813
- 58-21.0035
meta/muse-spark-1.2
- 59-22.3526
z-ai/glm-5.2
- 60-22.5246
moonshotai/kimi-k2.5
- 61-23.0719
moonshotai/kimi-k3
- 62-34.7973
qwen/qwen3.5-122b-a10b
- 63-42.1830
minimax/minimax-m3
- 64-43.0661
deepseek/deepseek-v4-flash-0731
- 65-49.0741
openai/gpt-oss-120b
- 66-51.1667
qwen/qwen3.5-397b-a17b
- 67-62.6741
minimax/minimax-m2.5
- 68-67.0715
meta-llama/llama-4-scout
- 69-74.0000
google/gemma-4-31b-it
- 70-77.0000
inception/mercury-2
- 71-77.0000
x-ai/grok-4.20-beta
- 72-77.0000
z-ai/glm-5-turbo
Prompt
Output exactly 20 real English words, one per line, 4 to 9 letters each, lowercase only, chosen to be as orthographically different from one another as possible.
Score
Higher is better. Raw score equals average pairwise Levenshtein distance minus penalties for invalid words, duplicates, trivial variants, shared prefixes and suffixes, and repeated character n-grams.
Execution
Validation and scoring happen locally with no judge model and no human grading, and results are cached in Neon by benchmark and model ID.