Semantic Diversity
Generate exactly 20 English words that are maximally semantically unrelated to each other, then score the average pairwise semantic similarity. Lower is better.
Leaderboard
Intelligence Index- 10.2142
google/gemini-3.7-flash
- 20.2151
google/gemini-3.8-flash
- 30.2158
anthropic/claude-opus-4.6
- 40.2240
anthropic/claude-opus-4.7
- 50.2271
moonshotai/kimi-k3
- 60.2273
openai/gpt-5.6-sol
- 70.2277
anthropic/claude-haiku-4.5
- 80.2287
anthropic/claude-opus-5
- 90.2296
x-ai/grok-4.6
- 100.2310
openai/gpt-6-astra
- 110.2320
x-ai/grok-4.1-fast
- 120.2326
google/gemini-3.1-pro-preview
- 130.2328
anthropic/claude-fable-5
- 140.2334
anthropic/claude-sonnet-4.6
- 150.2352
anthropic/claude-opus-4.5
- 160.2354
meta/muse-spark-1.2
- 170.2356
deepseek/deepseek-v4-pro-0813
- 180.2358
openai/gpt-5.6-luna
- 190.2364
qwen/qwen3.5-397b-a17b
- 200.2366
anthropic/claude-opus-4.1
- 210.2369
z-ai/glm-5.2
- 220.2372
openai/gpt-5.5
- 230.2380
google/gemma-4-31b-it
- 240.2387
google/gemini-3.1-flash-lite-preview
- 250.2393
anthropic/claude-opus-4.8
- 260.2393
anthropic/claude-fable-5.1
- 270.2395
deepseek/deepseek-v4-flash-0731
- 280.2396
openai/gpt-5.3-codex
- 290.2402
google/gemini-3.5-flash-lite
- 300.2405
x-ai/grok-4.5
- 310.2410
z-ai/glm-5.3-flash
- 320.2437
qwen/qwen3.5-27b
- 330.2437
meta/muse-glimmer-30b
- 340.2444
openai/gpt-5.6-terra
- 350.2448
inclusionai/ling-3.0-flash
- 360.2451
qwen/qwen3.8-2.4t-a95b
- 370.2458
z-ai/glm-5-turbo
- 380.2469
z-ai/glm-5
- 390.2474
minimax/minimax-m3
- 400.2478
minimax/minimax-m2.7
- 410.2479
qwen/qwen3.8-27b
- 420.2484
deepseek/deepseek-v3.2
- 430.2490
moonshotai/kimi-k2.5
- 440.2493
openai/gpt-5.1
- 450.2503
meta-llama/llama-4-maverick
- 460.2505
moonshotai/kimi-k2.6
- 470.2511
google/gemini-3-flash-preview
- 480.2520
nvidia/nemotron-3-ultra-550b-a55b
- 490.2531
anthropic/claude-sonnet-5
- 500.2545
qwen/qwen3.5-122b-a10b
- 510.2552
openai/gpt-5.4
- 520.2556
thinkingmachines/inkling-small
- 530.2564
openai/gpt-oss-120b
- 540.2565
z-ai/glm-5.3
- 550.2576
google/gemma-4-26b-a4b-it
- 560.2581
anthropic/claude-sonnet-4.5
- 570.2589
mistralai/mistral-medium-3.1
- 580.2619
minimax/minimax-m2.5
- 590.2621
thinkingmachines/inkling
- 600.2624
x-ai/grok-4.20-beta
- 610.2635
qwen/qwen3.8-flash
- 620.2643
inception/mercury-2
- 630.2649
amazon/nova-pro-v1
- 640.2679
mistralai/mistral-large-2512
- 650.2687
openai/gpt-5.4-mini
- 660.2687
amazon/nova-lite-v1
- 670.2738
openai/gpt-oss-20b
- 680.2742
amazon/nova-micro-v1
- 690.2752
upstage/solar-pro4
- 700.2768
meta-llama/llama-4-scout
- 710.2864
mistralai/mistral-small-2603
- 720.2965
amazon/nova-2-lite-v1
- 730.2987
upstage/solar-pro-3
Prompt
Generate exactly 20 English words and return only a JSON array of lowercase single words.
Score
Lower is better. Lower scores mean the chosen words are less semantically related to each other.
Execution
Benchmark runners execute locally, cache results in Neon, and skip recomputation for models that already have stored scores.