WeirdBench

Unconventional LLM benchmarks for the weird corners other evals skip. Definitions and runners live locally, scores are published openly.

Consolidated ranking6 benchmarks

WeirdBench Intelligence Index

A single ranking across every WeirdBench benchmark. Raw scores are normalized relative to each benchmark leader, then adjusted for coverage — so higher-is-better and lower-is-better benchmarks share one honest table.

Benchmarks

Full index

Regex Golf

Generate the shortest valid regular expression matching all 10 target strings while excluding all 10 distractor strings across 25 deterministically generated puzzles. Lower is better.

Leaderboard preview

1openai/gpt-6-astra5.360
2google/gemini-3.8-flash5.640
3openai/gpt-5.6-sol5.800
View benchmark

AI Writing Detection

Classify essays from a fixed balanced sample of 50 human-written and 50 AI-generated examples from the AI Generated Essays Dataset. Higher is better.

Leaderboard preview

1anthropic/claude-fable-5.11.000
2anthropic/claude-opus-4.11.000
3anthropic/claude-opus-4.71.000
View benchmark

Nutrition Prediction

Predict calories, protein, carbs, and fat from ingredient lists for a fixed 50-dish Nutrition5k sample. Higher is better.

Leaderboard preview

1qwen/qwen3.8-27b25.767
2anthropic/claude-opus-523.853
3mistralai/mistral-small-260323.823
View benchmark

Semantic Diversity

Generate exactly 20 English words that are maximally semantically unrelated to each other, then score the average pairwise semantic similarity. Lower is better.

Leaderboard preview

1google/gemini-3.7-flash0.214
2google/gemini-3.8-flash0.215
3anthropic/claude-opus-4.60.216
View benchmark

Orthographic Diversity

Search for 20 real English words that are maximally different in spelling under hard validity rules and deterministic penalties. Higher is better.

Leaderboard preview

1inclusionai/ling-3.0-flash6.067
2openai/gpt-oss-20b5.805
3openai/gpt-5.6-luna5.474
View benchmark

Wordle

Play 20 recent Wordle answers turn by turn with standard gray/yellow/green feedback. Invalid guesses still cost a turn, scores are capped at 10 turns per puzzle, and lower is better.

Leaderboard preview

1openai/gpt-5.3-codex3.600
2anthropic/claude-opus-53.650
3openai/gpt-6-astra3.700
View benchmark