WeirdBench
Unconventional LLM benchmarks for the weird corners other evals skip. Definitions and runners live locally, scores are published openly.
WeirdBench Intelligence Index
A single ranking across every WeirdBench benchmark. Raw scores are normalized relative to each benchmark leader, then adjusted for coverage — so higher-is-better and lower-is-better benchmarks share one honest table.
Benchmarks
Full indexRegex Golf
Generate the shortest valid regular expression matching all 10 target strings while excluding all 10 distractor strings across 25 deterministically generated puzzles. Lower is better.
Leaderboard preview
AI Writing Detection
Classify essays from a fixed balanced sample of 50 human-written and 50 AI-generated examples from the AI Generated Essays Dataset. Higher is better.
Leaderboard preview
Nutrition Prediction
Predict calories, protein, carbs, and fat from ingredient lists for a fixed 50-dish Nutrition5k sample. Higher is better.
Leaderboard preview
Semantic Diversity
Generate exactly 20 English words that are maximally semantically unrelated to each other, then score the average pairwise semantic similarity. Lower is better.
Leaderboard preview
Orthographic Diversity
Search for 20 real English words that are maximally different in spelling under hard validity rules and deterministic penalties. Higher is better.
Leaderboard preview
Wordle
Play 20 recent Wordle answers turn by turn with standard gray/yellow/green feedback. Invalid guesses still cost a turn, scores are capped at 10 turns per puzzle, and lower is better.
Leaderboard preview