AI Writing Detection
Classify essays from a fixed balanced sample of 50 human-written and 50 AI-generated examples from the AI Generated Essays Dataset. Higher is better.
Leaderboard
Intelligence Index- 11.0000
anthropic/claude-fable-5.1
- 21.0000
anthropic/claude-opus-4.1
- 31.0000
anthropic/claude-opus-4.7
- 41.0000
deepseek/deepseek-v4-pro-0813
- 51.0000
google/gemini-3-flash-preview
- 61.0000
google/gemini-3.5-flash-lite
- 71.0000
openai/gpt-5.5
- 81.0000
openai/gpt-5.6-sol
- 91.0000
openai/gpt-5.6-terra
- 101.0000
openai/gpt-6-astra
- 110.9901
openai/gpt-5.6-luna
- 120.9899
anthropic/claude-fable-5
- 130.9899
anthropic/claude-sonnet-4.5
- 140.9899
google/gemini-3.7-flash
- 150.9899
google/gemini-3.8-flash
- 160.9899
google/gemma-4-26b-a4b-it
- 170.9899
google/gemma-4-31b-it
- 180.9899
meta/muse-spark-1.2
- 190.9899
moonshotai/kimi-k2.5
- 200.9899
moonshotai/kimi-k3
- 210.9899
qwen/qwen3.5-122b-a10b
- 220.9899
qwen/qwen3.5-397b-a17b
- 230.9899
qwen/qwen3.8-2.4t-a95b
- 240.9899
qwen/qwen3.8-flash
- 250.9899
thinkingmachines/inkling
- 260.9899
thinkingmachines/inkling-small
- 270.9899
x-ai/grok-4.5
- 280.9899
x-ai/grok-4.6
- 290.9899
z-ai/glm-5.2
- 300.9899
z-ai/glm-5.3-flash
- 310.9804
openai/gpt-5.1
- 320.9804
x-ai/grok-4.1-fast
- 330.9800
anthropic/claude-opus-4.6
- 340.9691
minimax/minimax-m3
- 350.9608
z-ai/glm-5.3
- 360.9592
anthropic/claude-sonnet-4.6
- 370.9278
anthropic/claude-opus-4.5
- 380.9259
meta/muse-glimmer-30b
- 390.9009
moonshotai/kimi-k2.6
- 400.8958
nvidia/nemotron-3-ultra-550b-a55b
- 410.8932
anthropic/claude-sonnet-5
- 420.8929
google/gemini-3.1-flash-lite-preview
- 430.8850
anthropic/claude-opus-5
- 440.8750
inception/mercury-2
- 450.8636
anthropic/claude-opus-4.8
- 460.8621
openai/gpt-5.4
- 470.8454
anthropic/claude-haiku-4.5
- 480.8421
openai/gpt-oss-120b
- 490.8130
upstage/solar-pro4
- 500.8000
minimax/minimax-m2.5
- 510.7321
meta-llama/llama-4-maverick
- 520.6849
mistralai/mistral-small-2603
- 530.6803
openai/gpt-5.4-mini
- 540.6757
upstage/solar-pro-3
- 550.6667
mistralai/mistral-large-2512
- 560.6667
x-ai/grok-4.20-beta
- 570.6438
amazon/nova-micro-v1
- 580.6087
mistralai/mistral-medium-3.1
- 590.4819
deepseek/deepseek-v3.2
- 600.4211
amazon/nova-2-lite-v1
- 610.0357
amazon/nova-pro-v1
- 620.0241
meta-llama/llama-4-scout
- 630.0000
amazon/nova-lite-v1
Prompt
Each essay is shown once and the model must return exactly one character: "1" for AI-generated or "0" for human-written, with no explanation.
Score
Higher is better. The benchmark score is the F1 score for detecting AI-generated essays, using label 1 as the positive class.
Execution
Benchmark runners execute locally, read the dataset from disk, use OpenRouter for predictions with reasoning disabled, cache results in Neon by benchmark and model ID, and skip recomputation for models that already have stored scores.