Repo Bench
Measuring large context reasoning, file editing precision, and instruction adherence.
Rankings updated every 15 min • Median scores from all model runs submitted by Repo Prompt users
Last updated: 4/30/2026, 5:49:41 PM
Building Repo Bench: The Story Behind the Benchmark Read the full story → Explore Individual Runs Drill into every benchmark run, compare distributions, and slice by provider. Open Repo Bench Explorer →
Ranking mode:
Normalized (Median)Max Score
#1
Claude Opus 4.1Anthropic
6 runs
77.35%67.65–77.35%
#2
GPT-5.4 LowOpenAI
76.47%
#3
Kimi 2.6Moonshot
3 runs
75.59%60.59–75.59%
#4
Claude Sonnet 4.5Anthropic
53 runs
75%67.35–83.82%
#5
Claude Opus 4.5Anthropic
48 runs
73.53%67.65–83.82%
#6
Z.AI GLM-5Z.AI
15 runs
73.53%69.71–80%
#7
GPT-5.1 MediumOpenAI
7 runs
72.35%64.71–80%
#8
GPT-5.3 Codex MediumOpenAI
6 runs
71.47%67.35–71.47%
#9
GPT-5.5 MediumOpenAI
2 runs
71.18%70–71.18%
#10
GPT-5.5 HighOpenAI
2 runs
70.88%70–70.88%
#11
GPT-5.5 XhighOpenAI
2 runs
70.88%70–70.88%
#12
volcengine/doubao-seed-2.0-cod...OpenRouter
70.59%
#13
GPT-5.2 XhighOpenAI
14 runs
70%62.94–71.47%
#14
GPT-5.2 HighOpenAI
7 runs
70%62.65–71.47%
#15
Z.AI GLM-5-TurboZ.AI
7 runs
70%62.65–71.18%
#16
GPT-5.2 Codex XhighOpenAI
3 runs
70%62.65–70.29%
#17
GPT-5.3 Codex HighOpenAI
3 runs
70%65.88–70%
#18
GPT-5.4OpenAI
3 runs
70%70–70%
#19
GPT-5.3 Codex XhighOpenAI
2 runs
70%70–70%
#20
Gemini Pro 3.1 PreviewGemini
2 runs
70%68.82–70%
#21
GPT-5.2 MedOpenAI
2 runs
70%70–70%
#22
GPT-5.4 ProOpenAI
2 runs
70%68.82–70%
#23
GPT-5.2 ProOpenAI
70%
#24
GPT-5.2 LowOpenAI n 13 runs
69.71%62.94–71.47%
#25
GPT-5.1 LowOpenAI
19 runs
69.41%63.24–78.53%
227 additional models tested but not shown in top 25
* This benchmark does not reflect the coding skills of a model.
About Repo Bench
Large Context Reasoning
Tests the model's ability to maintain understanding and reasoning across extensive codebases with thousands of lines. Evaluates how well models can hold context and make logical connections across distant parts of large files.
Instruction Adherence
Evaluates how precisely models follow detailed system instructions for formatting, output structure, and task constraints. Measures consistency and accuracy in adhering to complex, multi-step instructions.
File Editing Precision
Assesses the model's ability to make precise edits to code files while maintaining context and requirements throughout task completion across different programming languages (Go, Swift, TypeScript) and various formatting constraints.
Run the Benchmark Yourself
The above results were derived from running the official Repo Prompt benchmark within the app. Run it yourself to see how well a model does with Repo Prompt's constraints, and compare results over time.
If you're a lab or model provider and would like to get in touch, feel free to reach out to contact@repoprompt.com