Repo Bench

Measuring large context reasoning, file editing precision, and instruction adherence.

Rankings updated every 15 min • Median scores from all model runs submitted by Repo Prompt users

Last updated: 4/30/2026, 5:49:41 PM

Building Repo Bench: The Story Behind the Benchmark Read the full story → Explore Individual Runs Drill into every benchmark run, compare distributions, and slice by provider. Open Repo Bench Explorer →

Ranking mode:

Normalized (Median)Max Score

#1

Claude Opus 4.1Anthropic

6 runs

77.35%67.65–77.35%

#2

GPT-5.4 LowOpenAI

76.47%

#3

Kimi 2.6Moonshot

3 runs

75.59%60.59–75.59%

#4

Claude Sonnet 4.5Anthropic

53 runs

75%67.35–83.82%

#5

Claude Opus 4.5Anthropic

48 runs

73.53%67.65–83.82%

#6

Z.AI GLM-5Z.AI

15 runs

73.53%69.71–80%

#7

GPT-5.1 MediumOpenAI

7 runs

72.35%64.71–80%

#8

GPT-5.3 Codex MediumOpenAI

6 runs

71.47%67.35–71.47%

#9

GPT-5.5 MediumOpenAI

2 runs

71.18%70–71.18%

#10

GPT-5.5 HighOpenAI

2 runs

70.88%70–70.88%

#11

GPT-5.5 XhighOpenAI

2 runs

70.88%70–70.88%

#12

volcengine/doubao-seed-2.0-cod...OpenRouter

70.59%

#13

GPT-5.2 XhighOpenAI

14 runs

70%62.94–71.47%

#14

GPT-5.2 HighOpenAI

7 runs

70%62.65–71.47%

#15

Z.AI GLM-5-TurboZ.AI

7 runs

70%62.65–71.18%

#16

GPT-5.2 Codex XhighOpenAI

3 runs

70%62.65–70.29%

#17

GPT-5.3 Codex HighOpenAI

3 runs

70%65.88–70%

#18

GPT-5.4OpenAI

3 runs

70%70–70%

#19

GPT-5.3 Codex XhighOpenAI

2 runs

70%70–70%

#20

Gemini Pro 3.1 PreviewGemini

2 runs

70%68.82–70%

#21

GPT-5.2 MedOpenAI

2 runs

70%70–70%

#22

GPT-5.4 ProOpenAI

2 runs

70%68.82–70%

#23

GPT-5.2 ProOpenAI

70%

#24

GPT-5.2 LowOpenAI n 13 runs

69.71%62.94–71.47%

#25

GPT-5.1 LowOpenAI

19 runs

69.41%63.24–78.53%

227 additional models tested but not shown in top 25

* This benchmark does not reflect the coding skills of a model.

About Repo Bench

Large Context Reasoning

Tests the model's ability to maintain understanding and reasoning across extensive codebases with thousands of lines. Evaluates how well models can hold context and make logical connections across distant parts of large files.

Instruction Adherence

Evaluates how precisely models follow detailed system instructions for formatting, output structure, and task constraints. Measures consistency and accuracy in adhering to complex, multi-step instructions.

File Editing Precision

Assesses the model's ability to make precise edits to code files while maintaining context and requirements throughout task completion across different programming languages (Go, Swift, TypeScript) and various formatting constraints.

Run the Benchmark Yourself

The above results were derived from running the official Repo Prompt benchmark within the app. Run it yourself to see how well a model does with Repo Prompt's constraints, and compare results over time.

If you're a lab or model provider and would like to get in touch, feel free to reach out to contact@repoprompt.com