Leaderboard preview Competition submissions are not open. Results below come from the Build-Bench paper.

Leaderboard

A transparent view of executable cross-architecture build repair. The public competition leaderboard is planned for September 2026.

Research baseline reference, not competition entries

These values are reproduced from Table 2 of the Build-Bench paper. Models were evaluated once per package under a fixed configuration with up to three repair iterations and 20 tool calls per iteration, using OBS for executable validation.

Read paper
268packages in view
7models in view
63.19%best success rate
Buildexecutable oracle

Main experimental configuration

Filter by migration direction. Rank is recalculated within the selected view.

Rank Model Direction Success Success rate Avg. time Avg. tokens
1GPT-5x86_64 → aarch64103 / 16363.19%31.18 min1,830.91K
2GPT-5-minix86_64 → aarch6447 / 16328.83%13.80 min1,683.95K
3Qwen3-maxx86_64 → aarch6428 / 16317.18%35.69 min505.39K
4GPT-4ox86_64 → aarch6422 / 16313.50%5.93 min541.66K
5Claude Sonnet 4.5x86_64 → aarch6416 / 1639.82%6.27 min328.76K
6DeepSeek V3x86_64 → aarch6413 / 1637.98%11.37 min235.53K
7Qwen2.5-3B-Instructx86_64 → aarch648 / 1634.91%27.65 min591.90K
1GPT-5aarch64 → x86_6431 / 10529.52%18.55 min1,518.66K
2GPT-5-miniaarch64 → x86_6428 / 10526.67%14.37 min1,894.60K
3GPT-4oaarch64 → x86_6413 / 10512.38%5.82 min614.12K
4Qwen3-maxaarch64 → x86_646 / 1055.71%27.46 min359.16K
5Claude Sonnet 4.5aarch64 → x86_646 / 1055.71%4.52 min332.99K
6DeepSeek V3aarch64 → x86_644 / 1053.81%19.27 min445.03K
7Qwen2.5-3B-Instructaarch64 → x86_642 / 1051.90%24.33 min373.81K

Success is the number of packages successfully built on the target architecture.

Average time and average tokens follow the definitions reported in the paper and are not proposed competition tie-breakers except where explicitly stated in the official rules.

The paper evaluates each model once per package under a fixed configuration. API and backend evolution may affect exact replication.

Build feedback changes the outcome

The research benchmark allows updated build logs and prior repair output to inform later attempts. The paper reports substantial cumulative gains across three iterations.

x86_64 → aarch64 · GPT-5
Iteration 136.81%
Iteration 248.47%
Iteration 363.19%

+26.38 percentage points from the first to the third iteration.

aarch64 → x86_64 · GPT-5
Iteration 111.43%
Iteration 216.19%
Iteration 329.52%

+18.10 percentage points from the first to the third iteration.

Public validation is planned for September 2026

Competition entries, submission limits, evaluator version, and final columns will be announced with the starter kit and public platform.

Read the Agent contract Explore the research code