+26.38 percentage points from the first to the third iteration.
Leaderboard
A transparent view of executable cross-architecture build repair. The public competition leaderboard is planned for September 2026.
These values are reproduced from Table 2 of the Build-Bench paper. Models were evaluated once per package under a fixed configuration with up to three repair iterations and 20 tool calls per iteration, using OBS for executable validation.
Main experimental configuration
Filter by migration direction. Rank is recalculated within the selected view.
| Rank | Model | Direction | Success | Success rate | Avg. time | Avg. tokens |
|---|---|---|---|---|---|---|
| 1 | GPT-5 | x86_64 → aarch64 | 103 / 163 | 63.19% | 31.18 min | 1,830.91K |
| 2 | GPT-5-mini | x86_64 → aarch64 | 47 / 163 | 28.83% | 13.80 min | 1,683.95K |
| 3 | Qwen3-max | x86_64 → aarch64 | 28 / 163 | 17.18% | 35.69 min | 505.39K |
| 4 | GPT-4o | x86_64 → aarch64 | 22 / 163 | 13.50% | 5.93 min | 541.66K |
| 5 | Claude Sonnet 4.5 | x86_64 → aarch64 | 16 / 163 | 9.82% | 6.27 min | 328.76K |
| 6 | DeepSeek V3 | x86_64 → aarch64 | 13 / 163 | 7.98% | 11.37 min | 235.53K |
| 7 | Qwen2.5-3B-Instruct | x86_64 → aarch64 | 8 / 163 | 4.91% | 27.65 min | 591.90K |
| 1 | GPT-5 | aarch64 → x86_64 | 31 / 105 | 29.52% | 18.55 min | 1,518.66K |
| 2 | GPT-5-mini | aarch64 → x86_64 | 28 / 105 | 26.67% | 14.37 min | 1,894.60K |
| 3 | GPT-4o | aarch64 → x86_64 | 13 / 105 | 12.38% | 5.82 min | 614.12K |
| 4 | Qwen3-max | aarch64 → x86_64 | 6 / 105 | 5.71% | 27.46 min | 359.16K |
| 5 | Claude Sonnet 4.5 | aarch64 → x86_64 | 6 / 105 | 5.71% | 4.52 min | 332.99K |
| 6 | DeepSeek V3 | aarch64 → x86_64 | 4 / 105 | 3.81% | 19.27 min | 445.03K |
| 7 | Qwen2.5-3B-Instruct | aarch64 → x86_64 | 2 / 105 | 1.90% | 24.33 min | 373.81K |
Success is the number of packages successfully built on the target architecture.
Average time and average tokens follow the definitions reported in the paper and are not proposed competition tie-breakers except where explicitly stated in the official rules.
The paper evaluates each model once per package under a fixed configuration. API and backend evolution may affect exact replication.
Build feedback changes the outcome
The research benchmark allows updated build logs and prior repair output to inform later attempts. The paper reports substantial cumulative gains across three iterations.
+18.10 percentage points from the first to the third iteration.
Public validation is planned for September 2026
Competition entries, submission limits, evaluator version, and final columns will be announced with the starter kit and public platform.