Reduit

An agent harness for the places the frontier models are not allowed to run.

Open weights, self-hosted, network access restricted to an allowlist of package mirrors. Not “air-gapped” — the benchmark’s own verification scripts fetch uv and pytest at runtime, and saying otherwise would be false. The constraint is not artificial: a bank, a hospital, an insurer or a cantonal administration already sits inside that box by law.

What we measured

78.3 % ± 7.6

GLM-5.2 off the shelf on Terminal-Bench 2.1 — 47 of 60 runs, 20 tasks stratified out of the 89, k=3, Terminus 2 unmodified, time limits untouched. Cost: $17.81. 95 % interval, clustered at task level: 63.3 to 91.7. Post-stratified to all 89 tasks: 82.1 %.

Z.ai claims 81.0 for this model and has never submitted it. Our interval contains that claim, so it is neither refuted nor confirmed — but the order of magnitude holds, and it sits far above the only verified open-weights entry on the board, which is 58.7 %.

The interesting part is not the number. It is that the open-weights slot at the top of this leaderboard is empty because nobody has submitted, not because it is hard. A full submission is 445 runs and costs a measured $132.

Why this is a selection problem

Across all 20 submissions on Terminal-Bench 2.1points
median gap between pass@5 and the hit rate15.0
range, without a single exception9.4 – 21.8
what pure harness ergonomics buys, measured on this benchmark0.2 – 8.1

The models already solve these tasks. They just do not hit reliably. Both numbers were recomputed from the raw submission JSON in the public leaderboard repository, not read off the website.

What did not work

Four of seven experiments refuted the premise they were built on, and they are published in full, limitations included. That is not a disclaimer, it is the point:

ExperimentWhat we expectedWhat happened
01 & 02an acceptance oracle can rank candidate solutionswrong in 2 of 5 measured cases, and the error is anticorrelated with correctness — the layer was dropped
06batching commands cuts turns and tokens by 40 %the model batched, and turns went up 1.19× and tokens up 1.75×
07a time budget rescues runs killed mid-stridein 17 runs the clock bound zero times — the failure mode was not there
the price listone published price describes what we paywrong three times running; five providers served one model on one key inside seven minutes at up to 2.002× each other’s price

How the numbers are kept honest

Where it stands

Submission pending. No entry on the leaderboard yet. The number above is our own measurement of the baseline model, not of this harness, and until a verified result exists this page will not claim one. A self-measured number sitting where a verified one belongs is exactly the practice this project criticises in others, so it does not happen here.

The commercial part, in one paragraph

The harness is and stays open source, under AGPL-3.0-or-later, with a commercial licence alongside for the cases where that does not fit. What is paid for is operation, adaptation and responsibility inside a regulated environment — not a binary. Everything measured stays published, because the whole standing of this project rests on being checkable.


Design partners

If you run agents where the data cannot leave the building — a bank, a hospital, an insurer, a public administration — and you would rather look at real numbers than a deck, write one line. No newsletter, no tracker, no scheduling link.

The form posts to this site and nowhere else. The page itself loads no external resource of any kind — no font, no script, no analytics, no cookie. Your message is stored in a database in Zurich and read by one person. Nothing else happens with it.