Skip to content
Pl. 11The engagement

Which model should we standardise on — proven on our code?

The harness is MIT and yours to run. When you would rather have the answer than the instrument, this is the engagement that produces it.

Not a subscription

Priced against the decision, not the compute.

An annual AI tooling budget is a six-figure line item chosen, almost everywhere, on marketing plus a feeling. This engagement replaces the feeling with a measurement taken on the only corpus that predicts your outcomes: your own repository.

It is deliberately a consulting engagement rather than a seat licence. One engagement at a time, run end to end, with the harness you can keep running yourself afterwards — because the harness is MIT and the method is published.

Later, if the pull is there: scheduled re-runs as models ship. “Did the new release actually get better on our code?” is the same question, asked on a cadence.

How it runs

Four stages. Your code never leaves.

  1. 01

    Scope

    Which repositories, which candidate configs, and — the part that matters — which decision the result has to settle. A model standardisation, a spend tier, an agent harness, a reasoning-effort default.

    A written scope and a candidate matrix.

  2. 02

    Mine & gate

    Guignet runs on your infrastructure. Your history becomes candidate tasks; the validity gate replays each one and discards everything that doesn’t reproduce. You see the soundness rate before a single agent runs.

    An admitted suite, with per-candidate discard reasons.

  3. 03

    Run & score

    N attempts per task per config, in disposable worktrees, with cost parsed from each harness’s own transcript. Contamination controls applied and reported, never quietly.

    Verdicts, costs, confidence intervals, flag rates.

  4. 04

    Readout

    The leaderboard, the recommendation, the tuned configs — and an explicit statement of what would change the answer. A number you can’t act on isn’t a deliverable.

    The report, the configs, and a live walkthrough.

Specimen — what you receive

One self-contained file, and its JSON twin.

Renders offline. No framework, no CDN. Regenerates from stored runs without re-executing a single attempt.

01Leaderboard
Every candidate config ranked on your admitted suite, with 95% Wilson confidence intervals rendered on every solve rate.
02$ per solved task
The executive number. Cost parsed from harness transcripts, never self-reported, with partial coverage marked as a lower bound.
03Taxonomy breakdowns
By type (bugfix / feature / refactor), by size bucket, and by path-based area tag — “wins on backend bugfixes, loses on UI features” is the actionable shape.
04Cutoff-split columns
Scores split by each model’s training cutoff, framed by your repository’s visibility, with regurgitation flag rates per config.
05Tuned configs
The run configs that produced the result, ready to keep — plus the prompt and effort settings that moved the numbers.
06The JSON twin
Every number in the HTML, also as JSON. Auditable, and yours to re-run whenever you like.

Every figure in the report is explained in the published methodology, including the limitations. If a number isn’t explained there, treat it as unexplained.

Start

Tell me the decision you’re trying to settle.

A short note about the repositories, the configs you are weighing, and the deadline is enough to know whether this is worth either of our time. If your history won’t support a sound suite, I will say so.