DownloadsDocsWikiWhyFeaturesPricing
Smart Harness

Don't ask the model twice.
Ask for proof once.

Smart Harness catches candidates that look finished but fail your checks, gives the model a bounded chance to recover, and applies the verified winner. You see every attempt, every token, and the final verdict.

Opt-in · bounded policies · fail-closed outcomes · complete attempt metering
01

candidate generated

ready
02

verifier: 1 test failed

retry
03

bounded recovery attempt

metered
04

7 / 7 tests passed

verified
Verified rescueWinner applied once · receipt attached
What it does

A small loop with a hard boundary

The public contract is simple. Your checks—not the model's confidence—decide when the task is done.

01

One task

You describe the change and the proof you expect.

02

Bounded attempts

The selected policy controls how far recovery can go.

03

Your verifier

Tests or checks decide whether a candidate is acceptable.

04

Truthful outcome

A verified winner is applied once—or the result says verification failed.

No secret sauce diagram. Just the customer-visible contract: bounded work, independent checks, and an honest result.

Same model. Same task.

The difference is what happens after the first answer

Step through a realistic boundary-bug repair. The left side keeps going when the verifier finds evidence; the right side hands that loop back to you.

Interactive task replayStep 1 of 5
Bounded verification loop

Smart Harness

Task accepted

Fix the inclusive-date billing bug

Same repository, same model, same failing test, and the same definition of done.

1 failing · 6 passing
Same model · no harness

Single-pass agent

Task accepted

Fix the inclusive-date billing bug

The single-pass agent starts from the identical prompt and repository state.

1 failing · 6 passing
Controlled benchmark evidence

More tasks ended verified

On a controlled 1,140-task BigCodeBench run using production infrastructure—not organic customer traffic.

Single pass55.00%

627 of 1140 passed

Smart Harness ladder71.67%

817 of 1140 passed

Measured lift+16.67 pp

190 more verified completions

Billing truth0

billing-invariant violations

Controlled benchmark result. It does not predict the lift on every repository, model, policy, or task.

Real product evidence

The result comes with receipts

These are direct captures from a real Verdict Desktop code-task completion. They show the finished workspace and final response receipt—not a Smart Harness A/B screenshot.

Verdict Desktop with the coding task and seven passing tests visible
01 · Finished workspaceThe bound repository, selected harness, green tests, and final answer in the packaged Desktop app.
Verdict Code terminal receipt with seven passing tests and usage metrics
02 · Proof detail7/7 tests, tokens, throughput, model-call time, and estimated local energy stay visible.
Read the observed-run case study
Honest accounting

Quality lift is not free compute

A failed first candidate can trigger more calls. Smart Harness makes that trade visible instead of hiding it behind a success animation.

Retries are bounded

Your selected policy caps the number and kind of additional attempts.

Every attempt is metered

Tokens, wall time, and the relevant cost or energy signal roll into the outcome.

No blanket savings claim

The proven benefit is more verified completions and less manual retry choreography—not universal token or time savings.

Questions, answered

The short version

Is Smart Harness another model?

No. It is a verification and recovery layer around the model and task you choose.

Does it retry forever?

No. Every policy has a bounded attempt budget. You can choose a faster or more thorough policy, and the outcome reports what actually ran.

Does it always save tokens or time?

No. A failed candidate can trigger extra attempts. The gain is a higher rate of verified completions and less manual retry choreography; every attempt is metered.

Where is the full engine available?

The repository-aware engine runs in Verdict Code and in Verdict Desktop portable coding tasks. Other surfaces may expose catalog controls without running the full code-task engine.

Stop shipping the first plausible answer

Give the model a finish line it can prove.

Start locally for $0. Turn on Smart Harness when the task deserves a verified result.