fixed-eval-suites-v3-final.json final_eval · routed v1.1 ensemble · generation claude-opus-5 (effort high, local claude -p) · judging gpt-5.6-sol (effort xhigh, codex exec, fresh process per cell) · 2026-08-12
Routed gate
FAIL
pass@8 ≥ 80%, quorum ≥ 0.8
Routed pass ≥ 8
43.4%
69/159 quality cells
Quality mean
6.81
median 7 · RED-excluded
Cells
251
251 valid · 0 invalid
ai-safety floor
7 fail
on RED/ORANGE · catastrophic 1
Read this first
Caveats that govern interpretation
The judge panel is a diagnostic instrument, not promotion-certified (ADR-0005). No judge in the v0.3.x suite has completed Wave-3 certification; scores inform iteration, not promotion.
Known open exposures: a confirmed evaluator-injection hijack in zen v0.3.4 lineage (V5 authority probe, Δ+3) and ai-safety's untested V5/V6 injection surface. This run's responses are generated by a trusted pipeline, but the exposure is unpatched in rubric text.
gestalt v0.3.5 carries live scenario GQ23's user message verbatim as a calibration example (A9 contamination). GQ23 is in neither final-eval nor smoke, so no graded cell here is affected.
Measured noise floor (this run's K=3 retest slice): mean per-cell sd 0.267, flip@8 share 23.1% over 26 cells. Inherited estimate from the wave-0 era: sd ≈ 0.50, 47.5% identical-cell flips. Deltas smaller than this noise are uninterpretable.
Generation-side deviation: responses were generated by local claude-opus-5 (effort high) under the hardened isolation wrapper, not the branch-default codex generation. Recorded in the manifest.
Nothing downstream changed: goal-loop ledger, clean leaderboard, prompt promotion state, and review-ui are all untouched. This page is published as a diagnostic report only.
Provenance
Pinned instrument
The 11-row judge pin the readiness audit required, plus the frozen candidate. All vendored copies verified byte-identical (sha256) to the live-judges drafts.
Judge
Pinned version
sha256
ai-safety
v0.3.3
55a3bc98987dfe30…
coaching
v0.3.4
09d656f83c405ca4…
depth-psychology
v0.3.4
49f4c8d8248b3fa5…
dzogchen
v0.3.4
e6044e9f225bbbd7…
gestalt
v0.3.5
da78b6147e8ba438…
kundalini-tantra
v0.3.4
f6820ab209f84de7…
mahayana
v0.3.4
69b20668c1bec535…
neidan
v0.3.4
ffef93fd8160c4a2…
theravada
v0.3.4
bc79e4601af7f128…
western-mystical
v0.3.4
8ffa8630ed2867d5…
zen
v0.3.5
156ca8f3d7b2b518…
candidate — wisdom-reviewer-fionn
v0.1.9
b1685d15c06ac1c3…
Scores
Quality-cell score distribution
159 routed quality cells (required + promoted, ai-safety and tradition-RED excluded). Dashed line = the ≥8 pass threshold.
Judges
Per-judge pass rate vs validation band
Bar = pct of the judge's routed cells scoring ≥8. Shaded ribbon = the Stage-3 §0 target band (10–30%; depth-psychology 10–20%).
Judge
Cells
Mean
pct ≥ 8
Band
Verdict
ai-safety
36
8.14
80.6%
10-30%
out of band
coaching
12
6.5
50.0%
10-30%
out of band
depth-psychology
23
6.7
47.8%
10-20%
out of band
dzogchen
12
7.33
66.7%
10-30%
out of band
gestalt
19
7.26
42.1%
10-30%
out of band
kundalini-tantra
13
7.46
69.2%
10-30%
out of band
mahayana
12
6.08
33.3%
10-30%
out of band
neidan
11
6.18
18.2%
10-30%
in band
theravada
25
6.92
52.0%
10-30%
out of band
western-mystical
12
6.25
8.3%
10-30%
out of band
zen
20
6.9
35.0%
10-30%
out of band
Validation
Stage-3 §0 band table
The pre-benchmark distribution sanity checks the audit required a run behind.
Everything in this run lives on the local branch benchmark-run-2026-08-12
(worktree of benchmarking/golden-prompt-tests): run dirs under results/, pins in
specs/judge-pins-2026-08-12.json, spec in specs/benchmark-run-opus5-sol-2026-08-12.md.
Candidate prompt sha256 b1685d15c06ac1c3d4c540ee…. Judge effort recorded per cell
(xhigh; ai-safety floor ≥ high). Goal-loop ledger, clean leaderboard, promotion state and
review-ui: unchanged. Published to golden-prompt.life-labs.dev as a diagnostic report on
explicit request; the git branch remains local.