Validation Lab — Dottie harness & training progress

Read-only scorecard for the Dottie orchestration stack: the live harness API and the distilled router model trained from run traces. Every number on this page is recomputed from committed event logs and reports — history, not live telemetry — and unmeasured quantities render as unmeasured.

Champion router model — orch-mlp-v1-v4

96.7% held-out tier accuracy (151 validation records)
NOT PROMOTED — champion measured accuracy 0.819672 does not strictly beat both baselines (freq prior 0.245902, heuristic 0.836066) on n=61 measured held-out records

Promotion is eval-gated and fail-closed: a champion ships to routing only after the gate passes on measured data. Until then it serves in an advisory role; the badge carries the gate's exact reason.

Hill-climb ladder — 8 variants, real CPU training runs

25%50%75%100%v1v1: 84.1% held-out tier accuracy — buckets 4096, embed 16, lr 0.001, 7.55s trainv2v2: 84.8% held-out tier accuracy — buckets 4096, embed 16, lr 0.003, 1.81s trainv3v3: 81.5% held-out tier accuracy — buckets 4096, embed 32, lr 0.001, 2.10s trainv4v4: 96.7% held-out tier accuracy — buckets 8192, embed 16, lr 0.001, 1.71s train96.7% — championv5v5: 84.8% held-out tier accuracy — buckets 4096, embed 16, lr 0.001, 1.50s trainv6v6: 90.7% held-out tier accuracy — buckets 4096, embed 16, lr 0.001, 1.52s trainv7v7: 81.5% held-out tier accuracy — buckets 4096, embed 16, lr 0.001, 1.63s trainv8v8: 81.5% held-out tier accuracy — buckets 8192, embed 32, lr 0.003, 1.89s train

Advantage-weighted training on the orchestration corpus (seed-pinned, ≤60s budget per variant; actual 1.5–7.6s). Bars share one hue — one measure across variants; hover any bar for its config.

Promotion gate — fail-closed

Measured held-out coverage (n ≥ 10)✓ PASS — 61 of 10 required measured held-out records
Beats frequency-prior baseline on measured hold-out✓ PASS — champion 82.0% vs baseline 24.6%
Beats heuristic router on measured hold-out✕ FAIL — champion 82.0% vs baseline 83.6%

Unmeasured renders as unmeasured, never as pass. The gate unlocks as real harness runs accumulate measured hold-out records.

Validation queue

MethodSiteStatus
Barlow Twinsresearchawaiting pilot
InfoNCE domain tuneresearchin execution

Candidates advance pilot → charter → deploy; verdicts publish here with their evidence references.

Training corpus — 1613 records by provenance

measured: 774 records mined from real run timelines and workflow journalssimulated: 839 records from the seeded synthetic battery
measured · 774simulated · 839

Provenance is labeled per record. The measured slice grows with every harness run; simulated records come from a seeded, deterministic battery (seed 20260809).

Corpus by routing tier

TierRecords
agentic_epic490
deep_research451
action_operator309
deterministic253
llm110

Label sources

Label sourceRecords
simulated820
measured-behavior752
measured-outcome41
✓ 10 measured hold-out records carry labels beyond measured-behavior and simulated — the gate's champion-vs-heuristic comparison is grounded in labels the heuristic did not produce

Each record's label provenance is tracked separately from its feature provenance. The promotion gate's champion-vs-heuristic comparison only becomes meaningful when the measured hold-out contains labels the heuristic did not make (measured-outcome or operator-corrected).

Correction queue — operator review

RunLabelSourceStatusCorrection command
20260813T173943611486Zaction_operatormeasured-behaviorokscout harness correct harness-run-20260813T173943611486Z <tier> --reason "…"
20260813T173940773648Zaction_operatormeasured-behaviorokscout harness correct harness-run-20260813T173940773648Z <tier> --reason "…"
20260813T173936967593Zagentic_epicmeasured-outcomefailscout harness correct harness-run-20260813T173936967593Z <tier> --reason "…"
20260813T173935751878Zaction_operatormeasured-behaviorokscout harness correct harness-run-20260813T173935751878Z <tier> --reason "…"
20260813T173934555198Zagentic_epicmeasured-outcomefailscout harness correct harness-run-20260813T173934555198Z <tier> --reason "…"
20260813T173932399458Zaction_operatormeasured-behaviorokscout harness correct harness-run-20260813T173932399458Z <tier> --reason "…"
20260810T121725288196Zagentic_epicmeasured-outcomefailscout harness correct harness-run-20260810T121725288196Z <tier> --reason "…"
20260810T121724195455Zagentic_epicmeasured-outcomefailscout harness correct harness-run-20260810T121724195455Z <tier> --reason "…"
20260810T121723037562Zagentic_epicmeasured-outcomefailscout harness correct harness-run-20260810T121723037562Z <tier> --reason "…"
20260810T121721866345Zagentic_epicmeasured-outcomefailscout harness correct harness-run-20260810T121721866345Z <tier> --reason "…"

The ten most recent measured harness runs from the committed corpus. A correction is ground truth: run the command with the tier that should have been routed (goal text via scout harness checkpoint show --run-id …; the corpus itself carries metadata only). Corrections land in label_corrections.jsonl and the nightly retrain consumes them as operator-corrected labels — the strongest signal against the gate's label ceiling.

Run scoreboard — mined from committed timelines

AgentRunsEventsOK ratep50 latency
builder122122100%0.02096000002893561 ms
critic122122100%0.10784499994542784 ms
executor124124100%0.04063200009341017 ms
mcp-operator252568%120.91145499994127 ms
operator3162100%0.013480000006893533 ms
planner124124100%0.024464000034640776 ms
researcher33100%30 ms
scout-lc-deep33100%55 ms
scout-lc-langchain33100%45 ms
scout-prime-coordinator33100%0 ms
strategist158158100%0.02600900000970796 ms

All runs to date succeeded — an OK rate of 100% is this period's data, not a claim; failure classes will appear here as they occur.

Venture artifacts — playbook-generated, provenance-checked

ArtifactPathClassificationGeneratedChecksum
validation/dataset-cardworkspace/artifacts/validation/dataset_card.mdREAL2026-08-10T02:28:03+00:00sha256:c45b2d9449f40044…
research/research-briefworkspace/artifacts/research/research_brief.mdREAL2026-08-10T02:28:02+00:00sha256:c0b675c7137d170e…
ops/changelogworkspace/artifacts/ops/changelog_entry.mdREAL2026-08-10T02:28:02+00:00sha256:b00597a9b069e96b…
ops/ops-digestworkspace/artifacts/ops/ops_digest.mdREAL2026-08-10T02:28:02+00:00sha256:2fefaf440f4f1f9f…
monitor/run-scoreboardworkspace/artifacts/monitor/scoreboard.jsonREAL2026-08-10T02:28:01+00:00sha256:9c7c2e015dcdaa1f…
monitor/run-scoreboardworkspace/artifacts/monitor/scoreboard.mdREAL2026-08-10T02:28:01+00:00sha256:44fbaaf7bb6fc015…

Generated by scripts/business/playbook.py run <venture> from committed sources only; each row's checksum is copied verbatim from the venture's own manifest, never recomputed here. Documents live under workspace/artifacts/ in the dottie repository.