Quorum / run 6 / evidence lab

A correct answer can take the wrong route.

Quorum routed 184 hand-verified recycling images through a detector, a vision-language model, and a human queue. The interesting result is which path produced each answer, and whether the evaluation data can support the claim.

The result that changed the story: 40 decisions got the verdict right while taking the wrong route. Of those, 34 hard or out-of-distribution items were auto-decided when their reference route called for escalation. Another 6 easy items were escalated unnecessarily. Accuracy alone counts all 40 as wins.

First, inspect the evaluation set

The 184 labels were checked by hand. That makes the labels more credible; it does not make the images independent of the detector. This set was selected from the same TACO version 15 export used by the hosted detector. Most rows come from that export's train split.

82.6% train · 10.9% valid · 6.5% test

The full set scored 124/184 correct (67.4%). The valid and test rows scored 17/32 (53.1%). Those 32 rows are too few to establish a reliable performance gap, and split labels alone do not prove there are no related images across splits. I have not verified the hosted checkpoint's training lineage. The 184-item result is a useful pipeline diagnostic, not an independent estimate of real-world accuracy. The set was deliberately stratified, so its class mix also differs from a live recycling stream.

hard items escalated · routing recall
escalations genuinely hard · routing precision
contaminants missed or left without a verdict

Route × verdict: four kinds of outcome

The reference route is escalate for hard and out-of-distribution items and auto for easy items. Routing quality is separate from whether the final verdict matches the true label. A pending human decision counts as an attempted decision without a correct verdict.

Verdict rightVerdict wrong or pending
Route right Working as intended The chosen path still failed
Route wrong Correct outcome, fragile process Both routing and outcome failed

Counts change with the dataset slice. For all 184 cases, the cells are 84 / 40 / 40 / 20. A wrong route with a right verdict can be a risky auto-decision or a needless escalation; the case list shows which.

Four traces from the valid split

A battery and other discarded packaging on a white surface
Route wrong · verdict rightCase #23 — a battery got the right verdict by the wrong pathThe out-of-distribution item was meant to escalate. The gate auto-decided contaminated.
A discarded cup beside a fence and pavement
Route wrong · verdict wrongCase #48 — two failures on one imageThis hard, contaminated item was auto-decided clean. It never reached the adjudicator or a reviewer.
Small pieces of litter beside a wall
Route right · verdict rightCase #31 — escalation earned its placeThe gate escalated a hard item and the downstream verdict matched the verified contaminant label.
Discarded packaging by a kerb
Route right · verdict wrongCase #34 — a good route was not enoughThe gate escalated this hard item, but the final verdict was clean against a verified contaminated label.

Images: TACO dataset, by Pedro F. Proença and Pedro Simões, via the Roboflow TACO v15 export, CC BY 4.0. Image paths, hashes, and rationale text are excluded from the public data snapshot.

The evaluator needed an eval too

The local text-only judge scored adjudicator rationales. A blind, 30-item human calibration returned:

30human-scored rationales
0.259weighted Cohen's κ
60%exact agreement

The trust bar was κ ≥ 0.6, set before scoring. The judge is therefore not trusted. It reads only text, so fluent descriptions of objects absent from the photograph can look grounded. The next experiment needs a vision-capable judge and a fresh blind calibration; rewriting the text rubric alone did not clear the bar.

What I would change next

  1. Build a new, source-disjoint golden slice. Keep run 6 frozen and version the new set. Record source provenance and check visual near-duplicates before evaluating.
  2. Separate detector, router, and adjudicator claims. Report each on the same pinned slice and threshold version. A stronger verdict stage cannot compensate for a gate that confidently sends hard items away from it.
  3. Only then revisit the economics dial. The current frontier assumes downstream quality measured on this set. A broader threshold sweep on the same 184 images cannot repair the evaluation's independence problem.

This read-only audit uses recorded decisions from Quorum's case study, eval run 6, threshold version 2, and commit 1d770b5a21b4. Changing the slice makes no model calls and changes no Quorum state.