Routeval · Python evaluation tool · 0.1.0

What did your escalations buy?

Replay examples through your AI pipeline. Compare answer quality, routing decisions, and cost in one report — for a model cascade, a human handoff, or a rules-and-LLM flow.

Python 3.10+ · no required package dependencies · three offline examples · no API key needed

Published on PyPI: install Routeval 0.1.0 with python -m pip install routeval==0.1.0. A fresh public install, the offline demo, and all three bundled fixtures passed verification. The trial kit also includes the same verified wheel and source archive.

See what answer accuracy leaves out

These are five synthetic cases, evaluated by the installed Routeval package. The second policy skips a required local review. Its final answer stays correct and its API cost stays the same — the review is free. Its escalation recall still falls.

correct answers / all five attempts
required escalations actually taken
synthetic API spend per correct answer
Answer correctAnswer wrong / failed
Route matches labelBoth checks passedChosen path failed
Route differsCorrect answer, policy mismatchBoth checks failed

These route labels express a reference policy. A mismatch does not prove an answer was accidental or that another path would have produced a better answer. The costs are illustrative; zero API spend excludes local compute and energy.

Get your first report

In a fresh folder, run:

python -m pip install routeval==0.1.0
routeval init
routeval run

The bundled demo needs no API key. Expect 20 attempts, 16 correct answers, one deliberately allowed failure, and $0.0613 synthetic total cost. Failed attempts remain in the accuracy denominator.

Try it with five cases

Download and unzip the trial kit. Open a terminal in its routeval-trial-kit folder, then run:

python -m pip install routeval==0.1.0
routeval --version
routeval run --config examples/local-hosted-cascade/routeval.json

You should see version 0.1.0, five correct answers, 100% escalation recall, $0.011 total synthetic cost, and $0.0022 per correct answer. The report also saves a complete JSON run artifact.

The kit includes three starting points. Each uses a lookup fixture; none calls a model or claims to implement a production integration.

Local / hosted cascade

A local fast path, a hosted escalation, and a free local review. Check that price never determines a route's role.

Support handoff

Self-service answers and cases that require a human. Keep policy labels separate from whether the final answer sounds right.

Rules / LLM extraction

Rules handle ordinary fields; an LLM handles ambiguous layouts. Check whether a policy change skips required fallbacks.

To reproduce all three fixtures and their deliberate routing regressions:

python run_trials.py

For your own pipeline, change one function and replace the five golden examples. Return the answer, actual route, and whole-decision cost, including calls made before escalation:

from routeval import Decision

def evaluate(request) -> Decision:
    result = your_pipeline(request)
    return Decision(
        answer=result.answer,
        route=result.route,
        cost_usd=result.total_cost_usd,
    )

Name the automatic and escalation routes in the supplied config. Add expected_route labels if you want routing metrics. Without those labels, answer quality, per-route cost and latency, and cost per correct answer still work. Missing cost stays unknown; cost gates require complete coverage.

Save a baseline, change one policy, and compare on the same five examples. An unavailable required metric, a changed measurement contract, or a routing regression returns a nonzero exit code.

Where this came from

In Quorum's verified routing audit, 40 of 184 attempts had correct verdicts with different route labels: 34 missed escalations and six unnecessary escalations. That is 21.7% of all attempts; it is a different quantity from the share of correct verdicts.

The result led me to separate answer quality, policy agreement, and economics. It did not establish that those answers were luck. The Quorum evaluation set is diagnostic evidence, not a verified independent estimate of production performance.

Routeval packages that separation as a small report and versioned artifact. Bring your own scoring function or judge. Existing evaluation tools can also accept custom scoring; Routeval is a focused reporting contract that can sit alongside them.

Help shape the next release

I'm looking for three independent trials: a cascade, a support handoff, and an extraction flow. The useful result is a repeatable report that changes a decision. Five cases are an onboarding test, not a production-quality estimate.

The kit includes FEEDBACK.md. Record time to your first useful report, wrapper effort, confusing labels, missing costs, and the decision the report changed. If you get stuck, that is useful feedback too.

Email me about a trial. Full artifacts can contain private answers and errors. The routeval summary command removes item contents and identifiers; review route and stage names before sharing. Nothing in this page or kit sends your results automatically.