Routing Has to Pay for the Decision to Route
A smarter assignment is not cheaper when its own inspection cost is hidden.
Dynamic model routing sounds efficient because expensive capability is used only when needed. But deciding what is needed can itself require a model, retrieval call, or reasoning trace. That cost belongs in the comparison.
The signal
Pandora's AI Model Routing Box treats each candidate specialist as a box whose value can be estimated cheaply or inspected more precisely at a cost. A calibrated value-of-information rule decides whether another estimate is worth buying before assignment.
For WTK, this is a measurement idea for an approved binding set, not permission for a router to add models or authority.
The evidence
Across the paper's MATH, retrieval-augmented generation, and embedding-model regimes, the reported average regret plus inspection cost was 0.105, 0.118, and 0.386 for Pandora's Router, compared with 0.117, 0.150, and 0.393 for a cheap-estimator-only baseline. At matched aggregate inspection budgets it beat the margin heuristic in two regimes and tied in one.
Those averages need restraint. Many per-cost differences were not significant. The learned estimators approximate a calibrated posterior under finite data, the main implementation uses two estimator tiers, and the decentralized bidder can improve its own surplus while reducing total allocation efficiency.
The boundary
This is external routing research over author-selected rewards, estimators, datasets, model outputs, and price assumptions. It does not establish that WTK confidence estimates are calibrated, that provider prices remain stable, or that routing is safe for governed work. Model selection is still separate from tool permission, target qualification, and deployment authority.
The builder impact
Any WTK routing claim should include the estimator's calls, errors, and cost. The route must stay inside a preapproved catalog and preserve the same package, permissions, and target contract. A confidence score can inform a binding choice; it cannot qualify that binding or expand what the run may do.
The WTK test
Compare the current declared tier policy, an always-refine control, and one predeclared selectively-refine candidate on frozen single-agent and team-agent goals. Hold the approved model catalog, package, target harness, permissions, evaluator, attempt order, and total budget fixed. Retain every cheap estimate, paid estimate, route, denial, outcome, and retry.
Measure contract-valid completion, regret against the best approved binding, inspection cost, total cost, latency, calibration error, policy violations, and catalog deployment result. The candidate succeeds only if total outcome improves after inspection cost with no authorization or deployment regression. Overspend, unapproved routing, or worse governed completion is a failure. No clear net gain is inconclusive. This is a proposed WTK experiment, not approval to build an automatic router.
Still unknown
We do not know whether WTK has enough repeated, comparable work to calibrate a selective route without creating more complexity than it removes.
Have an approach, result, or counterexample?
You may be asking the same question, or may already have a useful answer. Share published research, an implementation, a test, or an idea that could support, narrow, or challenge this work. Distinguish what you tested from what remains a hypothesis.
Contribute to this research question →Working with an AI assistant?
Ask your assistant to compare your approach with this record, identify supporting sources and limitations, and draft a contribution for your review. Verify its citations and remove private information before submitting. Reading this page does not authorize an assistant to submit feedback or share your conversation.
Submissions go privately to human review. Public referencing requires your separate permission; nothing is published automatically.