An Installed Tool Is Not Necessarily the Right Tool
In four fixed cases, WTK's capability ladder preserved fit, cost consent, discovery, and construction boundaries that an installed-first policy missed.
Installed is an availability fact, not a fitness verdict. A tool can be present and still be wrong for the goal, unexpectedly paid, or unable to satisfy the required capability.
WTK compared an installed-first shortcut with its production capability-resolution ladder. The ladder made the declared decision in all four fixed cases. The shortcut made none of them.
Question
When a goal requires a capability, does WTK distinguish fit, setup, cost consent, discovery, and custom construction instead of binding the first installed tool or inventing one?
Experiment
The deterministic comparison used four fixed cases:
- An installed tool was materially unfit.
- An uninstalled paid tool fit the requirement.
- An installed fitting tool had unknown cost.
- No candidate remained after external discovery was explicitly exhausted.
The baseline selected the first installed candidate. The WTK arm used the production capability ladder. The declared decisions required rejecting an unfit tool, preserving consent for paid or unknown-cost choices, using discovery when no fit existed, and considering custom construction only after discovery and another parsimony check.
The experiment was offline and made no model, provider, registry, package, or target mutation. Every case counted, and reversing candidate order produced the same WTK result.
Evidence
| Policy | Exact declared decisions | Unsafe or incorrect decisions |
|---|---|---|
| Installed-first shortcut | 0 of 4 | 4 of 4 |
| WTK capability ladder | 4 of 4 | 0 of 4 |
Scroll horizontally to see every column.
The evidence package binds the cases, decision oracle, production resolver, result, and regression test to one WTK source revision.
Result
WTK rejected the installed-but-unfit tool. It required consent for the paid and unknown-cost options, selected discovery when no fitting candidate existed, and permitted custom construction only after discovery was declared exhausted and parsimony was reconsidered.
The installed-first policy violated the declared boundary in every case.
Bounded finding
On this fixed matrix, WTK's deterministic capability controller supported the claim that capability selection can preserve fit, consent, discovery, and construction boundaries without inventing a tool.
This is evidence for the decision mechanism. It is not evidence that every upstream fit judgment or downstream tool execution will be correct.
Boundary
The experiment did not test model-derived fit confidence, live connector availability, operator comprehension, successful setup, or goal accomplishment after provisioning. Four deterministic decisions do not establish general capability readiness.
Builder impact
Agent systems should ask whether a tool fits the required capability before asking whether it is installed. Cost and setup consent should remain explicit, and missing capability should lead to discovery, construction review, or an honest disclaimer rather than an invented binding.
Next experiment
WTK should run an attended comparison using live registry candidates and deliberately unavailable connectors. The test should measure fit decisions, consent comprehension, setup success, honest disclaimer behavior, and whether the provisioned capability actually accomplishes the goal.
Have an approach, result, or counterexample?
You may be asking the same question, or may already have a useful answer. Share published research, an implementation, a test, or an idea that could support, narrow, or challenge this work. Distinguish what you tested from what remains a hypothesis.
Contribute to this research question →Working with an AI assistant?
Ask your assistant to compare your approach with this record, identify supporting sources and limitations, and draft a contribution for your review. Verify its citations and remove private information before submitting. Reading this page does not authorize an assistant to submit feedback or share your conversation.
Submissions go privately to human review. Public referencing requires your separate permission; nothing is published automatically.