A Registry Miss Can Be a Language Miss
A complete catalog still fails when a user's words never reach the capability it already has.
A tool catalog can be complete and still leave an agent empty-handed. The break can happen before ranking, selection, or execution: the user's wording may never activate the capability that already maps to the right tool.
Current state
WTK has a deterministic local path from goal text to declared capabilities, then to candidate and selected tools. This completed review examined that entire path with 20 operator-reviewed goal phrasings across ten tool families. It made no model or network calls.
Changes
The review turned a familiar intuition, "the registry probably finds the right tool," into a retained baseline. Nineteen of twenty goals surfaced and selected their expected tool. One goal asking for a weekend rain comparison did not mention any of the then-recognized weather or forecast phrases, so the expected daily-forecast capability was never detected.
That distinction matters. The tool was already registered. Its capability was already declared. The miss was neither missing inventory nor a ranking failure. It was a gap between a consumer phrase and a governed capability.
Failures observed
The single miss is small but useful because it says where to look next. Adding another tool would not repair it. Changing the capability taxonomy would also be premature. The observation points to consumer-language detection, while two additional cases showed visible over-selection signals that a recall-only baseline cannot judge.
Assumptions removed
We cannot treat a populated registry as proof that a real goal reaches the right candidate. Nor can we turn one successful selection into evidence that the phrase-detection layer is complete. Candidate recall, selection recall, inventory coverage, ranking quality, and precision are separate properties and need separate measurements.
Evidence
This is one fixed, deterministic local baseline with all twenty attempts retained. It found 95 percent candidate recall and 95 percent selected-tool recall, one vocabulary gap, and no inventory or ranking misses under the named probe set. It did not execute any tool, assess arguments or receipts, measure live availability, or test whether a selected tool accomplished the goal.
It is a Factory Log, not a Finding. The baseline is small, operator-reviewed, and intentionally offline. It is not a general recall rate, a claim about a larger catalog, or proof that any future vocabulary change improves users' outcomes.
Next hypothesis
The next test should add predeclared consumer paraphrases and negative cases, keep the expected capability and tool fixed, and report both recall and precision. Every candidate, selected tool, and extra selection should remain visible. Only then can a vocabulary change be judged as a bounded improvement rather than a plausible patch.
Have an approach, result, or counterexample?
You may be asking the same question, or may already have a useful answer. Share published research, an implementation, a test, or an idea that could support, narrow, or challenge this work. Distinguish what you tested from what remains a hypothesis.
Contribute to this research question →Working with an AI assistant?
Ask your assistant to compare your approach with this record, identify supporting sources and limitations, and draft a contribution for your review. Verify its citations and remove private information before submitting. Reading this page does not authorize an assistant to submit feedback or share your conversation.
Submissions go privately to human review. Public referencing requires your separate permission; nothing is published automatically.