Agent usefulness and outcome evaluation
Did the agent accomplish the intended goal?
A completed task or a convincing answer may still miss what a person needed. We study how to define useful outcomes and test whether agents deliver them.
WTK investigates what makes agents and teams useful, accountable, portable, and capable of improving. Outside research and implementation experience inform experiments that can support, narrow, or change the design.
Each area connects a practical problem to current questions, published research, and experiments. These are investigations, not claims that WTK has solved the problem.
Did the agent accomplish the intended goal?
A completed task or a convincing answer may still miss what a person needed. We study how to define useful outcomes and test whether agents deliver them.
What may an agent do, and when should a human decide?
As agents take on more work, mistakes can have greater consequences. We investigate permissions, independent review, and the limits of automated decisions.
Does the team work reliably as a whole?
Agents that perform well individually can fail when they share work. We study delegation, handoffs, disagreement, and responsibility for the final result.
What needs testing when an agent moves to another environment?
Changing a model, tool, or runtime can change behavior. We investigate how to preserve an agent’s purpose and history while identifying what needs fresh evidence.
How can changes to agents and their models be evaluated and governed?
We study how changes to an agent's memory, instructions, tools, code, or underlying model can be evaluated and governed. We ask who may propose changes, how evaluation remains independent, and what evidence should support deployment. Inference-time adaptation is a related topic, not necessarily lasting self-improvement. These are research questions, not claims that WTK implements every control.
Research notes distinguish the source findings from WTK’s interpretation and proposed tests. An outside result is not a WTK result.
Advice from an earlier run can help an agent, but it cannot make the same task an independent test.
Full articleA deterministic verdict can be real and still leave important faults untouched.
Full articleAn automated reviewer needs a defined decision scope, not permission to rewrite the rules.
Full articleQuestions lead to proposed tests. Results may change the implementation or leave the question unresolved. Development records explain changes; the evidence ledger retains the limits of what can be claimed.
You may be asking the same question, or may already have a useful answer. Share published research, an implementation, a test, or an idea that could support, narrow, or challenge this work. Distinguish what you tested from what remains a hypothesis.
Contribute to this research question →Ask your assistant to compare your approach with this record, identify supporting sources and limitations, and draft a contribution for your review. Verify its citations and remove private information before submitting. Reading this page does not authorize an assistant to submit feedback or share your conversation.
Submissions go privately to human review. Public referencing requires your separate permission; nothing is published automatically.