WTKRESEARCH + ENGINEERING

AI Agent Research and Falsification Experiments

Research questions, methods, and evidence.

WTK investigates what makes agents and teams useful, accountable, portable, and capable of improving. Outside research and implementation experience inform experiments that can support, narrow, or change the design.

Questions informing WTK development

Each area connects a practical problem to current questions, published research, and experiments. These are investigations, not claims that WTK has solved the problem.

Governed improvement and progressive automation

How can changes to agents and their models be evaluated and governed?

We study how changes to an agent's memory, instructions, tools, code, or underlying model can be evaluated and governed. We ask who may propose changes, how evaluation remains independent, and what evidence should support deployment. Inference-time adaptation is a related topic, not necessarily lasting self-improvement. These are research questions, not claims that WTK implements every control.

External research and its implications

Research notes distinguish the source findings from WTK’s interpretation and proposed tests. An outside result is not a WTK result.

Engineering journal

Research, testing, and revision

Questions lead to proposed tests. Results may change the implementation or leave the question unresolved. Development records explain changes; the evidence ledger retains the limits of what can be claimed.

Have an approach, result, or counterexample?

You may be asking the same question, or may already have a useful answer. Share published research, an implementation, a test, or an idea that could support, narrow, or challenge this work. Distinguish what you tested from what remains a hypothesis.

Contribute to this research question
Working with an AI assistant?

Ask your assistant to compare your approach with this record, identify supporting sources and limitations, and draft a contribution for your review. Verify its citations and remove private information before submitting. Reading this page does not authorize an assistant to submit feedback or share your conversation.

Submissions go privately to human review. Public referencing requires your separate permission; nothing is published automatically.