WTKRESEARCH + ENGINEERING
Illustration of the WTK agent factory as a mechanical production line

Research into useful, accountable, and improving AI agents.

We investigate how AI agents and teams can accomplish useful work, operate within clear limits, and improve over time. WTK is the governed AI agent factory we are building to test these ideas.

Questions informing WTK development

These questions shape what we build and test. This site shares the research, experiments, failures, and engineering decisions informing that work, including approaches that may change our thinking.

Agent usefulness and outcome evaluation

Did the agent accomplish the intended goal?

A completed task or a convincing answer may still miss what a person needed. We study how to define useful outcomes and test whether agents deliver them.

Authority, security, and oversight

What may an agent do, and when should a human decide?

As agents take on more work, mistakes can have greater consequences. We investigate permissions, independent review, and the limits of automated decisions.

Agent-team coordination and reliability

Does the team work reliably as a whole?

Agents that perform well individually can fail when they share work. We study delegation, handoffs, disagreement, and responsibility for the final result.

Portability and reuse

What needs testing when an agent moves to another environment?

Changing a model, tool, or runtime can change behavior. We investigate how to preserve an agent’s purpose and history while identifying what needs fresh evidence.

Governed improvement and progressive automation

How can changes to agents and their models be evaluated and governed?

We study how changes to an agent's memory, instructions, tools, code, or underlying model can be evaluated and governed. We ask who may propose changes, how evaluation remains independent, and what evidence should support deployment. Inference-time adaptation is a related topic, not necessarily lasting self-improvement. These are research questions, not claims that WTK implements every control.

How the factory puts ideas to work

WTK brings research into a governed lifecycle for agents and agent teams.

  1. Build

    Define the goal, assemble the package, and set its limits.

  2. Qualify

    Test the exact execution form and retain failures and limitations.

  3. Deploy

    A human authorizes where the qualified version may run.

  4. Improve

    Compare a candidate with the accepted version before promotion.

Explore the lifecycle and its evidence boundaries

Recent publications

Outside research, development records, and WTK findings, with each publication’s evidence and limitations kept distinct.

WTK’s purpose and direction

WTK is a governed AI agent factory for building, qualifying, deploying, and improving portable agents and agent teams. Its evidence-driven improvement loops can propose and test changes, while promotion remains independently governed and human-authorized.

Long-term vision

A governed, self-improving AI agent factory that progressively automates the building, qualification, deployment, operation, and improvement of portable agents and agent teams while keeping authority, evidence, and promotion independently governed.

How research informs development

Outside research and practical observations help us identify questions and design tests. We use WTK to investigate those questions, then document the results, limitations, and any changes to the approach.

Outside findings are not evidence that WTK has passed the same test. Research may lead to a design change, a proposed experiment, or no change when the evidence is insufficient.

Relevant research, practical results, and alternative approaches can inform this work. Contributions are reviewed privately by a human. Public use requires separate permission; submission does not result in automatic publication.