WTKRESEARCH + ENGINEERING

Why WTK Is Building a Governed, Self-Improving AI Agent Factory

Move humans from operating the factory to governing it.

As agents take on longer, more consequential work, model capability is only part of the system. Goals, authority, evidence, qualification, deployment, improvement, and recovery must survive changes in models and execution environments. WTK is building the factory around those responsibilities.

Purpose of the public site

This site is part of WTK’s research process, not just a record of completed work. We publish ideas, observations, failures, inconclusive results, and open questions so others can examine the evidence and challenge our assumptions.

External research and reader feedback can inform proposed experiments and design changes, or support keeping the current approach. The intended cycle is to publish what we know, review critiques, propose and test changes, and report what happened. We are developing this process; we have not established a reliable, fully automated feedback cycle.

Our learning agenda includes agent improvement, changes to underlying models, and adaptation during inference. We study these areas to understand what WTK may need to govern as it develops, not to claim that we already train self-improving models or control every improvement process.

Contributions go privately to human review. Feedback does not automatically change WTK or authorize an experiment or deployment. Experiment approval, evaluation, and deployment authorization remain separate decisions.

Explore the research questions or share evidence, a correction, or a counterexample. Read our research direction on governing improvement in agents and models.

Capability makes autonomous work possible. Evidence makes trust defensible.

Models and the software that runs them supply reasoning, tools, memory, and coordination. WTK studies how to keep an agent focused on its goal, limit what it may do, test its behavior, and make the supporting evidence available for independent checks.

The purpose of WTK is not to prove its architecture right. It is to learn what governed autonomous systems actually require.

Keep the original goal explicit through development.

An agent can satisfy a loosely stated request while violating the reason it exists. WTK treats explicit goals, delegated authority, evaluation obligations, and stop conditions as durable engineering artifacts.

Check important claims independently.

Model-level safety remains important. WTK asks which authorization, evidence, and qualification claims can also be checked outside the model and execution provider being evaluated.

Reassess controls when models or runtimes change.

Changing models or runtimes should not silently change the agent's goal, permissions, or evaluation requirements. WTK investigates which controls and evidence remain applicable, which cannot be preserved, and what must be tested again.

Improvement cannot authorize itself.

Failures and field outcomes should produce attributed repairs, comparable candidates, stronger tests, and reusable learning. Candidates remain provisional, failures remain visible, and an accountable authority controls qualification, promotion, rollback, and replacement.

Autonomy advances function by function.

WTK does not treat autonomy as one global switch. Research, construction, evaluation, deployment, monitoring, and remediation can move independently from human execution to assisted work, supervised execution, oversight, and bounded unattended operation.

Self-improvement may be automated. Self-authorization should not be.

Increasing automation changes where humans participate. It does not remove responsibility for goals, policy, security, risk acceptance, exceptions, suspension, or revocation.

Explore governed self-improvement

Understand the failure boundary before making the trust claim.

WTK approaches autonomous agent factories as systems that must withstand malformed inputs, excessive authority, unreliable evidence, drift, recovery failures, and governance mistakes. The work is not an argument for maximum process. It is an experiment in finding controls that measurably improve task outcomes, safety, or recovery without making the system too costly or cumbersome to use.

THE HONEST CAVEAT

WTK still needs to demonstrate that its controls improve outcomes.

WTK must demonstrate that its structures improve outcomes, safety, recoverability, or economic value under declared conditions. If the controls add process without measurable benefit, the architecture must change.

Submit an inquiry