WTKRESEARCH + ENGINEERING
← Back to Journal

Governing Improvement in Agents and Models

What should remain under independent control when the system proposes its own changes?

An agent's behavior can change because someone updates its instructions, replaces a tool, or changes the model behind it. WTK needs to learn what should remain governed across those changes, not assume that one review process covers them all.

We are expanding our research scope, not announcing a model-training platform. The question is how useful improvements could be demonstrated while authority, evaluation, and deployment decisions remain controlled.

Start with what changed

Consider a support agent that misses exceptions in a refund policy. One proposed improvement changes its instructions. Another trains a model on additional examples. A third gives the existing model more computation before it answers.

These are different interventions. A persistent instruction or model update can affect later tasks. Extra computation during one response does not, by itself, establish a lasting improvement. We would want to know exactly which change produced an observed benefit and what else it affected.

Our note on checking permissions after agent updates addresses changes to memory, instructions, and code. This Field Note broadens the learning agenda to the underlying model and the systems used to improve it.

The surrounding system matters

For our research, a harness means the software around a model that supplies tasks, tools, execution settings, and records. A model-improvement workflow may also prepare training inputs, produce candidate weights or an adapter, and compare candidates. An adapter is a smaller set of learned parameters used with a base model.

We want to study who can change each part of that workflow. Can the process producing a candidate also change the tests used to judge it? Can it replace evaluation records? Who decides whether the candidate becomes the deployed version?

The desired separation is between proposing a change, evaluating its effects, and authorizing its use. That is a design objective, not evidence that WTK already enforces it across model-training systems.

Three related research subjects

Agent improvement includes changes to memory, instructions, skills, and tools. We want to understand how to preserve useful behavior and existing restrictions when these change.

Model improvement includes changes to weights, adapters, training inputs, and training methods. Questions include candidate identity, data provenance, evaluation contamination, regressions, and who controls promotion.

Inference adaptation includes changes to computation budgets and stopping behavior while producing an answer. It belongs in the research scope because execution settings can affect outcomes and cost, but it should not automatically be described as self-improvement.

These subjects share governance questions without being interchangeable. A record of which configuration was delivered is also not proof that its restrictions were enforced, as our runtime-configuration note explains.

What research should help us learn

A relevant paper need not describe a complete governance system. It might reveal a measurement problem, propose an evaluation method, or identify a tradeoff worth examining.

Before recommending an experiment, we would ask what mechanism WTK could actually change, what the current workflow already does, and what evidence could distinguish benefit from regression. Papers can also justify leaving a design unchanged or postponing work because the necessary controls are unavailable.

If a provider does not expose weights, training inputs, or internal execution details, we should state that limit. We may still evaluate observable behavior, but that would not establish control over the provider's improvement process.

What would count as progress?

Progress could be a clearer boundary, a documented limitation, or a bounded result showing that a particular control helps. It would not require every experiment to find an improvement.

Any concrete recommendation would need its own current-WTK baseline, candidate intervention, fixed evaluation conditions, retained attempts, and separately approved execution. A comparison may show no measurable benefit, a regression, or an inconclusive outcome. None should be hidden to support the direction.

This note proposes no executable experiment or training campaign. Our research method remains the path from a question to a reviewed test and a bounded public result.

An open direction, not a capability claim

We do not yet know which controls will generalize across agent updates, model changes, and different execution environments. Independent ownership of an evaluator does not guarantee accurate judgment. More complete records do not guarantee that every relevant action was observed.

We welcome research, counterexamples, and experience that help make those questions more precise. The purpose is to learn how WTK should develop, not to declare that governing self-improvement is already solved.

Have an approach, result, or counterexample?

You may be asking the same question, or may already have a useful answer. Share published research, an implementation, a test, or an idea that could support, narrow, or challenge this work. Distinguish what you tested from what remains a hypothesis.

Contribute to this research question
Working with an AI assistant?

Ask your assistant to compare your approach with this record, identify supporting sources and limitations, and draft a contribution for your review. Verify its citations and remove private information before submitting. Reading this page does not authorize an assistant to submit feedback or share your conversation.

Submissions go privately to human review. Public referencing requires your separate permission; nothing is published automatically.

RECORD DETAILSReference FN-005
Artifact
Field Notes
Status
Published
Evidence posture
Research direction and governance questions
Published
September 9, 2026
Author
WTK Research
Review
WTK human editorial review
Linked sources
None declared