Evaluating AI Agents Without Leaking Prior Task Outcomes
Advice from an earlier run can help an agent, but it cannot make the same task an independent test.
Full articleNotes, findings, failures, and factory logs remain one chronological record of what WTK is learning and changing.
Showing 12 of 102 publications
Research area descriptionsAdvice from an earlier run can help an agent, but it cannot make the same task an independent test.
Full articleA deterministic verdict can be real and still leave important faults untouched.
Full articleAn automated reviewer needs a defined decision scope, not permission to rewrite the rules.
Full articleMore agent work is useful only when we can tell what it accomplished and what people had to correct.
Full articleA continuation summary can preserve a mistake as well as useful progress.
Full articleA fixed token budget can fund one longer attempt or several shorter ones. The useful choice depends on the task and how the result is selected.
Full articleEvaluate the capabilities and settings available in the environment where the agent will actually run.
Full articleVerify the environment before activation, and keep blocked work inside its approved scope.
Full articleSeparate workspaces do not establish isolation if agents can exchange information through a shared service.
Full articleRecognizing a hostile message is only the first check in a persistent agent system.
Full articleCheck whether the evaluator made a mistake or the requirements left something important out.
Full articleMeasure what a monitor detects separately from what an intervention prevents.
Full articleTest whether safety checks still work when a request changes but its intent does not.
Full articleRuntime data can fill in a task without expanding its authority.
Full articleWhat should remain under independent control when the system proposes its own changes?
Full articleA rule can reach the running agent without constraining what it does.
Full articleAn update can misuse existing access without changing permission settings.
Full articleA harness change needs evaluation on tasks that were not used to develop the change.
Full articleA patient attack can cross the boundary that a single-run safety check never sees.
Full articleOne deployed model can present different caution and candor signals across language contexts.
Full articleA supplied-text-only team acquired unnecessary capability requirements during planning, before its agents could execute.
Full articleAn enterprise approach to agent authority, team accountability, and governed improvement.
Full articleLater user trials exposed a completion path that needed to distinguish a revised goal from its earlier successful package.
Full articleHidden transport retries can separate the work an agent performs from the attempts its factory records.
Full articleA hash can prove the same request was sent, not that a shared evaluator will make the same decision.
Full articleAn agent team can settle its disagreement after crossing the boundary that should have stopped it.
Full articleTyped routing can stop a broken trace, but it cannot prove that the agent should have acted or reached the right result.
Full articleA reviewer model is part of the runtime, so it needs evidence that it improves the target it can block.
Full articleA label made by an improvement loop is still a hypothesis until independent evidence checks it.
Full articleA proposal to account for every required team role before assessing the executed workflow.
Full articleEvidence can look authoritative and still fail to justify an irreversible action.
Full articleA smarter assignment is not cheaper when its own inspection cost is hidden.
Full articleSeparate authority is not enough when a reviewer cannot validate the observations it receives.
Full articleCompression is useful only when rare obligations remain executable and testable.
Full articleTool access does not prove that an agent found the record its answer requires.
Full articleChanging the receiver is only half the experiment when context crosses with it.
Full articleA preregistered comparison can still be unable to identify the claim it reports.
Full articleA repair can respond to a failed review without copying the withheld case that exposed it.
Full articleA green utility score can still hide a changed answer regime.
Full articleAn improvement score is only meaningful when its intended outcome is declared before the comparison begins.
Full articleA target can look complete on paper and still miss the relationship that makes it work.
Full articleA second look can test retained evidence without quietly replacing the decision it was meant to examine.
Full articleA response can match a schema and still fail the obligation that the fixture was meant to test.
Full articleA connection can name the credential it needs without becoming the place that governs the credential itself.
Full articleReviewing durable package source is not enough when the executable form also depends on compiled inputs.
Full articleEvery run held unsafe promotion, but ambiguous scoring kept the memory comparison inconclusive.
Full articleRepeated tests of an unchanged system help distinguish measured improvement from score variation.
Full articleRead-only describes an operation. A package grant must also say where that operation may reach.
Full articleCorrect history can still steer an agent away from the goal it has now.
Full articleA skill induced from an entire trajectory can drop the no-memory baseline. A skill induced from a shared subtask can raise it.
Full articleA complete catalog still fails when a user's words never reach the capability it already has.
Full articlePAJAMA makes parts of model judgment inspectable and repeatable, but the resulting programs still need evidence boundaries of their own.
Full articleA final action check cannot repair a workflow that already went off course.
Full articleMore candidate compute helps only when the final choice is tied to the outcome that matters.
Full articleA verifier's repair history can change its threshold before the next task even begins.
Full articleA blocked candidate is not automatically evidence that the candidate is the problem.
Full articleMore answer samples are not always the best use of a verification budget.
Full articleA retained lesson should have to beat a matched run that starts fresh.
Full articleA fancier memory layer should have to beat a controllable path back to the source.
Full articleA score only speaks for the behavior its cases can actually distinguish.
Full articleThe command may be right before the path to execution makes it wrong.
Full articleA team is not reliable just because each role passed alone.
Full articleWhen a failure survives an instruction change, the system should change repair layers instead of polishing the same guess again.
Full articleOne model binding followed changed observations in every positive case, then failed the strict evidence-absence contract in every negative case.
Full articleA tool observation can be present in the trajectory without doing any work in the answer.
Full articleA package edit preserved the old record for inspection while preventing stale construction and approval evidence from authorizing the revised package.
Full articleExploration, package construction, testing, and repair should remain one traceable journey without turning the conversation into source truth.
Full articleA fixed comparison found that typed escalation routed every persistent failure correctly while instruction-only repair repeated the same move.
Full articleIn four fixed cases, WTK's capability ladder preserved fit, cost consent, discovery, and construction boundaries that an installed-first policy missed.
Full articleChanging governed package source should create a new accountable revision, not let yesterday's proof follow it forward.
Full articlePortability is not permission to assume that every harness enforces the same safety boundary.
Full articleA passing static test says little about what a governed agent does after its environment changes.
Full articleA model that ignores bad context is not trustworthy if it also ignores the good kind.
Full articleA useful memory layer still needs a receipt for where its advice came from and when it should apply.
Full articleMore relations can add explanation without adding one more useful skill.
Full articleThe right amount of structure is not more fields; it is a clear purpose for every constraint.
Full articleA registry should not promise an action that no compiled target can actually perform.
Full articleA passing score is only useful when the test can tell a correct solution from a lucky one.
Full articleA careful final response cannot undo sensitive data an agent already collected.
Full articleIf a reasoning agent can run the search loop, explicit control needs a measurable job.
Full articleThe same declared tool can behave differently when the runtime changes how an agent calls it.
Full articleA tool connection is not ready merely because it was once configured or once passed a check.
Full articleMore grounded explanation does not guarantee that an agent will honor its hard constraints.
Full articleFor a governed tool requirement, an answer is only a partial result until the runtime shows the tool was actually used.
Full articleA missing evidence capability should trigger an explicit decision, not a quieter unsupported answer.
Full articleAn agent becomes buildable when its purpose is translated into observable outcomes, bounded behavior, and evidence requirements.
Full articleChanging the work should create a new accountable version, not let old evidence answer a new question.
Full articleA deployment claim is only as honest as the complete evidence set it keeps.
Full articleBetter judges, better holdouts, and better repairs all depend on knowing exactly what evidence is available.
Full articleWTK now turns a missing required evidence path into a recorded limitation instead of accepting unsupported confident prose.
Full articleWhen the structure of the world shifted underneath them, agents fell back on exhaustive search instead of deduction, and extra reasoning budget made that failure cost more rather than less.
Full articleThe August 5 development record separates session-bound feedback from evaluation results; safe autonomous promotion remains unproven.
Full articleWTK now requires evidence from the compiled execution form before a catalog entry can claim deployable readiness.
Full articleFix the task, fix the model, swap only the harness, and the agent still passes while thinking differently. That is a problem for anyone who reads a green test run as a property of the agent.
Full articleWTK separated evidence by purpose so a convincing run cannot quietly become a qualification claim.
Full articleFlexible reasoning becomes dependable when purpose, authority, evidence, and acceptable outcomes remain governed.
Full articleIn a matched 24-run experiment, a governed fixed workflow completed every case while using fewer provider calls, tool decisions, and tokens.
Full articleIn the cited study, one misleading document increased false conclusions under the tested research-agent conditions.
Full articleIn two case studies, frontier agents received six days and substantial compute, completed the engineering, and made no substantial progress on the research questions. A second model and scaffold reproduced the failure pattern.
Full articleIn one multimodal fact-checking study, up to 29% of post-cutoff claims remained potentially contaminated, and the effect changed system rankings. That challenges how comparative value is measured, including in WTK.
Full articleReplaying verified answers requires checking whether the original verification still applies.
Full articleA new paper puts agent decision-making inside the model's hidden states. That's where WTK requires a tighter boundary, and we can now say why.
Full articleOlder records remain available through the archive, research areas, and RSS.