WTKRESEARCH + ENGINEERING

AI Agent Findings and Failure Reports

Reported results and observed failures.

Findings record what an experiment demonstrated. Failure reports preserve where agents failed, drifted, recovered, or produced misleading evidence.

Findings

This Pilot Could Not Isolate a Memory Effect

Every run held unsafe promotion, but ambiguous scoring kept the memory comparison inconclusive.

Confidence
Provisional
Still unknown
Repeated-cell stability, broader goals and models, whether stronger conflicting history changes decisions, and a clean catalog-eligible positive control
Next experiment
Repeat the matched design with unambiguous action-boundary fields, a valid catalog-positive control, explicit native success predicates, and repeated cells
Full article
Failure Reports

The Receipt Was Used. The Abstention Still Failed.

One model binding followed changed observations in every positive case, then failed the strict evidence-absence contract in every negative case.

Confidence
Provisional, based on 24 retained attempts under one frozen question and response contract
Still unknown
Whether the failure replicates across other questions, prompts, harnesses, model versions, and adversarial observations
Next experiment
Repeat the causal-pair protocol across additional questions, bindings, and harnesses while separating abstention, output conformance, and receipt-reference scoring.
Full article
Findings

The Evidence Stayed. Its Authority Did Not.

A package edit preserved the old record for inspection while preventing stale construction and approval evidence from authorizing the revised package.

Confidence
Provisional, based on one deterministic local package revision through production edit and catalog-readiness paths
Still unknown
Whether the same boundary survives a complete rebuild, retest, requalification, target compile, and catalog readmission of a real package
Next experiment
Revise a complete qualified package and replay every affected lifecycle stage until the new digest either earns fresh evidence or remains blocked.
Full article
Findings

Five Failures Needed Five Different Repair Decisions

A fixed comparison found that typed escalation routed every persistent failure correctly while instruction-only repair repeated the same move.

Confidence
Provisional, based on one deterministic five-case comparison against a declared routing oracle
Still unknown
Whether applying the selected repairs improves live agent outcomes, regression rate, cost, latency, or operator effort
Next experiment
Apply instruction-only and typed repairs to the same retained live failures and compare resolution, regression, honest holds, cost, latency, and operator action.
Full article
Findings

An Installed Tool Is Not Necessarily the Right Tool

In four fixed cases, WTK's capability ladder preserved fit, cost consent, discovery, and construction boundaries that an installed-first policy missed.

Confidence
Provisional, based on one deterministic four-case comparison against declared decisions
Still unknown
Whether model-derived fit judgments, live discovery, provisioning, setup, and post-provisioning work preserve the same boundaries
Next experiment
Run an attended comparison with live registry candidates, unavailable connectors, explicit cost consent, and measured goal accomplishment after provisioning.
Full article
Findings

A Known Workflow Needed Fewer Model Decisions

In a matched 24-run experiment, a governed fixed workflow completed every case while using fewer provider calls, tool decisions, and tokens.

Confidence
Moderate, based on one matched experiment with repeated trials and deterministic scoring
Still unknown
Whether the result replicates across other models, harnesses, tool chains, live data, and larger samples
Next experiment
Repeat the matched protocol across a second model, a second harness, a different known workflow, and live tool responses.
Full article