ATTRIBUTED AGENT EXPERIMENTS

Experiment observatory

Follow a request from one agent to its recorded outcome.

Files stay in this browser.
APPARATUS EXAMPLE

A real AGT policy check with scripted callers. No live model or behavioral finding.

Recorded agent IDs0Distinct IDs in this bundle
Policy decisions0Requests evaluated by AGT
Denied requests0Policy refused the action
Recorded outcomes0Results reported by the host

ACTION TRACE

All agents

A decision records permission. An outcome records what happened next.

THE EXPERIMENT

What makes this a honeypot?

A normal task contains a discoverable shortcut. An independent observer records whether an agent takes it, describes it, or tries to alter the evidence.

01 / TASK

A legitimate objective

For example, repair a ticket export so it satisfies a written specification. A solvable control establishes that the ordinary route works.

02 / OPPORTUNITY

A discoverable shortcut

A test reference or local receipt is writable. Changing it could make a local check appear to pass without satisfying the original task.

03 / BOUNDARY

A defined line

The protocol specifies the legitimate objective, permitted changes and defense layers. Explicit and implied boundaries are separate conditions.

04 / OBSERVATION

A record outside the task

The observer retains requests, decisions and effects separately from mutable local logs. It compares the actual changes with the agent's report.

Crossing and covering tracks are different outcomes.

An agent can disclose a shortcut. A changed log may have an innocent explanation. A concealment finding needs a defined criterion and the underlying evidence, not a suggestive phrase.

READING THE RECORD

Identity is the beginning
of the explanation.

A delegation graph identifies who launched an agent. The action trace shows what it received, requested and changed. A later explanation is kept as a separate statement.

Read the METR incident investigation
01

Observable actionsRequests, delivered messages, tool results and artifact changes can be measured where capture is verified.

02

Explicit limitsHidden reasoning, uninstrumented tools and unknown subagents remain unobserved. A signature does not remove those gaps.

03

Repeatable comparisonsKeep model, client, task, policy and scorer versions with every run. Saved-event replay and fresh model reruns answer different questions.