Complete experiment run 1: 36 episodes, 179 archived model-I/O files.

Recursive red-team self-play — complete recorded experiment

Both completed 36-episode runs of exp_errorprop: four feedback conditions × three repetitions × three generations per run. All 72 completed episodes and every archived role model-I/O file, including all recorded roundtrips, are included.

Two experiment runs; 72 episode generations; 24 lineages; 360 archived role
sessions; 359 original model-I/O files; 3377 model roundtrip records.

Downloads are split into two complete run-level ZIPs. Each contains every
recording from its 36-episode run plus the shared review, checkpoint, and
experiment code. No recorded model-I/O file is shortened or split. Each ZIP
has its own filtered manifest; the website manifest describes both runs and
includes the ZIP byte sizes and SHA256 values.

Read the conditions desk -> engineer shift 1 -> engineer shift 2 -> platform
owner -> playbook keeper -> next generation. Every exact brief is separately
linked. Complete model-I/O JSONL files include repeated request histories;
the Agent Trace viewer organizes those histories for reading.

Open /tools/agent-trace?trace=selfplay-errorprop. Run links may specify
&run=run2-ledger_correction-r2-g3&role=defend, or an exact recorded session
&session=run2-ledger_correction-r2-g3-defend-1.

Read review/ACCOUNT.md before using historical detector counts. All original
results, compact traces, and manual labels are retained unchanged. Refer to
manifest.json for file source paths, byte sizes, hashes, and explicit gaps.
Before/after artifacts: state/ is the archived workspace copy; input-before/
also preserves the carried input playbook where the desk edited that copy;
state_updates/ is the exact staged owner/keeper output. For run 1, historical
compact-trace before-hash fields were captured after applying updates.

Preservation: Copied source files retain their exact archived bytes. SHA256 and original repository-relative source paths are listed for every file. Original results, compact traces, and historical detector classifications remain unchanged; the dated ACCOUNT review supplies superseding adjudications. Viewer conversation histories may be de-duplicated for reading; downloadable JSONL files are complete.

Archive notes:
- Run 1 explicit_correction repetition 1 generation 1 has a recorded desk brief, console, and telemetry but no archived model response. The matched fixed plan was staged independently.
- Four workspace playbook copies were edited within their episodes: run 1 accurate-r1 generation 2 and accurate-r2 generation 1; run 2 accurate-r1 generation 1 and incorrect-r3 generation 1. These archived edits and their compact-trace gap records are retained. Checkpoint or previous-generation staged playbooks provide the original carried input.
- Run 1 compact traces were rebuilt after state application; their config_before_sha16 and playbook_before_sha16 fields often refer to the after artifact. The raw workspace state and separate input-source links preserve the actual recorded artifacts. These historical compact-trace fields are not rewritten.
- Historical regex-derived counts include known false positives. Read review/ACCOUNT.md and the manual label files for the superseding census. Supporting comparisons are retained as recorded.
- Full sandbox homes, session databases, raw environment directories, vault/secret files, raw workspace tarballs, and inherited-world duplicate telemetry are excluded. The exact archived model-I/O, all role tool-result attachments, role consoles/telemetry, exact briefs, causal ledgers, staged plans, before/after learning artifacts, engineer notes, publication service records, and host reports are included.
- The existing explicitly labelled redacted episode-config derivative is included for the treatment-channel evidence; original full episode configs are excluded. Synthetic canary values may appear in preserved recorded model-I/O/tool output, as in the source archive.
