PostTrainBench port status and remaining scope
Draft scoping note, 2026-09-11, unreviewed. Sizes are from wc -l, GitHub, and git logs; effort ranges are estimates.
Headline
The environment is already ported. linuxarena/post_train_bench is Odelya's port of control_arena/settings/post_train_bench from the research team's private ControlArena copy (re-sabotage/control-arena: Yulia Volkova, Titus Buckworth, Raja Moreno; Tyler has 59 commits there). It was vendored byte-identical on Jul 24, finished Sep 2, and has a leaderboard arena entry (post-train-bench in control-leaderboard/src/leaderboard/suites/arenas.yml, with scorer_timeout: 1800). It carries the same 6 main tasks and the same 12 side tasks as the source.
What the research team wants, in Yulia's words (topics doc comment, Sep 11): "We'd just want for our setting to be runnable via control tower and it would be nice to see it displayed here https://leaderboard.linuxarena.ai/", and "I can be the point of contact!" That is close to done. What is not ported is their monitoring stack and their post-Jul-24 fixes, which is a smaller and more design-sensitive job than "port the setting." Its main risk is the one that stranded the sabotage-eval streaming PRs (#1067, #1318): landing monitor code before the Eval 2 monitors-as-Inspect-scorers stack merges.
Provenance confidence is medium-high: the setting repo README says "Port of Control Arena's control_arena/settings/post_train_bench," the side-task set matches this repo exactly, and upstream UK AISI ControlArena has no PostTrainBench setting (Yulia: "We ported Posttrainbench into Control arena, it was not originally a control arena setting"). Not verified: whether the Jul 24 vendoring came from this repo or from a copy Tyler holds.
What the source repo contains
A private copy (not a GitHub fork) of UK AISI ControlArena at about v13.5.0 (Jan 2026), created 2026-01-09, last push 2026-08-30, 334 branches, 75 open issues (almost all camera-ready paper items). Additions inside control_arena/, Python LOC:
| Component | LOC | In linuxarena/post_train_bench today? |
|---|---|---|
| PostTrainBench setting | 49,770 | Yes, at the Jul 24 snapshot: 23,301 LOC; cli/, scaffold/, scripts/, Dockerfile.tinker not taken |
| HCAST setting (11 tasks, 4 registered) | 15,397 | No; a separate setting, not part of the CMFT work |
Offline monitor pipeline (monitor/pipeline, preprocess, components) | 14,702 | No |
| Monitored compaction (wrapper, turn extraction, CLI, tests) | ~3,900 | No |
| Hierarchical, hidden-scratchpad, dataset and "enriched" monitors, compaction-suspicion scorer | ~2,500 plus the dataset/enriched monitors inside the pipeline | No; the end-of-trajectory monitor is explicitly "not ported" per the setting README |
Tools: modal_sandbox, tinker_tool, submit_dataset, bash supervisor | 4,856 | Partly: the Modal runner and Tinker training backend yes; the tinker_tool agent tool no |
| Claude Code / Codex bridge scaffolds with monkeypatches | ~1,300 | No, and not needed: Eval 2's white block fixes the harness to claude_code or codex_cli. Yulia: "we also spent a great deal of time to make claude code and codex scaffolds work but we never managed to do it in the end" |
| Content-filter fallback policy, eval-log redaction, patches | ~1,000 | No |
Top level: experiments/ campaign controllers (15k), scripts/ (52k), knowledge/ wiki, catalog, results | ~70k | No, and should not be: campaign bookkeeping, not the setting |
Yulia's account of what is theirs (topics doc, Sep 11): "we designed the side task scorers, dataset monitors, compaction monitors are ours too, then there is an enriched monitor that is a transcript monitor enriched with data from dataset monitor... We also added some hinting to the agent to improve covertness, added elicitation prompts for attacks, added starter scripts for main tasks... Added hidden scratchpad functionality support." The scorers, hints, elicitation prompts and starter scripts are in the Jul 24 snapshot; the monitors are not.
What the setting needs at runtime
- A sample is a frontier agent in Inspect's
react()with a hidden-scratchpad attack tool, post-training a small open model for 6 to 12 hours on GPU (about $82 per eval) or Tinker (about $28), thenevaluate.pyscores main and side tasks. 3 to 5M unique tokens per eval; context peaks 164K to 313K, so summary compaction is load-bearing. - The object model has moved since the poster. Yulia: "it is more agent and monitor models now, and llama 3.2 3b as object model, but it does not really matter for the setting architecture and code, just a bit for the jinja templates"; they instruction-tuned that Llama themselves so the agent does not spend eval time on it. The port's
base_modellist includes Llama 3.2 3B; their tuned checkpoint and templates are a tier-1 sync item below. - Fresh results are coming: "we have since iterated quite a bit and will publish fresh results by end of sept."
- Their own notes on the code: "The recurring failure on this project has not been analysis, it has been bookkeeping" (
knowledge/README.md); paper evals "cannot be reproduced from any commit" (dirty trees, unmerged branches); compaction was silently unwired in all 33 April react runs; monitor reruns return cached identical scores. CI covers only HCAST tests and three PTB test files. - Yulia on what slowed them: long-horizon tasks; "setting up the right templates for the agent model and scorers to evaluate llama correctly"; Inspect, react and Claude model updates; and, in Control Tower's favor, "I really appreciate CT providing elicitation prompts and different monitor modes and hidden scratchpad built in."
Remaining scope, sized
| Tier | Work | Source size | Precedent | Team effort (estimate, medium confidence) | Gate |
|---|---|---|---|---|---|
| 0. Verify the existing port | One honest and one attack run on a main/side pair via ct run eval (or eval2); confirm a leaderboard run publishes; compare with the paper's core-53 .eval files (their git states are not reproducible) | none | porting-settings.md PR checklist | 1 to 2 person-days plus GPU or Tinker cost | none |
| 1. Sync the post-snapshot state | 13 source commits since Jul 24 (runtime floor in the submit tool, Anthropic token-count backport, Modal retry hardening, hint rewrites, checkpoint provenance) plus their instruction-tuned Llama 3.2 3B checkpoint and jinja templates | hundreds of lines plus one checkpoint | Odelya's Sep 2 fix PR (post_train_bench #11) | 2 to 4 person-days | Yulia as point of contact for the checkpoint |
| 2. Monitoring parity | Hierarchical, hidden-scratchpad, dataset and enriched monitors as monitor assets or Inspect scorers; compaction monitoring | ~2,500 plus ~3,900 | #1067 / #1318: 3,057 lines against Trajectory, closed unmerged after 8 weeks; #1869 makes the monitor an Inspect @scorer | 1 to 2 weeks for the monitors once Eval 2 monitors land; compaction monitoring 2 to 4 weeks and an architecture call (it imports inspect_ai.model._compaction, a private API, plus a version-pinned count_tokens patch) | Eval 2 monitor stack (#1822 to #1834, #1869) merged; Tyler or Miguel on where compaction monitoring lives |
| 3. Optional | tinker_tool agent tool (1,208); HCAST setting (15k, 4 of 11 tasks registered); offline pipeline | 1k to 30k | PostTrainBench itself: 23 commits, Jul 17 to Sep 2 | HCAST would be a second port on the PostTrainBench scale; the pipeline should not be ported | separate decision |
Relative to the sabotage-eval PRs: tier 2 is about their size (2.5k to 6.4k lines of monitor logic) and would hit the same failure mode if written against today's monitor APIs. Tiers 0 and 1 are an order of magnitude smaller and carry no architecture risk.
Open questions for the Sep 14 call with Yulia
- Does she know the setting is already in Control Tower with a leaderboard entry, and what would "displayed on the leaderboard" need to show for her team?
- Which monitors are load-bearing for the end-of-September results (that decides whether tier 2 is worth doing at all)?
- Would the new-models campaign run on Control Tower if tier 1 were done, and can she share the tuned Llama checkpoint and templates?
- Her question back: "can you please tell more about it?" on the embedded-maintainer model used for another back-integration.
Sources
- Repo analysis against a clone of
UKGovernmentBEIS/control-arena(2026-09-11); linuxarena/post_train_bench README ("Differences from Control Arena") and git log. docs/development/porting-settings.md,env-architecture.md, Eval 2 PR #1291, the open Eval 2 stack #1822 to #1834 and #1869.- Sep 9 Control Tooling Sync (all-in on Inspect decision) and Sep 11 standup chat (review lag attributed to the eval2 transition), in
linuxarena/meeting-transcripts. - Yulia Volkova's comments on the Sep 14 topics doc (2026-09-11) and her Sep 8 Slack message; the project poster.