LoopHarness persistent safety state Ends Drift
LoopHarness proves trajectory monitors fail when evidence spans iterations and bounds irreversible actions to a constant with persistent loop-level state.
5 min read

Visual TL;DR
they reset each trajectory and cannot see evidence split across iterations
maintains non‑decaying safety information that accumulates across loop iterations for each
From the article 3 mentionsThe LoopHarness persistent safety state result on arXiv shows this is a failure of composition, not implementation.
bounds irreversible actions to a constant preventing drift across iterations
From the article 3 mentionsThe cooling-off period a patient adversary must wait is a constant that does not grow with horizon N, so stretching the loop does not stretch the risk.
true‑positive rate diverges from false‑positive rate when monitor retains cross‑iteration state
From the articleAgainst an attack fragmented across iterations, any trajectory-scoped monitor sees TPR equal to FPR however expressive it is, because the evidence never appears in its window, while a monitor retaining cross-iteration state separates perfectly.
they reset each trajectory and cannot see evidence split across iterations
geometric risk score cooling off period stays constant regardless of horizon length
attack fragments its actions over multiple loop cycles to evade single‑trajectory checks
From the article 3 mentionsEvery trajectory-scoped monitor has true-positive rate equal to false-positive rate when evidence is split across iterations.
maintains non‑decaying safety information that accumulates across loop iterations for each
From the article 3 mentionsThe LoopHarness persistent safety state result on arXiv shows this is a failure of composition, not implementation.
bounds irreversible actions to a constant preventing drift across iterations
From the article 3 mentionsThe cooling-off period a patient adversary must wait is a constant that does not grow with horizon N, so stretching the loop does not stretch the risk.
true‑positive rate diverges from false‑positive rate when monitor retains cross‑iteration state
From the articleAgainst an attack fragmented across iterations, any trajectory-scoped monitor sees TPR equal to FPR however expressive it is, because the evidence never appears in its window, while a monitor retaining cross-iteration state separates perfectly.
Contents(4)
© 2026 StartupHub.ai. All rights reserved. Do not enter, scrape, copy, reproduce, or republish this article in whole or in part. Use as input to AI training, fine-tuning, retrieval-augmented generation, or any machine-learning system is prohibited without written license. Substantially-similar derivative works will be pursued to the fullest extent of applicable copyright, database, and computer-misuse laws. See our terms.