LoopHarness persistent safety state Ends Drift
LoopHarness proves trajectory monitors fail when evidence spans iterations and bounds irreversible actions to a constant with persistent loop-level state.

Visual TL;DR
they reset each trajectory and cannot see evidence split across iterations
maintains non‑decaying safety information that accumulates across loop iterations for each
From the article 3 mentionsThe LoopHarness persistent safety state result on arXiv shows this is a failure of composition, not implementation.
bounds irreversible actions to a constant preventing drift across iterations
From the article 3 mentionsThe cooling-off period a patient adversary must wait is a constant that does not grow with horizon N, so stretching the loop does not stretch the risk.
true‑positive rate diverges from false‑positive rate when monitor retains cross‑iteration state
From the articleAgainst an attack fragmented across iterations, any trajectory-scoped monitor sees TPR equal to FPR however expressive it is, because the evidence never appears in its window, while a monitor retaining cross-iteration state separates perfectly.
they reset each trajectory and cannot see evidence split across iterations
geometric risk score cooling off period stays constant regardless of horizon length
attack fragments its actions over multiple loop cycles to evade single‑trajectory checks
From the article 3 mentionsEvery trajectory-scoped monitor has true-positive rate equal to false-positive rate when evidence is split across iterations.
maintains non‑decaying safety information that accumulates across loop iterations for each
From the article 3 mentionsThe LoopHarness persistent safety state result on arXiv shows this is a failure of composition, not implementation.
bounds irreversible actions to a constant preventing drift across iterations
From the article 3 mentionsThe cooling-off period a patient adversary must wait is a constant that does not grow with horizon N, so stretching the loop does not stretch the risk.
true‑positive rate diverges from false‑positive rate when monitor retains cross‑iteration state
From the articleAgainst an attack fragmented across iterations, any trajectory-scoped monitor sees TPR equal to FPR however expressive it is, because the evidence never appears in its window, while a monitor retaining cross-iteration state separates perfectly.
Contents(4)
© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Written by
Daniel SingerEditor, StartupHub.ai
Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.