Steering LRMs Beyond Output Degradation

A new probe-based method, FPCG, distinguishes prediction from detection features to enable precise large reasoning models steering with minimal output quality degradation.

Abstract illustration of neural network nodes and data flow
Visualizing the internal representations of large reasoning models for controlled generation.
Visual TL;DR
LRM Output DegradationDriver
deployed large reasoning models often exhibit unpredictable behaviors
From the articleFPCG enables precise large reasoning models steering with remarkably little degradation in output quality, a significant improvement over previous methods.
Existing Steering MethodsDriver
rely on internal features detecting already generated text
From the article 5 mentionsPrior steering techniques inadvertently focused on features that signal existing behavior, which proved to be poor indicators of future actions.
Detection vs PredictionContext
distinguishing features that signal existing vs future behavior
From the article 4 mentionsThe core innovation presented by Kortukov, Komorowski, and colleagues in their arXiv preprint lies in identifying a critical distinction between detection and prediction features within LRM hidden states.
Activation ProbesCore
From the article 5 mentionsThis paper introduces activation probes trained to forecast future behavior likelihoods from intermediate reasoning steps.
Predicting Future BehaviorEffect
probes demonstrate significant accuracy from 64% to 91%
From the article 5 mentionsHowever, existing approaches often degrade output quality by relying on internal features that detect already generated text, rather than predicting future outcomes.
FPCG MethodCore
future probe controlled generation enables precise steering
From the article 5 mentionsFurthermore, FPCG demonstrates efficacy in steering scenarios where activation steering methods fail, underscoring its robustness and broader applicability.
Minimal Quality DegradationOutcome
From the articleFPCG enables precise large reasoning models steering with remarkably little degradation in output quality, a significant improvement over previous methods.

Deployed large reasoning models (LRMs) frequently exhibit unpredictable behaviors, a challenge that test-time steering methods have attempted to address. However, existing approaches often degrade output quality by relying on internal features that detect already generated text, rather than predicting future outcomes.

Unmasking Prediction Features for Control

The core innovation presented by Kortukov, Komorowski, and colleagues in their arXiv preprint lies in identifying a critical distinction between detection and prediction features within LRM hidden states. Prior steering techniques inadvertently focused on features that signal existing behavior, which proved to be poor indicators of future actions. This paper introduces activation probes trained to forecast future behavior likelihoods from intermediate reasoning steps. These probes demonstrate significant accuracy, ranging from 64% to 91%, in predicting the most probable behavior, thereby revealing a distinct set of internal prediction features.

Future Probe Controlled Generation: Precision Steering

Building upon these newly identified prediction features, the authors propose Future Probe Controlled Generation (FPCG). This novel text-level steering method enhances control by sampling multiple candidate sentences and selecting the optimal one based on a probe's prediction of future behavior likelihood. FPCG enables precise large reasoning models steering with remarkably little degradation in output quality, a significant improvement over previous methods. Furthermore, FPCG demonstrates efficacy in steering scenarios where activation steering methods fail, underscoring its robustness and broader applicability.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.