Governing LLM Reasoning with Formal Verification

EG-VAR introduces a Lean 4-based architecture for auditable LLM reasoning, achieving perfect accuracy and source fidelity on benchmarks by using formal verification as the sole claim issuer.

Diagram illustrating the EG-VAR architecture with the Lean kernel at its core.
EG-VAR: Evidence-Grounded Verified Agentic Reasoning Architecture.
Visual TL;DR
LLM Reasoning GapsDriver
outputs lack verifiable evidence, logical soundness, and provenance in high-stakes scenarios
EG-VAR SystemCore
introduces a Lean 4-based architecture for auditable LLM reasoning and verified claims
From the article 5 mentionsThe EG-VAR (Evidence-Grounded Verified Agentic Reasoning) system addresses this by leveraging the Lean 4 formal verification kernel as the exclusive minter of verified claims.
Lean Kernel ArbiterCore
exclusive minter of verified claims, ensuring structural guarantee of evidence descent
From the article 2 mentionsThrough tool-attestation axioms and declared source lifts, every verified output is structurally guaranteed to descend from an attested tool call and a chain of inference validated by the Lean kernel.
Tool-Attestation AxiomsContext
From the articleThrough tool-attestation axioms and declared source lifts, every verified output is structurally guaranteed to descend from an attested tool call and a chain of inference validated by the Lean kernel.
Abstain & Audit TrailOutcome
From the articleOutputs that cannot meet these stringent criteria are designated as 'Abstain' and are accompanied by a replayable audit trail, ensuring transparency.
Perfect AccuracyEffect
achieves perfect accuracy and source fidelity on benchmarks through formal verification
From the articleOn a subset of TableBench numerical reasoning tasks (n=120), it achieved a perfect 120/120 score, starkly contrasting with a 95% success rate for a same-tool baseline.
High-Stakes ApplicationEffect
enables LLM use in critical domains by ensuring verifiable, evidence-based reasoning
From the article 2 mentionsThis lack of verifiable provenance limits their application in high-stakes scenarios.
Contents(4)

The proliferation of LLMs in empirical reasoning tasks is hampered by a fundamental governance gap: tool access alone does not guarantee that outputs are evidence-based or logically sound. Claims may not descend from attested evidence, nor do deductions necessarily hold up under formal scrutiny. This lack of verifiable provenance limits their application in high-stakes scenarios.

Companies working on this

StartupHub profiles of the companies this article names, with funding and a one-liner from our database.

OpenAI
$852.0B
OpenAI is an AI research and deployment company dedicated to ensuring that artificial general intelligence benefits all of humanity.
Prometheus AI
$38.0B
Jeff Bezos-backed AI startup building physical AI systems for engineering and manufacturing.
Thinking Machines Lab Inc.
$12.0B
Frontier AI lab building multimodal AI systems with customizable, collaborative capabilities for enterprise applications, focusing on human-AI collaboration and personalized AI.
runway gen 5
$5.3B
Runway AI develops generative AI tools for creating videos, images, and multimedia content, focusing on video intelligence and world models.

The Lean Kernel as Sole Arbiter of Verified Claims

The EG-VAR (Evidence-Grounded Verified Agentic Reasoning) system addresses this by leveraging the Lean 4 formal verification kernel as the exclusive minter of verified claims. Through tool-attestation axioms and declared source lifts, every verified output is structurally guaranteed to descend from an attested tool call and a chain of inference validated by the Lean kernel. Outputs that cannot meet these stringent criteria are designated as 'Abstain' and are accompanied by a replayable audit trail, ensuring transparency.

Uncompromising Accuracy and Source Fidelity Under Stress

EG-VAR demonstrates a significant leap in reliability. On a subset of TableBench numerical reasoning tasks (n=120), it achieved a perfect 120/120 score, starkly contrasting with a 95% success rate for a same-tool baseline. Crucially, in counterfactual stress tests across five domains and two models, EG-VAR maintained 100% source fidelity, while the same-tool baseline faltered to 80-90% (and a no-tool approach scored only 50-80%). The system also achieves low semantic-formalization error rates when using LLMs like Sonnet (3.3%) and Opus (1.7%) as deployment-time formalizers.

A Formal Sidecar for High-Stakes Empirical Claims

EG-VAR is positioned as a critical technical-governance interface for high-stakes AI applications. By acting as a formal sidecar, it makes the target proposition, source scope, evidence boundary, proof obligation, and abstention condition explicitly auditable. This eliminates unsupported verified outputs today and transforms potential issues like formalization errors, disputes over source authority, ambiguities, and abstentions into explicit targets for auditing and improvement. The long-term vision involves integrating these typed sidecars into datasets, APIs, and public records, creating reusable infrastructure to amortize the formalization burden.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.

More from Daniel Singer