# Today in AI: GPT-6 Astra Reviews 41 Files in Minutes _Legora used GPT-6 Astra to complete a 41-document financial tie-out in minutes, catching all four planted errors._ **Published:** 2026-09-03 **Source:** https://www.startuphub.ai/ai-news/artificial-intelligence/2026/today-in-ai-gpt-6-astra-reviews-41-files-in-minutes --- Stockholm startup Legora ran 41 documents through [GPT-6 Astra](https://openai.com/index/legora-financial-statement-review-with-astra) in a single agent run and finished a financial statement tie-out in minutes that usually costs an evening. The test was not a demo file count but a full tie-out against trial balances, a consolidation schedule and prior year accounts, according to [OpenAI News](https://openai.com/index/legora-financial-statement-review-with-astra). ## Why 41 documents in one run matters Financial statement tie-out is grunt work that survives because mistakes are expensive. Legora legal engineer Percevale Perks describes it as work that can take an entire evening, sometimes days. The task is checking every figure in draft accounts against supporting schedules until each line agrees. With [Astra](/startups/astra), Legora's agent ingested all 41 documents at once, cross-checked every balance and logged each check for review. That scale matters because most legal AI demos still choke after ten files or lose citations halfway through. For builders the signal is context window plus structured output that auditors can actually audit. The gap is whether minutes holds when documents are scanned PDFs, not clean exports. ## How GPT-6 Astra scored on accuracy and completeness Legora tested [Astra](https://www.startuphub.ai/startups/astra) on its Benchmark for Agentic Reasoning, a suite of end-to-end legal tasks drawn from real workflows. On the tie-out workflow Astra scored nearly 40 percent higher than the previous model. Across the full BAR suite the average lift was only about 3 percent, which frames the headline number correctly. In the planted error test Astra caught all four planted breaks, including a £500,000 revenue note gap. It also retained every check the prior model got right and added around 50 additional correct checks. Perks calls that combination accuracy, completeness and reliability in one pass. That limitation is obvious: four synthetic errors do not equal a messy year end close with inconsistent naming. ## Human in the loop as Legora pushes beyond legal Legora is explicit that the agent does the exhaustive comparison but the lawyer makes the call. That human in the loop stance is now core to its push beyond contracts into audit, tax, compliance and risk. The company positions itself as an agentic operating system, not a point tool, used by more than 100,000 professionals across 1,800 legal departments in over 50 markets. That framing aligns with its recent moves to plug into enterprise stacks, from Intapp Walls to Google Cloud's Gemini Enterprise for Legal. It also lands just weeks after [OpenAI flagged critical cyber risks in the Astra model](/ai-news/artificial-intelligence/2026/openai-flags-critical-cyber-risks-in-astra-model), a reminder that broader agentic access raises exposure. For enterprises the temptation is to automate the first pass completely, but Legora's record every check design suggests it wants to be trusted as reviewer, not decider. The test ahead is whether clients will pay for that audit trail when the model itself keeps improving. ## Frequently Asked Questions ### What is a financial statement tie-out? It is the process of checking every figure in draft accounts against trial balances, consolidation schedules and prior year accounts until each line agrees. Legora says the work is tedious and error prone, which is why firms spend evenings on it. ### What is Legora's Benchmark for Agentic Reasoning? It is Legora's internal BAR suite that measures end to end performance on real legal workflows, not isolated Q and A. Legora reported a 40 percent gain on the tie-out task and about 3 percent on average across all BAR tasks. ### Does Astra replace lawyer judgment? No, Legora says Astra handles the exhaustive comparison and surfaces breaks while the legal professional retains final judgment on each result. --- Original analysis from [startuphub.ai](https://www.startuphub.ai), the #1 AI startup directory.