Basis AI accounting agents run 6x faster

Basis agents run 5+ hour accounting workflows 6x faster, but Cursor post shows vague context can silently alter behavior without changing the final output.

S
StartupHub.ai Staff
2 min read
Engineer revising behavior spec Markdown in Cursor alongside AI agent trajectory for accounting workflow
Basis agents complete Form 1065 returns in 6-7 hours versus 30-40 human hours, with behavior specs judged in Cursor.· Cursor Blog

Cursor Blog outlines how Basis AI accounting agents have slashed the time spent on Form 1065 partnership returns, dropping them from 30 to 40 human hours down to 6 to 7 hours. That's roughly a 6x speedup, and today the agents are running inside 40% of the top 25 accounting firms.

The system was built on Cursor from day one. Basis agents can run 5+ hours on a single deliverable in the background, handling month-end close, corporate and partnership tax, and audit planning and fieldwork. They return review-ready workpapers.

Affected systems aren't servers in the usual sense. They're the agents themselves and the artifacts accountants have to sign off on.

There's no remote exploit here. What an attacker needs is write access to context: the prompts, skills, tool descriptions, instructions, and memory that the model reads at runtime.

How the attack on Basis AI accounting agents actually plays out

Traditional code runs the same way no matter how you format the file. Language models don't work that way. Wording and organization change the next action, so a vague sentence, a buried exception, or a misleading example can redirect research, calculations, and tool calls.

Think of it like editing footnotes in a recipe. The dish title stays the same, but the steps quietly change and the final plate still looks correct.

Braintrust and Basis tackle this with behavior specs. A spec is a Markdown file not shown to the agent. It defines when a behavior applies, what evidence to inspect, and what failure looks like. A judge then scores the recorded trajectory as true, false, or NA.

Why this matters and what still needs fixing

For builders, context is a production input and has to be read and revised like code. Cursor is where Basis does that work, with Markdown preview, model switching, and side-by-side editing of specs and runtime prompts. It's a loop co-founder Mitch Troyanovsky described as the difference between having an agent window and actually being able to inspect and change the context.

Outcome evaluation alone won't catch a correct return produced without primary authority or without preserving source. The mitigation is process evaluation against explicit specs, then fixing the runtime wording and re-running until behavior holds. But gaps remain: there's no cheap ground truth for many accounting outcomes, trajectory review is expensive, and no static patch prevents a future vague instruction from reintroducing drift.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
S

Written by

StartupHub.ai Staff

Editorial team

The staff writers of StartupHub.ai, ranging from investment analysts to avid AI tool users, early adopters and critical enthusiasts. Backgrounds span engineering, business and the arts. We hold every piece to rigorous standards of research and review.

Startups in this story

Profiles for the companies named above.