Fine-Tuning LLMs: A Costly "Tech Debt," Says Data Scientist

Lease End's senior data scientist Dan Bjornn argues that fine-tuning LLMs creates "tech debt," advocating for agentic frameworks with prompt engineering as a more efficient and accurate alternative.

9 min read
Dan Bjornn speaking on stage at AI Engineer World's Fair
AI Engineer

Visual TL;DR. Fine-tuning LLMs leads to Creates 'Tech Debt'. Creates 'Tech Debt' manifests as Calcification Tax. Fine-tuning LLMs replaced by Agentic Framework. Agentic Framework enables Rebuild for Efficiency. Rebuild for Efficiency achieves Improved Performance. Improved Performance informs Fine-tune Only Necessary.

  1. Fine-tuning LLMs: initial approach for Lease End's customer interaction application via text
  2. Creates 'Tech Debt': accumulating technical debt associated with fine-tuning models, costly to maintain
  3. Calcification Tax: fine-tuning models become rigid, difficult to update with new information
  4. Agentic Framework: prompt engineering with LLMs as a more efficient and accurate alternative
  5. Rebuild for Efficiency: Lease End rebuilt their system using the agentic framework for better performance
  6. Improved Performance: new system generated $12 million revenue with a 50x ROI, more accurate
  7. Fine-tune Only Necessary: Bjornn's verdict: fine-tune LLMs only when absolutely essential for specific tasks
Visual TL;DR
Visual TL;DR, startuphub.ai Fine-tuning LLMs leads to Creates 'Tech Debt'. Fine-tuning LLMs replaced by Agentic Framework leads to replaced by Fine-tuning LLMs Creates 'Tech Debt' Agentic Framework Improved Performance From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai Fine-tuning LLMs leads to Creates 'Tech Debt'. Fine-tuning LLMs replaced by Agentic Framework leads to replaced by Fine-tuning LLMs Creates 'TechDebt' Agentic Framework ImprovedPerformance From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai Fine-tuning LLMs leads to Creates 'Tech Debt'. Fine-tuning LLMs replaced by Agentic Framework leads to replaced by Fine-tuning LLMs initial approach for Lease End's customerinteraction application via text Creates 'Tech Debt' accumulating technical debt associatedwith fine-tuning models, costly tomaintain Agentic Framework prompt engineering with LLMs as a moreefficient and accurate alternative Improved Performance new system generated $12 million revenuewith a 50x ROI, more accurate From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai Fine-tuning LLMs leads to Creates 'Tech Debt'. Fine-tuning LLMs replaced by Agentic Framework leads to replaced by Fine-tuning LLMs initial approachfor Lease End'scustomer… Creates 'TechDebt' accumulatingtechnical debtassociated with… Agentic Framework prompt engineeringwith LLMs as a moreefficient and… ImprovedPerformance new systemgenerated $12million revenue… From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai Fine-tuning LLMs leads to Creates 'Tech Debt'. Creates 'Tech Debt' manifests as Calcification Tax. Fine-tuning LLMs replaced by Agentic Framework. Agentic Framework enables Rebuild for Efficiency. Rebuild for Efficiency achieves Improved Performance. Improved Performance informs Fine-tune Only Necessary leads to manifests as replaced by enables achieves informs Fine-tuning LLMs initial approach for Lease End's customerinteraction application via text Creates 'Tech Debt' accumulating technical debt associatedwith fine-tuning models, costly tomaintain Calcification Tax fine-tuning models become rigid, difficultto update with new information Agentic Framework prompt engineering with LLMs as a moreefficient and accurate alternative Rebuild for Efficiency Lease End rebuilt their system using theagentic framework for better performance Improved Performance new system generated $12 million revenuewith a 50x ROI, more accurate Fine-tune Only Necessary Bjornn's verdict: fine-tune LLMs only whenabsolutely essential for specific tasks From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai Fine-tuning LLMs leads to Creates 'Tech Debt'. Creates 'Tech Debt' manifests as Calcification Tax. Fine-tuning LLMs replaced by Agentic Framework. Agentic Framework enables Rebuild for Efficiency. Rebuild for Efficiency achieves Improved Performance. Improved Performance informs Fine-tune Only Necessary leads to manifests as replaced by enables achieves informs Fine-tuning LLMs initial approachfor Lease End'scustomer… Creates 'TechDebt' accumulatingtechnical debtassociated with… Calcification Tax fine-tuning modelsbecome rigid,difficult to update… Agentic Framework prompt engineeringwith LLMs as a moreefficient and… Rebuild forEfficiency Lease End rebuilttheir system usingthe agentic… ImprovedPerformance new systemgenerated $12million revenue… Fine-tune OnlyNecessary Bjornn's verdict:fine-tune LLMs onlywhen absolutely… From startuphub.ai · The publishers behind this format

Dan Bjornn, a senior data scientist at Lease End, a company specializing in connecting auto lease customers with financing options, shared a cautionary tale about the perceived benefits and hidden costs of fine-tuning large language models (LLMs). In a presentation titled "Your Fine-Tuned Model Is Tech Debt: A 50x ROI House of Cards," Bjornn detailed Lease End's journey from an initial RAG-based system to an LLM application that generated $12 million in revenue with a 50x ROI, only to reveal the accumulating technical debt associated with fine-tuning.

Fine-Tuning LLMs: A Costly "Tech Debt," Says Data Scientist - AI Engineer
Fine-Tuning LLMs: A Costly "Tech Debt," Says Data Scientist — from AI Engineer

The Initial Approach and Its Pitfalls

Lease End's first LLM application, built in late 2024, aimed to facilitate customer interactions via text, allowing them to ask questions, schedule calls, and receive reminders. The initial solution employed a workflow-based approach using a Retrieval Augmented Generation (RAG) system. This system searched a vector database of past messages to classify customer intent. Bjornn explained that while this approach worked, it struggled with the nuance of human language, leading to inaccuracies.

Driven by the need for better accuracy, lower costs, and reduced latency for real-time responses to thousands of daily messages, Bjornn championed the idea of fine-tuning. He believed fine-tuning was ideal for the narrow, structured task of classifying customer intent into six categories. Furthermore, he felt it would provide greater control over model providers and allow for model agnosticism. Bjornn described building a pipeline for data collection, LLM-as-judge classifications, manual review, and fine-tuning, calling it "a data scientist's dream." The success was evident, with the application contributing $12 million in revenue and a 50x ROI within a year.

The "Calcification Tax" of Fine-Tuning

However, Bjornn revealed that this success came with hidden costs, which he termed the "calcification tax." The fine-tuning process proved complex and time-consuming. It involved gathering and synthesizing examples, validating them, labeling them, and then iterating through fine-tuning and evaluation cycles. Crucially, Bjornn noted that they rarely achieved desired results on the first iteration, often fixing one problem only to create regressions elsewhere, turning the process into a "whack-a-mole" game that could take up to a week to deploy a fix.

This complexity led to "model lock-in," as switching between model providers or even versions within a provider required significant rework due to differences in data structure and interaction methods. The architecture also became rigid, preventing Lease End from adopting newer, more performant AI approaches. Bjornn highlighted two specific examples of how the fine-tuned model failed: "The Confused Confirmer," where the model mistakenly initiated a call after a customer confirmed an appointment, and "The Overeager Puppy," where the model immediately offered a call upon a simple greeting.

The "Aha Moment" and the Agentic Framework

The turning point came earlier this year when Lease End began using cloud code for their coding tasks. Bjornn observed that they could achieve better results simply by adjusting the context, skills, and resources passed to the models, without needing to fine-tune. This realization led him to question why the same principle couldn't be applied to their messaging application. Bjornn admitted it was difficult to admit this, as he had been a proponent of fine-tuning.

Fortunately, Lease End was able to integrate this new approach into an existing project. They migrated their workflow-based system to a series of skills, tools, and resources that could dynamically load context. This new "agentic framework" was tested as one of their first production applications.

The Rebuild: Efficiency and Improved Performance

Bjornn then compared the "before" and "after" of this rebuild. The "before" state involved a week-long, costly process for retraining and fixing issues, often requiring a critical mass of problems before addressing them. The "after" state, using the agentic framework, simplified the process dramatically. When a problem was identified, the fix involved adjusting the system prompt or the affected skill, validating performance, and then deploying the update by simply uploading MD files to an S3 bucket. This entire process, from discovery to deployment, was reduced to less than an hour, allowing for a far more reactive and customer-centric approach.

While the per-message API costs increased due to the use of better models, Bjornn stated that overall costs decreased due to the significant reduction in maintenance time. Accuracy also improved substantially compared to the fine-tuned model. The agentic framework's model-agnostic nature provided the desired freedom from vendor lock-in, allowing them to switch between providers like OpenAI and Anthropic as needed, with the crucial element being the quality of the provided context.

The Verdict: Fine-Tune Only When Necessary

Bjornn concluded by advising others to reconsider fine-tuning. He summarized the key reasons why they moved away from it:

  • Better Accuracy: The rebuilt system outperformed the fine-tuned model.
  • Lower Cost at Volume: While per-message costs increased, overall costs decreased due to reduced maintenance.
  • Lower Latency: Marginal gains from smaller fine-tuned models were negligible in practice.
  • Narrow/Structured Task: Even seemingly straightforward tasks like intent classification became "tech debt" when fine-tuned.
  • Vendor Control: Fine-tuning created lock-in, hindering flexibility.

Bjornn suggested that fine-tuning might still be relevant in specific scenarios requiring privacy and data control or offline solutions, but cautioned that these should be carefully evaluated against the "calcification tax" and potential long-term issues. His final advice: "Fine-tune only when you literally cannot call a frontier model, and even then, your decision still has to beat the tax."

© 2026 StartupHub.ai. All rights reserved. Do not enter, scrape, copy, reproduce, or republish this article in whole or in part. Use as input to AI training, fine-tuning, retrieval-augmented generation, or any machine-learning system is prohibited without written license. Substantially-similar derivative works will be pursued to the fullest extent of applicable copyright, database, and computer-misuse laws. See our terms.