# Databricks Speeds Up Open-Source LLMs _Databricks enhances open-source LLM performance with automatic prompt caching, reducing latency and boosting throughput without user configuration._ **Updated:** 2026-08-22 **Published:** 2026-05-22 **Source:** https://www.startuphub.ai/ai-news/technology/2026/databricks-speeds-up-open-source-llms --- Databricks is rolling out automatic prompt caching for open-source large language models (LLMs) on its platform. This feature, previously available for proprietary models, aims to accelerate LLM inference by reusing identical prompt prefixes across requests. According to [Databricks](https://www.databricks.com/blog/accelerating-llm-inference-prompt-caching-open-source-models-databricks), this can dramatically cut down on wasted compute cycles, reduce latency, and increase overall throughput. Repeated LLM PromptsDriver common in chatbots and batch tasks, leading to wasted compute cyclesFrom the article 3 mentionsRepeated prompts are common in many LLM applications, from chatbots using consistent system messages to batch processing tasks with identical initial instructions.Databricks PlatformCoreFrom the article 6 mentionsDatabricks is rolling out automatic prompt caching for open-source large language models (LLMs) on its platform.No User ConfigurationContextfeature works automatically without requiring manual setupFrom the article 2 mentionsThis includes models like GPT-OSS 20B and 120B, Gemma 3 12B, and various Llama 3.1 and 3.3 configurations.Automatic Prompt CachingContextreuses identical prompt prefixes across requests, skipping prefill stageFrom the article 5 mentionsThe caching is entirely automatic; users do not need to configure any settings for Databricks Prompt Caching to function, similar to how other solutions like Tensormesh exits stealth with $4.5M to slash AI inference caching costs operate.Feature AvailabilityContextnow extends to open-source LLMs, not just proprietary onesFrom the article 2 mentionsThis feature, previously available for proprietary models, aims to accelerate LLM inference by reusing identical prompt prefixes across requests.leads toReduced LatencyEffectfaster response times for LLM inference, improving user experienceFrom the article 3 mentionsAccording to Databricks, this can dramatically cut down on wasted compute cycles, reduce latency, and increase overall throughput.Increased ThroughputEffectFrom the article 3 mentionsThis directly translates to lower latency and higher throughput, allowing more tokens to be processed per unit of compute. The core idea behind prompt caching is simple: why reprocess the same initial instructions or system prompts repeatedly? When a prompt prefix matches a cached entry, the LLM can skip the initial computation phase, known as the 'prefill' stage. This directly translates to lower latency and higher throughput, allowing more tokens to be processed per unit of compute. ## Why It Matters Repeated prompts are common in many LLM applications, from chatbots using consistent system messages to batch processing tasks with identical initial instructions. Without caching, these repeated computations are a significant performance bottleneck. Prompt caching allows for the cost of long, domain-specific system prompts to be amortized across many queries, effectively boosting model quality in specific contexts without a proportional increase in inference cost. This is particularly relevant as research shows open-source models can now rival proprietary models in enterprise tasks through techniques like automated prompt optimization. ## Feature Availability and Security Databricks has extended its built-in prompt caching to a range of open-weights models available through their Foundation Model APIs (FMAPIs). This includes models like GPT-OSS 20B and 120B, Gemma 3 12B, and various Llama 3.1 and 3.3 configurations. The feature is available for batch inference, pay-per-token, and provisioned-throughput workloads, and it implicitly powers higher-level services like Agent Bricks and Genie. Security remains a priority, with prompt caches isolated to volatile memory and never persisted. The caching is entirely automatic; users do not need to configure any settings for Databricks Prompt Caching to function, similar to how other solutions like [Tensormesh exits stealth with $4.5M to slash AI inference caching costs](/ai-news/ai-research/2025/tensormesh-exits-stealth-with-45m-to-slash-ai-inference-caching-costs) operate. ## Real-World Performance Gains In early production tests on GPT-OSS models, Databricks observed substantial improvements. One large-scale batch inference pipeline saw a 2.5x increase in per-replica input-token throughput and a 3x reduction in P50 latency, even with a relatively modest cache hit ratio of 30%. This demonstrates the tangible impact of efficient caching. By automatically reusing KV caches for identical prompts, Databricks enables faster, more cost-effective, and secure operation of open-source LLMs. This enhancement can significantly improve inference pipelines for various applications, from real-time chat to large-scale document processing. --- Original analysis from [startuphub.ai](https://www.startuphub.ai), the #1 AI startup directory.