Databricks Opens Up On-Prem Storage

Databricks unveils a new storage ecosystem using OpenSharing to connect its AI platform directly to on-premises data, eliminating migration needs.

Databricks logo with abstract network connections representing data storage.
Databricks announces its new storage ecosystem, connecting its platform to diverse data locations.
Visual TL;DR
Data Can't MoveDriver
sovereignty, gravity, latency prevent cloud migration
From the article 3 mentionsFurthermore, edge applications in retail, manufacturing, and telecommunications demand low-latency access to local data that cloud round-trips can’t provide.
Databricks Storage EcosystemCore
From the article 5 mentionsThe company today announced the Databricks storage ecosystem, a new approach designed to bring its Data Intelligence Platform directly to data residing in private clouds, on-premises data centers, and edge environments.
OpenSharing ProtocolCore
open-source standard for secure data connection
From the article 6 mentionsThe core of this new offering is the OpenSharing protocol, an open-source standard that enables Databricks to connect securely to these disparate data sources without requiring any data movement.
Launch PartnersContext
companies enabling this new storage approach
From the articleBy leveraging the OpenSharing protocol, storage partners can expose their data estates directly to Databricks' serverless compute, generative AI capabilities, and Large Language Models (LLMs).
No Data MigrationContext
eliminates need to move data from sources
From the article 9+ mentionsThe sheer economics of storing and transferring petabytes or exabytes of data also make cloud migration a non-starter for some, leading companies to even repatriate workloads.
Direct AccessEffect
Databricks platform accesses data where it lives
From the article 2 mentionsEverpure (formerly Pure Storage): In Private Preview, Everpure provides an OpenSharing connector for secure, gated access to object storage data within Databricks workspaces without replication.
Unlock On-Prem DataOutcome
enables AI on previously inaccessible datasets
Contents(4)

Databricks is breaking down the barriers that have kept vast amounts of enterprise data locked away on-premises. The company today announced the Databricks storage ecosystem, a new approach designed to bring its Data Intelligence Platform directly to data residing in private clouds, on-premises data centers, and edge environments.

StartupHub data

Companies working on this

Profiles of the companies named in this story, with founding year, headquarters, and a short description from our database.

Qumulo is a data storage and cloud infrastructure company focused on managing data at enormous scale for global businesses.

Founded
2012
Location
Seattle, Washington, USA
Valuation
$1.1B

A unified data analytics and AI platform built on the lakehouse architecture.

Founded
2013
Location
San Francisco, United States
Valuation
$190.0B

Defining the future of data management by reshaping storage.

Founded
2011
Location
Santa Clara, California, USA
Valuation
$11.0B

A universal data platform for AI and analytics, offering high-performance flash storage and a data lakehouse architecture.

Founded
2016
Location
New York, United States
Valuation
$30K

The core of this new offering is the OpenSharing protocol, an open-source standard that enables Databricks to connect securely to these disparate data sources without requiring any data movement. This addresses a critical pain point for organizations constrained by data sovereignty regulations, massive data gravity, or the latency demands of edge computing.

The 'Data That Can't Move' Problem

For years, the prevailing strategy was to migrate all data to the cloud. However, this approach is increasingly unfeasible for many enterprises.

Strict data sovereignty and regulatory mandates, like GDPR and HIPAA, often prohibit moving sensitive datasets. The sheer economics of storing and transferring petabytes or exabytes of data also make cloud migration a non-starter for some, leading companies to even repatriate workloads. Furthermore, edge applications in retail, manufacturing, and telecommunications demand low-latency access to local data that cloud round-trips can’t provide.

These aren't niche concerns; they represent a fundamental shift from "migrate everything" to "govern everything."

Enter the Databricks Storage Ecosystem

The new Databricks storage ecosystem is built to serve this evolving landscape. By leveraging the OpenSharing protocol, storage partners can expose their data estates directly to Databricks' serverless compute, generative AI capabilities, and Large Language Models (LLMs).

This means organizations can now run Databricks workloads, including training AI models on classified data or analyzing network telemetry, directly on their existing on-premises infrastructure. The benefit is immediate: isolated data becomes active, AI-ready assets without the cost, complexity, and compliance risks associated with data migration or duplication.

This integration provides a unified catalog across hybrid environments, allowing customers to use Databricks Serverless Compute, Genie, and AgentBricks to query and reason over data that never leaves its original location. This is not a future vision; these integrations are available today.

Key Launch Partners

Databricks has launched with support from several major storage providers:

  • MinIO: Offering General Availability, MinIO AIStor natively implements OpenSharing, connecting Databricks to on-premises Apache Iceberg™️ and Delta tables under Unity Catalog governance.
  • Everpure (formerly Pure Storage): In Private Preview, Everpure provides an OpenSharing connector for secure, gated access to object storage data within Databricks workspaces without replication.
  • Qumulo: Slated for Private Preview in July 2026, Qumulo's NeuralSearch integrates with OpenSharing to allow secure, non-replicating sharing of data across core, cloud, and edge environments.
  • VAST Data: Expected in Private Preview in August 2026, VAST Data's AI Operating System will support OpenSharing, bridging Databricks workflows with hybrid infrastructure data without massive movement.

This initiative marks a significant step in Databricks' strategy to unify data governance across the entire enterprise data estate, regardless of where the data physically resides.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.

More from Daniel Singer