Uber Optimizes Data Exports
Uber slashes data export costs by combining Hudi column stats and table sorting, preserving cloud storage tiering.

Visual TL;DR
DSARs and compliance requests trigger full-table scans on massive historical archives
From the article 2 mentionsConsequently, even with a secondary index, the query engine still had to inspect a substantial set of files, incurring high metadata and file-scan costs.
leverages Apache Hudi's built-in metadata to quickly identify relevant data files
From the article 4 mentionsColumn stats provide the mechanism for pruning, and sorting enhances its precision.
organizes data physically on disk, reducing the need to scan irrelevant blocks
From the article 7 mentionsThe solution, detailed on the Uber Engineering blog, combines Apache Hudi™ column statistics with table sorting.
even selective queries on cold data inflate storage, retrieval, and egress costs
From the article 2 mentionsHowever, export workloads disrupt this by keeping all data "hot" through repeated scans.
not suitable for broad historical scope and small, specific record lookups
From the articleWhile secondary indexes can identify candidate files, they proved insufficient for Uber's specific export workloads.
Hudi column stats and table sorting work together to optimize data access
avoids rehydrating cold data, keeping costs low on Google Cloud Storage
From the articleKey lessons from the initiative include: access patterns dictate storage costs; file layout is as important as storage tiering; metadata-driven pruning is essential; and a one-time rewrite cost can yield long-term savings.
significantly reduces storage, retrieval, and egress expenses for export workloads
From the articleUber engineers have developed a method to significantly reduce the cost of running specific data export workloads.
Contents(5)
© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Written by
Daniel SingerEditor, StartupHub.ai
Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.