Iceberg Unites Snowflake and Google Cloud

Snowflake and Google Cloud are integrating via Apache Iceberg, creating a federated, AI-ready data lakehouse with unified governance and programmatic access for intelligent agents.

Abstract representation of interconnected data clouds for Snowflake and Google Cloud
Apache Iceberg facilitates data interoperability between Snowflake and Google Cloud platforms.· Snowflake
Visual TL;DR
Data SilosDriver
From the article 9+ mentionsThis move aims to dissolve data silos by allowing different analytics engines to access the same data without costly duplication.
Apache IcebergCore
open table format developed at Netflix, now an Apache Software Foundation project
From the article 6 mentionsSnowflake and Google Cloud are pushing for a borderless data lakehouse, betting on Apache Iceberg as the common language to bridge their respective platforms.
Snowflake + Google CloudCore
integrating platforms to create a federated, AI-ready data lakehouse
From the article 9+ mentionsStartupHub.ai data shows Snowflake with a score of 72/100, while Google Cloud scores 17/100, highlighting Snowflake's stronger position in data management infrastructure.
Unified GovernanceContext
managing metadata, access control, and concurrent writes across platforms
From the article 3 mentionsThis is where the catalog comes in, acting as the lakehouse's governance layer.
Compute to DataContext
From the article 9+ mentionsThe core idea, rooted in the principle of data locality, is to bring compute to the data rather than moving massive data volumes.
Border-less Data LakehouseOutcome
dissolving data silos by allowing different analytics engines to access same data
From the article 4 mentionsSnowflake's Model Context Protocol (MCP) and Google Cloud's BigQuery MCP server provide universal interfaces for AI agents to query lakehouse data programmatically.
Snowflake StrongerContext
From the article 9+ mentionsStartupHub.ai data shows Snowflake with a score of 72/100, while Google Cloud scores 17/100, highlighting Snowflake's stronger position in data management infrastructure.
AI-Ready DataEffect
programmatic access for intelligent agents to derive actionable insights
From the article 9+ mentionsThe outcome is an AI-ready open lakehouse where data is interoperable, governed, semantically rich, and accessible to both humans and AI agents.
Contents(4)

Snowflake and Google Cloud are pushing for a borderless data lakehouse, betting on Apache Iceberg as the common language to bridge their respective platforms. This move aims to dissolve data silos by allowing different analytics engines to access the same data without costly duplication.

The core idea, rooted in the principle of data locality, is to bring compute to the data rather than moving massive data volumes. Apache Iceberg, an open table format developed at Netflix and now an Apache Software Foundation project, serves as this shared language. It's supported by engines like Spark, Trino, Flink, BigQuery, and Snowflake, eliminating the need for proprietary adapters. StartupHub.ai data shows Snowflake with a score of 72/100, while Google Cloud scores 17/100, highlighting Snowflake's stronger position in data management infrastructure.

The Catalog: Governance for Open Data

While an open format solves interoperability, managing metadata, access control, and concurrent writes requires a central authority. This is where the catalog comes in, acting as the lakehouse's governance layer. It tracks tables, schemas, access policies, and file locations.

To enable cross-vendor compatibility, catalogs need a standardized integration protocol. The Iceberg REST Catalog (IRC) specification provides this, allowing engines to communicate with any catalog. This standard, combined with vended credentials, enables federation, where one platform's query can access data managed by another's catalog. For example, Snowflake queries can reach into Google Cloud's Lakehouse catalog, and vice versa, with each catalog issuing time-limited access tokens.

Federation: Bridging Lakehouses

This federated model allows for bidirectional access to Iceberg tables across different catalogs. Organizations can choose between self-managed catalogs, like Apache Polaris, or managed solutions. Most opt for managed catalogs to offload operational overhead.

Google Cloud's Lakehouse runtime catalog is a serverless metastore supporting multiple engines. Snowflake Horizon Catalog integrates Polaris and offers enterprise governance features like RBAC and data lineage. These managed catalogs can coexist via federation. Snowflake connects to Google Cloud's catalog via a Catalog-Linked Database (CLD), presenting Google Cloud's Iceberg tables as native Snowflake tables. Conversely, Google Cloud services connect to Snowflake's Horizon catalog via its IRC endpoint.

From Open Lakehouse to AI-Ready

Interoperability is only the first step. Making data useful for AI requires addressing gaps in programmatic access, semantic grounding, and contextual intelligence.

Snowflake's Model Context Protocol (MCP) and Google Cloud's BigQuery MCP server provide universal interfaces for AI agents to query lakehouse data programmatically. This ensures AI applications, like Gemini Enterprise, can access governed data through standardized protocols.

Semantic models are crucial to prevent AI hallucinations by defining business logic, metrics, and relationships at the data layer. Snowflake's Semantic View Autopilot and Google Cloud's Universal Semantic Layer aim to automate the creation and maintenance of these models. This ensures AI systems inherit consistent, correct business logic, grounding their responses in factual data.

Contextual intelligence adds surrounding knowledge, descriptions, data quality signals, and usage patterns, to help AI interpret data correctly. Snowflake Horizon Context and Google Cloud's Knowledge Catalog provide this layer, ensuring AI agents operate with trusted information.

Agentic AI: Actionable Insights

With governed data and context, AI agents can act with high accuracy. Snowflake's CoWork and CoCo agents offer personal productivity and development workflow assistance, respectively. They leverage semantic grounding and governance across human and AI actions.

Google Cloud's Gemini, with its multimodal understanding, can process structured and unstructured data from the lakehouse. Gemini Enterprise connects via MCP, routing queries through AI agents and semantic models for grounded responses. The outcome is an AI-ready open lakehouse where data is interoperable, governed, semantically rich, and accessible to both humans and AI agents.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.

More from Daniel Singer