#Data Engineering
50 articles with this tag

Databricks Lakehouse: The AI Context Layer
Databricks' Data Hub creates a governed AI context layer by unifying R&D data, prioritizing context coverage as a quality metric for both human and AI users.

Dow's Carbon Ledger Accelerates Sustainability
Dow leverages Databricks to create a Carbon Footprint Ledger, drastically cutting PCF calculation time and boosting sustainability transparency.

Databricks Lakehouse: Unified Data AI Platform
Databricks' Lakehouse Architecture unifies data warehousing and data lakes for analytics and AI, emphasizing governance and open sharing.

Dmitry Petrov: Bridging Agents and Physical Data
Dmitry Petrov of DataChain explains how specialized 'data harnesses' are crucial for enabling AI agents to effectively process unstructured physical data like videos and sensor logs.

Pinterest's Medic AI Tames Spark Failures
Pinterest's Drasko Profirovic details Medic, an AI agent designed to diagnose and fix Apache Spark job failures, highlighting the evolution from prototype to a sophisticated multi-agent system.

Databricks Builds AI Soccer Coach
Databricks' new 'La Pizarra' app turns massive soccer match data into real-time tactical insights, leveraging AI for scouting and opponent analysis.

Spark 4.2: AI-Native Analytics Arrive
Apache Spark 4.2 integrates AI-native analytics, enhances remote access via Spark Connect, and streamlines data processing with Auto CDC.

Netflix's Real-Time Service Map Pipeline
Netflix’s real-time service map pipeline uses a streaming-first, three-stage architecture to visualize complex service dependencies at scale.

UltraX: Redefining LLM Data Refinement
UltraX redefines LLM data refinement by introducing function-calling for fine-grained editing, achieving superior performance with fewer training tokens.

Synapse to Databricks: The Migration Playbook
Moving from Azure Synapse to Databricks offers a unified Lakehouse, streamlining analytics and AI workloads while cutting costs and complexity.

Databricks Auto Upgrades debut
Databricks Auto Upgrades automates the deployment of new lakehouse table features, enhancing performance and reliability without manual intervention.

Databricks Vibe Auto-Generates Data Models
Databricks unveils Vibe Data Modeling, an LLM agent that auto-generates Silver-layer data models from plain English in hours, not months.
Databricks tags dbt pipelines
Databricks Query Tags enhance dbt pipelines with granular cost attribution and performance insights, making resource usage transparent.
Databricks redefines databases
Databricks' new LTAP architecture decouples database storage and compute, enabling real-time analytics on fresh data without impacting transactional workloads.

Data Platform, Not Model, Drives Legal AI
The data platform, not the model, is the key to effective and responsible legal AI, enabling context-aware reasoning and compounding institutional intelligence.

Snowflake Streamlines Real-Time Data
Snowflake's Snowpipe Streaming, now enhanced with AI coding agent CoCo, simplifies real-time data ingestion and analysis.

Snowflake Embraces AI Agent Discovery
Snowflake is adopting the open Agentic Resource Discovery Specification to make AI agents easily discoverable and usable across organizations.
Databricks Free Edition Gets Major Upgrade
Databricks Free Edition now includes Genie Code, serverless GPUs, Lakebase, Agent Bricks, and Lakeflow Designer, offering a complete toolkit for data and AI projects.
Databricks Expands Genie Code
Databricks unveils major updates for Genie Code at Data + AI Summit 2026, including a new command center and scheduled autonomous tasks.
Databricks Refines Partner Framework
Databricks updates its Partner Well-Architected Framework with AI-ready guidance, Dev Kit, and open-source Firefly to accelerate partner innovation.
Data Pipeline Architecture Explained
Understand the core layers, common patterns like ELT and Medallion, and best practices for building robust data pipelines.
Databricks Automates Data Ops
Databricks launches Genie ZeroOps, an AI agent that automates data and AI operations, monitoring, diagnosing, and fixing production issues securely.
Databricks Automates Data Ops with Genie ZeroOps
Databricks introduces Genie ZeroOps for automated data operations, alongside expanded connectivity and real-time processing capabilities.
Databricks Indexes Speed Up Text Search
Databricks introduces beta full-text search indexes to accelerate text queries on large datasets by up to 100x, without application changes.
Azure Databricks embraces agentic era
Databricks unveils major Azure updates, integrating AI agents into productivity tools, real-time data processing, and a new CDP for the agentic era.

Snowflake AI Platform Aims for Easier Data Migration
Snowflake launches AIM, an AI-powered platform simplifying enterprise data migration and modernization with dual paths for full transformation or rapid virtualization.
Databricks Rethinks Data Migration
Databricks advocates for a new approach to data migration, prioritizing parallel progress and early value realization over traditional phased methods.

AI Reshapes Data Engineering
AI is fundamentally redefining data engineering, shifting roles from manual tasks to strategic oversight and enabling agentic AI to construct complex data pipelines.
Databricks Lakebase Branches Up Databases
Databricks Lakebase introduces copy-on-write database branching, making isolated developer environments a reality and reshaping the DBA role.
Databricks Touts AI Outcome Acceleration
Databricks launches its Forward Deployed Engineering (FDE) organization to accelerate customer AI outcomes through embedded engineering and a unified platform.
Databricks Hits Petabyte Scale Ingest
Databricks Zerobus Ingest achieves petabyte-scale data ingestion at 12 GB/s per table, eliminating infrastructure management.
Databricks Taps Students for AI Future
Databricks launches its first Student Fellows program, empowering students to lead in data and AI through campus initiatives and real-world application.

Uber's Data Abstraction Layer
Uber's Data Abstraction Layer (DAL) simplifies data access, drastically cutting report generation time and enabling more sophisticated advertiser tools.

Snowflake Supercharges AI Data Engineering
Snowflake enhances its platform with AI-driven tools for data engineering, aiming to accelerate pipeline creation and improve reliability.

Snowflake Simplifies Python Deployment
Snowflake's CoCo agent now streamlines the deployment of Snowpark Python pipelines with a single prompt, simplifying production workflows for data engineers.
Databricks Unlocks Database Evolution
Databricks Lakebase's new database branching capabilities make evolutionary database development principles a reality at scale.
Spark's Real-Time Mode Powers Gaming
Databricks' Apache Spark Real-Time Mode with transformWithState now enables sub-second latency for gaming session tracking, eliminating complex architectures.

Snowflake Streams for Real-Time AI
Snowflake enhances its platform with Datastream for Kafka-compatible streaming, AI-powered tools, and expanded data integration capabilities to fuel agentic AI.

Snowflake's Adaptive Compute
Snowflake's new Adaptive Compute technology dynamically scales resources for data workloads, promising higher performance and reduced operational complexity.

Snowflake CoCo Goes Everywhere
Snowflake's AI coding agent, CoCo, is expanding beyond its data cloud with desktop, mobile, and Slack integrations, aiming to embed governed AI development everywhere.
Databricks Lakebase: Database Branching Reimagined
Databricks Lakebase's new database branching feature allows developers isolated, production-like environments, streamlining database evolution.
Databricks Shines at SIGMOD 2026
Databricks' Enzyme engine and Spark Declarative Pipelines are showcased at SIGMOD 2026, simplifying complex data engineering tasks.

Cloudflare's AI Data Agent
Cloudflare unveils Town Lake, a unified data platform, and Skipper, its AI agent for natural language data querying, enhancing internal data access and governance.
Databricks Unifies Operational Data
Databricks' new Lakebase Change Data Feed simplifies operational data integration into the Lakehouse, enabling direct streaming and unified governance.
Octopus Energy Slashes Costs 50x
Octopus Energy slashed data engineering costs by 50x to meet UK's MHHS regulation, processing 98.8% fewer data rows and improving freshness.
LinkedIn Sales Navigator Search Speed Boost
LinkedIn engineers drastically cut Sales Navigator's search data processing time by optimizing its Spark pipeline, enabling faster results for users.
Databricks Unifies Data, Analytics, and AI
Databricks aims to simplify data operations with its unified Lakehouse Platform, integrating data warehousing, analytics, and AI development.

Snowflake Turbocharges dbt with Fusion
Snowflake integrates dbt Fusion, boosting compilation speeds for complex projects with no added cost or reconfiguration required.

Databricks Launches Analytics Engineer Path
Databricks launches a new learning pathway for SQL practitioners to become analytics engineers, covering data modeling, pipelines, and AI agent deployment.

Snowflake Simplifies Data Pipelines
Snowflake introduces DCM Projects and Cortex Code for declarative data pipelines, simplifying workflow management and reducing manual coding.