Iceberg vs Delta Lake: Choosing a Table Format

Iceberg vs Delta Lake: features have converged, so CDC tooling and catalog lock-in now decide the format. Compare commits, deletes, and upkeep.

Summarize with AI:

The Iceberg vs Delta Lake decision stopped being a feature comparison once Iceberg spec v3 and Delta Lake 4.x converged on the same deletion-vector encoding. What still separates them is where lock-in lives. Your CDC tooling can rule a format out before any feature matters, and the catalog you pair with it decides whether your policies and lineage move with your bytes or stay behind. Teams that pick on features and defer those two questions end up paying for them at migration time, when tables are large and pipelines are already in production.

TL;DR

  • Both formats now store Parquet data files and share a RoaringBitmap deletion-vector encoding, so the feature checklist rarely decides the choice.
  • Your CDC tooling can rule out a format before you start comparing features.
  • The catalog is the second lock-in surface: Delta plus managed Unity Catalog ties format to policy, while Iceberg plus Apache Polaris keeps them separable.
  • Databricks teams running CDC upserts must choose between efficient Delta writes with deletion vectors and UniForm access for Iceberg readers, because one table cannot have both.

Try Airbyte Flex

How Do Iceberg and Delta Lake Differ Under the Hood?

Iceberg and Delta still diverge in how they commit, how much metadata they carry, and how they change layout after the fact. Apache Iceberg 1.11.0 shipped spec v3 as production-ready on May 19, 2026, with deletion vectors, VARIANT, and row lineage. Iceberg tracks table state in a metadata tree and commits through a catalog pointer swap, while Delta Lake, created by Databricks and now governed by the Linux Foundation, tracks state in a transaction log.

Commit Protocol Decides Multi-Writer Safety

Iceberg checks for data conflicts before commit, then has the catalog perform an atomic Compare-And-Swap (CAS) of the metadata pointer. Universally Unique Identifier (UUID) names on metadata files mean no writer overwrites another's, and no object store PutIfAbsent is required, as Jack Vanlightly's consistency analysis traces step by step.

Delta appends JSON records to _delta_log, and before version 4.0, that append depended on a Java Virtual Machine (JVM) lock on a single Spark driver. This architecture prevented optimistic concurrency across clusters until Delta Lake 4.0 (June 9, 2025) shipped coordinated commits, backed by a DynamoDB commit coordinator. Delta Lake 4.1.0 (March 1, 2026) added conflict-free feature activation, so activating deletion vectors or column mapping no longer blocks concurrent writers.

Metadata Shape Sets the Planning Cost

Iceberg's immutable Avro manifests carry partition stats, so a planner skips whole manifests without opening them. The tree costs disk space: DataVidhya's format comparison reports 2–4 GB of Iceberg metadata for hourly commits over 12 months, versus 200–500 MB for Delta's log for the same activity.

Delta keeps its footprint small with Parquet and Delta 3.0 V2 checkpoints. Both formats record row history: Iceberg v3 stamps _row_id and _last_updated_sequence_number on each row, and Delta exposes Change Data Feed. Which one your engines read shapes how you carry lineage across deployments.

Layout and Schema Changes Follow Different Models

Iceberg stores partition functions in metadata, so switching from day(ts) to hour(ts) is a metadata-only spec change, and both specs coexist in one table with no rewrite.

Liquid Clustering lets you redefine Delta clustering keys without an immediate rewrite, but it reorganizes data through background compute and is generally available (GA) on Databricks Runtime 15.4 LTS and above.

Schema changes split the same way. Iceberg tracks columns by ID, so add, drop, rename, and reorder all run without a rewrite. Delta's column mapping supports dropping and renaming columns from Delta Lake 1.2 onward, requires reader version 3 and writer version 7, and type widening reached GA in Delta 4.0.

Both formats time-travel by snapshot or version and by timestamp, and both need pruning: Iceberg through expire_snapshots, Delta through VACUUM with a retention window. These internals stay invisible until a second engine reads the same table, and row-level deletes are where that sharing breaks first.

Where Do Row-Level Deletes and Interoperability Break Down?

Iceberg v3 and Delta write the same deletion-vector bits, yet cross-format reads still fail on equality deletes, UniForm's constraints, and hidden partitioning. These compatibility limits make client access part of data governance.

Delta has written RoaringBitmap deletion vectors since Delta 3.1, and Iceberg v3 adopted binary compatibility with that encoding, storing vectors in Parquet files with at most one per data file. The Iceberg spec requires each new deletion vector to replace every earlier position delete for its data file, so readers can ignore them.

Equality deletes exist only in Iceberg, identify deleted rows by column value, and force a reader to evaluate that predicate against every base file in the snapshot. PyIceberg raises a ValueError with the message "PyIceberg does not yet support equality deletes."

The official UniForm docs state that Iceberg clients get read-only access, that UniForm does not expose Change Data Feed to them, and that UniForm does not work on tables using deletion vectors. Deletion vectors are the Delta path for efficient updates, so a Databricks team running CDC upserts cannot have both on one table.

Hidden partitioning cuts the other way. Jack Vanlightly's review of format interoperability shows that when Iceberg is the primary format, and Delta is a secondary target, the table has to give up hidden partitioning. Every cross-format path therefore costs a feature on one side, which is why the pipeline writing the table matters as much as the format itself.

How Does Your Ingestion Pattern Decide the Format?

Kafka Connect writing to object storage with no Flink cluster cannot yet carry mutation workloads into Iceberg, because the v1.10 Apache Iceberg sink for Kafka Connect does not support UPSERT. The Debezium route runs into the other Iceberg-only delete type, because Debezium Server's Iceberg consumer writes equality deletes in upsert mode, a delete type PyIceberg rejects.

Snowflake closes that path from the warehouse side. Its interoperability page lists equality-delete read and write for Iceberg tables as "Not planned" as of March 2026. A Debezium-to-Iceberg pipeline therefore needs rework before format selection means anything. Flink CDC 3.4.0 added an Iceberg pipeline connector in May 2025, which provides an alternative route.

Choosing among Kafka Connect, Debezium, and Flink makes format selection partly a data integration tooling decision.

The write pattern still punishes wide tables. In an example the Apache Fluss team published on the small-file explosion, 500 MB of CDC data became 2 million files before compaction and slowed queries 10–100x. Delta's Change Data Feed has a longer production track record than Iceberg v3 row lineage, and that maturity favors Delta for high-frequency CDC upserts on Spark. Iceberg's counterweight is recovery, since incremental reads between two snapshots let a failed downstream job resume from the last committed snapshot rather than reprocess the table.

Why Is the Catalog the Lock-In Decision?

The catalog decides lock-in because it holds the policies, lineage, and access rules that the table format leaves out. Atlan's review of Unity Catalog limitations states the problem for governed data: "The bytes are portable. The policies and the meaning are not."

Databricks open-sourced Unity Catalog in June 2024, but managed lineage and auditing stay proprietary. The same source notes that the Iceberg REST catalog API is Iceberg-specific, so Delta and Hudi tables do not join its credential-vending, Role-Based Access Control (RBAC) model natively. Delta plus managed Unity Catalog is therefore compound lock-in at two layers.

Iceberg plus Apache Polaris separates storage from governance concerns, and Polaris graduated to an Apache Top-Level Project in February 2026. Version 1.4 (April 2026) added federation to Hive Metastore, AWS Glue, and external Iceberg REST catalogs. AWS extended its own Glue catalog federation to GovCloud in August 2026, which matters when workload placement spans regulated regions.

The audit primitive lives either in your boundary or in a vendor runtime. Iceberg row lineage puts per-row provenance in the table format itself, so, as securitydataworks puts it in its catalog decision piece, "swapping Polaris for Nessie doesn't cost you the audit trail." Delta keeps equivalent history in Change Data Feed, which only Delta readers see. Once lineage and policy placement are settled, the remaining difference between the formats is the operating bill.

What Does Ongoing Maintenance Cost for Each Format?

Self-managed Iceberg costs more operator time than managed Delta on Databricks. Managed Iceberg services reduce that burden by handling compaction and snapshot expiration.

IOMETE's production anti-patterns guide walks through a streaming pipeline committing every second: 86,400 commits per day, 432,000 new files per day at five files per commit, 13 million files after a month. Once compaction falls behind, rising file counts can overwhelm maintenance. Large-volume deployment practices must keep compaction ahead of file creation.

Iceberg needs compaction for both data files and manifests, plus rewrite_position_delete_files for dangling deletes, and Starburst warns that deleting old snapshots does not replace ongoing data-file compaction. Iceberg 1.6+ adds a fault-tolerant partial progress mode for long-running compaction. In a snapshot-expiration bug report filed as GitHub issue #10907, an expire_snapshots run completed successfully while the snapshot count rose from 2,130 to 2,164, and the run removed no S3 files.

Delta requires maintenance for both file layout and deletion vectors. File compaction handles file layout, while compacting deletion vectors into base files needs a separate REORG TABLE ... APPLY (PURGE). On the managed side, Amazon S3 Tables and Snowflake Iceberg Tables handle compaction and snapshot expiration, and Databricks Predictive Optimization manages compaction, file layout, and statistics.

Which Table Format Should You Choose?

Pick Iceberg for new builds and multi-engine or multi-cloud environments, Delta when Databricks is your platform and CDC upserts dominate, and change the pipeline before choosing either when Kafka Connect alone carries your CDC.

The format decision is independent of your warehouse hosting model; it tracks your engines and catalog instead. Warehouse support has moved toward Iceberg, with Snowflake making Iceberg v3 GA on May 7, 2026. The table below matches the recommendation to your platform and ingestion path.

Your situationPickWhy
Databricks is your platform and stays that wayDelta LakePredictive Optimization for compaction, file layout, and statistics; Unity Catalog governance
Multiple engines or vendors query the same tablesIcebergOpen REST catalog spec; Iceberg support on Snowflake, BigQuery, Trino, DuckDB, and AWS services, with read/write depth varying by engine and catalog
Snowflake is your primary warehouseIcebergIceberg v3 GA; Delta readable only through UniForm
BigQuery is your primary warehouseIcebergBigQuery reads Delta but does not write it
High-frequency CDC upserts on SparkDelta LakeDeletion vectors plus Change Data Feed with more production mileage
CDC through Kafka Connect with no Flink clusterNeither, yetAdd Flink or another UPSERT-capable path before selecting the format
Already on Delta, need Iceberg readers, no deletion vectorsDelta Lake with UniFormOne-directional read compatibility, GA since Delta 3.2
Already on Delta with deletion vectors, need Iceberg readersManaged Iceberg tables or migrationDelta tables with deletion vectors cannot use UniForm
Regulated or multi-cloud, audit primitive must outlive the catalogIcebergRow lineage lives in the format; REST catalogs federate (Polaris 1.4, AWS Glue)
Starting fresh with no platform commitmentIcebergBroadest engine support and an open REST catalog spec

The rows that point to Iceberg share one property: they keep lineage and policy outside any single vendor's catalog, which is what cloud data governance programs audit. That leaves the pipeline landing the data as the last place where ownership can leak.

How Does Airbyte Flex Land Lakehouse Data in Your Boundary?

Airbyte Flex keeps the replication layer under the same ownership model as the tables it writes. Airbyte operates the control plane while the data plane, along with your records, credentials, keys, and compute, runs in your environment, though some metadata such as cursor and primary-key values sits in the control plane. Replicated records land in Iceberg tables in your own object storage or in Delta tables through the Databricks destination, so the format and catalog decision above stays yours.

Flex runs CDC replication across 700+ connectors without locking production tables, and the same catalog is available in every deployment model. The connectors and data movement engine sit on an open-source foundation you can inspect, and Flex adds RBAC plus organization-level audit logging on supported paid tiers, for events such as connection, permission, and source changes.

Where Should You Start?

Decide which layer holds your lineage and policies first. Parquet bytes move between formats with little friction, and you pay for the governance around them, so settle the catalog and CDC path before the format.

Airbyte replicates operational data into Iceberg and Delta destinations, and Airbyte Flex keeps those pipelines inside your boundary when your lineage and access policies need to stay there too.

Get a demo to see how Airbyte Flex lands governed replication into your chosen table format in your boundary.

Frequently Asked Questions

Can You Migrate Delta Lake Tables to Iceberg Without Losing History?

Yes. Apache Iceberg 1.11.0 documents a Delta Lake migration procedure whose snapshot migration mode carries data history into the Iceberg table. It also writes table properties recording the source, so you can trace the migration afterward.

Does Apache XTable Remove the Need to Choose?

Not yet. Apache XTable (incubating, version 0.3.0) offers omnidirectional metadata translation across Delta, Hudi, and Iceberg, unlike UniForm's one-directional conversion. It remains in incubation, and translation lacks write-side features specific to the secondary format, including hidden partitioning.

How Does Microsoft Fabric Handle Both Formats?

OneLake generates Iceberg-compliant metadata on demand when an Iceberg engine reads a Delta table, with no data movement. Microsoft made automatic translation of Iceberg metadata into Delta metadata generally available for all Fabric engines on September 16, 2025.

Can You Trust Published Iceberg vs Delta Benchmarks?

Treat them as historical. The most-cited results come from 2022 vendor runs on Iceberg 0.13.1 with a configuration other vendors disputed, and no independently reproduced benchmark exists for Iceberg 1.11.0 and Delta 4.x. PureStorage's open LakeBench-k8s repository supports current versions if you want to run one on your own hardware.

Integrate with 700+ apps using Airbyte

Move data from 700+ sources into warehouses, lakes, and beyond. Set up pipelines in minutes with pre-built connectors and the Connector Builder.