Coderhouse goes global

🌍

Read more

What Is Apache Iceberg and Why AI Teams Are Standardizing on Open Table Formats

AI Engineering

Build and ship real applications with LLMs, agents and automation.

8 weeks · Live online · View course →

Data Analytics

Go from zero to job-ready analyst with SQL, Python and BI tools.

10 weeks · Live online · View course →

Dan Patiño

AI Strategy & Innovation at Coderhouse

Data

What Is Apache Iceberg and Why AI Teams Are Standardizing on Open Table Formats

Apache Iceberg is an open table format for huge analytic datasets. It sits between your files (usually Parquet in object storage like Amazon S3 or Google Cloud Storage) and the engines that query them, adding a metadata layer that turns a pile of files into a real table: with a schema, snapshots, ACID transactions, and a reliable list of which files belong to which version. Spark, Trino, Flink, Snowflake, BigQuery, Databricks, and DuckDB can all read and write the same Iceberg table without copying data into a proprietary format first.

Why it matters now: AI teams want one copy of governed data that analytics, feature pipelines, model training, and agents can all use. The Apache Iceberg project was built for exactly that multi-engine world, and adoption shows it. A survey of 252 senior data leaders running Iceberg in production, published by Ryft, found that 95% are using or planning to use Iceberg for AI or ML workloads and 79% plan to move their remaining data to it within a year.

What an Open Table Format Actually Does

A classic data lake is a folder structure. The “table” is whatever files happen to be in a directory, and the Hive Metastore tracks partitions by path. That works until two jobs write at the same time, someone renames a column, or a query has to list millions of files just to plan its scan. An open table format fixes this by tracking state explicitly in metadata files instead of relying on folders.

Iceberg keeps a tree of metadata: a table metadata file points to snapshots, each snapshot points to manifest lists, and manifests point to the actual data files along with column-level statistics. Every write creates a new snapshot atomically. Readers always see a consistent version, and the query planner can skip files using those statistics without touching storage.

The Features AI and Data Teams Rely On

Time travel and rollback. Because every commit is a snapshot, you can query a table as it was at a specific point and roll back a bad write. For ML, that means you can reproduce the exact training set behind a model version.

Schema evolution. You can add, drop, rename, or reorder columns without rewriting data, because Iceberg tracks columns by ID rather than by name or position.

Hidden partitioning. Iceberg derives partitions from column values (for example, a day from a timestamp), so users write normal filters and the engine prunes partitions for them. Partition layouts can also change over time without breaking old queries.

Row-level changes. Merges, updates, and deletes work at row level, which matters for GDPR deletion requests and for applying streams from change data capture (CDC).

The newer v3 specification goes further. As Databricks describes in its Iceberg v3 announcement, v3 adds a VARIANT type for semi-structured data, deletion vectors for cheaper row-level deletes, and row lineage for incremental processing — all useful when AI pipelines ingest messy JSON from APIs and event streams.

Iceberg vs Delta Lake vs Hudi

Iceberg is not the only open table format. Delta Lake (created by Databricks) and Apache Hudi solve similar problems. Delta has deep integration with the Databricks platform; Hudi grew up around streaming upserts. Iceberg’s advantage has been neutral community governance under the Apache Software Foundation and wide engine support. That is why many companies choose it as the default: Grab’s engineering team says community governance, engine compatibility, and long-term flexibility drove its decision to move from Hive Parquet to Iceberg. Interoperability layers and v3’s convergence with Delta are also shrinking the cost of picking “wrong.”

Where Iceberg Fits in an AI Data Stack

Iceberg is the storage layer behind most modern data lakehouses. Above it you need a catalog (such as Apache Polaris, AWS Glue, or Unity Catalog) that tracks where each table’s current metadata lives and enforces access rules. Pair that with a business-facing data catalog and data lineage so people and agents can find and trust tables. On top, the same Iceberg tables can feed BI dashboards, a feature store, and training jobs, without separate copies drifting apart.

Common Mistakes When Adopting Iceberg

The format solves consistency, not operations. Streaming writes create many small files, and snapshots pile up. Without scheduled compaction, snapshot expiration, and orphan-file cleanup, query performance degrades and storage costs climb; the Ryft study flags this operational maturity as the next big challenge for most teams. Other common mistakes: running several catalogs that disagree about the current table version, migrating everything at once instead of starting with high-value tables, and assuming time travel is a backup strategy when expired snapshots are gone for good. Start with one domain, set retention and compaction policies from day one, and choose a single source-of-truth catalog.

If you want to build the data foundations AI products depend on, structured practice helps. Explore Coderhouse’s courses below.

FAQ

Is Apache Iceberg a database?

No. Iceberg is a table format: a specification for how data files and metadata are organized. You still need a query engine (Spark, Trino, Snowflake, and others) and a catalog to use it.

Do I need Iceberg if I already have a cloud data warehouse?

Not always. Iceberg pays off when several engines or teams need the same data, when storage costs at scale matter, or when you want to avoid locking data into one vendor’s format. Many warehouses now read and write Iceberg tables directly.

How does Iceberg help reproduce ML experiments?

Every write creates a snapshot. Recording the snapshot ID used for training lets you query that exact version of the data later, as long as your retention policy keeps it.

Dan Patiño

I'm Dan Patiño, head of AI Strategy & Innovation at Coderhouse. My day-to-day work involves merging the tactical management of e-commerce (CRO, Email Marketing and SEO) with the development of disruptive solutions. I specialize in building internal AI-powered apps to automate tasks and boost innovation within the team. I firmly believe that technology is strategy's best ally. To dive deeper into my professional journey, I'll be waiting for you on my LinkedIn profile.

© 2026 Coderhouse. All rights reserved.

© 2026 Coderhouse. All rights reserved.

© 2026 Coderhouse. All rights reserved.