hyperlakeDiscuss a deployment ↗
Blog · · 5 min read

Open Table Formats Explained for Data Lakes

Open table formats add ACID transactions, time travel, schema evolution, and efficient query planning to data lakes. Compare Iceberg, Delta, Hudi, and Paimon.

Video thumbnail: Open Table Formats Explained
Watch: Open Table Formats Explained (2:03) · Video page

Open table formats are metadata layers that make files in object storage behave like managed database tables. They record committed changes, coordinate concurrent writers, preserve historical states, track schema changes, and help query engines find relevant files. Apache Iceberg, Delta Lake, Apache Hudi, and Apache Paimon implement these capabilities with different architectural priorities.

This matters because object storage is durable and economical, but files and directories alone do not provide reliable updates, transactional consistency, or a trustworthy view of past table states. Open table formats add those controls without requiring data to live inside a traditional database. The video above walks through the core ideas.

What are open table formats?

An open table format defines how data files, metadata, snapshots, and table changes are organized on object storage. It sits between the storage layer and compute engines such as query, batch processing, or streaming systems.

A plain data lake might contain thousands or millions of Parquet files arranged into directories. Without an additional metadata layer, engines may need to infer table state from filenames and folder layouts. Updating records, coordinating multiple writers, or reconstructing an earlier version becomes difficult.

A table format turns those files into a logical table by defining:

  • Which data files currently belong to the table.
  • Which schema applies to a particular table version.
  • How writes become visible to readers.
  • How historical snapshots and deleted files are tracked.
  • What metadata engines can use to avoid unnecessary file scans.

The format does not replace object storage or the compute engine. It provides a shared table contract between them, separating durable storage from the systems that read and write the data.

How do open table formats provide ACID transactions?

Open table formats publish changes through atomic metadata commits. A set of new, replaced, or removed files becomes part of the table only when its corresponding metadata change commits successfully.

The exact mechanism differs by format. Delta Lake uses a transaction log, Apache Iceberg advances table metadata and snapshot pointers, and Apache Hudi maintains a timeline of table actions. Conceptually, each creates an authoritative history of committed table states.

A typical write follows this sequence:

  1. A compute engine reads the current table state.
  2. It writes new data files without immediately exposing them as current data.
  3. It prepares metadata describing the intended table change.
  4. It atomically commits that metadata or retries if another writer created a conflict.

Readers therefore see a valid state before or after the commit rather than a partially published update. This foundation enables ACID behavior on object storage and supports concurrent workloads more safely than direct, uncoordinated file writes.

Diagram: A writer reads table state, writes files, prepares metadata, and atomically commits the change.
Data files become part of the current table only after the metadata commit succeeds.

What capabilities come from table metadata?

Committed metadata enables time travel, schema evolution, and efficient query planning. These features are related because they all depend on knowing exactly what the table contained at each committed state.

Time travel lets an engine inspect or query an earlier snapshot. This is useful for reproducibility, debugging, auditing, and recovering from an incorrect data change, subject to the table’s history-retention and file-cleanup policies.

Schema evolution records how columns and data types change over time. Rather than treating every directory as an unrelated collection of files, the format associates schema information with table metadata and historical states.

Efficient planning lets query engines use metadata about files, partitions, and value ranges to determine which files may contain relevant records. Engines can skip unrelated data instead of listing or scanning every file. Apache Iceberg snapshots provide one example of this metadata-driven approach; see how Iceberg snapshots work for more detail.

Which open table format should you choose?

Choose the format whose write patterns, engine ecosystem, and interoperability model fit how the data moves. Feature checklists matter less than the architectural trade-offs surrounding batch processing, streaming ingestion, record-level updates, and multi-engine access.

The main options discussed in the video are:

  • Apache Iceberg was designed as an engine-agnostic table specification. Its broad engine support makes it a strong candidate when several compute systems must work with the same tables.
  • Delta Lake originated around Apache Spark and retains deep integration with that ecosystem. It commonly fits teams whose data processing architecture centers on Spark and Delta-compatible engines.
  • Apache Hudi was designed around incremental processing and record-level changes. Its indexing approaches help locate records that need to be updated rather than treating every change as a full table rewrite.
  • Apache Paimon is a newer format oriented toward streaming workloads. Its LSM-tree architecture is designed for frequent writes, including pipelines built with systems such as Apache Flink.

Selection should also account for catalog support, writer concurrency, maintenance operations, and every engine that must read or modify the table. A format supported by one engine may not expose identical behavior through another, so teams should validate the complete workload rather than only the storage specification. For a focused comparison, see Apache Iceberg vs. Delta Lake.

Diagram: Iceberg and Delta emphasize engine ecosystems, while Hudi and Paimon emphasize changing and streaming data.
The right format depends on engines, interoperability, update patterns, and streaming requirements.

Key takeaways

  • Open table formats add a reliable table abstraction to files stored in data lakes.
  • Atomic metadata commits prevent incomplete changes from becoming the current table state.
  • Historical metadata enables time travel, schema evolution, and more efficient query planning.
  • Iceberg, Delta Lake, Hudi, and Paimon prioritize different engines and data-movement patterns.
  • Format selection should reflect actual readers, writers, update frequency, and interoperability requirements.

How Hyperlake helps

Hyperlake can assemble governed data foundations using Iceberg and object storage with a fitting query engine, including Trino where appropriate for the workload. It combines data services with identity, access policy, observability, and lifecycle controls in infrastructure that the organization or its client controls. To discuss the engines, deployment environment, and policies your workload requires, talk to our team.

Frequently asked questions

Can several query engines use the same open table format?

Yes, if each engine supports the chosen format, catalog, storage system, and required operations. Read compatibility may be broader than write compatibility, and support for newer format features can vary by engine version. Teams should test schema changes, concurrent writes, deletes, and maintenance procedures across every engine that will access the table.

Do open table formats replace object storage or databases?

No. An open table format organizes data and metadata stored in systems such as object storage, while compute engines execute queries and writes. It can provide database-like table behavior for analytical data, but it does not replace every capability of an operational database, including low-latency transactions tailored to application workloads.

Can a data platform change table formats later?

Yes, but migration is a data-platform project rather than a metadata toggle. The team may need to rewrite or register data, translate schemas and partitioning, preserve required history, update catalogs, and validate every reader and writer. Choosing around real workload and engine requirements reduces the likelihood of a disruptive migration.

Start with a workload. Build the environment around it.

Explore example deployments, or see how the platform assembles, deploys, governs and operates the stack.