
Apache Iceberg snapshots are immutable records of a table’s exact state, represented by metadata pointers rather than full copies of its data. Each snapshot identifies the manifests and data files that belong to the table at commit time, enabling snapshot-isolated reads, safe concurrent updates, historical queries, and efficient file pruning.
This matters because the snapshot model supports reliability and query performance without requiring a full copy of the table for every change. It also provides the foundation for several capabilities associated with modern lakehouse architecture. The video above walks through the core ideas.
What is an Apache Iceberg snapshot?
An Apache Iceberg snapshot is a metadata record that identifies the exact files representing a table after a committed data change. It points through a metadata hierarchy to those files instead of copying or modifying the complete dataset.
Each successful append, overwrite, delete, or merge produces a new logical table state. That state remains separate from earlier states because Iceberg writes new files and metadata rather than changing existing files in place during the commit.
A snapshot therefore acts like a consistent version of the table. Readers using the same snapshot see the same table state even if another process commits new data while their query is running. This model is a central feature of data lakehouse architecture.
Snapshots are lightweight compared with full table copies, but they still depend on the underlying data and metadata files remaining available. Retention and snapshot-expiration policies determine how long historical states can be queried.
How does Iceberg create a snapshot?
Iceberg creates a snapshot by writing new data files, describing them in manifests, grouping those manifests in a manifest list, and recording a snapshot that points to that list. Existing committed data files do not need to be modified.
The simplified commit sequence is:
- Write data files. The writer creates files containing the records added or rewritten by the operation.
- Create or update manifests. A manifest identifies data files and records metadata about their contents, including useful column statistics.
- Build the manifest list. The manifest list identifies the manifest files associated with the snapshot.
- Commit the snapshot. A snapshot entry is added to the table metadata state and points to the manifest list.
The resulting hierarchy runs from table metadata to a snapshot, from the snapshot to its manifest list, and from manifests to data files. Iceberg can also reuse existing manifests and files when they remain part of the new table state, reducing unnecessary metadata work.
Because commits publish a new metadata state rather than editing data files in place, readers can discover the table consistently through its metadata pointer. This separation of metadata from immutable files also distinguishes Iceberg from storage layouts that depend on directory listings or mutable file sets; see Apache Iceberg versus Delta Lake for a broader format comparison.

Why are concurrent Iceberg reads and writes safe?
Concurrent operations are safe because readers remain attached to one immutable snapshot while writers prepare and commit a different snapshot. A reader’s view does not change partway through a query simply because new data arrives.
Writers typically create their files before attempting to publish the new table metadata. Iceberg uses optimistic concurrency controls during the commit: if another writer has changed relevant table state, a conflicting operation may need validation, retry, or reconciliation rather than silently overwriting the other writer’s work.
This design provides snapshot isolation for reads. A long-running query can continue using the files referenced by its starting snapshot, while newer queries can use the newly committed state. Readers and writers therefore avoid interfering through in-place edits to shared files.
Operationally, file cleanup must respect snapshot retention. Files referenced by retained snapshots cannot be removed merely because they are absent from the newest table state.
How do snapshots enable time travel and file pruning?
Snapshots enable time travel by retaining pointers to earlier valid table states, while their manifests enable pruning by describing files before a query opens them. Both capabilities follow from the same metadata hierarchy.
For time travel, a query selects a historical snapshot by identifier or timestamp, depending on the query engine. This is not a restore operation: the engine reads the historical state directly from the files referenced by that snapshot. Historical queries remain possible only while the snapshot and its required files are retained.
For query planning, manifest metadata can include column-level statistics such as lower and upper value bounds for individual data files. When a query contains a predicate, the engine compares that predicate with available statistics and excludes files whose ranges cannot contain matching values.
For example, if a file’s recorded range cannot satisfy a date or numeric filter, the engine can skip that file without reading its rows. On a large, well-organized table, this metadata-driven pruning can remove many irrelevant files from consideration before data scanning begins. Actual pruning effectiveness depends on the filters, data distribution, file layout, available statistics, and query engine.

Key takeaways
- An Iceberg snapshot is a metadata pointer to an exact table state, not a complete copy of the data.
- New snapshots preserve immutable table states, allowing readers and writers to operate without in-place file conflicts.
- Historical snapshots support direct time-travel queries while their referenced files remain retained.
- Manifest statistics help query engines eliminate irrelevant files before reading table data.
- Snapshot expiration and file cleanup must be coordinated so retained historical states remain valid.
How Hyperlake helps
Hyperlake can assemble a governed data foundation using Iceberg and object storage with a fitting query engine, alongside operational, search, vector, graph, and streaming services when the workload requires them. Teams can deploy and manage that foundation in their own infrastructure or a client’s environment, with shared controls for identity, access, observability, lifecycle, and audit; operational procedures vary by engine and solution pack. To discuss an Iceberg-based data environment, talk to our team.
Frequently asked questions
Does every Iceberg snapshot contain a separate copy of the table?
No. A snapshot is primarily a metadata structure that identifies the manifests and files belonging to one table state. Multiple snapshots can reference some of the same immutable data files, while newer snapshots add or replace file references as data changes.
Can Apache Iceberg snapshots replace data backups?
No. Snapshots support table history and recovery from some logical mistakes, but they depend on the underlying metadata and data files remaining intact. Backups, replication, access controls, and tested disaster-recovery procedures address infrastructure loss, accidental deletion, and failures beyond normal table history.
What happens when two writers update an Iceberg table at once?
Each writer prepares new files and attempts to commit a new metadata state. Iceberg validates the commit against the current table state; compatible operations may succeed, while conflicting changes can require a retry or fail validation. This prevents one writer from silently replacing another writer’s committed work.


