# Data Lakehouse Architecture: The Four Layers

> Data lakehouse architecture uses four layers—object storage, table formats, metadata catalogs, and compute—to deliver open, governed analytics.

Source: https://hyperlake.cloud/blog/data-lakehouse-architecture-explained-the-four-layers-that-make-it-work
Published 2026-10-07 · by Hyperlake Team · Hyperlake

Video: [Watch: Data Lakehouse Architecture Explained   The Four Layers That Make It Work (2:12)](https://www.youtube.com/watch?v=JQJ8Xzb1diw)

Data lakehouse architecture combines object storage with database-like table management and independent compute. Its four layers are object storage, an open table format, a metadata catalog, and compute engines. Together, they keep data in open files while adding transactions, schema tracking, governance, historical versions, and flexible analytics or machine learning.

This separation matters because a lakehouse only delivers flexibility when each layer has a clear responsibility and interoperates with the others. Poor choices can instead add operational complexity, duplicate data, or tie workloads to a single engine. The video above walks through the four-layer design.

## What are the four layers of data lakehouse architecture?

A data lakehouse consists of four layers arranged from durable physical storage to workload-specific processing. Each layer adds capabilities without requiring the layer below it to understand every engine or application above it.

1. **Object storage** holds the underlying data as files, commonly using Apache Parquet for columnar analytical data.
1. **The table format** adds transactional metadata, schema management, table versions, and historical snapshots to those files.
1. **The metadata catalog** records table names, schemas, locations, and access information so engines can discover data logically.
1. **The compute layer** runs SQL queries, large-scale transformations, machine learning, and other processing workloads.

This layered model is what distinguishes a lakehouse from a basic data lake. A data lake may contain the same Parquet files but lack consistent table semantics, transaction handling, or a shared catalog. For broader context, compare the roles of a [data lake, data warehouse, and data lakehouse](https://hyperlake.cloud/blog/data-lake-vs-data-warehouse-vs-data-lakehouse-which-one-do-you-actually-need).

![Diagram: Four lakehouse layers from object storage through table formats and catalogs to compute engines](https://hyperlake.cloud/blog/img/production/877e293c2e6227564f122cf45cbafea0906432d1-1200x750.png?w=1600&fit=max&auto=format)

*Each layer adds management or processing capabilities without replacing the layer below it.*

## Why is object storage the foundation of a lakehouse?

Object storage provides the durable physical location for lakehouse data. It separates stored data from the compute engines that read and transform it, allowing compatible engines to work with the same underlying files.

Analytical datasets are commonly stored in Apache Parquet because its columnar layout supports efficient reads of selected columns. Open file formats also reduce dependence on a proprietary storage layer and make it practical to retain large volumes of structured or semi-structured data.

Storage alone, however, does not create a lakehouse. A directory of Parquet files does not inherently provide transactions, reliable schema evolution, table versions, or a consistent definition of the current table state. Those responsibilities belong to the table format above it.

## What does an open table format add to object storage?

An open table format turns collections of files into managed tables that query engines can interpret consistently. This is one of the most consequential architecture choices because it defines how writes, schema changes, snapshots, and concurrent operations are coordinated.

Formats such as Apache Iceberg, Delta Lake, and Apache Hudi maintain metadata that identifies which files belong to a table and which version represents its current state. Depending on the format and engine support, this layer can provide:

- ACID transactions for reliable reads and writes.
- Schema evolution without treating every change as a new dataset.
- Version tracking across table updates.
- Time travel queries against retained historical snapshots.
- Safer coordination between concurrent readers and writers.

The table format does not replace object storage; it manages the files stored there. Format selection should therefore consider engine compatibility, governance requirements, operational tooling, and portability. A closer [Apache Iceberg and Delta Lake comparison](https://hyperlake.cloud/blog/apache-iceberg-vs-delta-lake) explains how two common approaches differ.

## How do metadata catalogs and compute engines work together?

The metadata catalog gives engines a stable, logical way to find and interpret tables, while the compute layer performs the actual processing. Together, they prevent every engine and user from having to manage physical file paths directly.

A catalog typically tracks which tables exist, their schemas, their underlying storage locations, and information used for access governance. Instead of querying a path such as a specific object-storage prefix, an engine requests a named table. The catalog resolves that logical name to the relevant metadata and files.

This indirection matters when tables move, files are compacted, partitions change, or snapshots advance. Without a shared catalog, each engine could develop its own view of the dataset, creating inconsistency and fragile dependencies on storage layout.

Above the catalog, independent compute engines can serve different workloads. SQL query engines can support interactive analytics, Spark can perform large-scale transformations, and machine learning frameworks can consume prepared data. Because compute is separated from storage, compatible engines can scale independently, run concurrently, or be replaced without copying the entire dataset.

The architecture still requires coordination. Engines must support the selected table format and catalog behavior, while identity, authorization, and audit controls must be enforced consistently at integrated access points.

![Diagram: A catalog connects stable lakehouse storage and tables to independent compute engines](https://hyperlake.cloud/blog/img/production/c872323743501114f56aa91d8b6fea2f353aead6-1200x750.png?w=1600&fit=max&auto=format)

*The catalog lets compatible engines find governed tables without depending on physical file paths.*

## Key takeaways

- Data lakehouse architecture separates storage, table management, metadata discovery, and compute into four distinct layers.
- Object storage holds open data files, while a table format adds database-like transactions and version management.
- A metadata catalog gives engines a governed logical interface instead of exposing physical file paths directly.
- Independent compute engines can process the same data for SQL, transformation, and machine learning workloads.
- Interoperability between the table format, catalog, engines, and access controls determines whether the architecture remains flexible.

## How Hyperlake helps

Hyperlake can assemble a governed data foundation using Iceberg and object storage with a fitting query engine, alongside operational, vector, graph, search, and streaming services when the workload requires them. It runs in customer-controlled infrastructure and provides shared controls for identity, access, observability, lifecycle operations, lineage, and audit, with procedures varying by engine and solution pack. To discuss a lakehouse architecture for your environment, [talk to our team](https://hyperlake.cloud/contact).

## Frequently asked questions

### Can multiple query engines use the same lakehouse data?

Yes, multiple compatible engines can access the same lakehouse tables through a shared catalog and supported table format. For example, one engine might run interactive SQL while another performs batch transformations. Compatibility still matters: each engine must correctly understand the table format, catalog implementation, transaction semantics, and current metadata version.

### What happens if object storage has no table format?

The storage remains a data lake containing files, but it lacks a shared mechanism for managing those files as reliable database-style tables. Readers may need to infer schemas and locate paths themselves, while concurrent writes, schema changes, and file replacement become harder to coordinate. A table format supplies the metadata and transaction model needed for consistent table state.

### Is a data lakehouse the same as a data warehouse?

No. A lakehouse generally stores data in open files on object storage and places independent table, catalog, and compute services above them. A traditional warehouse more often combines storage and compute within a managed system. Both can support governed analytics, but their storage models, engine flexibility, operational responsibilities, and workload tradeoffs differ.

### Why is the metadata catalog important for governance?

The catalog creates a logical inventory of tables, schemas, and storage locations that users and engines can reference consistently. It can also participate in access governance by connecting table identities with authorization controls and audit processes. Enforcement must still occur at the relevant query, storage, or service access points rather than relying on metadata alone.
