Data Lakehouse definition
A data lakehouse is a data architecture that combines the low-cost, flexible storage of a data lake with the reliability, performance and governance features of a data warehouse. It stores data in open file formats on object storage and adds a table layer, such as Delta Lake, Apache Iceberg or Apache Hudi, providing ACID transactions, schemas and fast SQL.
How does a data lakehouse work?
A lakehouse keeps data as files, typically Parquet, in cloud object storage. An open table format sits on top and tracks which files belong to each table version in a transaction log or metadata tree. This allows multiple engines to read and write safely, with ACID transactions, schema enforcement and evolution, time travel to earlier versions, and efficient updates and deletes, which are hard to achieve with plain files in a traditional data lake.
Query engines such as Databricks SQL, Spark, Trino, Snowflake, BigQuery and Amazon Athena can read these tables directly. A governance layer, such as Unity Catalog or a similar catalog service, manages permissions, lineage and discovery, so the same storage supports BI dashboards, SQL analytics, data science notebooks and machine learning training without copying data between separate systems. That single copy simplifies security reviews too.
The medallion architecture
Lakehouses commonly organize data into three quality layers, an approach popularized by Databricks as the medallion architecture. Each layer builds on the previous one, and transformations between layers are automated, tested and monitored so problems are caught early. The names matter less than the discipline of separating raw, cleaned and business-ready data. Gold tables should map to real business questions.
- Bronze: raw data ingested from sources, kept for replay and audit.
- Silver: cleaned, validated and conformed data joined across sources.
- Gold: aggregated, business-level tables for reporting and ML features.
Benefits of a data lakehouse
The main benefit is one platform instead of two. Organizations that previously copied data from a lake into a separate warehouse maintained duplicate pipelines, inconsistent numbers and two sets of costs. A lakehouse reduces that duplication. Open formats also reduce lock-in, because the same tables can be read by different engines, and storage stays on inexpensive object storage. Data science and BI teams work from the same governed data, which shortens the path from analysis to production models.
Trade-offs and challenges
Lakehouses require solid data engineering. Teams must manage file sizes, compaction, partitioning, table maintenance and catalog configuration, which managed warehouses largely hide. For small, mostly structured datasets used only for BI, a conventional cloud warehouse may be simpler and cheaper to operate. Choosing between table formats and catalogs also matters, though interoperability between Delta Lake and Iceberg has improved as vendors have converged on open standards. Budget time for table maintenance jobs from the start.
Example of a lakehouse
An ecommerce company ingests clickstream events, orders and product catalog data into bronze Delta tables. Silver tables join sessions to orders and clean product attributes. Gold tables feed revenue dashboards in Power BI and a recommendation model trained in the same platform, both using identical customer definitions. Nexzem designs lakehouse platforms on Databricks and open table formats when clients need analytics and machine learning on the same governed data. Adding a new gold table takes hours rather than a new pipeline project.