Coderhouse goes global

🌍

Read more

What Is the Medallion Architecture and Why AI Teams Organize Data Into Bronze, Silver, and Gold

AI Engineering

Build and ship real applications with LLMs, agents and automation.

8 weeks · Live online · View course →

Data Analytics

Go from zero to job-ready analyst with SQL, Python and BI tools.

10 weeks · Live online · View course →

Dan Patiño

AI Strategy & Innovation at Coderhouse

Data

What Is the Medallion Architecture and Why AI Teams Organize Data Into Bronze, Silver, and Gold

The medallion architecture is a way of organizing data in a lakehouse into three layers of increasing quality, usually called bronze, silver, and gold. Raw data lands in bronze exactly as it arrived, gets cleaned and conformed in silver, and is shaped into business-ready tables in gold. Each step adds structure, validation, and meaning, so anyone reading a gold table knows it has already passed through the checks below it.

Why it matters now: AI projects pull data from more sources than ever, from application databases and event streams to documents and third-party APIs. When everything lands in one undifferentiated pile, teams cannot tell which tables are trustworthy enough to train a model, feed a dashboard, or ground an AI agent. The Databricks glossary describes the pattern as a data design for logically organizing data in a lakehouse, with the goal of incrementally and progressively improving its structure and quality as it flows through each layer. That simple promise, clear stages with clear guarantees, is why the pattern has spread well beyond the platform that popularized it.

How the Medallion Architecture Works

The architecture sits on top of a data lakehouse: cheap object storage plus open table formats that add transactions, schema enforcement, and time travel. Data moves from one layer to the next through pipelines that can run in batch or as streams.

Bronze: Raw and Complete

The bronze layer stores data as close to the source as possible: full API payloads, database change records, log lines, and files, along with ingestion metadata such as load time and source system. Little or no transformation happens here. The point is to keep an append-only history you can replay. If a bug in a later step corrupts a table, you can rebuild it from bronze instead of asking the source system to resend months of data. Techniques like change data capture often feed this layer.

Silver: Cleaned and Conformed

In silver, data is deduplicated, typed, validated, and joined into a consistent model of the business. Customer records from three systems become one customer table; timestamps share one time zone; invalid rows are quarantined instead of silently dropped. Microsoft’s guidance for Azure Databricks recommends that silver tables contain validated, enriched records and serve as the source for downstream analysis, while keeping at least one non-aggregated representation of each record.

Gold: Business-Ready

Gold tables are modeled for a specific use: aggregated metrics for a sales dashboard, a feature table for a churn model, or a curated view an AI agent is allowed to query. They are typically denormalized, documented, and optimized for reads. Gold is where business definitions live, such as what counts as an active user or how revenue is recognized.

Why AI and Data Teams Use It

Trust you can explain. When a model or an analyst consumes a gold table, everyone knows it went through the same deduplication and validation steps. That makes results easier to defend.

Reproducible training data. Because bronze keeps raw history, teams can rebuild a training set as it looked at a given point and retrain or audit a model with the same inputs.

Safer AI agents. Agents that write SQL or call data tools are far less likely to produce nonsense when they are restricted to curated gold tables with clear column definitions, rather than raw event dumps.

Clear ownership and debugging. When a number looks wrong, the layers narrow the search: was the raw data wrong, the cleaning logic, or the business aggregation? Pairing the layers with data observability checks at each boundary catches problems before they reach a dashboard or a model.

Open formats, many engines. Storing each layer in an open format such as Apache Iceberg or Delta Lake lets SQL engines, notebooks, and machine learning frameworks read the same tables without copying them.

Common Mistakes When Adopting Medallion Layers

The first mistake is treating the three names as a rigid rule. Some teams need an extra staging layer or several gold tables per domain; others can skip straight from bronze to a single curated layer for simple sources. The labels matter less than the guarantees each layer provides. The second mistake is transforming data in bronze, which destroys the replayable history that makes the layer valuable. The third is letting gold multiply without governance, so five tables define revenue five different ways. Agreeing on expectations with source owners, for example through a data contract, and documenting every gold table keeps the architecture from turning back into a swamp.

A Simple Example

Imagine an online store. Bronze receives raw order events from the checkout service and nightly exports from the payments provider. Silver joins them into one clean orders table, removes duplicate events, converts currencies, and flags orders whose payment never settled. Gold publishes daily revenue by country for finance, a customer lifetime value table for marketing, and a feature table that a demand forecasting model reads every morning. If finance spots an anomaly, engineers trace it back through silver to the exact raw events in bronze.

If you want to build the SQL, data modeling, and pipeline skills behind architectures like this, and learn how AI systems consume curated data, structured practice helps. Explore Coderhouse’s courses below.

FAQ

Is the medallion architecture only for Databricks?

No. Databricks popularized the name, but the pattern works on any lakehouse or warehouse that can hold layered tables, including platforms built on Apache Iceberg, Delta Lake, or Apache Hudi.

Is medallion the same as ETL?

Not exactly. ETL describes the process of extracting, transforming, and loading data. The medallion architecture describes where data lives at each stage of that process and what quality each stage guarantees.

Do I always need exactly three layers?

No. Three is a common starting point, but the right number depends on your sources and use cases. What matters is that each layer has a clear purpose and consistent quality rules.

Dan Patiño

I'm Dan Patiño, head of AI Strategy & Innovation at Coderhouse. My day-to-day work involves merging the tactical management of e-commerce (CRO, Email Marketing and SEO) with the development of disruptive solutions. I specialize in building internal AI-powered apps to automate tasks and boost innovation within the team. I firmly believe that technology is strategy's best ally. To dive deeper into my professional journey, I'll be waiting for you on my LinkedIn profile.

© 2026 Coderhouse. All rights reserved.

© 2026 Coderhouse. All rights reserved.

© 2026 Coderhouse. All rights reserved.