dbt + DuckDB in 2026: You Don't Need a Warehouse for That
Everyone asks which warehouse to buy. In one project, the right answer was none.
EB
Egor Burlakov
••6 min read
A lot of data projects begin with the same meeting. Someone draws a few boxes on a whiteboard, someone else starts naming platforms, and before long the conversation turns into a familiar question: “Snowflake or Databricks?”
It is a reasonable question, but it often arrives too early. By asking which warehouse or lakehouse to choose, you have already assumed that you need one. In one project we worked on, that assumption turned out to be wrong: the transformation layer ended up as a small dbt project and one scheduled container. There is no permanent analytical cluster and no warehouse waiting between runs.
The stack is deliberately unexciting. dbt contains the logic, DuckDB executes it, and Postgres serves the results. The point is not that DuckDB is cheaper; it is that some workloads do not need a permanent analytical platform at all.
The Use Case: Small Data, Real Stakes
The project came from a physical operations environment with several automated and semi-automated processes. The systems produced event logs describing process starts and stops, step transitions, faults, some states and manual interventions. Volumes were modest at the point in time, and months of history could be processed comfortably on one machine. However, the growth plans were ambitious and we had to take this into account.
The business wanted a web-application with a consistent view of throughput, cycle time, equipment availability and quality metrics. Operations, engineers and management all needed the same numbers for different reasons, which meant a disagreement between the portal and a spreadsheet quickly became a trust problem.
So the architecture problem was not “how do we process enormous amounts of data?” It was “how do we turn messy operational data into metrics that are correct, testable and calculated the same way everywhere?”
Why We Did Not Start With Databricks or Snowflake
Both Databricks and Snowflake are excellent products, and in 2026 both can handle small workloads efficiently. The more useful question is: What would Databricks or Snowflake actually give us that we need?
The DuckDB file is scratch space; durable inputs stay in the lake, and durable outputs go to Postgres. A model is ordinary SQL:
-- models/daily_jobs.sql
select
finished_at::date as day,
count(*) as finished_jobs,
median(date_diff('second', started_at, finished_at)) as median_seconds
from read_parquet('s3://example-lake/clean/jobs/*/*.parquet',
hive_partitioning = true)
where status = 'finished'
group by all
DuckDB treats many Parquet files as one table, while Postgres remains the application database. That separation matters: DuckDB is good at scanning and aggregating a batch; Postgres is good at staying online and serving many small queries.
The bigger benefit is that the team gets one place to define business metrics. “Throughput” otherwise tends to reproduce: one version in a dashboard, another in a spreadsheet, a third in Python. Each may have started out correct, but after a few changes nobody knows why the same day has three different answers.
The rule we used was simple: every metric is defined once, in the transformation project. The portal displays it, reports read it, and ad-hoc analysis starts from the same result. If a definition changes, it changes in one place, together with the tests that describe its expected behaviour.
A dbt data test can be just a query that should return no rows:
-- tests/daily_jobs_are_sane.sql
select day
from {{ ref('daily_jobs') }}
where median_seconds < 0
or day > current_date
dbt unit tests can also define fixed inputs and expected outputs, turning vague metric logic into something much closer to an executable contract.
The Hard Part Was Not DuckDB
The biggest problem in this project was not query performance or memory limits. It was the meaning of the source events.
At one point, an operation could be aborted and later restarted from a step in the middle of the process. The event stream did not record enough context about where the restarted run began, so a run that completed only the final part of the process could look like a surprisingly fast complete run. No database engine can recover information that the source system never recorded.
That was the more important lesson: teams often spend more time choosing the system that calculates a metric than asking whether the source events contain enough information to calculate it correctly.
The stack still has clear limits: each run must fit on one machine, many tiny object-storage files can slow scans, and scheduling or alerting still has to come from somewhere else. If those constraints become painful, a larger platform may be justified.
When dbt + DuckDB Is the Right Call
Most architecture debates become easier when you replace product names with concrete questions:
I would use this setup when a batch fits comfortably on one machine, transformations run on a schedule, and the application needs prepared facts rather than arbitrary analytical queries. I would choose something larger when many people need interactive access, governance spans many teams, or one machine is genuinely no longer enough. At that point, Snowflake or Databricks is solving a problem that has actually appeared.
Start Small, but Keep the Exit Open
A small architecture is only attractive if it does not trap you. Most of the logic here is ordinary dbt SQL, so a later move to Snowflake or Databricks would require engine-specific changes, but the metric definitions, dependencies and much of the test suite would survive.
That is the principle I keep coming back to: a data architecture should become complicated because the problem became complicated, not because the industry has normalized complicated architecture.
If the data fits on one machine and the hard part is producing correct, consistent and testable metrics, a scheduled dbt + DuckDB job is a strong place to start. Buy the shared warehouse or lakehouse when the workload asks for it.
Engineering and Science Leader with experience building scalable data infrastructure, data pipelines and science applications. Sharing insights about data tools, architecture patterns, and best practices.
Explore Further
Dive deeper into the tools and categories mentioned in this article.