The DuckDB Dilemma: AWS Just Bought the Team, But Who Really Owns the Duck?
The news hit the data engineering world like a caffeinated duck to the face: DuckLabs, the fiercely independent, bootstrapped Amsterdam company behind DuckDB, is joining AWS. The announcement came with all the expected reassurances, MIT license stays, DuckDB Foundation keeps stewardship, the team remains intact in Amsterdam. Andy Warfield, AWS’s Distinguished Engineer and VP, even called the DuckLabs crew “one of the most technically deep, humble, and high-velocity teams” he’s ever worked with.
But let’s be honest about what just happened. The company that built the most beloved analytical database of the past decade, the one that told venture capitalists to get lost and bootstrapped its way to over a million daily downloads, just sold itself to the world’s largest cloud provider. The “revolutionary” open-source pledge mostly revolutionized paperwork.

The Bootstrapped Bastion That Couldn’t Scale Itself
Here’s the part that stings. DuckLabs didn’t need money. They turned down VC funding. They grew to 30+ people in Amsterdam, funded entirely through support contracts and feature-development work. Mark Raasveldt and Hannes Mühleisen built something genuinely rare: a profitable, founder-owned open-source company with a product developers actually adore.
But reading between the lines of their announcement, the founders were hitting a wall that cash alone couldn’t solve. They explicitly worried about becoming “a bottleneck for the project.” Building a larger sales, support, and operations organization would have pulled them away from the technical work. In other words: being a good open-source steward and being a good company were pulling in opposite directions.
There’s a certain irony here. DuckDB’s entire pitch was that analytics shouldn’t require a data warehouse. It runs in-process, queries Parquet files directly, and gives you warehouse-level performance on a laptop. DuckDB in production has cut costs by 70% for teams that adopted it, outperforming Spark clusters for the sub-terabyte queries that dominate real-world analytics.
Key Stats
- 70% cost reduction in production
- 1M+ daily downloads
- 30+ team members in Amsterdam
- 0 VC funding taken
And now the “anti-warehouse” database is going to be owned by the company that sells the warehouses.
What AWS Actually Bought (Hint: It’s Not the Code)
AWS was careful to say they’re not acquiring the DuckDB open-source project. The official announcement makes clear that DuckDB will remain under the independent DuckDB Foundation, MIT license intact. The CWI representative on the foundation board, Peter Boncz, emphasized that the foundation “holds all IP of open-source DuckDB, and will continue to do so.”
Here’s what AWS did buy: the people who know how to build DuckDB.
MotherDuck CEO Jordan Tigani, who built BigQuery at Google and now runs the company closest to DuckDB’s commercial ecosystem, put it bluntly: “That’s Amazon’s playbook, after all, wait until an open source project gets big enough, then launch it as a service.” He also notes Amazon now has a “financial reason to keep the project open and healthy”, because if DuckDB becomes the standard, it drives compute consumption on AWS infrastructure.
That last point is the real strategic calculus. AWS doesn’t need to own DuckDB’s code. They need to own its roadmap, its performance priorities, and its integration depth. When DuckDB’s development team is on AWS’s payroll, guess which cloud’s S3 endpoint gets optimized first? Which object-store quirk gets the dedicated code path?
The S3 Trojan Horse
The AWS blog post from Mai-Lan Tomsen Bukovec drops some genuinely revealing numbers. Amazon Quick, the company’s internal dashboarding tool, has processed 2.5 billion queries using DuckDB integrations since October 2025. Those integrations reduced average query latency by 30%. This isn’t speculative technology, DuckDB is already critical infrastructure inside AWS.
The Allen Institute provides the customer-facing example: they’re using DuckDB to analyze terabytes of scientific data stored in S3, with queries that took minutes now returning “in less than a second.” DuckDB also runs in-process to AWS Lambda functions, which is a hilariously cost-effective pattern, you pay only for the milliseconds of compute while Lambda executes your query.
But read the strategic framing carefully. Bukovec describes the plan as combining “the superpower of DuckDB at everyday queries of a terabyte or less with the proven exabyte-plus enterprise scale of S3 and our AWS analytics services.” That’s a division of labor that conveniently positions DuckDB as the front-end query engine and S3 as the permanent storage substrate. Every DuckDB query against S3 generates egress and request costs. Every DuckDB deployment on Lambda generates compute revenue.
This is not a charitable endeavor. This is AWS integrating a beloved open-source tool into its cost recovery machine, and DuckDB’s role is to make S3 the frictionless default for analytics.
The Governance Question Nobody Can Answer
The DuckDB Foundation is getting a technical advisory board. Extensions signed by outside developers will be allowed to run. These are genuine improvements that suggest the community’s concerns were heard.
But the uncomfortable truth is that open-source governance under corporate ownership has a mixed track record. Developer sentiment on forums captures the anxiety well: AWS has a history of acquiring projects, integrating them into their stack, and then letting them wither when strategic priorities shift. The comparison to Google’s or Microsoft’s acquisition graveyards gets thrown around a lot, and while it’s reductive, it points at a real risk.
There’s also the question of what “community” means when the project’s core maintainers are now AWS employees. DuckLake SDK decoupling from DuckDB showed that the ecosystem can fork or diverge when dependencies become problematic. But DuckDB’s value proposition is precisely its tight kernel optimization, you can’t easily replace the team that has spent a decade tuning its vectorized execution engine.
Then vs. Now
- Before: Independent, bootstrapped, community-focused
- After: AWS-owned subsidiary, corporate-aligned priorities
- License: MIT (unchanged, but direction uncertain)
- Risk: Strategic alignment with AWS’s data formats
The project itself remains MIT-licensed, which is the right license for maximum adoption and minimum friction. But licenses don’t guarantee development direction. They don’t guarantee that the next performance optimization targets S3 over MinIO or Azure Blob Storage. They don’t guarantee that AWS doesn’t eventually do what dbt Labs did with ELv2.
The MotherDuck Elephant in the Room
The most fascinating subplot here is MotherDuck. The Seattle-based company sells a cloud service built on DuckDB and was founded in partnership with the DuckLabs team. Three of its engineers are among the top 10 outside contributors to the DuckDB project.
Tigani’s public response is diplomatic: “We welcome the competition.” He notes that DuckLabs will remain a wholly owned subsidiary with its organization intact, and that the DuckDB Foundation has what he calls “iron clad control over the DuckDB IP.”
But MotherDuck is also now offering enterprise support for DuckDB, something it previously avoided to not compete with DuckLabs. Mühleisen and Raasveldt gave their “explicit blessing” for this shift. You don’t need a business degree to understand what’s happening here: MotherDuck is positioning itself as the independent commercial layer on top of DuckDB, just in case AWS’s stewardship goes sideways.
This mirrors the Fivetran-dbt merger dynamics we’ve seen elsewhere in the data stack, consolidation creates both risk and opportunity for adjacent players. And just like dbt Core v2’s open-source reversal, the real test is whether the community’s trust survives the corporate transition intact.
What This Means for Your Data Stack
If you’re a DuckDB user (and if you’re in data engineering in 2026, you probably are), here’s the practical reality:
Short term
Nothing changes. The MIT license means you can keep using DuckDB exactly as you are today, forever. The 1.0 API remains stable. The extension ecosystem is opening up, which is strictly better for users.
Medium term
Expect deeper AWS integration. DuckDB will get first-class treatment for S3 Tables, likely better performance profiles for querying data lake formats, and possibly tighter Lambda integration. If you’re on AWS, this is probably good news.
Long term
The risk isn’t that DuckDB gets closed, it’s that it becomes strategically aligned with AWS to the exclusion of other platforms. If you’re on Azure or GCP, or if you’re running on-prem, you may find that DuckDB’s performance advantages increasingly favor AWS object stores. This is how cloud lock-in works: not through force, but through convenience.
The Postgres data lakehouse trajectory shows what happens when beloved open-source databases start absorbing adjacent functionality. DuckDB’s path under AWS will similarly be shaped by where AWS wants analytics to go.
The Verdict: Cautious Optimism, Careful Skepticism
The DuckDB team has earned an extraordinary amount of trust over the past decade. They took the road less traveled, stayed independent, and built something genuinely beloved. Their commitments to keeping the project open source appear sincere, and the governance improvements (technical advisory board, open extension signing) are concrete rather than cosmetic.
But the SQLMesh vs. dbt battles and the broader open-source data tool consolidation we’ve watched over the past few years should temper our enthusiasm. Corporate stewardship of open source is a feature, not a bug, as long as the incentives align. And right now, AWS’s incentive is clear: make DuckDB the default query engine for data stored in S3, then charge for the compute that comes with it.
For the rest of us, the relational AI workloads that DuckDB powers will continue to work exactly as before. The question is whether the duck stays free to swim wherever it wants, or whether it becomes a very well-fed, but ultimately captive, bird in the AWS pond.
Watch the extension registry, watch the S3 integration roadmap, and watch whether DuckDB’s performance optimizations start favoring AWS’s data formats. Those will tell you everything you need to know about who really owns the future of this project.




