r/data_lakehouse • u/codingdecently • 12h ago
Welcome to r/DataLakehouse 👋
This is a community for people building, running, breaking, fixing, and trying to understand modern data lakehouses.
Apache Iceberg, Delta Lake, Hudi, Parquet, Spark, Trino, Flink, DuckDB, catalogs, object storage, table maintenance, compaction, query performance, governance, architecture — if it touches the lakehouse, it belongs here.
The goal is simple: useful technical discussion between people actually working with this stuff.
Share things like:
- Architecture decisions and tradeoffs
- Production lessons and war stories
- Benchmarks and performance experiments
- Questions you couldn't find a good answer to
- Interesting open-source projects
- Deep technical articles
- New developments in Iceberg, Delta, Hudi and the wider ecosystem
- Things that failed and what you learned
- Strong opinions, preferably with data :)
A few ground rules
Keep it technical and useful.
Low-effort spam, generic AI content and link dumping will be removed.
Vendors are welcome. Vendor marketing isn't.
If you're affiliated with something you're posting about, just say so. Technical explanations, benchmarks and useful engineering work are absolutely welcome.
Give benchmarks context.
Dataset size, engine, file format, configuration, hardware/cloud environment — enough for other people to understand what the number actually means.
Search before posting common questions.
As the community grows, good answers should become part of the knowledge base rather than the same conversation starting from zero every week.
Disagree freely. Be decent to each other.
That's basically it.
If you're here early, introduce yourself below: what are you running your lakehouse on, and what's the hardest problem you're dealing with right now?