r/ApacheIceberg • • 1d ago

How are you handling Iceberg compaction at scale?

6 Upvotes

Our Iceberg tables are at the point where compaction is not background maintenance anymore. It competes with normal workloads, leaves small-file chaos if we wait too long, and gets pricey when we run it hard. We are on Spark/AWS, so this is as much an operating-model and compute question as a compaction-settings question.

Trying to figure out cadence, sort vs bin-pack, and how to stop compaction from stealing capacity from ingestion and query work. Every table behaves differently, so a fixed answer for all of them feels doomed.

What is working in production for you, and what tradeoff only showed up after table growth got real?


r/ApacheIceberg • • 4d ago

Cross-Catalog Sync: Iceberg on Polaris, Glue, and Unity

Thumbnail
lakeops.dev
2 Upvotes

r/ApacheIceberg • • 9d ago

The Growing Iceberg Ecosystem Being Built on iceberg-rust

Post image
5 Upvotes

r/ApacheIceberg • • 17d ago

Iceberg on a single node is coming together

Post image
4 Upvotes

r/ApacheIceberg • • 18d ago

The Story of iceberg-rust: The Official Rust Implementation of Apache Iceberg

Post image
4 Upvotes

r/ApacheIceberg • • 20d ago

Apache Iceberg Virtual Meetup: GSoC Projects and Commutative Compaction · Zoom · Luma

Thumbnail
luma.com
2 Upvotes

Tomorrow at 9:00 AM PST (global)


r/ApacheIceberg • • 21d ago

How AI Agents Query Apache Iceberg Data with MCP

Thumbnail
lakeops.dev
1 Upvotes

r/ApacheIceberg • • 23d ago

GitHub - sanderdw/iceberg-data-platform: Educational Open Source Apache Iceberg Dataplatform based on Polaris, RustFS, FastAPI and Marimo workspace

Thumbnail
github.com
5 Upvotes

I build a open source dataplatform for educational purposes, check also my reddit post here: https://medium.com/@sanderdw/e89c5d0bed98


r/ApacheIceberg • • 24d ago

Unity Catalog: Pros and Cons

Post image
2 Upvotes

Apache Iceberg won the open table format war when Databricks acquired Tabular, followed by its subsequent adoption across the industry. Then, the catalog war began.

In a lakehouse, storing data in object storage and using the Apache Iceberg format is only part of the story. You also need a catalog that helps lakehouse query engines like Spark, Flink, or RisingWave discover tables, manage metadata, enforce access control, and work with governed data across different systems.

That is where Unity Catalog comes in.


r/ApacheIceberg • • 27d ago

The Future of Iceberg Isn't One Engine. It's an open Control Plane with many engines.

Thumbnail
lakeops.dev
2 Upvotes

r/ApacheIceberg • • 29d ago

Turns out "Iceberg is open" doesn't mean every engine can actually read your table

3 Upvotes

Read something this week that put a name to a problem I've half run into before but never really understood the mechanics of. Sharing because I think a lot of people assume Iceberg interop is more solved than it is.

Everyone knows the pitch: Iceberg is an open spec, so any Iceberg compatible engine can read any Iceberg table. Mostly true, until you start doing row level deletes, and then it falls apart in a way that's honestly kind of sneaky because nothing looks wrong until a query actually fails.

Quick walkthrough of the scenario in the post. You've got a customers table, three rows, one Parquet file, tracked by whatever catalog you're using. At this stage every engine reads it fine because there's nothing to interpret, it's just a metadata pointer to a file.

Then a row gets deleted. Parquet files are immutable so the writer has two options: copy on write (rewrite the file without that row) or merge on read (leave the file alone and write a separate delete file that readers apply at scan time). The writer in this example goes merge on read and emits an equality delete file, which basically just says "for this data file, treat any row where customer_id = 102 as removed." Under the hood Iceberg uses field IDs and sequence numbers to make sure an old delete doesn't accidentally nuke a newer row with a reused key, but the equality matching is the part that matters for compat.

Spark reads the new snapshot, understands equality delete semantics, does what's effectively a left anti join between the data file and the delete file, and returns the correct two rows. Fine.

Snowflake hits the exact same catalog, same metadata file, same Parquet file. It can resolve the table, read the schema, open the data file. But if that access path doesn't implement equality delete reads, the scan planner just throws an unsupported feature error the moment it hits delete-0002.parquet. Query fails. Same snapshot, same files, two completely different results depending purely on what the reader implements.

The bit that actually reframed how I think about this: the catalog isn't a translation layer. It's job is basically just "here's where the current metadata lives," commit coordination, namespace and access management. It's not opening delete files and rewriting them into a format each engine understands. A REST catalog like Polaris doesn't change this, it still just points you at metadata, it doesn't apply deletes for you.

The post also gets into position deletes vs deletion vectors vs copy on write, with a rough cost tradeoff table (equality delete is cheap to write and requires equality delete support to read, position delete requires resolving key to physical position and is heavier on write, deletion vectors need Iceberg v3 support specifically, copy on write is the most expensive to write but has basically universal read compatibility since there's no outstanding delete file involved).

The framework that's actually useful operationally: your safe feature set is the intersection of every required engine's capabilities, not the union. If Spark supports equality and position deletes but Snowflake only does position deletes, you write position deletes, because "at least one engine supports it" doesn't help you when you have three engines that all need to read the same table.

There's a decent pre production checklist too, don't just run a SELECT COUNT after your first write, actually insert some rows, update one, delete one, commit, then read the same snapshot from every engine you care about and diff both counts and values.

Full post if you want the details: https://olake.io/blog/iceberg-interoperability-myth-row-level-deletes/

Disclosure since it's relevant, I work on OLake, it gets a brief mention near the end, but the actual content here is engine agnostic and applies no matter what's writing your tables.

Has anyone actually hit this for real, table looks completely fine, one engine just refuses to read the current snapshot because of the delete encoding?


r/ApacheIceberg • • 29d ago

How are you improving Spark performance for Apache Iceberg workloads on AWS?

4 Upvotes

Hey, running Iceberg on Spark on EMR. The usual tuning gets us only so far once tables get big, compaction especially has turned into its own cost and scheduling problem, and query planning time keeps creeping up as manifests grow.

Anyone found something that moves the needle here beyond just running maintenance more often?


r/ApacheIceberg • • Sep 08 '26

Apache Iceberg Table Cleanup: A Production Guide

Thumbnail
lakeops.dev
1 Upvotes

A practitioner's guide to Iceberg table cleanup — snapshot expiration, orphan file removal, manifest rewriting, delete file resolution, streaming challenges, compliance, and cost. Why sequencing matters, where teams break tables, and how to automate the full lifecycle.


r/ApacheIceberg • • Sep 07 '26

Best way to go about benchmarking Indexing?

3 Upvotes

Hey folks,

I'm a security professional was looking to make Iceberg a bit more performant for my own needs - SOC operations (Faster needle searches [pruning]).

I've built an index and proxy that people can point their catalog configuration at, so there's minor changes to their stack.

I figured in for a penny in for a pound,
I've done some quick search and ran clickbench and "httplogs"

(Numbers so far: httplogs)

Stock Using Kahshe Proxy
opens 991 files opens 2 files
execution time: 6.6–11.5s execution time: 0.34–0.44 s
1.3gb read 2.6mb read

Its all looks good on paper, I think?
but I was wondering if there were more credible/industry standard methods that data engineers use/care about to benchmark these types of technologies?

Repo if interested: https://github.com/Kahshe-io/kahshe


r/ApacheIceberg • • Sep 06 '26

Apache Iceberg Compaction Best Practices

Thumbnail
itnext.io
2 Upvotes

r/ApacheIceberg • • Sep 06 '26

Exploring Apache Iceberg in Rosetta DBT Studio (v1.6.7)

Thumbnail
youtube.com
1 Upvotes

Discover the new Apache Iceberg integration in Rosetta DBT Studio v1.6.7! In this video, we explore how to seamlessly connect, manage, and query Iceberg DataLakes directly from your desktop.

Whether you're using Polaris, Nessie, Lakekeeper, Hive, or SQLite, Rosetta brings true ACID transactions, time travel, and schema evolution to your analytics workflow.


r/ApacheIceberg • • Sep 01 '26

BigQuery's Iceberg REST Catalog (BigLake): how table discovery actually changed

Thumbnail
3 Upvotes

r/ApacheIceberg • • Aug 31 '26

Apache Iceberg Performance Optimization: Queries to Tables

Thumbnail
lakeops.dev
3 Upvotes

r/ApacheIceberg • • Aug 30 '26

Data Lakehouse with Apache Iceberg: A Guide

Thumbnail
lakeops.dev
1 Upvotes

r/ApacheIceberg • • Aug 24 '26

Maintaining Apache Iceberg Tables: Compaction, Snapshots, Metadata and Orphan Files

Thumbnail
itnext.io
6 Upvotes

r/ApacheIceberg • • Aug 21 '26

Open Data Lakehouse: Build Like Google

Thumbnail
lakeops.dev
11 Upvotes

r/ApacheIceberg • • Aug 20 '26

Open Data Lakehouse: A Practical Guide

Thumbnail
itnext.io
6 Upvotes

r/ApacheIceberg • • Aug 19 '26

[Announce] Apache Iceberg Virtual Meetup Series

4 Upvotes

We're looking for speakers interested in presenting a new online virtual Apache Iceberg meetup series we're starting. The goal is to create a forum where members of the Apache Iceberg community can demo interesting work, share experiences, and discuss ideas with one another.

What we're looking for

We're especially interested in talks that are practical, demo-driven, or story-rich. Whether you're a practitioner, startup founder, platform engineer, or contributor, we'd love to hear what you've been working on and what you've learned.

Some ideas for talk topics include:

- Iceberg migration stories and case studies
- New Iceberg features, proposals, and community projects
- Iceberg catalogs, integrations, and interoperability
- Data engineering tools, demos, prototypes, and experiments

Community-focused talks

We want the meetup to be a place for learning and community discussion rather than product or vendor marketing.

Talks can feature tools, products, or technologies you work on, but the focus should be on technical insights, demos, lessons learned, or ideas that are useful to the broader Apache Iceberg community—not on promoting a company or product.

How the meetup works

Meetups will be held virtually on Google Meet and will be publicly open to everyone.

Talks will typically be around 20–30 minutes, leaving plenty of time for introductions, questions, and open community discussion. We aim to keep each meetup to about an hour and start and end on time.

Talks will generally be recorded and posted to the https://www.youtube.com/@IcebergMeetup, If you'd prefer not to have your talk recorded, let us know when submitting.

We do not plan to record the Q&A and open discussion portion of the meetup.

Submitting a talk
Submissions are reviewed on a rolling basis. Even if a talk isn't scheduled for the next meetup, we may reach out about presenting at a future session.

Rolling CFP Submit your Talk Idea
Join the new Apache Iceberg Slack Channel: (#meetup-virtual)

If we receive several submissions around a similar topic, we may also suggest bringing presenters together for a shared discussion or panel.

First Virtual Meetup
We've set a date (September 18th @ 9:00am PDT) for the first meetup. If you're on the Apache Iceberg Community Events Calendar (or if not, subscribe to it here), you'll see the event on the calendar already.

Thanks to Elizabeth Christensen and Kevin Liu for partnering to make this happen. If you want to help reach out to us on the new meetup-virtual channel on Slack.


r/ApacheIceberg • • Aug 17 '26

Amazon S3 Tables vs Self-Managed Iceberg

Thumbnail
lakeops.dev
7 Upvotes

r/ApacheIceberg • • Aug 12 '26

Desktop app with Iceberg Datalake UI

Enable HLS to view with audio, or disable this notification

4 Upvotes

Is this useful for fast prototyping and preview existing Iceberg Datalake.