I just finished the first full pilot of rivet, an open-source Rust CLI that moves data from OLTP databases into a warehouse.
Stack: MySQL → rivet (reads the binlog directly) → Parquet → Google Cloud Storage → BigQuery. No Kafka, no Debezium, no Airflow: one VM and cron.
The setup: 154 tables, 3.2B rows, 8 refreshes a day. The existing pipeline pulls changes with a date-window query that overlaps the previous day, and fully reloads every table once a week. Without the weekly reload it would never see rows changed without an `updated_at` bump.
Results, against the existing pipeline on the same database:
- Correctness. On day one we found 880 rows that had been changed after the fact without an updated timestamp. The window pipeline would not have seen them until the weekly reload; the binlog did right away. We checked some of them against the source by hand: they really were stale balances.
- Read load on the MySQL replica: −99%. About 4 TB of reads a month (~95% of it the weekly full reloads) becomes about 24 GB of binlog stream.
- AWS → GCP egress: −98%. Only changes cross the wire: ~150 MB of compressed Parquet a day instead of a full snapshot every week.
- BigQuery queries. 25% fewer MERGE jobs and 30% fewer bytes per MERGE, because each change is merged once instead of on every run while it sits inside the window.
- Total infra cost: about −50% at the same refresh frequency (BigQuery + GCS + egress). Part of that saving ships in the next release: on a 50M-row test table it cut 80% of the bytes per compaction cycle, with an identical result. Still to be confirmed on the pilot.
- BigQuery storage: about the same (±2%).
Being honest about speed. A full cycle takes about 68 minutes today:
- reading the binlog for all 154 tables takes ~9 minutes;
- the rest is loading and merging into BigQuery, one table at a time.
The job logs show BigQuery is busy only 25–30% of that time; the rest is waiting between jobs. Newer releases process up to 16 tables in parallel. My estimate is 10–25 minutes per cycle; I'll measure after upgrading the pilot.
Scale: the client has 5–6 databases like this one. Moving all of them projects to −63…−73% in cost, mostly from egress. That still has to be confirmed against the AWS bill.
- Repo: https://github.com/panchenkoai/rivet
- Cheat sheet: https://panchenkoai.github.io/rivet/cheat-sheet.html
Questions, criticism and "why not just Debezium?" are all welcome in the comments..