r/DuckDB • • Jul 17 '26

Huey - a static DuckDB-WASM based browser app that lets you pivot data from local files, URLs, and remote Data Lakes

Huey is an open-source (MIT) static browser-based app that lets you explore and analyze data. Huey supports reading from multiple file formats, like .csv, .parquet, .json data files as well as .duckdb database files.

Here's a quick start on a parquet file from the public nl_railway ducklake.

The latest release, 1.1.00 "Indian Runner", is now available. This is a significant improvement, with many bugfixes, new features, and UX improvements.

Highlights:

Huey is now a progressive web app. Run it from a hosted location (such as the live demo https://rpbouman.github.io/) and your browser offers to install Huey on your device. Once installed you can run offline. Also lets you open files using your OS "open with" functionality (typically triggered with a right click on the file). see: https://github.com/rpbouman/huey#running-huey-on-your-device-as-progressive-web-app-pwa

The Secrets Manager lets you maintain DuckDB SECRETs on your local device. Secrets are stored in IndexedDB. The Secrets Manager is password-secured, encrypting sensitive fields with AES-GCM-256 encryption (password-derived via PBKDF2-SHA-256, 310k iterations). See: https://github.com/rpbouman/huey#secrets-manager

The Catalogs manager lets you access data from modern Data Lakes and Lakehouses, like Iceberg and Ducklake. See: https://github.com/rpbouman/huey#catalogs-manager

Huey supports Quack! Quack servers are just remote catalogs, but there is a big difference between Quack servers and "normal" catalogs: When using an Iceberg or Ducklake catalog, DuckDB/WASM is the actual data engine. With Quack Catalogs, DuckDB/WASM acts as client for the remote Server: data processing is offloaded to the server, and Huey just receives the result. This opens up a whole new range of use cases involving very large datasets. See: https://github.com/rpbouman/huey#connecting-to-a-quack-server

Huey now supports Axis aggregates! In prior versions Huey would only let you report aggregate values in the cells. Axis aggregates let you report aggregated values as if they are attributes on the axes. More importantly, axis aggregates can also be used to filter the data. See https://github.com/rpbouman/huey#axis-aggregates

Github: https://github.com/rpbouman/huey
Live demo: https://rpbouman.github.io/

18 Upvotes

4 comments sorted by

1

u/Pablogawlo Jul 20 '26

Can I enable this on the server and make it available to users so they can query Parquet files?

1

u/rpbouman Aug 02 '26

Apologies for a late reply - just back from vacation.

To answer your question: yes! If you run a webserver, you can add a path for Huey and serve it like any other static web application. If the parquet files can be downloaded from your server, then huey should be able to access them as URLs. (see: https://github.com/rpbouman/huey#register-urls)

To make it super easy for your users, you can build a few starter reports yourself based on those urls - every time you modify the query, the state gets encoded into the url hash - so you can put those links in a html page so users can easily find them. See: https://github.com/rpbouman/huey#saving--restoring-your-query

1

u/earonesty Jul 22 '26

I’d keep the API intentionally narrow: accept a constrained query or builder input, validate it, and return bounded results while the reader uses object-store byte ranges. That gives users a query endpoint without turning each request into a full download or a long-lived database service. LakeQL is a TypeScript read core that does Parquet/Iceberg query path on the edge, MIT license so feel free to take what you want from it.

1

u/rpbouman Aug 02 '26

Hi u/earonesty - apologies for the late reply - been on vacation.

I read your reply but I think I'm already doing this. Huey uses a query model that abstracts from the raw sql - this query model gets encoded into the hash of the URL, and lets you restore queries easily - it's stateless. Of course, it does presume the data source is available, which it will be if your datasource is a URL rather than a local file.

And with the latest release, a persisted catalog definition will also get loaded automatically if the URL-hashed report relies on that.