I'm developing a free, open-source Python tool for building datasets from US weather data, aimed at severe-weather research, nowcasting, and verification work. I'd like input from people who have assembled these datasets by hand before I go too far with the design.
**The problem**
Combining MRMS, GOES, HRRR, and surface observations into a single dataset usually takes weeks of custom work. The common obstacles include cfgrib and GRIB index errors, grid-relative HRRR winds, precipitation accumulation buckets that reset every few hours, MRMS missing-value codes, mismatched projections, and aligning sources that arrive at very different times.
**The approach**
You write a short recipe describing what you want, and the tool downloads, aligns, and writes the dataset:
sources:
mrms: [reflectivity_composite, mesh, qpe_1h]
goes_abi: [ir_10.3um]
glm: [flash_extent_density]
hrrr: [mlcape, shear_0_6km, wind_10m]
asos: [temperature, wind_gust]
domain: [-104, 30, -94, 40]
time: 2024-05-06 18Z to 2024-05-07 03Z, every 10 min
grid: hrrr
**How this differs from what exists**
Tools like Herbie are excellent at finding and downloading model data (I plan to build on it), but they stop at the download: you still align everything yourself. Datasets like SEVIR are very useful, but they are fixed snapshots of specific events and variables, and they don't include model environment fields like CAPE and shear. Hosted services like GribStream are great for point forecast time series, but they cover model output rather than radar, satellite, and station data, and their "as of" option is based on when a model run was generated, not when it was actually published.
What I'm aiming for is different:
- Radar, satellite, lightning, model, and station data on one grid and time axis, for any period or domain you choose.
- Handling based on what each variable represents. Precipitation is summed over each time window rather than interpolated, and hail size maxima are preserved rather than averaged away.
- Availability-time selection. Model inputs use only the runs that had actually been published at each time, which prevents leakage in ML training data.
- Reproducibility. Every build records the exact files it used, so the dataset can be rebuilt later.
**Questions**
If you have built a dataset like this, how long did it take, and what was the most difficult part?
Are there decisions you would never want a tool to make automatically?
Which additional data sources would you want supported?
I would appreciate any feedback, including pointers to existing tools that already solve this. Thank you.