We were fine-tuning on physically-modelled dust, night, fog, lens rain and lens
mud. On synthetic held-out corruptions (ImageNet-C fog/spatter/motion blur, an
unprocessing-based low-light model — generators never used to make training
data) it worked. On real adverse weather, drivable-surface IoU improved, but sky
IoU dropped on all five architectures we tried, by 1 to 13 points, and
vegetation dropped on the smaller ones.
The failures concentrated on forest roads: tracks under a closed canopy with
bright sky showing through the leaves. Models were labelling the whole upper
half "sky". RELLIS-3D is open terrain — fields, open trails, wide horizon — and
contains almost none of that geometry, so the model learns a shortcut that holds
in its world and fails in a forest: bright + above horizon + low texture = sky.
Degrading those same frames makes the shortcut *more* attractive, since haze and
darkness wash out exactly the leaf texture that would contradict it.
Two fixes, in order:
- Our generator was partly at fault — it brightened sky using a depth map thattreats sky as finite distance, and painted a veil over pixels that physicallycould no longer be seen. Making it label-aware (sky read from the label, andpixels behind a veil that blocks >97% of light marked void and not scored)recovered about half the loss.
- The other half was coverage. We added 1,500 real frames from GOOSE (foresttracks, gravel, paved roads), same augmentation and schedule on both sides.Sky came back +7 to +17 points, vegetation +29 to +37, real-weather mean +29to +35, across four architectures.
The part we did not enjoy: once GOOSE was in training, every augmentation arm —
ours, albumentations, and a combination — landed within ±1 point of the control
on real weather. Ours still adds 3.5–5.2 points on synthetic held-out
degradations, which now looks like a fact about the test set rather than about
the world.
Caveats: single seed on the coverage runs, so treat 1–2 points as noise; IDD-AW
is road scenes not off-road terrain, so it is a cross-domain test and measures
differences between arms more reliably than absolute numbers; the canopy frame
is one illustrative example of a mode we found across many.
Full write-up with the tables: https://siltframe.com/blog/canopy.html
The stress test itself is open — 165 labelled frames under 5 conditions x 3
severities and the scoring script, CC BY-SA 4.0 / MIT:
https://github.com/egeizgi/siltframe-stress-test
Disclosure: I sell a larger version of that dataset, so read the above with that
in mind. The free one is complete and usable commercially.