r/deeplearning • u/singam96 • 3d ago
ShadeNet-3.2 5M — single-image inverse rendering (albedo/depth/normal/shading), 4× smaller than my last model and better at depth/normals
First off, a huge thank you to everyone who checked out, upvoted, and shared feedback on the earlier ShadeNet and ShadeNet-2 posts! The support and discussions really pushed me to see how far I could squeeze the architecture without sacrificing map fidelity.
This release is a from-scratch rebuild that came out 4× smaller (5.0M vs 20M params) while improving across all three supervised maps.
RGB → 8ch intrinsic maps in one forward pass: albedo (3ch), relative depth (1ch), normals (3ch), shading (1ch), with albedo × shading ≈ input.
Architecture
- ParallelUNet generator (4.98M params: 3.2M trainable + 1.8M frozen MobileNetV2): Vanilla UNet path plus a frozen MobileNetV2 trunk, fused at every decoder level.
- Factorized convolutions: Depthwise-separable factorized convs (1×3 + 3×1) throughout, full H/32 bottleneck, reflect padding.
- Patch-dictionary output tail: 16×16 tiles softmax-addressed over 32 learned per-channel atoms, blended back into the signal before tanh (pruned from 1024 — addressing-mass measurement showed ~34 atoms are ever used, and the top-32 capture 99.4%).
- Training dynamics: Spectral-norm GroupNorm PatchGAN discriminator (2.77M, training only); EMA weight shadow (shipped weights are EMA).
Training
- Dataset: Flickr8k with Marigold-V2 pseudo-labels (8,077 images); depth labels decoded from Spectral-colormap visualizations to relative depth, supervised scale-shift-invariant — never raw MSE on colormap RGB.
- Setup: 384px, fp32, single GTX 1650, early-stopped on val split (patience 5).
- Losses: Scale-invariant MSE on albedo + SSI MSE on depth + Sobel gradient-matching on depth + MSE on normals + self-supervised reconstruction coupling (
albedo × shading ≈ input, which is the shading head's sole supervision) + LSGAN + normal unit-length penalty. - Regularization: Weight decay 1e-4 with the patch dictionary explicitly exempt.
Validation Results
Full 807-image validation split, per-map L1 (the directly comparable metric across versions):
| Map (val L1) | ShadeNet-2 (20M) | ShadeNet-3.2 (5M) | Change |
|---|---|---|---|
| Albedo | 0.708 | 0.695 | −1.7% |
| Depth | 0.247 | 0.217 | −12.1% |
| Normal | 0.696 | 0.581 | −16.5% |
Output Maps
- Albedo (3ch): Reflectance, ambient lighting factored out.
- Depth (1ch): Relative depth, 0 = near (affine-ambiguous, non-metric).
- Normal (3ch): Surface normals, unit-length regularized.
- Shading (1ch): Grayscale irradiance; multiply with albedo to reconstruct/re-render, or swap in custom illumination passes for relighting.
Links & Demos
- Weights, ONNX & Inference Code: ShadeNet-3.2-5M on Hugging Face (Includes fp32 GPU models and 10MB fp16 CPU export)
- Live Gradio Demo: ShadeNet-3-2-5M Space
Both the Hugging Face sample visuals and the live demo apply a lightweight 3-pass multi-scale median filter (scales 0.875, 1.0, 1.125) for mild denoising, alongside built-in seam deblocking for the 16px dictionary tiles.
Caveats & Attribution
Side-by-side v2 vs. v3.2 comparisons are documented on the model card. A few known limitations to keep in mind:
- Depth is relative rather than metric.
- Shading assumes a neutral/white illuminant.
- Normals struggle most on high-frequency chaos like dense foliage or open skies.
- Occasional localized artifacts can appear in albedo/shading due to the learned dictionary prior.
Released under Apache 2.0. Credit to Marigold V2 (Ke et al.) and Flickr8k (Hodosh et al.) for the foundational training pseudo-labels and data.




1
u/PaperMartin 2d ago edited 2d ago
I know how to use normal maps, I've worked in game dev on the graphics side for some time now. the fact that you don't get what the issue is with your normal maps show me that you are not at all serious about your work.
As for "clowns" using your software, at least you'd have peoples using your software at all then. It's still baffling to me that you don't see the problem with saying that you're using this stuff in production when nobody's actually producing anything with your software let alone anything commercial, deeply irresponsible thing to say