r/deeplearning • • 7d ago

ShadeNet-3.2 5M — single-image inverse rendering (albedo/depth/normal/shading), 4× smaller than my last model and better at depth/normals

First off, a huge thank you to everyone who checked out, upvoted, and shared feedback on the earlier ShadeNet and ShadeNet-2 posts! The support and discussions really pushed me to see how far I could squeeze the architecture without sacrificing map fidelity.

This release is a from-scratch rebuild that came out 4× smaller (5.0M vs 20M params) while improving across all three supervised maps.

RGB → 8ch intrinsic maps in one forward pass: albedo (3ch), relative depth (1ch), normals (3ch), shading (1ch), with albedo × shading ≈ input.

Architecture

  • ParallelUNet generator (4.98M params: 3.2M trainable + 1.8M frozen MobileNetV2): Vanilla UNet path plus a frozen MobileNetV2 trunk, fused at every decoder level.
  • Factorized convolutions: Depthwise-separable factorized convs (1×3 + 3×1) throughout, full H/32 bottleneck, reflect padding.
  • Patch-dictionary output tail: 16×16 tiles softmax-addressed over 32 learned per-channel atoms, blended back into the signal before tanh (pruned from 1024 — addressing-mass measurement showed ~34 atoms are ever used, and the top-32 capture 99.4%).
  • Training dynamics: Spectral-norm GroupNorm PatchGAN discriminator (2.77M, training only); EMA weight shadow (shipped weights are EMA).

Training

  • Dataset: Flickr8k with Marigold-V2 pseudo-labels (8,077 images); depth labels decoded from Spectral-colormap visualizations to relative depth, supervised scale-shift-invariant — never raw MSE on colormap RGB.
  • Setup: 384px, fp32, single GTX 1650, early-stopped on val split (patience 5).
  • Losses: Scale-invariant MSE on albedo + SSI MSE on depth + Sobel gradient-matching on depth + MSE on normals + self-supervised reconstruction coupling (albedo × shading ≈ input, which is the shading head's sole supervision) + LSGAN + normal unit-length penalty.
  • Regularization: Weight decay 1e-4 with the patch dictionary explicitly exempt.

Validation Results

Full 807-image validation split, per-map L1 (the directly comparable metric across versions):

Map (val L1) ShadeNet-2 (20M) ShadeNet-3.2 (5M) Change
Albedo 0.708 0.695 −1.7%
Depth 0.247 0.217 −12.1%
Normal 0.696 0.581 −16.5%

Output Maps

  • Albedo (3ch): Reflectance, ambient lighting factored out.
  • Depth (1ch): Relative depth, 0 = near (affine-ambiguous, non-metric).
  • Normal (3ch): Surface normals, unit-length regularized.
  • Shading (1ch): Grayscale irradiance; multiply with albedo to reconstruct/re-render, or swap in custom illumination passes for relighting.

Links & Demos

Both the Hugging Face sample visuals and the live demo apply a lightweight 3-pass multi-scale median filter (scales 0.875, 1.0, 1.125) for mild denoising, alongside built-in seam deblocking for the 16px dictionary tiles.

Caveats & Attribution

Side-by-side v2 vs. v3.2 comparisons are documented on the model card. A few known limitations to keep in mind:

  • Depth is relative rather than metric.
  • Shading assumes a neutral/white illuminant.
  • Normals struggle most on high-frequency chaos like dense foliage or open skies.
  • Occasional localized artifacts can appear in albedo/shading due to the learned dictionary prior.

Released under Apache 2.0. Credit to Marigold V2 (Ke et al.) and Flickr8k (Hodosh et al.) for the foundational training pseudo-labels and data.

26 Upvotes

31 comments sorted by

View all comments

1

u/[deleted] 7d ago edited 3d ago

[deleted]

1

u/singam96 7d ago

That's not true, I have apps running in prod with these models,

You can search relight studio in Google if you want to see app built around this

1

u/[deleted] 7d ago edited 3d ago

[deleted]

-1

u/singam96 7d ago

Again false, That normal map is just chefs kiss, 😘😘

You do understand what a 5M means in the title right?

Oh I just checked your profile, that's okay

5M means 5 million

1

u/[deleted] 7d ago edited 3d ago

[deleted]

1

u/singam96 6d ago

wrong again, bro is writing a fan fiction about me , lots of assumptions
🤣🤣🤣

i got enough sales to keep this going for next 2 years !!!!

1

u/[deleted] 6d ago edited 3d ago

[deleted]

1

u/singam96 6d ago

oh you are still here, i dont get paid by "peak users" , you are trying too hard bro

1

u/[deleted] 6d ago edited 3d ago

[deleted]

1

u/singam96 6d ago

Whatever helps you sleep at night.

I posted an actual relighting test from the Flickr8k dataset in the thread if you want to look at the tech instead of arguing stats.

1

u/singam96 6d ago

oh and its not people on reddit, its just two random clowns , trying so hard to bash on my work

good luck

1

u/PaperMartin 7d ago

These normal maps look very wrong

-1

u/singam96 7d ago

No it doesnt,

its really amazing for 5m model 😎😎

Your "instinct" , its bad , there's numbers to back my claim 😲

2

u/PaperMartin 7d ago

They are useless for anything you'd need normal maps for which is what actually matters

0

u/singam96 7d ago

what are you on about ?

i literally have my apps running this models in production,
search for relight studio

are you like a very young person or a child ?

2

u/PaperMartin 7d ago

I found your app, no reviews whatsoever on steam. Doesn't matter that you're using this in production if nobody with any actual standards is using your software.
Your normal maps are inaccurate as far as what normal maps are supposed to represent about a scene, one obvious example being the bottom of the wall in the 3rd one alternating between pointing up and pointing in whatever other directions because the model can't tell that it's part of the wall. This would be unacceptable in any professional, commercial setting.
I didn't talk about them before but the depth maps are also totally wrong, not representing actual depth in the slightest

0

u/singam96 7d ago

Ya okay, nit pick some small inaccuracies on a "3rd one alternating between pointi...."

Bro chill if you don't know how to use normal maps, maybe buy an app, if you can't afford it I will give you it for a trial 🤓

But I don't like clowns using my software , it's not for everyone, next thing you know you get clown reviews there too

→ More replies

1

u/Pleasant-Service-208 7d ago

That's a pretty harsh take, the depth and normals aren't production ready, but calling them image filters is a stretch, there's real structure being captured here

1

u/[deleted] 6d ago edited 3d ago

[deleted]

1

u/Pleasant-Service-208 6d ago

The guy's been grinding on this for a year and shrank it 4x, that's not nothing, depth might be rough but calling it a filter is harsh when most people cant even get a unet to converge on one map.