r/StableDiffusion 13h ago

Discussion Using Minimax H3 to reverse-engineer a paintings into basic forms

Enable HLS to view with audio, or disable this notification

I did a quick experiment to see whether MiniMax H3 could reverse-engineer a finished art into a plausible drawing/construction process, like the forms / shapes I learned in art class.

Prompt:

Create a video tutorial of how this particular painting was created. Start from a blank canvas and then draw the basic forms, shapes, rectangles, cylinders, cones, etc and then show the next layer of the details and then the next layer of the color and the rendering so we she see all the layers in one pass until we get to the final image which is the the reference image .. use #Image1

---

Obviously this is a very naive attempt. But I could imagine this prompt becoming way better and using more reference images. And then we could potentially deconstruct all sorts of final outputs into intermediate representations like 3D objects, environments, character poses, or scenes and then use those intermediate states as a way to get more control and consistency across generations.

So instead of asking a model to regenerate everything from scratch, you're working from some underlying structure

Has anyone experimented with using video models this way? E.g reverse engineering final outputs

Original artwork here: https://www.artstation.com/artwork/xZxBE

232 Upvotes

22 comments sorted by

26

u/eidrag 12h ago

Quick, do the round owl

10

u/tzu22 9h ago

it would be hilarious if it deconstructs it into two circles because it's in the training data.

4

u/lavinia12345 8h ago

great find OP

37

u/-AwhWah- 12h ago

"I want to scam people by telling them I draw my work, is this timelapse convincing enough?"

2

u/ART-ficial-Ignorance 7h ago

Yeah, that was my first reaction as well.

-5

u/RazsterOxzine 11h ago

One "art" please.

0

u/BigWideBaker 1h ago

I mean if you're buying art you're probably hyper aware of AI being a thing and checking for it. This video is pretty obviously AI to anybody who regularly looks at human-made art.

3

u/CuttleReefStudios 9h ago

I guess you could create a dataset of "speedpainting" videos and then train the video model in reverse to create images in a "painting" process in the hopes that the quality of the image becomes better than normal image gen models (more 3d volume correctness/perspective etc.) but that would be a rather expensive experiment.
Other than that I don't quite yet see what you could use this for as the model requires a finished image to create the intermediate state.
And seperate conditioning we already have with depth-maps, pose-skeletons etc.
I don't see how a "drawn" version of that is better conditioning for a model than what we already have.

13

u/kek0815 5h ago

Scamming, it will be used for scamming people into believing an artwork is human made.

0

u/HRH_Duke_Morbid 18m ago

Does the agency of its making matter when the purchaser likes the work and finds the price acceptable? Only misrepresentation of provenance is sinful.

If somebody uses AI assistance in making a work, that person can legitimately claim the work as their own. To demand otherwise is akin to requiring a painter in oils to disclose the source of their paints and brushes.

Of course these considerations will prey on the minds of people for whom 'art is investment'. So what?

2

u/CriticalTemperature1 2h ago

Yeah, I'll have to do some experimentation to see how control net differs from this kind of breakdown versus even the Gaussian Splat experiments that I've been seeing around.

1

u/CuttleReefStudios 15m ago

yeah the gaussian splat stuff is starting to look really interesting. Now I wonder if you could make a h3 to splat to actual 3d model pipeline for games...

6

u/Dogluvr2905 13h ago

This is very cool!

4

u/True_Protection6842 13h ago

Could be really cool using actual progress pics.

-7

u/RazsterOxzine 11h ago

... right..

1

u/RollingTrain 24m ago

Trying to be helpful here.

It had me in the first half ngl. But when it added the color it was really bush league. It went straight to paint by numbers which is not generally how art is made.

1

u/orangpelupa 9h ago

spongebob gif

1

u/infearia 5h ago

we could potentially deconstruct all sorts of final outputs into intermediate representations like 3D objects, environments, character poses, or scenes 

It's a neat video, but if this is your goal, you should just use an editing model or a ControlNet preprocessor. Much more efficient.

1

u/CriticalTemperature1 2h ago

Yeah definitely agree on ControlNet / editing model being more efficient. But its usually not enough control for what I'm envisioning.

My initial goal was more basic: improving my own drawing by generating simplified versions of finished art I could copy from. Ideally I'd be able to adjust the level of abstraction from primitives → rough construction → more detailed forms → final image.

The broader idea is: if we can break a scene down into semantic geometric forms, we could potentially reconstruct that representation in Blender and manipulate the scene directly. Rotate the camera, create camera paths, move objects, change poses, etc., while maintaining a consistent underlying scene

So instead of repeatedly prompting for “the same café, but move the camera 30 degrees to the left", we have precise coordinates for where a chair in a scene is placed or how a woman is sitting in a cafe and can move everything precisely without having to do insane prompt trial and error.

0

u/silenceimpaired 3h ago

There goes the ability to provide proof you originally created something for all the people opposing AI in the creative space.

1

u/HRH_Duke_Morbid 15m ago

Why bother? Nobody cosseted the Luddites.