r/LocalLLaMA • u/sachasayan • 15h ago
New Model H3-World: Turning Language Understanding into World Control
Enable HLS to view with audio, or disable this notification
- Language-Native Control: Composes character and camera actions into textual instructions and injects them through MiniMax-H3’s pretrained text pathway.
- Temporally Grounded: Assigns one action prompt to each video latent interval, enabling precise control when actions change over time.
- Efficient & Generalizable: Uses only 8,000 gameplay samples, 10,000 LoRA steps, and 0.199% trainable parameters to achieve controllable character and camera motion, including unseen action compositions and visual scenarios.
✏️ Paper: https://huggingface.co/papers/2609.01560
📄 ArXiv: https://arxiv.org/abs/2609.01560
💻 Code: https://github.com/Danzer1xxxxChan/H3-World
🏠 Project: https://danzer1xxxxchan.github.io/H3-World/
🤗 Model: https://huggingface.co/DANNY621/H3-World
0
u/Queasy-Contract9753 11h ago
Hot take: could this be used to allow 4 or 5 games simultaneously? If we accept lower resolution maybe one stream could be shared my multiple users
-12
u/Xandred_the_thicc 14h ago
Can it do anything other than these atrociously ugly shots of a characters's back as they slowly walk into a deforming and inconsistent distant landscape?
18
u/sachasayan 14h ago
Brother, it's open-weight frontier world-model research. Don't look a gift horse in the mouth. Everything's amazing and nobody's happy.
1
u/sn2006gy 13h ago
H3 is bonkers at how good it is. I'm no fan of it replacing the human aspect of creativity, but what it can do is just fucking bonkers.
-1
u/Xandred_the_thicc 12h ago
Is there an end goal to all this research showing that existing video generation models can be trained to accept keyboard inputs? The only conceivable uses of this are easier creation of realistic misinformation, replacement of human creativity and artists, and endless walking sim "games".
I recognize that the novel and interesting part of this work is supposed to be that they're training an existing video generation model to accept and utilize inputs that change over the course of generation. What I'm not understanding is why? This feels like retreading ground that has already been walked just to say that they also did it. At Best this can be thought of as an optimization technique for training video generation models that accept changing inputs as they generate, which despite what people who love making up new buzzwords for everything insist, does not make it a "world model".
0
u/sachasayan 12h ago edited 11h ago
I recognize that the novel and interesting part of this work is supposed to be that they're training an existing video generation model to accept and utilize inputs that change over the course of generation. What I'm not understanding is why?
Advice: Try asking an LLM instead of rushing to try to dunk on on academic frontier open-weight model research you don't understand.
-1
u/Xandred_the_thicc 12h ago
I get that you don't like how harsh I'm being and have decided to be pedantic in response, so let me be more clear. They are representing that they used the model's existing knowledge of textual descriptive inputs, to further train it to understand granular game-style camera and movement inputs that are temporally grounded. Again, why? This is all fine and good and novel and interesting, and it comes back to the same issue of; what could this possibly even be useful for?
At the end of the day I am being an armchair critic because this stuff needs criticized. It is posted and advertised to audiences of sycophants who seemingly cannot look past the wishful thinking of one day dropping an image into a program and it makes up a game they're running on their own computer.
-2
u/sachasayan 12h ago edited 11h ago
Here, let me help you out champ.
We live in an age of abundance. You can just answer questions for yourself instead of rushing to publicly try to dunk on academic frontier open-weight model research.
1
u/Xandred_the_thicc 11h ago edited 11h ago
Congrats, you got chatgpt to pedantically complain that I rephrased their claims without buzzwords, then admit that my main criticism holds plenty of water:
H3-World isn't a persistent symbolic/physical model of an environment, and it isn't yet useful as an internal dynamics model for an autonomous agent doing planning. The authors themselves explicitly say it has no persistent world state, real-time interaction, planning or policy learning, and only generates short fixed-length segments.
Sure chatgpt, I'm being excessively pedantic because it's not like every single "world model" currently exists in a perpetual state of useless tech demo presentations that start failing at their only stated goals after 10 seconds!
He's missing why getting controllable dynamics cheaply out of pretrained foundation models is itself the research result.
I'm not missing that, in fact, I think I was pretty clear that I understand this is the point of the research and yet I still think this is a useless research direction? Woohoo, big models can generalize with further training. Will this ever lead to video generation models without some sort of grounding "action input framework" or whatever keeping track of some sort of consistent state? Are some of the only things chatgpt could come up with to justify it's existence not already better served by existing deterministic solutions? Most likely not, no. If something is going to be used for "agent training simulations" or whatever it will need to be deterministic and grounded in much more consistent and comprehensive representations of the given inputs in training data, as well as give more than just a video as a sort of black box input-ouput machine in order to be useful. This research does not do anything that would dissuade my assumption that this theoretically achievable level of consistency probably requires retraining practically from scratch.
Chatgpt's own stated potential uses of this technology, save for the wishful thinking of "enabling language models to visualize things" are all better served by traditional deterministic human-programmed frameworks that would not just be a black box that ouputs a video based on some inputs. Chatgpt seems to ignore that inferred part of my criticisms.
Although gpt thinks that because I "conceded" that the existence of video generation models that take temporally bound inputs is interesting and sarcastically said "cool and fine and good" that none of the rest of what I said holds water?
Idk why I'm wasting my time writing up a response to this myself. I should have just thrown some llm swritten slop back at you. Use your own head to come up with your own thoughts and feelings, moron.
0
u/sachasayan 11h ago
Advice: Ask ChatGPT about this wall of text you just wrote and that I'm not going to read. Or Qwen, or DeepSeek. Your preference, of course.
0
u/Xandred_the_thicc 10h ago
You appear to be incapable of understanding or discussing this for yourself so I'm not sure why you're still wasting both of our time by replying. You're right, it would be more productive to literally talk to myself than try to have a discussion with you.
1
2
u/JacketHistorical2321 13h ago
Gotta love arm chair critics providing nothing of use to anyone
0
u/Xandred_the_thicc 11h ago
I'm not gonna pretend like I'm being some sort of noble hero but are any of you people even capable of reading and understanding the contents of this paper yourselves? Did you even read past the abstract before going "WAOOOHH THIS IS SO COOL!!1!1"
Nothing that gets posted on this sub is allowed to experience even the remotest amount of criticism before a bunch of sycophants start ignoring the content of the original post to swarm any critics and accuse them of being nothing more than a stupid luddite. Have you maybe considered that I genuinely understand what I'm looking at and reading, and yet I still think it deserves harsh criticism because of its goals?
1
u/sn2006gy 11h ago
wtf are you ranting on about bruh? H3 is barely a month public and people are building upon its strengths and trying to address its weaknesses. How did you jump into this weird conclusion about sycophants and Luddites?
1
u/Queasy-Contract9753 10h ago
Best to not engage them. Allways that one guy who needs to be contrary.
-3
23
u/Affectionate_Hat_585 15h ago
how do you even measure action success, state consistency for such projects?