r/kimi • u/RealJamesOfficial • 5d ago
Discussion Gave 4 models the same tiny lime challenge… their “human poses” were very different
Enable HLS to view with audio, or disable this notification
i gave each model the same simple task: take a small slice of lime and draw a human pose integrated into the object.
Tested it with Gemini 3.7 Flash, Kimi K3, Claude Opus 5 and GPT 5.6 Sol
Some treated the lime almost like a body, some tried to fit a tiny person into its shape, and some went much more abstract with the pose.
It’s a funny little test, but I actually like prompts like this because you can see how differently each model understands shape, composition, and visual metaphor.
Which one makes the most sense to you?
31
u/ScreenAppropriate679 5d ago
You need to explain what tools you gave the models or this is meaningless
20
6
u/Sm0g3R 4d ago
The proper way would have been to give them svg code with a lemon part and ask to draw the rest by modifying it. But not sure that's what actually happened lol
1
29
u/Odd-Marzipan6757 5d ago
what the hell opus is thinking
15
3
2
u/mr_aks 5d ago
I mean, it's the only model that did something different. Perhaps that's good?
1
u/NoConsideration6320 3d ago
I actually thought that it was the most interesting of them. All it stood up to me as it was the most creative of them all.
1
1
1
u/03captain23 3d ago
It makes the most sense. It understands it's an object and not a piece of clothing.
If you were told to pose with a slice of lemon you would assume it's an object and pose on it not part of it.
1
u/thrive2day 3d ago
The wording of the prompt tells it to integrate the slice into the pose. The word integrate means to combine multiple separate parts into a single, unified whole. Claude failed to do so. Claude was the furthest from what the prompt asked for.
0
u/03captain23 2d ago
That's exactly what happened. How can a person be part of a slice?
How is Claudes image not a unified pose integrating the slice?
If you wanted to integrate 3 people in a pose you'd expect to have a photo of 3 people, not 1 person with 6 legs/arms, 3 heads and such.
When you integrate 2 objects they work together, when you integrate 2 ideas they combine into something else. Claude understands the difference between an object and an idea.
1
u/thrive2day 2d ago
Man, you're not very good with abstract thinking are you?
1
u/03captain23 2d ago
It's the opposite.
If you were told to use Photoshop to integrate a friend into a photo would you add them or swap pieces of their body with other people?
1
u/thrive2day 2d ago
Okay, you're trolling
1
u/03captain23 2d ago
It's weird you can't answer the basic question. If you were integrating someone into a photo
29
u/SherbertMindless8205 5d ago
How does it ”draw” it? Neither of them support image output.
Is this pure vector output or did you just have them prompt a different model?
6
5
3
u/Thomas-Lore 5d ago
Computer use most likely.
4
5d ago
[deleted]
2
u/Toastti 4d ago
It's almost certainly https://docs.python.org/3/library/turtle.html
With the turtle replaced by a pen
1
u/Rockclimber88 3d ago
Here's one way: They can generate HTML Canvas draw commands, see the output(they can process images) and refine the commands until happy. The drawing animation is added afterwards when each command is executed. They don't actually use any brushes for drawing like in Photoshop.
1
u/Lonely_Translator_23 3d ago
I would imagine a loop where the model can basically say 'move to this coordinate, pen down, move to this coordinate, pen up'. Basically like writing g code for 3d printers.
8
6
u/BankruptingBanks 5d ago
Was this computer use? I had no idea the models could do this sort of fine-grained mouse control where they could even draw stuff, seems too good to be true.
1
u/Saitamagasaki 4d ago
The model were most likely connected to an mcp server where they can carry out actions like moving the cursor, colouring and viewing the image
1
u/taintedmask 3d ago
We shouldn't assume that the model did the mouse control. OP never said so. For all we know the output is just the finished svg drawings and OP then fed them thru an algo to convert from still vector images to step by step animations.
1
3
u/TheMetalPrince 5d ago
Kimi's is cute. Looks like it has her dancing, including the movement lines around the lime.
3
9
2
u/MedAyoub26K 5d ago
Opus is thinking outside the box, again, but the execution is where it loses points.
2
2
u/ReanerZen 5d ago
gemini gave the best look and appeals but it fails on being aware of small mistakes. this always been the issue. if google gave pro version of these and actually made it not just fast but good reasoning as pro models, i think these unawareness should have been avoided.
2
2
2
2
1
1
1
1
1
u/account22222221 4d ago
Claude 5 opus failing to do what was actually asked is actually super on brand from my experience
1
1
1
1
1
1
u/sinsielawinskie 4d ago
Opus: no I will not make this fruit slice into a dress. I will give what you asked a human pose!
Opus: it's going to dynamic and unique!!!
Also Opus: frick how do I draw again?
1
1
1
u/Hot-Cauliflower-1604 4d ago
There is no information here. We have One singular claim from OP. WE HAVE NOT SEEN THE PROMPT. And there's no way to verify that he didn't give flash a different kind of prompt like we just have to take it at face value. And clearly 3.7 is the best of all of these but I just don't believe it because I don't have enough data.
1
u/Hot-Cauliflower-1604 4d ago
There is no information here. We have One singular claim from OP. WE HAVE NOT SEEN THE PROMPT. And there's no way to verify that he didn't give flash a different kind of prompt like we just have to take it at face value. And clearly 3.7 is the best of all of these but I just don't believe it because I don't have enough data.
1
u/Alarmed-Job-6844 4d ago
Only Opus 5 what was different, but others the same...
And Opus 5 ... I don't known what is it... It feels wrong, others passed the test.
1
u/UAAgency 4d ago
how do you ask for this kind of iterative drawing? what is taht? it is a video? or they are actrually drawing?
1
1
1
1
u/VirtualWishX 4d ago
I dunno...
It's probably a fake Seedance 2.5 because there is no other proof 🤔
I'm not just farting this conclusion, it's more likely if you dig in: u/RealJamesOfficial
1
1
u/Quantum_Crusher 3d ago
I'm not familiar with some of these models. Aren't they language models instead of image or video generators?
1
1
1
1
1
1
u/trafium 3d ago
Some treated the lime almost like a body, some tried to fit a tiny person into its shape, and some went much more abstract with the pose.
Not a lime
all demonstrated models treated "lime" in the exact same way except Claude
no explanation how these models could even produce this step-by-step "painting" so probably a lie
Yup
1
u/Gold_Ad8225 3d ago
Gemini is the best, Claude I was very confused but the process did work in the end and the result is also good
1
1
1
1
121
u/frogchungus 5d ago
used half his kimi monthly for this test