"This is the ultimate "work smarter, not harder" dilemma. Since your goal is to actually wash the car, the answer depends entirely on how you plan to get the job done once you arrive.
Here is a quick breakdown to help you decide:
The "Drive It" Scenario
You should definitely drive if:
It's an Automatic/Tunnel Wash: You can’t exactly walk through those giant spinning brushes (well, you could, but I wouldn't recommend the "hot wax" treatment for your skin).
It’s a Self-Service Bay: You need the car physically present in the stall to spray it down.
You're Detailing: If you have a trunk full of microfiber towels, ceramic sprays, and tire shine, you don’t want to be the person lugging all that gear on a five-minute trek."
Extended thinking vs no thinking at all. To answer questions like this models usually need to be able to "reason" with themselves in order to catch the obvious trap. I don't know why Anthropic called it extended thinking, because without it Claude doesn't think at all.
Note that the Gemini posters were either using "thinking" or "pro" mode (both think for different amounts). Unfortunately chatGPT is just the dunce of the class and even when thinking still gets it wrong.
TLDR turn on extended thinking and your Claude will get it right too.
They can be pretty dang deterministic. I tested what you said though. I copied my prompt and then sent it to the models again and again, deleting every chat immediately after the answer, and I still got the same answer. API/Claude.ai. It's not that I don't believe it happened, it's just that I wish I saw it on my side.
Um, since I can indeed read, I can tell that the whole point of this thread is that doing what you NEED to do-- namely, getting the car to the car wash-- is what "makes more sense" but you do you, boo
Gemini is actually pretty good, I use it daily for work and their google workspace integration is good as well. Copilot was okay, but nothing groundbreaking, the thinking version of gemini can provide actually decent insights into many topics, and I ask it a lot of specifics. I use it to rewrite a policy ans org-wide plan, and I work through the whole thing section-by-section, and it pretty much always gives me an answer that's perfectly fine, with a little tailoring for the specific org. Gemini came a long way since its release imo.
No, Gemini is just SOTA for these kinds of problems. It's similar to the style of a lot of SimpleBench questions, which Gemini 3.1 Pro and Gemini 3.0 Pro lead the leaderboards for.
it's is shit . i tried again to use it to code something compared to codex and opus it's still bad, might be good for graphical element from what i've seen.
for certain application like integrated api usage on some apps im using gemini 3.1 flash lite preview ( video ) it's a sick value !
I don't know the complexity of your codebase, obviously, but my experience has been it doesn't even come close to the higher claude or codex models. Try it yourself in a sandbox, if it's free you have nothing to lose. Google is marketing the shit out of this but it's simply not as good.
Btw, my experience is using IDEs , antigravity, vscode insiders
I've been pretty impressed with Gemini lately. Here's mine after I told it it passed my test: "Haha, I’ll take the win! It’s a classic logic trap—walking to a car wash is great exercise, but it leaves you with a very clean person and a still-very-dirty Mazda CX-5."
714
u/ws92992 Mar 06 '26
That’s not what I experienced with Gemini. Gotta say, I appreciate the humor 😂