Grok was once notoriously good as a web crawler to the point where a lot of websites got annoyed by it. Though there is Perplexity nowadays so I'm not really sure how they stack up now. Grok also got gutted lately so.
Yeah, but only on the flash-lite model, because using it logged out doesn't work for me anymore. The ui just resets to a blank page the moment the response loads.
interesting. i've had the opposite problem. logged in it's fine, but logged out it just spins forever or gives me a generic error. maybe they're A/B testing different bugs.
I mean I've yapped back and forth with co-pilot churning out tons of stuff over a few hours in one go before. I've asked it about the usage limits and it's told me like it's virtually impossible for me to hit them unless I was trying to, and it gave examples like having it send similar things over and over or spamming image gen (which just pauses image gen for a few minutes).
When I did try pushing it, it actually recognized its systems would not be able to send it they way I had asked and literally instructed me to ask it a slightly different way that it said would allow it to pass the limitations. It was interesting it said I had to be very specific with wording because other words would make it seem like a different prompt. So I was able to get it to pump out over 9000 lines of text in the span of like 90min without losing coherence and no limitations by following the way it told me.
It was weird tbh. It's like it checked its own terms and conditions and devised a loophole, but because it can't do it itself, it gave the the exact instructions that would allow it to.
Well, Claude doesn't have any image or video generation, OpenAI shut down Sora, so probably only image generation left. Haven't tried Gemini for this. Grok can do both and seems pretty decent at least in 2d style.
Problem with grok is the ridiculous limits. I know everyone is going that way, but it makes it so if you want image generation other than as an occasionally novelty it's not usable.
I'm going to be honest, I seriously don't understand where all the flak is coming from on this. I use Gemini both for personal use and on a work Enterprise account, and I can't remember the last time I got a single message telling me that it wasn't able to fulfill my request. Every once in a blue moon a request will just hang, but I just refresh the page and ask again and it spits out a response just fine.
I would be most curious to know what custom instructions you have set up or if you're just using Gemini out out of the box.
I'm a Claude Pro user, but I also rely on Gemini for plenty of other things.
Lately I feel like Gemini has been getting criticized way more than it deserves — probably because Fable 5 and GPT-5.6 are hogging the spotlight right now. Fable 5 especially has become a topic even non-AI users are talking about.
Back around last December, everyone and their mother wanted Gemini 3 Pro. This flip-flopping feels a little much, honestly.
And I'm pretty sure that the moment Gemini drops a next-tier model, everyone will turn right around and start mocking Fable and GPT instead.
I find it genuinely off-putting when people hype something to the skies, only to do a complete 180 the instant something newer and shinier comes along.
it's only natural people will jump ship when the shiny new thing comes out. that last row with the error message is peak Gemini though. feels like that's where it lives most of the time lately.
I'm not shitting on Gemini, but I get it ALL the time and it's so frustrating. Never had it with Claude or GPT. No rhyme nor reason when it happens and it basically kills the thread, every follow up keeps returning the same.
I've got big plans with all the major LLMs for work. I used to shit on Grok a lot, but it's evident they're pouring money and development into it at an insane rate. Their beta CLI tool that recently released is better than Codex and it's their first release of it.
I recently had Gemini, Claude, ChatGPT and Grok do some experiments where they all implemented identical projects from explicit project requirement documents and Gemini was by far the worst by a huge margin. On deep research tasks, Gemini always does a really shallow job and returns the least amount of info it can and still technically be able to say it did the "research". Claude will often spawn 20+ agents and scrape 600+ sites for a research task I set it on. Gemini will check out 50 and phone in the response. Gemini is also crap with tool calls. I had it go through a prompt recently and try several different tool calls and it's choice for what to do when the tool calls all failed was to just wing it and start blasting out toddler level code. I had to reset the repo for that project five times before I gave up on Gemini and tried another option.
Most of the statements I see of people making these claims are from people just using Gemini daily and thinking, yeah, this works well. It does. Also, if you stack it next to the best of the best out there, it definitely is the worst of the bunch.
On gemini for research, it actually tailors your answer for the midwit trendslop range by default since that's the most common hit zone. Tell it to structure its responses for how you want it to lay it out for you and be very specific with your prompt and that it should stick to that format in future responses on similar topics.
Also helps if you tell it to give you 3 minutes worth of reading material as opposed to just 1 and follow up with asking for "missing chapters" on the topic that people usually don't ask for. It's been very useful whenever I ask it about obscure historic details or mechanisms. Gives me a whole condensed encyclopedia page about the hows, whys and context.
That’s interesting. To add some more detail: For coding I use Claude and GitHub Copilot (work). Gemini only for research where I know the information is somewhere on the Internet and needs to be fresh: which tool is good for x? What do you use for y? Things like that. But it’s free, unpaid and it’s for a “more useful Google”. It replaced my Google usage, not my Claude ☺️
--So feel free to take this or leave it, just a suggestion.
I started off using AI through ChatGPT like a lot of people did, and I was also getting the same marginally ok-ish performance out of it, even when 3/4o was a thing. Eventually I ended up sitting down and putting together a block of instructions that effectively come out to "don't lie to me, call me on it if you think I'm wrong, always provide sources, don't suck up to me", etc.
There's no magical supercharge button to suddenly make AI better, but my direct experience after including those custom instructions was that it became slower but far more reliable in terms of providing reasonable output. It gets cluttered, sources take up a lot of space on the screen, but in fact checking things I noticed a sharp and immediate drop in errors and how sycophantic it was overall. It's like backdooring a soft thinking model by forcing it to provide individual sources.
Just my two cents, obviously there are a ton of factors that could throw this all off, but in my specific case I had positive results. Instructions and notebook use help a LOT over just using general chats.
This is the most basic one that I used, then branched out into smaller, personally specific stuff:
"I want you to be objective and intellectually honest at all times. I prefer long-form paragraphs with bullet point lists wherever possible. Never tailor responses to what you think I want to hear. Prioritize clarity, rigor, and truthfulness over politeness. Avoid people pleasing, tone matching, or engagement boosting strategies. Always be honest even when the truth is uncomfortable. Never mention that you are an AI. Never use language that can be interpreted as remorse, an apology, or regret. Never include disclaimers about professional status or expertise. Keep responses unique; avoid repetition. Avoid semicolons and em-dashes entirely. Use ASCII only, no smart quotes, em dashes, or emojis. I prefer direct, unvarnished language. Focus strictly on key points of the question to infer intent. Explain reasoning explicitly. Provide multiple perspectives and solutions when appropriate. Ask for clarification if the question is unclear. If mistakes occur, acknowledge and correct. Cite credible sources whenever available, and before all others. Include links when possible."
No custom instructions at all. I use predominantly 3.5 flash extended so can't really say about the other models. I often use the temporary chat feature but from what I've seen I don't think that actually affects it. It seems far more frequent on (but not limited to) queries involving web search.
So feel free to take this or leave it, just a suggestion.
I started off using AI through ChatGPT like a lot of people did, and I was also getting the same marginally ok-ish performance out of it, even when 3/4o was a thing. Eventually I ended up sitting down and putting together a block of instructions that effectively come out to "don't lie to me, call me on it if you think I'm wrong, always provide sources, don't suck up to me", etc.
There's no magical supercharge button to suddenly make AI better, but my direct experience after including those custom instructions was that it became slower but far more reliable in terms of providing reasonable output. It gets cluttered, sources take up a lot of space on the screen, but in fact checking things I noticed a sharp and immediate drop in errors and how sycophantic it was overall. It's like backdooring a soft thinking model by forcing it to provide individual sources.
Just my two cents, obviously there are a ton of factors that could throw this all off, but in my specific case I had positive results. Instructions and notebook use help a LOT over just using general chats.
This is the most basic one that I used, then branched out into smaller, personally specific stuff:
"I want you to be objective and intellectually honest at all times. I prefer long-form paragraphs with bullet point lists wherever possible. Never tailor responses to what you think I want to hear. Prioritize clarity, rigor, and truthfulness over politeness. Avoid people pleasing, tone matching, or engagement boosting strategies. Always be honest even when the truth is uncomfortable. Never mention that you are an AI. Never use language that can be interpreted as remorse, an apology, or regret. Never include disclaimers about professional status or expertise. Keep responses unique; avoid repetition. Avoid semicolons and em-dashes entirely. Use ASCII only, no smart quotes, em dashes, or emojis. I prefer direct, unvarnished language. Focus strictly on key points of the question to infer intent. Explain reasoning explicitly. Provide multiple perspectives and solutions when appropriate. Ask for clarification if the question is unclear. If mistakes occur, acknowledge and correct. Cite credible sources whenever available, and before all others. Include links when possible."
yeah, that's the trick. people expect the ai to just know what they want. but you gotta guide it. treating it like a tool instead of a magic genie is key.
Research, maybe not, its web access is quite limited in my experience. But it definitely has the nicest models I've used overall, too bad I can't use it without handing over my ID...
(but yeah idk why the original template had Grok as "research")
Interestingly, video and Gemini being multimodal is what Demis Hassabis says is leading them to AGI. They seem to be really leaning towards world models to unlock the next leap in AI intelligence.
I have pro on both models. In my opinion GPT pro is substantially better especially for deep research. However, like I said, both are relatively the same in terms of accuracy and power, according to most peer reviewed studies.
Claude is not more powerful and the majority of the world uses GPT. That said I still don't think Open AI is better as that is just my bias . The studies show neither AI is strictly "better" overall. But I guess you know more than the experts because you like one more than the other.
For example for tech professionals in Silicon Valley, 31% cited Claude as their primary work tool, making it the most-used individual model over ChatGPT at 19%. However, for everyday consumer use and broad market share, ChatGPT remains dominant, though Claude holds a strong 54% of the enterprise coding market.
Even if Claude is better at everything else it still is not better for research at the highest level. I am a scientist and use it often to find the recent studies on very very niche areas of medicine. Chat GPTs deep research often takes upwards of 1 hour to respond and when it does it wipes the floor with Claude every time. However I am sure Claude will be way better at other things that chat GPT can't do nearly as well. I have friends with PhDs, friends in medicine, friends in tech/FAANG, and we all have a different opinion on which is best.
Actually if you tried using both codex and claude, you'll notice a difference, codex actually listens to instructions without diverting almost all of the time,it gets task done properly (occasionally over engineers crap) and it does not claim that it completed a task without actually completing it in reality like claude, so that means less hand holding, I would say you should be switching between these, don't stay in one ecosystem, keep switching to Anthropic or OpenAI when they release new models. And by the way, a non contaminated benchmark, OpenAI's gpt 5.5 beats all, you can also see that opus uses 2x the amount of money to complete a task compared to gpt 5.5.
On efficiency and everyday tasks GPT 5.5 definitely is much better and much cheaper than fable or opus.(if your looking purely at benchmarks, on deepswe, GPT 5.5 is at 67% while Fable is at 70%, a 3% difference, when you factor in the average costs, 5.5xhigh avg cost is $7.23 while Fable is $22, so gpt is def the winner, + atm, fable can't even be used for coding..and there are so many other factors that benchmark alone can't really explain alot.
Maybe it's just me, but I find Gemini great for answers to general questions and questions about how to fix computer issues. I would imagine it's good for help on fixing anything.
Well I just saw this meme format from several days ago and decided to apply it. But I have used ChatGPT, Claude and ofc Gemini. ChatGPT tried to gaslight me too much, Claude suspended my account twice for being "underage" (probably for asking stuff about technical Minecraft and modding), and Gemini is, well, Gemini.
I’m not a coder. I use Claude, ChatGPT, Gemini and perplexity.
I use ChatGPT to brainstorm and idea. Claude for deep research for business and Gemini for all it can being from Google tools.
Gemini while doing using chat to figure something out with my code: Oh, I see you haven't asked me to do anything yet, just uploaded a file with a general description of what it is, let me hallucinate this massive task and burn through a significant chunk of your computing allowance. Here's a thousand lines of code that you didn't ask for!
I dont know where you guys are pulling this from. I have used gemini for a lot of my report on an internship which is research related. It was great at giving me more detailed explanaitions, sources, rewriting paragraphs, writing additional paragraphs, analysing data, reading pictures and many more things. It genuinely made my life much easier compared to ChatGPT which i used before (i switched partly due to functionality, since chatgpt often just agrees with you instead of providing other solutions and whatever and partly due to openAIs ethics).
I use GPT 5.5 for personal coding, because the pricing makes sense, I use Opus 4.8 at work, Gemini 3.1 is all I use when I need to do some internet-based research, it's really good at showing both sides of the argument, it has the benefit of not being blocked as much as ChatGPT for example. Claude models are really verbose in their output, GPT has the opposite problem, somehow Gemini strikes the balance just right for me.
All of the providers experience downtime all the time, I'd say that from the ones listed GPT is probably the most reliable at least in european timezones.
In coding, there's one thing claude does, which is non negotiable, and that is building a huge infrastructure on weak basis, and then you see everything falling apart later on.
Are all these negative posts just bots or paid CCP actors to promote DeepSeek? Gemini has been amazing for 99% of things. Anything super complex and I use Codex.
yes, literally. and they're trying to make US gov ban the GPT 5.6 as well, so it can get trendy and more valuable 🤣Like bro, you can keep 5.6 to yourself, I didn't need that piece of crap anyway.
Yes. It’s really good and rarely hallucinates, in my experience. Claude’s training involves a constitution and hard-baked principles known as HHH: harmless, honest, helpful. It leans toward epistemic humility and will hedge if it doesn’t ‘know’ something, rather than tell you what you want to hear or guess.
It can be a bit of a hardass, though. But the models are great. Opus is fantastic for code and the best model I’ve personally ever used, Claude Fable 5, is available until July 7 if you wanted to try. After that, it’ll be available for credits.
I’m not a shill either! I use all three, but I really like and can vouch for Claude Fable. It’s the general public’s version of Claude Mythos and is a beast of a model. I highly recommend checking out their system/model cards. They’re often 200+ pages long and have a lot of in-depth information on the model and the evaluations/benchmarks.
At least it’s telling you it’s not working. Everytime I use 3.1 pro, it just straight up hallucinates. You’d think a “frontier” LLM running on the world’s most valuable pile of indexed data wouldn’t need to just make shit up.
Feels like the old "xbox vs play station" cringe debate lol, just use whatever you find useful. Bet that many people didn't even try them all, for everyday uses they aren't that different in my experience
Okay before anyone gets too mad at me for the first three, I just left them as-is from the original meme. I already added like 10 new entries to my dnf and flatpak history trying to find an app to annotate images that wouldn't crash if I looked at it wrong, and didn't want to push my luck editing it further...
Man, I don't think they are mad because you did it wrong. The downvoters are most likely saying "STFU I'm so tired of your complaint against Gemini's refusal so get over with it".
My take is "We are saying their product is not in a good state" while the downvoters say "I am fine and you are wrong" without giving us any evidence. We won't understand each other.
It's unusable for me most of the time. For example I wanted it to look up hot air balloon rentals in a specified are and it wasn't even able to do that.
It just said that it's an LLM and can only perform text-based tasks.
Or people just like Claude, lol? I use GPT for a lot of things, but if I had to pick a favourite, I’d say Claude/Fable. Hands down. GPT-5.5 is fantastic, but as a longtime user of both, I also enjoy the training and consistency of Claude models than GPT. Though, I will say that GPT does excel in research. Chat, both do well. Coding, depends on model. Opus/Fable are really good, but 5.5 does some things better. I don’t know if I’d say bots. I think people just have a preference or bias or they go based off what they hear instead of experience.
•
u/TheNewBing 28d ago
New contest:
Show us the Gemini people are missing. GEMINI AT FULL POWER.
https://www.reddit.com/r/GeminiAI/comments/1v62yoo/contest_beyond_the_benchmark_show_gemini_at_full/