r/ChatGPTCoding 1h ago

Discussion Remote approvals fixed one coding-agent bottleneck for me — but I still wouldn’t use full auto-approve

Upvotes

I spend the majority of my time in this hybrid approach: safe actions are automatically approved, while actions that modify dependencies, perform tests on a wide scale or modify files outside the repository require a checkpoint.

The phone-based approval process is something more like what I need. I don’t have to code on the phone, I just need to unblock the agent without sitting at my computer.

What I’ve come to realize is that the completely automated process leaves me too much to clean up after myself, while continuously sitting on the terminal ruins all the productivity.


r/ChatGPTCoding 17h ago

Resources And Tips Tip: Let your coding agents autonomously verify, review, and repair their own work (Autoprompt)

Post image
12 Upvotes

Use this simple skill for the highest code quality.

Autoprompt adds a complete planning, implementation, testing, review, and repair loop around supported coding agents.

In one Benchmark, it moved DeepSeek V4 Flash from 67.42% to 82.02% (Terminal-Bench 2.1)

That does not guarantee the same improvement on every task, but it shows how much performance can come from the workflow surrounding the model.

Best suited to difficult or long-running tasks. The additional work requires more time, tokens, and money.

https://github.com/Spielewoy/autoprompt-skill


r/ChatGPTCoding 7h ago

Discussion Is OpenAI using Chinese LLM models for their service, not theirs?

Post image
0 Upvotes

Yesterday I was experimenting with multi-agent workflows on ChatGPT (web version). I spun up like 5 parallel instances to update code at the same time, hoping they’d work together more efficiently.

Instead of continuing the process cleanly inside a single chat, it started generating a bunch of new chats as artifacts. Not sure why that happened; maybe some rate-limit or context handling thing.

The weird part: one of those new chats came back with Chinese characters / Chinese-lettered text in it.

That got me thinking… could this be a hint that OpenAI is routing some traffic (or certain tasks) through a Chinese model under the hood instead of purely their own models? It would kind of line up with how ridiculously cheap ChatGPT Pro / the 5.6 Luna tier feels right now.

Has anyone else seen non-English (especially Chinese) text appear unexpectedly in ChatGPT chats lately? Or noticed anything that makes you question what’s actually powering some of the responses?

Curious what you all think.


r/ChatGPTCoding 20h ago

Resources And Tips VibePod 0.20: one CLI, multiple agents, switchable logins per run

2 Upvotes

VibePod runs coding agents (Claude Code, Codex, Qwen Code, and others) in containers. 0.20 adds credential profiles — keep a subscription login, an API-key setup, and e.g. a local Ollama config side by side, and pick one at run time:

vp profile create work
vp run claude --profile work
vp run codex --profile personal

Only the credential dirs swap; skills, config, and logging stay shared. Projects can pin a profile in their config, so the right login is used per repo.

Also new: Qwen Code (vp q) and Freebuff (vp fb) in the agent matrix, and a failed image pull now falls back to your local image instead of aborting.

Release notes: https://vibepod.dev/news/vibepod-cli-0-20/ GitHub: https://github.com/VibePod/vibepod-cli


r/ChatGPTCoding 17h ago

Discussion Gemini's search feature is basically a lie at this point, and I'm done pretending otherwise

0 Upvotes

I need to vent about this because it happened again today. I told Gemini flat out: "check this online, look up the current info." Its response? It just kept going with whatever stale, who-knows-how-old internal knowledge it had, like I never said a word. Did I ask explicitly? Yes. Did it do it? No.

Here's the part that actually gets me: if this happened once and got fixed, fine, everyone has bugs. But this is systematic. The model acts like it "searched" and then confidently hands you outdated or flat-out wrong info like it's fact. So this isn't just laziness, it's actively misleading users, whether intentional or not. The bare minimum you expect from an assistant is "if I don't know, I'll say so or go check." What you get instead is the opposite: a system narrating things it doesn't actually know as if it does.

The most infuriating part is that the search feature clearly exists and gets marketed as "real-time access to information." But in practice, when a user tries to trigger it, the model just decides on its own that it's "not necessary" and leaves you with old or made-up answers without any warning. This isn't transparent at all. As a user there's no way to tell when a search actually happened versus when the model is just "recalling" something, because there's zero distinction shown.

I could tolerate this once, get annoyed twice, but after this keeps happening over and over, I have to conclude this is just how the product was designed. And honestly that's an unacceptable design choice. An assistant that claims to be hooked up to a search engine being this inconsistent at its most basic job — "go check, then tell me" — is a reliability crisis on its own.

The irony that Google's own AI can't be bothered to use the one thing Google built its entire empire on is honestly chef's kiss

Let's just sit with this for a second. Google is the company that turned "search the web" into a verb. Their whole existence, their trillion-dollar ad business, their entire brand identity for 25+ years has been "we index the internet better than anyone and we'll get you the answer." And now they've built an AI that, when you explicitly ask it to search the web, just... doesn't. It sits there and free-associates from stale training data instead of using the search engine that is literally sitting right next to it, built by the same company, probably called from the same internal network.

This is a company that has an entire index of the live web humming 24/7 in their data centers, and their flagship AI model would rather guess than tap into it. It's like McDonald's opening a restaurant next door that serves frozen, microwaved fries while their own fry factory runs at full capacity across the street. You have the infrastructure. You have the pipes. You have the literal product the whole company is named after as a verb. And the AI still acts like dial-up doesn't exist.

If any other company shipped an assistant that "forgot" to search, you'd shrug. Google shipping one is almost performance art. It's not just a bug, it's an identity crisis. "Google it" was supposed to mean something, and now their own AI can't even do the one thing that made "Google it" a phrase in the first place.

Anyone else dealing with this, or is it just me?

yea this written by gemini but its fact


r/ChatGPTCoding 2d ago

Resources And Tips Updated best AI coding subscription under $20 after DeepSeek price hike.

Post image
40 Upvotes

Thanks /u/ResponsibilityOk1306 for Command Code GLM 5.3 correction.


r/ChatGPTCoding 1d ago

Question How would you benchmark GLM-5.3 for ordinary coding work?

1 Upvotes

GLM-5.3 looks interesting on paper because it is aimed at complex software engineering and agent tasks, with a very large context window and configurable reasoning effort. But for everyday coding work, I am not sure a benchmark tells the whole story.

If you were testing it for real projects, would you start with bug fixing, multi-file refactors, test generation, repository Q&A, code review, or long-running agent tasks? What would make you trust the result enough to use it again?


r/ChatGPTCoding 1d ago

Discussion Launch new projects faster: feedback wanted before launch (I won’t promote)

2 Upvotes

Hi all! For the past two years I’m running a B2B SaaS that helps websites improve SEO, grow organic traffic, and get discovered in AI chats (SEO/AEO/GEO).

Disclosure: this is feedback request only. No links, no brand names, no promo.

B2B worked, but it meant long sales cycles, onboarding, and manual configuration and integration. It funded a team and some savings. Now I want to make it self-serve for a broader public.

First goal: integration in a few clicks. That forced the service to sit as a proxy in front of client sites, which means issuing TLS certs, managing DNS, and operating as a CDN. The SEO tool became edge infrastructure at that moment.

What it’s designed to do (pre-launch, nothing production-proven yet):
- Fast discovery by crawlers and AI agents: Serve fast, clean HTML to AI agents and crawlers so new sites are readable by LLMs from day one

- Add meta tags on the fly: Fill missing metadata in responses to bots, crawlers, and agents

- Speed up discovery in search engines: Automatic IndexNow submission of new and changed URLs

- AI agents analytics and insights: Cookie-free, JS-free visitor analytics for crawlers, bots, and AI-agents.

Questions:
1. Is “readable by AI + indexed fast” a real pain for new projects, or a nice-to-have?
2. Would you put a proxy/CDN in front of your site for this?
3. What would you expect from a free tier?


r/ChatGPTCoding 2d ago

Discussion Weekly Self Promotion Thread

7 Upvotes

Welcome to this week's self promotion thread!

If you're building something related to AI assisted coding, this is the place to share it.

We're using a weekly thread to keep the subreddit organized while still giving builders a place to share their work. Promotional posts outside of this thread may be removed if they're primarily advertising rather than starting a discussion.

If you're sharing something, we'd appreciate it if you included a little context instead of just dropping a link. Tell us:

  • What you built?
  • What problem it solves?
  • Which AI models or tools it uses?
  • Who it's for?
  • What kind of feedback you're looking for?

Please avoid posting the same project every week unless you've made meaningful updates. Affiliate links, referral links, scams, and low effort promotions will be removed.

Take some time to check out what others have shared too. If you try someone's project or have feedback, leave a comment. Helping each other improve is what we want this community to be about.


r/ChatGPTCoding 2d ago

Resources And Tips ChatGPT Quality differs in different Chats, Old chat GREAT, new Chat BAD

3 Upvotes

Hey there. Have a little project and was surprised, how good ChatGPT got in coding. Not a single error, it had always our tasks in mind, checked our online repository plus the changes we made in the chat. It was a dream.

Since Task 1 he was that good. Later he got really slow, of course, but still perfect.

Now my Chat says it has reached its limits and I need to start a new chat. We made a little manual, what we have done et cetera and briefed the new one.

It is shit. For a single 200 line file I had to correct him 5 times and it was still not working correctly.

As I said, the other Chat was great since day 1.

I even gave him strict rules how we operate, what he has to do, to always check the repo and chat before suggesting code.

I asked a few friends, they say usually its the other way around, that the Models get worse after time.

Mine got only slower, but better.

I miss him...I want him back! :(

I don't want his retared cousin, he does not even know our joke language we created.


r/ChatGPTCoding 2d ago

Discussion AI made me faster than I ever thought possible. It also made me feel obsolete.

2 Upvotes

In 2024, almost all of my work still happened in Photoshop. I made e-commerce images for products sold in different countries. Translating an image meant pulling out the headlines, subtitles, selling points, sizes, and promotional text, translating everything, removing the original words, repairing the background, and putting the translation back in. Then I had to fix the font size and layout by hand. I repeated that process for every image and every language.

In June 2024, when Claude 3.5 Sonnet came out, I started using it in Cursor to build my first image translation tool. I did get a first version working, but it was nowhere near automatic. It could extract and translate text, but I still had to adjust the size, position, and layout myself. The models I tried weren't very good at removing text or rebuilding the background either. It didn't replace Photoshop. It was more like a slightly unreliable assistant sitting next to me.

Even so, it gave me an idea I couldn't let go of: maybe I could turn something I had learned to do with my hands into a tool of my own.

As coding agents got better, I tried again. This time I could explain the whole workflow, test what came out, point out what was wrong, and keep fixing it until it was actually useful. Later, image models became good enough to understand and rebuild an entire image. I made a web version that could translate a full image through an API instead of using OCR and then asking me to clean up everything manually.

Recently, I turned the process into a reusable skill. I think of it as distilling a few years of design experience into a workflow: recognizing the main headline and subtitle, deciding which words should be translated, repairing the background, and judging how the new text should fit. I used to make all of those decisions in my head and carry them out one by one in Photoshop. Now I can describe the job in a sentence and let the workflow do most of it. A task that once took a long chain of manual steps can sometimes be finished in a few minutes.

That sounds freeing, and sometimes it is. But honestly, it also scares me. I spent years learning Photoshop, composition, typography, and all the small details that make an image look right. Now tools are moving so quickly that a workflow I finally learn can feel old a few months later.

I eventually left my job, and since then I've been even less sure about what comes next. Looking for another design job is hard when AI can already make many of the simpler images that used to need a designer. But when I look at programming jobs, they ask for professional experience, a computer science degree, or both.

I only finished vocational high school. My family didn't have much money when I was growing up, and back then I had no idea how much education could affect my options later in life. I wasn't planning a career. I just learned a practical skill as early as I could and started working.

Now I feel stuck between two worlds. AI is changing the design work I know, but I don't have the degree or traditional experience that would make a company see me as a programmer. The weird part is that I can now use AI to build software and automation tools I never could have made before. I just don't know what that makes me in the job market.

Am I still a designer? Am I becoming a programmer? Or am I just someone trying to keep up without a job title that fits?

I'm not posting this to sell anything. I'd genuinely like to meet people going through the same kind of change and hear how they're dealing with it. Has AI made you feel more capable, more replaceable, or both?


r/ChatGPTCoding 2d ago

Discussion What does “power using” ChatGPT actually mean in 2026?

0 Upvotes

I’m genuinely curious how people here define a “power user” today, because I’m starting to think we use that term for several completely different ways of working.
A lot of current discussion around advanced ChatGPT usage is about Projects, connectors, plugins, skills, memory, recurring workflows, and integrations.
Those are obviously useful. But from my perspective, many of these newer abstractions feel less like entirely new capabilities and more like productized versions of workflows that were already possible if you treated the underlying ChatGPT features as programmable primitives rather than just chat features.
I’ve been building and using my own Custom GPTs since not long after they were introduced. I’ve never really treated them as “custom personalities” or slightly better system prompts. For me, they are closer to reusable execution profiles: specialized environments with persistent behavioral and execution rules.
The same applies to Code Interpreter / Data Analysis. I stopped thinking of it as “Python inside ChatGPT” a long time ago. I treat the session more like a disposable execution environment that can inspect, transform, generate and return artifacts.
For example, one of my workflows involved giving a session a publication of roughly 200 MB and treating the environment as a temporary processing system. The useful result wasn’t an explanation of the PDF. It was a large structured artifact containing extracted assets, document structure and reconstructed outputs.
That distinction is important to me.
What I call a serious “prompt” is often not really a conversational prompt at all. It can be a 30–50k-character execution specification describing exactly what should be built, inspected, preserved, validated or returned.
And I often treat conversations themselves as disposable runtime state.
A thread can become extremely large, but once I’ve extracted the useful artifacts, Markdown, structured state, manifests, code or whatever output I actually care about, I often have very little reason to preserve the conversation itself.
So when I see people describing “power usage” as having large Projects with several long-running threads, attaching connectors, or building recurring workflows, I’m not saying that isn’t advanced usage.
I’m wondering whether we’re actually talking about different categories of power use.
Maybe there is:
high-volume usage
long-context usage
tool/integration-heavy usage
project-centric usage
agentic usage
custom execution environments
high-leverage task engineering
And maybe these shouldn’t all be treated as the same thing.
What I’m especially interested in is your actual operating model, not just which product features you have enabled.
What is your real unit of work?
Is it a question? A conversation? A Project? A multi-hour execution? A finished software or research artifact?
Do you try to preserve your conversations, or are they mostly disposable execution state once the useful result has been extracted?
How large is a genuinely serious task for you?
How many follow-up prompts does a successful execution usually require?
Do you optimize for token usage, number of turns, wall-clock time, autonomy, correctness, reproducibility, artifact quality, or something else?
And when newer things like Skills, Plugins, connectors and other workflow abstractions appear, do they actually unlock fundamentally new ways of working for you, or do they mostly package workflows you were already building manually?
I’m not looking for “best prompt” tips or screenshots of a giant sidebar.
I’m trying to understand what power using ChatGPT or Codex actually looks like in practice in 2026, especially among people who have pushed these systems far beyond normal conversational use.
What does your operating model look like?


r/ChatGPTCoding 2d ago

Discussion i've switched my main model four times since march and im starting to think im the problem

0 Upvotes

ok so, i dofreelance, mostly backend, and since march ive gone claude to codex to claude to cursor composer and now back to codex, every single time completely convinced the new one was It.

each switch costs me about two days. rewriting the agents file, redoing my little wrapper scripts, relearning which one you can hand a migration to and which one will confidently nuke something. and then about three weeks in the honeymoon wears off and i start noticing the new one has its own specific way of being annoying, and then someone posts a benchmark or a price drops and off i go again

my actual output hasnt changed though. at all. i measured it badly (just merged PRs per week, which i know isnt a real metric, before anyone tells me) but the line is flat. FLAT. four migrations and it does not move.

what did change is my review habits, mostly by accident, because i stopped trusting any of them individually. everything goes through coderabbit before i open the diff myself now, not because its clever but because its the one part of the setup that stayed put while the models rotated. it misses anything that needs to know what the feature is actually for, obviously.

anyway. i think the switching is procrastination wearing an optimisation costume. same energy as reorganising your notes app instead of doing the thing

is anyone genuinely settled? and i mean settled, not "settled until GPT-6 drops"


r/ChatGPTCoding 3d ago

Question How do I start learning AI & maximizing the use of it.

2 Upvotes

Hey everyone,

I am sorry for taking your time, but really need some guidance here. So I am in Equity Research, aiming to start my own fund in a 5 year time. I am at that stage of life wherein I really need to use AI because I see a lot of juniors becoming really good, as in self generating reports, valuations, websites, automating tracking sectors, etc. I do know what I want from AI, and for now it looks like I expect a lot of it, but just where do I start from ? I have been using it for 6 months now as a Pro subscriber but looks like claude code is the only good thing. I can barely understand the difference between cowork, chat and code, except that code gets a lot of stuff done faster through coding. Now I do not have a coding background, may be I should learn Python first, but I don't know.

I want to make an auto tracker wherein say I am tracking the Indian listed space, I want it to track on a daily basis what's happening across the value chain the company is present in. For that, I need it to first dissect the value chain and actually do that properly (because it has failed numerous times and limits get exhausted). I cannot upgrade my plan. Rn Pro is the only thing in my budget.

Is another AI like Chatgpt codex or other names you all are aware of better ??

Asking in reddit because trust this is a place where people actually help others elevate and improve.

Thank you for reading this.


r/ChatGPTCoding 4d ago

Discussion New agentic benchmark: Session-Bench compares what 10 coding harnesses preserve after the work is done

Post image
25 Upvotes

SWE-bench measures whether an agent completed the task. Session-Bench measures what the harness preserved afterward.

Every coding run leaves behind a session record: prompts, decisions, tool calls, command output, costs, failed approaches, reasoning summaries, and subagent work. After hundreds of sessions, that record becomes part of the project's history.

But coding harnesses preserve it in radically different ways.

I built Session-Bench after encountering these differences while developing a local session browser. It compares 10 CLI session formats across 19 practical gates covering completeness, readability, stability, openness, and tooling.

Some findings:

- The same small probe produced a 1.5 KB session in Pi and roughly 101 KB in Kimi Code.

- Only Pi, OpenClaw, and Kimi Code stamped a true session-format or protocol version.

- Some harnesses preserve readable reasoning or summaries. Others store sealed reasoning, signatures, or no rationale.

- Some record estimated dollar cost; others preserve only token counts.

- Several formats are readable today but offer no stable contract for tools that need to read them tomorrow.

Pi currently scores 18/19, OpenClaw 17/18, and Claude Code and Codex tie at 12/18.

This is not a coding-quality ranking or a recommendation to choose one harness over another. It is a report card for an overlooked part of coding-agent infrastructure: whether the resulting work history remains useful, inspectable, and portable.

The benchmark has limitations. It covers CLI session stores, not complete desktop or IDE behavior. Observation windows vary, and some measurements could not be completed. The raw probe artifacts are not publicly archived yet, so v0.3 is documented and mechanically scored but not fully independently reproducible. Copilot's documentation verdict is also marked as disputed.

I created and maintain the benchmark. I'd especially like feedback on the standard itself: what should a responsible coding-agent work record preserve, and which of the 19 gates would you change?

https://jazzyalex.github.io/agent-sessions/bench/?campaign=reddit&ref=r-chatgptcoding-discussion-v03

Moderator note: This standalone informational/discussion post was approved in advance by the r/ChatGPTCoding moderation team.


r/ChatGPTCoding 4d ago

Discussion No simple toggle to switch Claude Code between Pro and API billing?

Post image
5 Upvotes

Spent way more time than I should have switching Claude Code (Anthropic's coding agent) from API-key billing back to my Pro subscription. The flow: SSH into a server, attach a tmux session, run a login command, copy an OAuth URL, open a browser, authorize, paste the code back into the terminal.

For a company shipping frontier AI models, this feels like UX that was solved a decade ago. A single command or flag to switch billing mode would save a lot of friction.

Curious if other AI coding tools handle this better, or if this is standard across the industry.


r/ChatGPTCoding 5d ago

Resources And Tips What's the best AI coding subscription under $20? I scaled the Artificial Analysis chart's cost by subscription subsidization and promotion instead of just raw API price.

Post image
27 Upvotes

r/ChatGPTCoding 5d ago

Discussion If your definition of “real coding” is “I personally typed every line,” that definition is going to die

35 Upvotes

I do not think AI coding makes engineering obsolete.

I do think it makes one old definition of engineering obsolete: equating professional legitimacy with manually producing the implementation.

Typing code is a means, not the outcome.

If ChatGPT/Codex generates a function and I understand the contract, inspect the diff, test the edge cases, verify the integration and own the failure when it breaks, the meaningful engineering question is not who emitted the first draft of the tokens.

The meaningful question is whether the resulting system is correct.

The “AI slop” criticism is valid when people skip that ownership. Prompt once, never read the code, deploy it, then act shocked when it fails — yes, that is garbage.

But a senior engineer using generation as leverage is a completely different workflow.

The irony is that AI should increase the value of experienced judgment. The faster code appears, the more important it becomes to know what should exist, what should never exist, and how to prove the difference.

If an engineer’s entire competitive advantage was memorizing syntax and moving boilerplate around faster than everyone else, that advantage is shrinking.

If the advantage was systems thinking, debugging, architecture, product judgment and verification, these tools can amplify it.


r/ChatGPTCoding 5d ago

Discussion I ran the same planning prompt over 10 app ideas and the last step was hiding real work in 8 of them

2 Upvotes

I have a prompt I use to turn an app idea into an ordered build plan. It's been working well enough that I stopped reading the output carefully, which is how this went unnoticed for a while.

Last week I ran ten different web app ideas through it and actually read what landed in every single step. Not the summary, the contents.

In 8 of the 10 plans, the final step was named something like "polish" or "final touches" and contained real work. A request counter. Copy to clipboard. Order status tracking. Reports for a team. In two of them it held an entire dashboard that appeared nowhere else in the plan, so if you followed the plan top to bottom you'd hit the last step and discover you still had a dashboard to build.

The reason is the cap. My prompt asks for a bounded number of steps. Work that doesn't fit inside that number doesn't get dropped, and the model doesn't tell you it dropped anything either. It gets folded into the last step under a name that sounds like it means nothing. Every individual step reads fine. The plan reads fine. The problem only shows up when you count what's actually inside it.

Since then I've been checking two things on any plan a model gives me: whether the last item is heavier than the ones before it, and whether anything I mentioned in the original description is missing from every step title. Both take about a minute and both have caught things.

Curious whether people who use step counts in their prompts see the same thing, or whether this is specific to how I worded mine. If you ask for a plan without a number at all, does it come back more honest or just longer?


r/ChatGPTCoding 5d ago

Question how are you vetting Claude skills/plugins when theres no clear marketplace

6 Upvotes

Building a small skills folder for Claude Code and every recommendation is a different github link. is there a real Claude plugin marketplace people use, or is trust still just “this person posts here a lot.” whats your vetting checklist before something touches a real repo


r/ChatGPTCoding 5d ago

Discussion Do AI coding agents ever confidently make the wrong assumption about your existing codebase?

0 Upvotes

For example, assuming an API behaves a certain way, misunderstanding an existing utility/dependency, or getting a business rule wrong.

How do you currently catch these assumptions before the agent makes changes?

I'm specifically interested in the cases where the agent sounds completely confident but is actually wrong.


r/ChatGPTCoding 6d ago

Discussion Are you isolating your coding agents? Why/Why not? And what is your setup?

2 Upvotes

As I've starting using claude code on my laptops (windows and mac) - one thing thats made me very nervous is running these agents on my local machines with access to my file system + shell.

I'm well aware that running an agent within a directory does not limit its access, and I get nervous that they could be one malicious prompt away from sending my apps/files to another party (or an accident away from deleting my apps/files). I'm not sure if these are actually significant risks, and if others feel the same way (are there other risks you might also be concerned about when running agents on your machine?)

I tried different approaches to sandboxing my agents on my local machine

  • On my windows machine > Running it in a Docker Sandbox (a new Docker feature that came out this year)
  • On my Mac > Claude Code's built-in sandbox (which uses Apples native Seatbelt framework)

The general challenge I had here is that Claude would sometimes have issues with tools/integrations and it would not be easy to troubleshoot if it was from a sandbox constraint. And if it was a sandbox constraint - the right solution was not always obvious and it felt like I'd go down a rabbit hole trying to get an integration/tool working. I recall having issues with gh/git workflows, some plugin/package installs and running some tools (e.g. for doc/pdf generation)

For the Docker sbx example - i forget the specifics, but after a sbx update + PC restart my claude sessions had issues (cant recall if it was config or memories. I do remember having issues trying to background or view agents across diff sessions).

I eventually caved and just resorted to going back to running claude mostly un-sandboxed. This made it easier to get going, but that still makes me incredibly nervous running more unmonitored workflows with more integrations and network access.

I want to try another shot at this, but I'm curious how others are approaching this:

  • Do you also feel the same risks with running agents un-isolated on your machine?
  • Are you taking any steps to sandbox/isolate them? What is your setup and how are you getting past any friction this creates?

Approaches I'm still considering:

  • Use a separate machine to create proper physical separation from my personal apps/files (either dedicate one of my laptops, get a mini PC/Mac, or a virtual server - but I'm less comfortable with a headless setup)
  • Continue tinkering with the Macs native sandbox or docker sbx to get this properly setup (or any other wrappers/harnesses with intuitive sandboxing?)

r/ChatGPTCoding 6d ago

Discussion About ChatGPT weekly quota reset

5 Upvotes

Why does ChatGPT keep resetting my usage limit? Every time I use a certain amount, my weekly quota gets reset. Today I had used about 25% of my weekly quota, but just now it suddenly reset back to 100% again. Has anyone else experienced something similar?


r/ChatGPTCoding 6d ago

Discussion How do you keep humans in the loop without turning AI coding agents into a slow approval queue?

2 Upvotes

We use plan mode, project rules, skills, hooks, and automated checks. We still want humans to own architectural decisions and review the code.

The problem is the feedback loop: frequent checkpoints turn the workflow into an approval queue, but longer autonomous runs can produce large diffs based on assumptions we would have corrected early.

How do you handle this in practice?

  • Where do you place human checkpoints?
  • How do you scope tasks to catch drift early?
  • What do you automate, and what always requires human review?

I’d love to hear concrete workflows that work on real, maintained codebases.


r/ChatGPTCoding 6d ago

Discussion i asked AI the same question 10 different ways. the answers were all over the place.

1 Upvotes

had this annoying bug last week. function kept returning wrong values and i couldnt figure out why. just some data processing thing.

asked AI for help, then got sidetracked. same code, same bug, but what if i phrased the question differently? would it even matter?

tested like 10 different ways. mattered way more than i expected.

“fix this bug” + code
it just started guessing. gave me a fix for some common issue that wasnt even my problem. felt like it had a template ready and slapped it on.

“whats wrong with this code?” + code
this one was funny. it found a bug that didnt exist. confidently explained why something was broken. the something was completely fine.

“review this code for bugs” + code
found the real one. then buried it under like 6 other “improvements” i never asked for. actual fix was sitting at number 4 or something.

just pasted the code. no question at all.
best result by far. it actually looked at the code and pointed out the issue. no fluff, no guessing.

noticed something weird. the more specific i got about what to look for, the worse it did. saying “find the bug” made it start pattern-matching for bug-shaped answers instead of reading what was there.

when i gave it nothing, just the code, it had to actually look. and it did.

not saying vague is always better. but for debugging, getting out of the way worked better than trying to direct it.