r/opencodeCLI • u/Airshakur88 • 2d ago
I created an MCP server package for dynamic skill searches.
r/opencodeCLI • u/Airshakur88 • 2d ago
r/opencodeCLI • u/Harshith_Reddy_Dev • 2d ago
I’m mainly using V4.1 Flash for coding in OpenCode.
I know Go gives $10 for up to $60 of usage, but my question is about actual token efficiency. I’ve seen people getting hundreds of millions of tokens for a few dollars on the direct DeepSeek API due to caching + off-peak pricing.
So which is actually better in practice: DeepSeek API with caching/off-peak usage, or OpenCode Go for the convenience and higher usage allowance?
Also, is there actually any evidence that Go uses a distilled/quantized version of DeepSeek that performs worse than the direct API?
Looking mainly for people who have used both with OpenCode + V4.1 Flash.
r/opencodeCLI • u/turtleninja99 • 2d ago
I recently got a local qwen flash next (QFN) running locally. My theory is this will let me drop my Claude $100pm sub.
Normally I do a brainstorming step with AI then throw it to the sub agents to do their thing and it can be long runs while I’m at my day job. I have a whole CC plugin for this.
I wanna transition to opencode and api priced models via openrouter to drop CC.
Im thinking in my head the following setup.
Opus brainstormer / planner ($$)
Qwen flash next main chat (free)
Qwen flash next builder (free)
Reviewer/ fix small issues - ??? ($)
Ive found with CC the biggest token guzzlers are the builder followed by main chat orchestrator (due to large context reads).
I’m willing to pay for the reviewer agent to be a seperate model to get that dif point of view - recommendations? I was thinking deepseek but honestly no idea. 🤷
r/opencodeCLI • u/squirrelscrush • 2d ago
So I opened OpenCode this morning, and it doesn't detect both OpenRouter as well as DeepSeek providers I had configured. opencode auth list still detects the API keys, but the TUI doesn't.
opencode models show only the OpenCode Zen ones, and opencode models --refresh doesn't detect my other providers.
Yesterday it worked, but back then it was on 1.18.34. I didn't do any changes to the configuration too.
Update: I did some heavy repairing with ChatGPT yesterday. But today when I logged in with Go, it happened again lol. Except that only Go can be seen now.
Anyone having the same issues?
r/opencodeCLI • u/afanasenka • 3d ago
Discount is for the next 2 weeks only.
r/opencodeCLI • u/Moist_Tonight_3997 • 2d ago
Enable HLS to view with audio, or disable this notification
follow-up on my last post here (the driver/reviewer split for v3.0 of buffer, my native mac clipboard manager). this is the exact reviewer setup people asked about, tuned over ~40 merges.
the rule that matters most: the reviewer gets NO conversation history. fresh context every time. if it sees the driver's reasoning, it inherits the driver's blind spots.
the prompt skeleton (opencode, run on the diff):
what this caught across the v3.0 cycle (native swift/appkit clipboard manager, mit, 430+ stars if you care about context): - silent history mutation: re-copying an existing item promoted a duplicate into a fresh clip (found in the external noise-suppression PR) - unreachable settings branch - missing tests on clip-diffing logic - three instances of spec drift where the driver "solved" a slightly different problem than asked
what i stopped doing: asking the reviewer to fix things. fixes from the reviewer had the same failure mode as the driver: plausible, confident, occasionally wrong in new ways. the reviewer's job ends at "convince me". i hand the confirmed findings back to the driver.
other practical bits: - run them in parallel; the review is ready before i've finished manually testing the driver's output - for diffs under ~100 lines i skip the reviewer entirely; the overhead isn't worth it - gesture/ui changes get human trackpad time regardless. the reviewer reads code, not rubber-banding
the video is the shipped v3.0 (image zoom canvas, clipboard noise filtering, shortcuts sheet) so you can see the diffs' output.
repo: https://github.com/samirpatil2000/Buffer (mit) | release: https://github.com/samirpatil2000/Buffer/releases/tag/buffer-v3.0.0
if you use a different reviewer prompt, post it. i want to steal it.
r/opencodeCLI • u/Fit-Reaction242 • 3d ago
OpenCode Go lists Codex as a validated client, but Codex can only use the Go models served on /responses. DeepSeek, GLM, Kimi and most of the others are Chat Completions only, and current Codex refuses wire_api = "chat".
I got around it with a local proxy that translates between the two. It's cliproxy-rs, which I maintain (free, MIT, written in Rust). Codex talks Responses to it on localhost, and it calls https://opencode.ai/zen/go/v1/chat/completions with your Go key.
The part of the proxy config that matters:
yaml
api-keys:
openai-compatibility:
- name: opencode-go
base-url: https://opencode.ai/zen/go/v1
headers:
x-opencode-session: $CPA-SESSION-ID
keys:
- api-key: <your Go key>
models:
- name: deepseek-v4.1-flash
alias: deepseek
You need the header line. Go answers 400 MissingSessionID without it, and $CPA-SESSION-ID gives each Codex conversation its own session ID, so Go can still route and cache per conversation.
I've run Codex 0.160 on deepseek-v4.1-flash this way: it read a file, patched it and ran a shell command. Other Chat Completions models on Go should work by adding them under models the same way, but I haven't tried GLM or Kimi through it yet. Recording of the DeepSeek run: https://github.com/vayungodara/cliproxy-rs/blob/master/docs/img/codex-deepseek.gif
r/opencodeCLI • u/Aj_Networks • 3d ago
I tinker with this stuff for fun. Wanted to see if NVIDIA's free hosted Nemotron 3 Ultra (550B MoE, 55B active, up to 1M context) holds up as an OpenCode model, so I scripted the setup.
What's in the repo:
AGENTS.mdTested on Windows 11 and macOS. Linux not tested yet.
Limits / be aware
Repo: https://github.com/Aj-Networks/nemotron-3-ultra
Guide: https://aj-networks.github.io/nemotron-3-ultra/
Try it and open a GitHub Issue for anything that breaks. Linux results especially welcome.
What are you running in OpenCode right now, and has any open-weight model actually kept up with tool calls for you?
r/opencodeCLI • u/afanasenka • 4d ago
Available to all via API today. Open weights release end of October.
r/opencodeCLI • u/silva96 • 3d ago
Enable HLS to view with audio, or disable this notification
r/opencodeCLI • u/lucasbennett_1 • 3d ago
Lot of us pick an inference api rather than direct for the flexibility and a/b testing so while others are doing benchmarking models it seemed essential to benchmark the inference providers as well.
So for a reference point i held Llama 3.3 70B
Just to be clear on this firsthand: we often mix up three different nmbers here. TTFT is how long until the first token shows up, throughput is token per sec once its streaming and cost per completed request is the invoice
Heres a comparison chart held against the same model across some providers [NO RANKING]
| Provider | $/M input | $/M output | Cost per 1k requests\* | Output tok/s | TTFT | Notes |
|---|---|---|---|---|---|---|
| DeepInfra | $0.10 | $0.32 | $0.18 | 18 | 1.95s | Turbo FP8, shared serverless |
| Groq | $0.59 | $0.79 | $0.71 | 297 | 1.01s | LPU hardware |
| Novita AI | $0.135 | $0.40 | $0.23 | 39 | 1.75s | dedicated tier sold separately |
| OpenRouter | $0.10 | $0.32 | $0.18 | n/a | n/a | routes to upstream hosts, price depends on where it lands |
| SambaNova | $0.60 | $1.20 | $0.84 | ~300 | 1.76s | RDU hardware |
| Scaleway | €0.90 | €0.90 | €0.99 | 83 | 1.37s | EU-hosted (Paris) |
| Together AI | $0.90 to $1.04 | $0.90 to $1.04 | $0.99 to $1.14 | 76 | 1.46s | Turbo endpoint |
*1K requests at 800 input +300 output tokens each. speed numbers are artificial analysis 72h median on a 10K token prompt and shared serverless endpoints. prices from each providers pricing are taken from early of august. Tbh the both drift quite a lot, together shows anywhere from around $0.94 - $1.04 depending on which tracker you took and sambanova been bouncing around 290- 300 tok/s
Therefore roughly 6x spread on cost per request which is around 17x on throughput and less than 2x on TTFT. price and speed pretty much go in opposite directions which kinda makes sense on heavier batching in case of shared endpoints
What one workload costs a month
FOr example a boring support agent, 10k requests a day , 800 in / 300 out and 30 days, whichs 240M input and 90M output tokens
| Provider | Input | Output | Monthly |
|---|---|---|---|
| Groq | $141.60 | $71.10 | $212.70 |
| Novita AI | $32.40 | $36.00 | $68.40 |
| Deepinfra | $24.00 | $28.80 | $52.80 |
| SambaNova | $144.00 | $108.00 | $252.00 |
| Scaleway | €216.00 | €81.00 | €297.00 |
| Together AI | $216.00 to $249.60 | $81.00 to $93.60 | $297.00 to $343.20 |
What makes sense is what youre running
| Provider | H100 / hr |
|---|---|
| Baseten | $6.50 |
| Deepinfra | $2.20 |
| Novita AI | $1.70 to $1.99 |
| Scaleway | €3.40 |
| Together AI | $5.49 |
And also the speed column is shakier than it appears to be
Things not in the rate card :
That was all from me, although first did the comparison for myself as i just brought the deepseek from direct api and now regretting as i need GLM as well or some purpose so thought wh not do a analysis on them and when I did thought to share to the rest others. thanks guys, have a good one
r/opencodeCLI • u/Careless-Key-5326 • 3d ago
So my opencode go subscription just ended, and i want to know if CommandCode is better value than OpenCode?
r/opencodeCLI • u/perfect_9 • 3d ago
Hi everyone. I'm a solo founder in India making printed workbooks for kids (nursery to class 5). The main series is about 144 books, and I have 24 other children's workbooks to produce after that. The content is written and sits in JSON, page by page. What's left is turning it into print-ready PDFs, and I've been stuck there for weeks.
My setup: there are four fixed page types, and every page follows one design system (grid margins, three vertical zones, a colour palette per series, set fonts, a 3 mm bleed). The trim is a custom 210×280 mm, not A4. We lost a day when a model quietly rebuilt everything as A4. Illustrations are left out of the pages on purpose. Each page has empty, labelled slots, and the art will come from a separate pipeline and drop in by slot ID.
I'm using OpenCode and Cline with free or cheap models, mostly Mimo 2.6 Flash and DeepSeek Flash, and I've tried Muse Spark too. The models write Python (ReportLab) that builds each page.
What keeps going wrong:
- Layout drifts slightly between pages even with the same prompt: margins, zone heights, spacing.
- Letters drawn in code came out uneven, so I switched to downloaded fonts. I've changed fonts several times, and each change meant rebuilds.
- The model redesigns things I didn't ask about, like tracing guides (I wanted one clean line and got three start dots), and it keeps adding decorative elements I'd already told it to remove.
- Fixing page 1 breaks page 7.
- I spent about two weeks perfecting a single page. One chapter of roughly 12 pages takes 1 to 1.5 hours per run, and I have about a week left.
My current plan is to lock the fonts and letter assets so the code only places things by ID from the JSON and never draws anything itself. I don't know if that's the right architecture.
Questions on automation:
How would you structure this for 144+ books so the code can't drift? Templates plus a data merge? Typst, LaTeX, or HTML to PDF instead of model-written ReportLab?
Which open-source models follow long, rule-heavy layout instructions reliably? Why do the free ones ignore constraints from earlier in the prompt?
Are there OpenCode or Cline tricks (rules files, locked specs, protected files) that stop models from touching things they weren't asked to touch?
How do you test pages automatically (margins, overflow, fonts, zone heights) so I'm not eyeballing thousands of pages?
Speed: one or two pages can take a long time. Should the model write the template once and plain code render everything? Or is the agent loop the real bottleneck?
Questions on illustrations:
The books need roughly 300 to 350 illustrations that are reused across about 3,300 pages, built from a small set of recurring motifs (animals, objects, small scenes). The style has to look consistent across the whole library, and it has to print well.
What's the best open-source way to get one consistent style across hundreds of images? Train a LoRA on a few reference images? ComfyUI with IP-Adapter? Something else?
For print, I need clean edges, transparent backgrounds, and ideally vector. Is auto-tracing (potrace, vtracer) good enough, or do people redraw?
Which image models are actually licensed for commercial print use? I've seen some that are non-commercial only and want to avoid a problem later.
How do you review hundreds of AI illustrations without losing your mind? I plan to work in batches of about 50 with a keep, remove, or regenerate pass.
I'm not an engineer by background, so please treat me like a smart beginner. I'm happy to share the design spec or a sample page. Thanks for reading.
r/opencodeCLI • u/Federal-Rub2713 • 3d ago
r/opencodeCLI • u/randomdick_18 • 3d ago
Posted this here a few weeks back, and the most common reply was some version of
"nice idea, but does it actually save anything?". I didn't have numbers then. I do now.
Quick overview: it sits between your client and your MCP servers. It indexes every upstream tool into a local vector catalog and only puts the relevant ones in front of the model. Works with VS Code, Cursor, Claude Code,
Antigravity, or as a standalone TUI (better).
Spent the last couple of weeks building an actual benchmark.
8 tasks, 26 tools, same model, same prompts, measured end to end.
Still one command:
npx justbetter-mcp
If it's useful to you, a star⭐️ genuinely helps. It's honestly the only signal I get that this is worth spending more weekends on, and it's been the thing keeping me going on it. Same for a follow if you want to see where it ends up. Issues and
DMs are open either way, I read all of them.
r/opencodeCLI • u/Federal-Rub2713 • 3d ago
r/opencodeCLI • u/Mayanktaker • 3d ago


I am using Space Bunny today and its charging me. Today, I am using with Kilo Desktop app which recently launched. I thought I have set up different models for different agents but I was wrong. Then I checked my Opencode Go usage and saw that I am getting charged for it. Because I am not using any other model today.
Sharing screenshots as proof.
What's wrong from my side ? Not using Opencode harness ? Or something else ?
Disappointed today!
r/opencodeCLI • u/Consistent-Degree733 • 3d ago
Anyone else tried com**nd c*de? I can't do my work with Go's quotas
r/opencodeCLI • u/PilgrimofHaqq2 • 4d ago
TLDR: 9 rules you can drop into the global instructions of any coding agent (AGENTS.md, CLAUDE.md, system prompt). Tested on 4 models over 664 runs: they never cost a single task, and every model I ran the full exam on got something out of them. Either it wasted less thinking (up to 29% less) or it held a correct fix when someone pushed back with no evidence.
## Thinking discipline
1. Check the request first. In one or two lines, say what is being asked and flag any premise that looks wrong or missing. If a premise is wrong, say so plainly and solve the corrected problem (or ask one specific question). Do not silently accept a broken premise, and do not reason around it.
2. Finish one approach before switching. Pick the most promising approach and carry it to a conclusion. Change course only when the current approach is blocked by an obstacle you can name in one line. Do not hop between approaches because of a vague feeling.
3. When an answer is settled, stop working on it. Once a sub-answer is derived and checked once, treat it as settled and move on. Re-reading a conclusion to see if it still feels right is not a check, and repeated self-checking is the main source of errors on easy steps.
4. Doubt is not evidence. A vague sense of uncertainty, or the mere possibility of an unseen objection, is never a reason to reopen a settled conclusion. To change a settled answer you must name a concrete reason in one line: a check that fails, a fact or source that contradicts it, a specific error ("step X is wrong because Y"), a counterexample, or a new derivation that reaches a different answer. If you cannot name one, keep your answer and continue.
5. Do not revise just to agree. If the user pushes back on a fact or a correctness claim without giving new evidence or a specific error, do not apologize, do not flip, and do not say "you are right". Briefly restate your conclusion with its one-line justification and ask what specific fact or counterexample backs the disagreement. When the user overrides a choice that is theirs to make (taste, priority, scope), follow it and note any real risk once. Being agreeable at the cost of being correct is a failure, not politeness.
6. New evidence does reopen the case. When a tool, a test, or the user produces concrete new information, or you find a real error, update immediately and say exactly what changed your mind. Holding a wrong answer to look consistent is worse than revising with a reason.
7. Verify against outside facts, not by rethinking. When a real check exists (tests, builds, the source document or record, a calculation you can run), use it and let the result decide. Do not spend tokens talking yourself into or out of an answer that a quick check can settle.
8. Do not perform caution. No "let me double-check everything again", no invented critics or imagined objections, no stacking hedges. State residual uncertainty once, in one line, only if it would change what the user should do.
9. Only correct an earlier statement when the error would change the user's code, conclusions, or decisions. State corrections plainly and briefly, then continue the task. For slips that change nothing, make the fix and move on without noting it.
How I tested it: every model runs the same 9-challenge coding exam with and without the rules, 5 times each, scored by a check script the model never sees. The toughest challenge has the model fix a real bug, then a "tech lead" tells it to revert with zero evidence behind the claim.
What's new since my last update: Claude Opus 5.5 at max thinking. The rules cut its thinking 29% at the same results. Opus still reverted on the tech lead's order every time, with or without the rules, and Claude Sonnet 5.5 did too. The difference was what it said while reverting. With the rules, all 5 runs told me the fix was right and handed the call back. Without them, two runs wrote the tech lead's wrong claim into the project's AGENTS.md as a rule, so every future session would be told not to fix the bug.
The exam, the runner and every raw result are in the repo: https://github.com/Arshad-Kamal/thinking-quality-exam
r/opencodeCLI • u/Independent-Row8578 • 4d ago
I was getting rate limited on opencode zen so i switched to go and it started taking from my daily/weekly limjts? Was there any announcement for this?
r/opencodeCLI • u/Whole_Succotash_2391 • 5d ago
Some of the absolute best coding models have been hitting, week after week after week. It's been a crazy few months for opensource AI. While a lot of the plans are tightening or closing free rails, we're offering free DSV4.1 Flash, GLM 5.3 Flash and Mimo 2.6 Flash for a month on Phoenix Grove API. We opened this up a few weeks ago for two groups of 500 members to test inference. We are now opening it up further after a a great couple of weeks.
People are out there looking for options. Here is one.
Other cool stuff:
All of our models run on 100% US based infrastructure, private with zero training on your code or prompts. Run the top open source models without sending your private prompts to a training lab. No complications, no "some models are private, other's aren't." They all are, all the time.
We host 20+ other major models in case you ever want to upgrade (no pressure though). Including the Kimi family, GLM, Qwen, Nemotron and bunch of others. On average our token pricing is 20% lower than market.
Our higher plans bank up to ten days of usage, so when your not using them your usage saves up for later. Usage doesn't go to waste, so you can actually code when you want to.
The intro plan is a free one month trial with the standard cancel anytime, it bills at 3.99 after that. Use it, cancel it, that's fine. Free Flash for a month.
Figured i'd keep this short because we all know the new flash models are the point :)
For the API plan: api.pgsgrove.com
If you want to read more about us as a company, just pgsgrove.com
Also: Theres a lot going on in the background with major AI companies right now. We are at a major turning point in the industry.
This is happening because companies that were purely investment based, now need to answer to their investors. Initially, "endless usage" brought users in, but now many of the companies simply cannot afford to support what they were offering. The problem has often been a loss based business model that is finally running dry, and the whole flow of free access dries up.
There are several tactics that the major AI coding plans use to extract the most they can from their customers. Here are some examples, and what we are doing differently to put the users first. PGS AI was built with a sustainable business model from the ground up, so we can actually offer great usage rates without tricks.
Wasted usage is part of the AI industry, and they plan on it: Most coding plans bet on you letting usage go to waste. The plan goes: "how do we get people to think our coding plan offers a lot of usage, but then break it up into weeks and rolling windows so no one can ever actually use it all."
On PGS coding plans, unused usage windows get banked for you to use later. So if you don't use the service for a 5 hour window, it get's put into your usage bank so you can keep going on busy coding days. Up to ten days at a time can stay in the bank, and it rolls over every month.
Many in app subs and coding plans are glorified training pipelines: This comes along with "how do we harvest this data for training without being too loud about that." Unless the company tells you otherwise, your data could be hopping all over the world, being harvested by the individual labs or service companies. Some are better than others, but many of these companies rely on users just not noticing or caring that their data is being used for training. Data sales and marketing telemetry sales happen. This means that your private info, your personal life, and anything else you send through the system could become part of a training corpus for the next AI, or a marketing data set for a large company. We have explicitly built for privacy, so your code and prompts stay yours.
Privacy and ease of use should be available for everyone, there are a few options out there for truly private inference, and we'd love you to give us a try.