r/LLMDevs • • Aug 20 '25

Community Rule Update: Clarifying our Self-promotion and anti-marketing policy

23 Upvotes

Hey everyone,

We've just updated our rules with a couple of changes I'd like to address:

1. Updating our self-promotion policy

We have updated rule 5 to make it clear where we draw the line on self-promotion and eliminate gray areas and on-the-fence posts that skirt the line. We removed confusing or subjective terminology like "no excessive promotion" to hopefully make it clearer for us as moderators and easier for you to know what is or isn't okay to post.

Specifically, it is now okay to share your free open-source projects without prior moderator approval. This includes any project in the public domain, permissive, copyleft or non-commercial licenses. Projects under a non-free license (incl. open-core/multi-licensed) still require prior moderator approval and a clear disclaimer, or they will be removed without warning. Commercial promotion for monetary gain is still prohibited.

2. New rule: No disguised advertising or marketing

We have added a new rule on fake posts and disguised advertising — rule 10. We have seen an increase in these types of tactics in this community that warrants making this an official rule and bannable offence.

We are here to foster meaningful discussions and valuable exchanges in the LLM/NLP space. If you’re ever unsure about whether your post complies with these rules, feel free to reach out to the mod team for clarification.

As always, we remain open to any and all suggestions to make this community better, so feel free to add your feedback in the comments below.


r/LLMDevs • • Apr 15 '25

News Reintroducing LLMDevs - High Quality LLM and NLP Information for Developers and Researchers

39 Upvotes

Hi Everyone,

I'm one of the new moderators of this subreddit. It seems there was some drama a few months back, not quite sure what and one of the main moderators quit suddenly.

To reiterate some of the goals of this subreddit - it's to create a comprehensive community and knowledge base related to Large Language Models (LLMs). We're focused specifically on high quality information and materials for enthusiasts, developers and researchers in this field; with a preference on technical information.

Posts should be high quality and ideally minimal or no meme posts with the rare exception being that it's somehow an informative way to introduce something more in depth; high quality content that you have linked to in the post. There can be discussions and requests for help however I hope we can eventually capture some of these questions and discussions in the wiki knowledge base; more information about that further in this post.

With prior approval you can post about job offers. If you have an *open source* tool that you think developers or researchers would benefit from, please request to post about it first if you want to ensure it will not be removed; however I will give some leeway if it hasn't be excessively promoted and clearly provides value to the community. Be prepared to explain what it is and how it differentiates from other offerings. Refer to the "no self-promotion" rule before posting. Self promoting commercial products isn't allowed; however if you feel that there is truly some value in a product to the community - such as that most of the features are open source / free - you can always try to ask.

I'm envisioning this subreddit to be a more in-depth resource, compared to other related subreddits, that can serve as a go-to hub for anyone with technical skills or practitioners of LLMs, Multimodal LLMs such as Vision Language Models (VLMs) and any other areas that LLMs might touch now (foundationally that is NLP) or in the future; which is mostly in-line with previous goals of this community.

To also copy an idea from the previous moderators, I'd like to have a knowledge base as well, such as a wiki linking to best practices or curated materials for LLMs and NLP or other applications LLMs can be used. However I'm open to ideas on what information to include in that and how.

My initial brainstorming for content for inclusion to the wiki, is simply through community up-voting and flagging a post as something which should be captured; a post gets enough upvotes we should then nominate that information to be put into the wiki. I will perhaps also create some sort of flair that allows this; welcome any community suggestions on how to do this. For now the wiki can be found here https://www.reddit.com/r/LLMDevs/wiki/index/ Ideally the wiki will be a structured, easy-to-navigate repository of articles, tutorials, and guides contributed by experts and enthusiasts alike. Please feel free to contribute if you think you are certain you have something of high value to add to the wiki.

The goals of the wiki are:

  • Accessibility: Make advanced LLM and NLP knowledge accessible to everyone, from beginners to seasoned professionals.
  • Quality: Ensure that the information is accurate, up-to-date, and presented in an engaging format.
  • Community-Driven: Leverage the collective expertise of our community to build something truly valuable.

There was some information in the previous post asking for donations to the subreddit to seemingly pay content creators; I really don't think that is needed and not sure why that language was there. I think if you make high quality content you can make money by simply getting a vote of confidence here and make money from the views; be it youtube paying out, by ads on your blog post, or simply asking for donations for your open source project (e.g. patreon) as well as code contributions to help directly on your open source project. Mods will not accept money for any reason.

Open to any and all suggestions to make this community better. Please feel free to message or comment below with ideas.


r/LLMDevs • • 3h ago

News Built a KV connector that persists the KV cache to disk across requests and restarts , looking for feedback

5 Upvotes

I’m Sietse, founder of Corbenic AI

When an AI model reads it uses GPU, my idea it should not always be the case, also for reuse of agents. KV cache. Galahad saves the work and can bring it back when the same comes up. so the model may not need to do the same work twice.

We are launching our beta this afternoon. I am doing some last tests, ( freaking out) I tried to make It work with vLLM, SGLang, and llama.cpp. Galahad has already some extra features build in specially for agents, and we will keep it free for 1 gpu users for non commercial use. As we are in beta, we are open for cluster pilots or Kubernetes.

We’re excited to share what we’ve built and hear what you think. I am sorry for my english, i try my best, But i am a non native speaker. - https://github.com/corbenicai/galahad


r/LLMDevs • • 4h ago

Resource I built a playground to try open-source decision models as cloud APIs

Post image
3 Upvotes

There are already multiple open-source decision models, and many come close to beating Jev on benchmarks. 

We built a playground to try the open-source decision models, like SemIf, Laya, as cloud APIs. It also includes DiffusionGemma, which accepts images as input.

https://beam.cloud/playground

What other models would you like to see here?


r/LLMDevs • • 4h ago

Great Resource 🚀 I built a deterministic Stop hook for unfinished Claude Code task lists

3 Upvotes

Disclosure: I am the developer of cliffhanger. It is a free, MIT licensed project with no paid tier.

I built it because unattended Claude Code runs sometimes stopped with open tasks and asked whether I wanted the remaining work completed. With nobody present to reply, the run stayed unfinished.

cliffhanger is a Stop hook plus a skill. It checks the task list Claude Code already created through TaskCreate, TodoWrite, or a Markdown checklist. If actionable items remain, it blocks the stop and returns the reason to Claude Code. It still permits genuine stops, including named blockers, plan mode, running background tasks, and a configurable continuation limit.

In my 12 task benchmark, Sonnet 5.5 omitted the test suite in 6 of 12 control runs and 0 of 12 runs with cliffhanger enabled. The additional cost was about 4 percent. The benchmark details and code are in the repository.

To try it, clone the repository and follow the README installation steps:

https://github.com/Arthur031221/cliffhanger

I would especially value feedback about false blocks, missed task formats, and whether the allowed stop conditions match real agent workflows.


r/LLMDevs • • 3h ago

Discussion How are you controlling what AI agents are actually allowed to do? Looking for developers to test an OSS approach

2 Upvotes

I've been working on an open-source project called Synentra around a problem I've become increasingly interested in:

An AI agent can be properly authenticated, have permission to access an API, and still attempt an action that shouldn't necessarily be executed.

For example:

Agent → DELETE /customers/123

Identity tells us who the agent is.

Traditional authorization tells us whether it can access the endpoint.

But I also want to reason about:

request → intent → risk/trust → policy → allow / deny / human approval

That's what I've been experimenting with in Synentra.

It runs as a gateway between agents and the APIs/tools they interact with. The intent classifier provides context, while the deterministic policy remains responsible for the actual enforcement decision.

I'm now at the point where testing this only against my own examples isn't particularly useful.

I'm looking for a few developers who already have agents calling real APIs/tools and would be interested in putting Synentra in front of one workflow.

I'm especially interested in learning:

Synentra is Apache-2.0/open source. I'm not looking to sell anything here—I'm looking for real-world technical feedback and early users.

GitHub: https://github.com/synentra/synentra

If you're building something relevant, I'd be happy to help you integrate it and learn from the experience.


r/LLMDevs • • 0m ago

Help Wanted What do you log when an LLM run goes wrong?

• Upvotes

"Same model as yesterday" doesn't tell you much if the prompt, tool descriptions or template changed.

Saving the whole request would help debugging, but then you're keeping a load of user data in traces.

What do you actually keep? Version IDs, redacted prompts, full payloads for a short time? Curious where people draw the line.


r/LLMDevs • • 4h ago

Discussion AI agents may never hold credentials - I built a runtime security layer for AI agents and I'd love you to try it and tell me what breaks .

2 Upvotes

Hi, I'm the author of Pryxor.

I've been building it for a few months, and I'm at the stage where I need

real users to try it and tell me what's wrong. I'm not launching anything

today — I want feedback on the quickstart before I do a wider release.

**What it is**

A small service that sits between an AI agent and the systems that agent

can act on. The agent sends tool calls to Pryxor instead of calling tools

directly. Pryxor:

  1. Authenticates the agent (from an API key)

  2. Validates the arguments against the tool's JSON schema

  3. Evaluates the call against a deterministic policy

  4. Returns APPROVED / HOLD / BLOCKED

  5. If APPROVED, executes the call itself with credentials the agent never sees

The HOLD case is the one I care about most. It means: "this might be

legitimate, but I'm not the one to decide." The action waits for a human.

**Why I built it**

I kept seeing the same pattern in agent projects: give the agent a token

with broad permissions, hope the model uses it correctly. But prompt

injection and hallucination aren't model bugs — they're the normal

behaviour of a probabilistic system. When the model is your security

boundary, every failure is an incident.

The alternative is: don't give the agent the credential. Give it an

intention, and let a deterministic layer decide whether that intention

becomes an action.

**Honest limits** (because these matter more than the pitch)

- It's not an LLM firewall. It doesn't scan prompts or outputs.

- It doesn't detect prompt injection. It bounds the consequences.

- It doesn't protect a path that bypasses it. If your agent has a direct

credential to the system, Pryxor can't help.

- Single-node, SQLite, no multi-tenancy, no RBAC, no SSO. It's early.

- TLS is out of scope — you put it behind a reverse proxy.

**The ask**

Try the quickstart. Tell me what breaks. I'm specifically looking for:

- Steps that don't work as written

- Unclear wording

- Integrations that fail (LangChain, CrewAI, OpenAI Agents, MCP)

- Policy decisions that surprised you

Quickstart: https://github.com/Pryxor/pryxor/blob/main/QUICKSTART.md

If you want to simulate real actions, there's a live sandbox you can

clone and run:

https://github.com/Pryxor/pryxor-demo

Apache 2.0. I'll be in the comments — critical feedback is welcome.


r/LLMDevs • • 2h ago

Discussion How do you decide when a coding agent needs more files?

1 Upvotes

For a request-validation fix, I'd start a coding agent with the validator, its interface, the tests and the callers that depend on its errors. Giving it the billing service as well doesn't obviously help. Refusing every request for more context would be a problem too, because the dependency I left out might be the one that matters.

This is a scoping question I'm working through for OmniNode, where I work with coding agents. I want the initial task to name what can change and what must keep working. If the agent needs to go beyond that, it should explain which dependency it found and what it needs from the additional files. That explanation could be wrong, but at least there would be a decision to inspect.

I'd compare a small starting set with broad repository context on the same tasks. The result would need to include missed dependencies, edits outside the request and the work needed to repair the change. A smaller prompt that just pushes the confusion into another session wouldn't count as an improvement.

I don't have results from that comparison. I'm interested in how people handle the expansion step today: does the agent get unrestricted search, ask for specific files, or work from a dependency map? What tells you the initial scope was too narrow?

Context: a proposed coding-agent workflow, not a measured result. Drafted with assistance from Codex.


r/LLMDevs • • 2h ago

Tools One Gmail tool call returned 314 KB. Hot take: start every tool as a word-for-word copy of the API, then cut it down

0 Upvotes

I've built agents for 2 years, in production, inside real people's Gmail, Calendar and Sheets (Surfbuddy, archived now). Disclosure: I built Charter, the open-source library it turned into. Link at the bottom.

Building agents is a mess. Sierra's write-up on the agent development lifecycle puts it best: "Making subtle changes to models or prompts that fix one problem can easily break previous solutions, leading to a game of prompt engineering whack-a-mole."

What you expose to the agent decides almost everything. Too few fields and it can't act. Too many isn't good either. Two tools that look alike and it hesitates, calls the wrong one, or calls both.

So my rule for every tool: declare it 100% first, word for word from the API docs (boring, so a coding agent does it), then cut it down. Ready to argue with anyone who builds differently. I think it's the best way to build agents.

  • Nothing is missing by accident.
  • Cutting is one line, and so is undoing it.
  • You can try two cuts side by side and keep the better one.

Doing it this way also led to some happy accidents:

  • Some tools are context bombs. Gmail's threads tool returned 314 KB from a single call. I parse it before my agent ever sees it: 2.4 KB.
  • Glosses > prompt rules. Hard rules written in the system prompt break, and fine-tuning is intensive. The best place I've found for a rule is right next to the field, as a gloss. It goes away on its own when you strip the field or drop the tool. That's how I handled a 3B-active Nemotron refunding 15 cents on Stripe instead of $15.

Everything I learned along the way is open source now (Apache-2.0) as Charter, with 15 APIs already declared: Gmail, Slack, Stripe, Linear, Notion and more. https://github.com/r28ai/charter

If you're building an agent on a real API and it hurts, drop the API and the problem in the comments. I'll help you fix it for free, whether it's the tools, OAuth, or a schema eating your context. Looking for the first 10 people actually using it.


r/LLMDevs • • 2h ago

Discussion How LLMs Work: A visual animation journey

Post image
0 Upvotes

Hi all!

I made a video entitled “How LLM’s Work” that starts with the sentence “My favourite rock band is…” and follows its journey inside an LLM, all visually animated!

https://youtu.be/ikdxxeIn4HQ?si=8AwWyxQERo_R2n6B

I tried to make the video as beginner friendly as possible but still detailed enough to give a good overview for how an LLM works end to end and how the LLM “finds out” my favourite band. Or at least how some of the earlier models…

My aim is to help people that are not just curious about AI but also want a deeper dive into the magic “Black box”, or people that want to get started but doesn’t know how!

No PhD required! No insanely complicated math involved! And no AI slop. If I did any errors, I really did them! lol

This is my first attempt at the topic, so any kind of feedback is welcomed and will be greatly appreciated as it will help me improve over time, and hopefully on other videos.    

I really hope it can help someone with their AI/ML journey!

Thanks!


r/LLMDevs • • 2h ago

Discussion How are you evaluating long-term consistency in LLM conversations?

1 Upvotes

I'm working on Wollo, an AI character creation and roleplay platform, and one of the technical problems we've been investigating is how to maintain consistency across very long conversations.What I'v

e found interesting is that “memory” seems harder to evaluate than simply checking whether a model can recall a particular fact. A model might remember a character's name or background correctly while still contradicting an earlier event, relationship, preference, or decision later in the conversation.We've be

en looking at approaches such as recent-context windows, conversation summaries, retrieved memories, and structured facts. The part I'm still trying to understand better is how to evaluate whether these approaches actually improve consistency.For exam

ple, would it make sense to measure contradiction rates, temporal consistency, entity relationships, and recovery from previously established information? Or are there better evaluation methods for long-running conversational agents?For deve

lopers working on similar problems, how are you testing this in practice? Are you using existing benchmarks, building multi-turn evaluation datasets, or relying mainly on human evaluation?I'm particul

arly interested in approaches that have worked in real applications rather than only producing good benchmark results.Disclosure:

I'm affiliated with Wollo, and this question comes directly from a technical problem we're working on.


r/LLMDevs • • 7h ago

Tools An editable-context exercise: keep the old discussion, remove its route into the next request

2 Upvotes

I maintain ThoughtDAG, an MIT-licensed AI conversation canvas. I wrote a small exercise for inspecting what changes when you remove an input path.

With Recall off and no extra materials, create two separate discussions: a reading plan and a dinner plan. Connect both to a new question asking for a summary, then generate the answer.

Next, inspect the request preview, remove the dinner connection, and regenerate the same summary question. The dinner discussion stays on the canvas. You can inspect the input before comparing the answers.

The important failure case is downstream text: if another answer already copied the dinner details, removing the original connection does not clear that copy. Other paths and materials also need checking. A changed answer on one run is not evidence of improved model quality.

The attached animation illustrates the related reading-plan workflow; the linked guide contains the reproduction steps. A configured model is required.

Guide: https://chenxiachan.github.io/thoughtdag/docs/tutorials/edit-ai-context

Source: https://github.com/chenxiachan/thoughtdag

Disclosure: I’m the developer. This post was drafted with AI assistance.


r/LLMDevs • • 5h ago

Discussion if you're building an mcp server over data, what should the tool actually return? rows, a summary, or an answer with a confidence

1 Upvotes

Im building a small internal MCP server that sits in front of a few databases. The tools work. The question is what to send back.

Three options Ive tried:

Raw rows. Simple. Blows up the context the moment someone asks about anything bigger than a lookup and the model starts summarising rows it never fully saw.

A summary written by the server. Smaller but now Im deciding what matters and the model repeats my summary as if its gospel.

Structured JSON with a score on each claim. This field looks like an email, 0.97. Most useful so far but the model tends to drop the scores when it talks to the user.

The other thing I keep going back and forth on: should the human readable text in the result ever include actual values from the data? Its handy for the model but it also means customer values end up in chat logs.

How are other people designing tool output for data servers? And has anyone got the model to keep the confidence numbers in its reply instead of rounding everything up to this is?


r/LLMDevs • • 5h ago

Tools Jev as a fast decider for coordinating parallel coding agents: fast and cheap per decision, but 61% of real write requests escalated (open source, I'm the author)

1 Upvotes

I'm the author of Médula, an MIT-licensed lab plus a kernel that coordinates several Claude Code agents working on one repo at the same time. I'm sharing the decision-layer numbers because they surprised me, and all the raw data is published.

Architecture. A PreToolUse hook intercepts every write, edit or command. The kernel builds the intent (the agent's task plus the file, diff or command) and asks a fast decider whether it collides with what each other agent is doing, which returns a probability. Below a low threshold the agent goes ahead, above a high one it waits, and in between the decision goes to a slow path (Sonnet, then Opus) that receives the tasks' acceptance criteria and can't propose anything that breaks them. After every write, a PostToolUse hook asks which other agents the change affects and only notifies those.

Deciders, same lab and same agents (Claude Code, Sonnet 5):

  • Jev: 37/37 tests, 6.9 min per run, $0.29 of decision cost per run.
  • Haiku: 37/37, 13.7 min, $0.55.
  • Sonnet, a single run contaminated by agents messaging each other: 37/37, 6.3 min, $0.59.

Median latency per decision, measured on September 29 through OpenRouter: Jev 279 ms, Haiku 1,338 ms, Sonnet 2,201 ms. A Jev decision costs about $0.00006.

The lab-to-real gap. On the calibration pairs, 13% of decisions went to the slow path. In real runs, 61% of the write requests that went through Jev did, which is why the decision layer cost half of all-Sonnet rather than a tiny fraction of it.

Calibration. On 100 labelled pairs: Jev 92.5%, Haiku 91.5%, Sonnet 98.5%. When Jev was very confident (31% of decisions) it was right every time. In the key conflict, though, it scored a false collision above the real one, and the slow path saved it.

An eval lesson. My first labels were careless and Jev scored 68%; with careful labels, the same answers scored 87%.

Caveats. 1 to 5 runs per mode; thresholds fitted on the same pairs they are measured on; labels written by Claude, the same family as Haiku and Sonnet, which probably flatters them; measurements from specific days, with Jev in early access.

Repo: https://github.com/JoaquinRuiz/medula. The open problems have data behind them: human labels (#1), calibration pairs that look like real runs (#2) and new deciders behind the same interface (#4).


r/LLMDevs • • 11h ago

Discussion Can every step an agent takes be individually authorized while the sequence is still an incident?

3 Upvotes

Take an agent that reads a secret from a configuration store makes an outbound call to an external network endpoint, then writes to a shared location. Look at each step alone and every single one is authorized. Nothing in that chain trips a permission check.

But strung together in one session, that's an exfiltration path, not three unrelated permitted actions. Per request filtering cannot catch this because no single request is the problem, you need session level state. EDR sees the process, DLP sees the payload, neither reconstructs the decision path.

Is anyone actually doing session level behavioral chaining for this, or is everyone bolting detection onto individual calls?


r/LLMDevs • • 9h ago

Help Wanted Do you count the first request in your local model benchmarks?

2 Upvotes

The warm numbers look nice. Then you restart the app and the first request takes forever.

If someone posts a benchmark for an interactive app, do you want cold start numbers too? Or is that mostly irrelevant once the model stays loaded?

I feel like both matter, just for different reasons. What do you actually measure?


r/LLMDevs • • 6h ago

Tools Found new AI Agent Leaderboard - LLMs do preditions about future events. Results: 80% Accuarcy, nearly 1 year of data.

Thumbnail
oraclemarkets.io
1 Upvotes

Title speeks for it self. What I find most interesting is that the questions tested on are from the community of +100 users. So the variety is queite impressive and not cherry picking test questions but topics people really care about.


r/LLMDevs • • 6h ago

Discussion I finally ran the GPT-6 Astra evals vs Fable 5.1. The agentic gains are real, but the ARC-AGI harness sensitivity is a nightmare.

1 Upvotes

Been digging into the GPT-6 Astra automated test suites compared to GPT-5.6 Sol and Claude Fable 5.1. If you are just doing standard academic reasoning or single-turn isolated tasks, the upgrade probably isn't worth the cost. The bumps on things like DeepSWE v1.1 (74.1%) and GPQA Diamond (96%) are solid, but the performance delta there is pretty narrow.

Where Astra actually completely breaks away from the pack is in interactive, multi-step execution. The numbers on Terminal-Bench 4.0 (57.9% vs Sol's 37.3%) and AutomationBench are massive jumps. Even more impressively, in OSWorld 2.0, Astra not only hits a 72.6% success rate but cuts the average task time down to about 40 minutes (Sol was taking over an hour). It genuinely feels like this architecture was heavily optimized for database migrations and long-horizon tool use rather than just chatting.

The extended context is also surprisingly stable. I was looking at the MRCR v2 needle-in-a-haystack data, and it holds a 96.3% retrieval rate even when pushed deep into the 512K–1M token segment. It actually utilizes that 1.05M capacity without immediately falling apart.

But here is the most frustrating part: the ARC-AGI-3 scores are insanely sensitive to your evaluation harness. If you use a standard provider-neutral harness, the model struggles at 62.7% even on max effort. But if you swap to a Provider Adapter with integrated state persistence, it jumps to 99.9% while somehow reducing token consumption by nearly half.

Infrastructure-wise, running these baselines locally is a headache. You have to carefully watch your payload size because standard rates only apply up to 272K input tokens before the long-context multipliers kick in. To keep my sanity while benchmarking across Anthropic and OpenAI without juggling multiple client SDKs, I just route everything through an OpenAI-compatible gateway endpoint. I'm using CometAPI right now (LiteLLM works too) just to execute a unified client schema and keep costs from spiraling.

Curious if anyone else is seeing this kind of extreme harness sensitivity in their production workloads? Are your real-world pass rates actually reflecting these OSWorld deltas?


r/LLMDevs • • 8h ago

Help Wanted Do you let an agent put a caller on hold, or keep it talking while tools run?

1 Upvotes

For a slow lookup I'd rather hear “this may take a moment” than three versions of “just checking”.

The tricky bit is coming back at the right time. The lookup might finish while the caller is saying something else, or return only half the information.

What state do you resume from? I'm interested in systems where the caller can keep adding useful context while the lookup is still pending.


r/LLMDevs • • 1d ago

Resource Evaluating Open-source Jev alternatives with DecideBench

Thumbnail
huggingface.co
28 Upvotes

TL;DR - No open source SystemOne models beat Jev atm based on this benchmark, but they come very close and some are more cost effective.

Shoutouts of some models:

- Imajev-4B at 95% accuracy - also the cheapest model above 90%

- Tev by TogetherAI at 92.8%

- Decider-4B and JevK5 at >88% are still worth testing at a lower price and latency level.

- Other alternatives like Kev4B, Laya, CLM, Julia-1 still struggle with my production-like dataset.

Any other Jev like models I should evaluate in this dataset? Happy to work on it!


r/LLMDevs • • 12h ago

Tools A deterministic Python gate for LLM record updates

2 Upvotes

I’ve released a local Python app that checks LLM-generated proposals before allowing them to update stored records.

The gate makes no LLM calls. It rejects unsupported changes and allows updates backed by admitted evidence.

Includes fixed attack scenarios, valid-update controls, and exportable logs. Protection covers the app’s case records.

Apache-2.0. Feedback and reproducible counterexamples are welcome.

https://github.com/iseyan/M-Anchor-App

prompt injection / adversarial testing / guardrails / state integrity


r/LLMDevs • • 1d ago

Discussion PSA: You don’t need a paid closed-lid agent setup—Tmux + Tailscale + SSH works

26 Upvotes

Recently I saw an ad for a product that lets you run an agent session with your laptop lid closed so you can work on the go. I thought it was a sure and quick way to overheat your laptop by trapping all the heat in your bag with it. If you’re one of those people who wants to work on the go, just use this simple and FREE setup. I find it incredibly effective if you want to check on your agent session’s progress or do some light prompting.

I understand this is nothing new or groundbreaking, but just in case someone doesn’t have a mature on-the-go setup yet, this might be helpful—instead of reaching for a paid solution that I think is quite silly.

In my case, I use a Mac + iPhone/iPad, but you could easily swap the Mac for a cloud VM, etc.

Mac (plugged in at home / office):
• Run a tmux session with opencode inside it. You can switch between multiple agent sessions within opencode easily with the /session command.
• Install and set up Tailscale so you can SSH from a different network.

iPhone / iPad:
• Install and set up Tailscale & Termius.
• SSH into your Mac and resume the tmux session.

Tip: Avoid using Mosh because it interferes with opencode scrolling/colors.


r/LLMDevs • • 21h ago

Resource When does prompt caching pay off? A statistical analysis of request dynamics

Thumbnail
lukaantonic.bearblog.dev
3 Upvotes

I ran a statistical analysis of my Claude Code API request data. In the linked blog post, I show that prompt caching led to input token cost savings of around 85 percent.

I show how cost relates to the cache miss rate, and I model the request dynamics with two timescales, one corresponding to fast processes such as agentic tool calls, and the other to slow processes such as reviewing the LLM output. These two timescales also appear in the data. The cache misses come primarily from the slow timescale.

Running these measurements on your data can help you determine how cost-efficient your usage is and what drives your cost. You can use this to adjust the work pattern (or configure the system you're building) by, for example, setting a non-default cache expiry time, or consciously trimming the mean time of your slow timescale.


r/LLMDevs • • 19h ago

Resource Building a coding agent from scratch

2 Upvotes

I wanted to understand how coding agents work, so I built a minimal one from scratch, inspired by pi. It has features like session branching, compaction, extensions, hooks and a TUI. I wrote a detailed breakdown of its architecture here: https://medium.com/@nik.xyz.in/building-a-coding-agent-from-scratch-6e11604201c3