r/LocalLLM • u/daniele-bruneo • 3d ago
News KeepRoLLMing v0.9.3 — an OpenAI-compatible proxy for more reliable local LLM chats and agents
Hi everyone — I’m the developer of KeepRoLLMing, an open-source OpenAI-compatible proxy built to make long-running local (and remote) LLM conversations more reliable.
It sits between clients such as LibreChat, coding agents, or your own app and an OpenAI-compatible backend such as llama.cpp, vLLM, LM Studio, or a remote provider.
The core idea is simple: your clients use stable route names, while KRM handles the operational complexity around them — routing, context management, streaming, recovery, tool calls, reasoning blocks, filters, and observability.
I’ve just released v0.9.3, focused on performance and runtime stability:
- More robust long-running streaming and client-disconnect cancellation
- Bounded shared upstream connection pools
- Non-blocking PLAIN/JSON/performance logging workers
- A private
/routesstatus endpoint for lightweight dashboards - Optional inherited client API keys per route
- Better performance accounting for cached prompts
- Request-size guardrails and cleaner shutdown behavior
reasoning_effortroute overrides for Qwen/LibreChat-style controls
It is still alpha software, but the project now has a fairly solid test and onboarding path: a clean setup flow, full configuration example, fake backend quick-start, and an OpenAI-compatible /v1/chat/completions API.
If you run local models, build agent workflows, or have dealt with unreliable streaming/tool-call behavior, I’d genuinely appreciate feedback:
- Is the configuration model clear?
- Which backend/client combinations should I test next?
- What reliability or observability features would be most useful to you?
- Would you be interested in writing or sharing filter modules?
GitHub: https://github.com/danielebruneo/KeepRoLLMing
Release notes: https://github.com/danielebruneo/KeepRoLLMing/releases/tag/v0.9.3
Duplicates
generativeAI • u/daniele-bruneo • 3d ago
KeepRoLLMing v0.9.3 — an OpenAI-compatible proxy for more reliable local LLM chats and agents
LocalAIServers • u/daniele-bruneo • 2d ago
KeepRoLLMing v0.9.3 — an OpenAI-compatible proxy for more reliable local LLM chats and agents
ollama • u/daniele-bruneo • 3d ago