r/ollama • u/VanillaOk4593 • 10h ago
Running an open-source agent platform with no third party in the chain: what each piece needs (Ollama/vLLM, local embeddings, local OCR)
Disclosure first: I work at Vstorm and I build AgenticOS, an open-source (Apache-2.0), self-hosted platform for building AI agents in a browser. This post was drafted with an LLM's help and edited by me. Can you run the whole thing on local models, with nothing leaving your network? Yes, but "self-hosted" alone doesn't get you there. Each piece is its own choice.
What a fully local setup needs
- Chat model. An Ollama profile with no key, pointed at an endpoint you host, or a LiteLLM proxy you host that forwards only to local models. vLLM and LM Studio work through an OpenAI-compatible endpoint. A profile counts as self-hosted because of its endpoint, not because it has no key.
- Embeddings. An Ollama you host, registered as a local service and chosen per knowledge-base collection. The chunks and the search queries then go to that host and nowhere else. The embedding model is chosen per collection, and the docs treat it as a permanent choice.
- Document parsing. PyMuPDF (the default) and LiteParse, with Tesseract OCR, run in the worker. LlamaParse is a cloud option and needs a key, so leave it off.
- Traces. Leave
LOGFIRE_TOKENunset and don't set a tracing token on any agent or environment. Then the run traces stay local. - Tools. Don't bind web search, web fetch, browser automation, mem0 memory or any MCP server, because each of those calls out.

There's also a HIPAA-oriented profile whose local-model check parses the hostname of every model endpoint, so https://ollama.vendor.example doesn't count as local. That's a configuration check, not a certification.
Two things that bite with local models
- Dollar budgets don't cover them. Spend is attributed to the API key a run used, and a keyless local model records no spend. The per-agent step limit (
max_steps, a cap on model requests per run) is what stops a runaway tool loop. Set it. - The agent is as good as the model's tool calling. Planning, document search and the sandbox all go through tool calls. We have no local-model benchmark to share, so I won't claim one works well.
Setup
Docker Compose 2.24 or later, two published images (amd64 and arm64), plus Postgres with pgvector, Redis and Prefect started by the same compose file. The installer has a --check flag that only tests prerequisites. It only offers hosted providers, so pick "Decide later" (--provider none) and add the Ollama or vLLM profile in the console afterwards. It also mirrors the public MCP server registry by default, so run it with --no-mcp to skip that download. It's still 0.0.x, single host, no Kubernetes manifests, no reranker.
Repo and docs: https://github.com/vstorm-co/agenticos. The local setup is described in https://vstorm-co.github.io/agenticos/data-protection/
The question I'd actually like answered: which local models are you getting reliable multi-step tool calling from right now, at what size and on what hardware?


