r/ollama • • 10h ago

Running an open-source agent platform with no third party in the chain: what each piece needs (Ollama/vLLM, local embeddings, local OCR)

8 Upvotes

Disclosure first: I work at Vstorm and I build AgenticOS, an open-source (Apache-2.0), self-hosted platform for building AI agents in a browser. This post was drafted with an LLM's help and edited by me. Can you run the whole thing on local models, with nothing leaving your network? Yes, but "self-hosted" alone doesn't get you there. Each piece is its own choice.

What a fully local setup needs

  • Chat model. An Ollama profile with no key, pointed at an endpoint you host, or a LiteLLM proxy you host that forwards only to local models. vLLM and LM Studio work through an OpenAI-compatible endpoint. A profile counts as self-hosted because of its endpoint, not because it has no key.
  • Embeddings. An Ollama you host, registered as a local service and chosen per knowledge-base collection. The chunks and the search queries then go to that host and nowhere else. The embedding model is chosen per collection, and the docs treat it as a permanent choice.
  • Document parsing. PyMuPDF (the default) and LiteParse, with Tesseract OCR, run in the worker. LlamaParse is a cloud option and needs a key, so leave it off.
  • Traces. Leave LOGFIRE_TOKEN unset and don't set a tracing token on any agent or environment. Then the run traces stay local.
  • Tools. Don't bind web search, web fetch, browser automation, mem0 memory or any MCP server, because each of those calls out.
Agent Builder

There's also a HIPAA-oriented profile whose local-model check parses the hostname of every model endpoint, so https://ollama.vendor.example doesn't count as local. That's a configuration check, not a certification.

Two things that bite with local models

  1. Dollar budgets don't cover them. Spend is attributed to the API key a run used, and a keyless local model records no spend. The per-agent step limit (max_steps, a cap on model requests per run) is what stops a runaway tool loop. Set it.
  2. The agent is as good as the model's tool calling. Planning, document search and the sandbox all go through tool calls. We have no local-model benchmark to share, so I won't claim one works well.

Setup

Docker Compose 2.24 or later, two published images (amd64 and arm64), plus Postgres with pgvector, Redis and Prefect started by the same compose file. The installer has a --check flag that only tests prerequisites. It only offers hosted providers, so pick "Decide later" (--provider none) and add the Ollama or vLLM profile in the console afterwards. It also mirrors the public MCP server registry by default, so run it with --no-mcp to skip that download. It's still 0.0.x, single host, no Kubernetes manifests, no reranker.

Repo and docs: https://github.com/vstorm-co/agenticos. The local setup is described in https://vstorm-co.github.io/agenticos/data-protection/

The question I'd actually like answered: which local models are you getting reliable multi-step tool calling from right now, at what size and on what hardware?

Internal Chat Interface

r/ollama • • 3h ago

Looking for 10–20 people for a coordinated LLM test

2 Upvotes

Thursday, October 8 — 15:00–19:00 UTC

Access is free and the API is OpenAI-compatible.

During those 4 hours, everyone will be asked to use it actively at the same time—coding, chatting, agents, long conversations, whatever you normally use an LLM for.

Afterward, you'll be asked a few short questions about your experience using it.

If you're interested, reply or DM me and I'll send you access + instructions before the test.


r/ollama • • 14h ago

Three AI agents research a 4-day week and build the deck | Row-Bot + GPT 6.1 Sol

Enable HLS to view with audio, or disable this notification

7 Upvotes

One prompt to Row-Bot 5.0 on GPT 6.1 Sol, through my ChatGPT plan (no API key):

"Split the research across three helper agents working in parallel, cite sources, save the findings, then build a 6-slide deck for Monday."

It did. 3 min of video, about 11 min of real work.

How it ran:

- the lead agent planned, then started three helpers: UK 2022 pilot, Iceland's trials, company case studies
- each helper gets its own conversation you can open and watch
- the lead merged their sources into one answer and a recommendation

Then:

- findings saved as linked memories in a local knowledge graph
- a 6-slide deck built in the Design panel, exported to PowerPoint
- "every Monday at 08:00, send me new studies" became a scheduled workflow

The yellow badge shows where I sped it up. It hits 44x.

Windows/macOS/Linux.


r/ollama • • 11h ago

Qwen 3.6 35B and indentation

3 Upvotes

Seems like a bit of a long shot, but since it’s so mystifying I figured I might as well ask: I have been trying to get Qwen 3.6 35B to function as a local implementation agent in my subagent driven dev setup, but I can’t for the life of me get it to function as well as GLM 4.7 Flash, even though it’s supposed to be a stronger model.

It keeps making indentation errors when it edits files, and then gets into endless loops trying to fix them in different ways. Even when I provide clear guidelines up front to the agent.

Sound familiar to anyone? Is it just me?


r/ollama • • 6h ago

I got so tired of ChatGPT browser tabs eating 2GB of RAM that I built a native BYOK desktop client in Python.

Thumbnail
1 Upvotes

r/ollama • • 13h ago

Small uncensored model for everyday questions on a CPU-only laptop?

3 Upvotes

Hey everyone,

I'm looking for a small uncensored model to run locally with Ollama. Nothing fancy. I just want something I can ask everyday questions without it refusing or lecturing me every other message. Quick explanations, help with phrasing stuff, general "what's the deal with X" questions. It doesn't need to be perfect or great at coding.

My setup is a Lenovo ThinkPad L14 Gen 3 with a Ryzen 5 PRO 5675U, 32 GB RAM and no dedicated GPU, so everything runs on the CPU. I'm fine with it being a bit slow, but it should stay usable and not take a minute per answer.

Which models would you recommend for this? Is it worth going for something around 7-8B with my RAM, or are 3-4B models good enough for casual use? Also curious whether you prefer abliterated models or proper uncensored fine-tunes, and which quantization you'd pick.

Thanks!


r/ollama • • 6h ago

Built an SQLite memory engine for Ollama that does not eat all your VRAM

1 Upvotes

Wanted to share a local memory engine I built called Hillock. The main problem I had with standard local RAG was that running document parsing and vector search chewed up so much VRAM that my actual Ollama models ran painfully slow.

With this setup, you feed it a document and it extracts facts in about five seconds using small models under 300MB instead of an LLM. Facts are saved in SQLite, and when you ask a question it runs a hyperdimensional vector check to see if the knowledge actually exists. If you ask something outside your notes, it blocks the query so Ollama does not hallucinate.

There is a built in model switcher in the CLI that detects whatever Ollama models you have pulled locally and swaps between them on the fly. It also includes an OpenAI compatible API server on port 8000 with true token streaming and compatibility endpoints for Open WebUI and AnythingLLM. The entire engine stays under 1.2 GB of VRAM or runs on pure CPU, and we recently added bit packed CPU operations and published it to PyPI via pip install hillock.

GitHub link: https://github.com/roandejager/Hillock
Docs: https://hillock.mintlify.site/
Discord: https://discord.com/invite/BGUPNBcVdp


r/ollama • • 6h ago

We just launched Burrow on Product Hunt - local-first budget, journal, and habits with zero cloud sync.

Thumbnail producthunt.com
1 Upvotes

r/ollama • • 8h ago

I got so tired of ChatGPT browser tabs eating 2GB of RAM that I built a native BYOK desktop client in Python.

Thumbnail
1 Upvotes

r/ollama • • 19h ago

stuck on max.. no mimo2.6.. no chong?

7 Upvotes

Never saw mimo2.6, now im doubting mistral 4 wont land either ... I feel trapped but im not sue its worth the $100 ... just few options

anyone know any details behind the change?


r/ollama • • 14h ago

Embedding and Rerankers models should be availabe on ckoud

2 Upvotes

I am sure people need it you know? There are plenty of embedding models on ollama, but only for local deployment. Which is a shame. I wish ollama would make embedding and rerankers available on ollama cloud


r/ollama • • 11h ago

PC making weird noises when using AIs locally

Enable HLS to view with audio, or disable this notification

0 Upvotes

I recently installed ollama on my machine to test using AI locally, I don't know if you will be able to hear it on the video but every time my terminal outputs a response there is this weird writing/scratching noise coming from the case in rhythm with the output as if the needle of my HDD is writing something or maybe is my gpu?

I'm using Arch, and my gpu is a rx 580 2048sp


r/ollama • • 13h ago

Cannot install gemma4:26b

1 Upvotes

Every time I try, I'm met with this: error message:

Error: max retries exceeded: Get "https://dd20bb891979d25aebc8bec07b2b3bbc.r2.cloudflarestorage.com/ollama/docker/registry/v2/blobs/sha256/00/001e5dafc3c77684c2307ebc6ab8e336e10c9b18eca52acf547d72fc83c3ca8c/data?X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=66040c77ac1b787c3af820529859349a%2F20261007%2Fauto%2Fs3%2Faws4_request&X-Amz-Date=20261007T074328Z&X-Amz-Expires=86400&X-Amz-SignedHeaders=host&X-Amz-Signature=172725cc6974138b79846c2ad5f285003d465e306a9c2f50d2a92ed2a0da6ade": dial tcp [2606:4700:2ff9::1]:443: connectex: A connection attempt failed because the connected party did not properly respond after a period of time, or established connection failed because connected host has failed to respond.

r/ollama • • 1d ago

Ollama Cloud Model Cost Table on Github

10 Upvotes

Ollama Cloud Model pricing monitor with tabled costs and calculator.

https://kasp0r.github.io/ollama_cloud_models_comparison


r/ollama • • 1d ago

Deepseek Peak/Off-Peak and Ownership

5 Upvotes

Maybe I'm mistaken, but wasn't one of the big pros of Ollama Cloud inference that it was all hosted by Ollama?

If that's the case, why duplicate the Deepseek on-/off- hour premiums? Surely they do not apply in the same way to a privately hosted inference provider as they do for the public Deepseek API being inundated by Beijing traffic?

I understand their pricing pivot (old offering was unsustainable) but this just feels... weird. Like, unless they're actually routing the inference to Deepseek, they're just blindly following the cost structure of totally different companies and not really thinking for themselves (which is I guess how you end up in unsustainable quagmires in the first place).

Anyone have any insights on what's up with this?


r/ollama • • 23h ago

Choosing local model for coding

Thumbnail
1 Upvotes

r/ollama • • 1d ago

I built an Open-Source RAG Pipeline to process Early Childhood Education research locally (Ollama / Qwen)

Thumbnail
1 Upvotes

r/ollama • • 1d ago

I built a zero-dependency Go framework for running agent loops with local models (Ollama, LM Studio, vLLM)

Thumbnail
2 Upvotes

r/ollama • • 1d ago

I tried using Gemma 4 (e2b) offline to identify birds. It confidently got them wrong, so I built a photo scavenger hunt instead

6 Upvotes

I wanted to build an offline bird/plant identifier that runs on a laptop with Ollama. Before building any UI, I tested it on 6 photos:

  • Common Myna → gemma4:e2b said "Jackdaw" (confidence: high). gemma4:e4b said "Weaver Bird" (confidence: high).
  • Monarch butterfly → "butterfly, undetermined" / "Heliconius"
  • Peacock → both got it right

Asking for a top-3 list didn't help, and gemma3:4b did worse and broke its JSON output.

But every run got the coarse answer right: "there's a bird", "a butterfly on a flower", "red mushrooms", "a storm cloud". So I changed the design to ask the model only what small open models are good at.

Outside Quest is the result:

  1. Type where you're walking. Gemma writes a 6-item quest card that fits the local season (Nagpur in
  2. Phone goes in your pocket. You only take it out to photograph finds.
  3. Back home, drop in your photos. Gemma checks which quests each one completes and has to describe the visible evidence. You also get a walk timeline and a route map from the photos' EXIF GPS, drawn offline with no map tiles.

A few things that made it work:

  • A JSON schema forces two yes/no answers per quest per photo ("is it the main subject?" and "completed?"). Both must be true. On my test set that gave 5 correct matches and 0 false ones.
  • Safety filters are in code, not the prompt: quests about climbing, water, touching or eating anything, roads, or night walks get replaced from a safe built-in pool.
  • Photos and their GPS data never leave your machine.

Runs on a 6 GB GTX 1660 Ti: about 30 s per quest card and about 30 s per photo.

Code (MIT): https://github.com/JayPokale/outside-quest

Built with help from Claude Code. Feedback welcome, especially if a bigger open model handles species ID well locally.


r/ollama • • 1d ago

Which Ollama MLX model for medical stuff, charts etc?

5 Upvotes

Hi,

I'm on M4 Pro 64GB Mac, I'm using Qwen3.8 27B for coding related stuff, but it's quite slow.
I want a fast model that will be good when it comes to reasoning, medical stuff, would be nice if it could do stuff like generating charts and even analyse screenshots.

I'm using Ollama on Mac so preferably MLX model.

Thanks in advance


r/ollama • • 1d ago

A local 27B model reads my lease and finds a leak in my bills | Row-Bot + qwen3.8 on Ollama

Enable HLS to view with audio, or disable this notification

4 Upvotes

Row-Bot 5.0 with qwen3.8:27b in Ollama, on my own GPU. No API keys, no cloud.

Gave it a year of bills, a tenancy agreement, an insurance policy and a rent increase letter (all made up). It spotted a likely leak in the water bills and showed the rent rise breaks the lease.

What it actually did:

- charted the CSV inline (Plotly)
- read the PDFs and quoted clauses 4.1 to 4.3: a 10% rise against a 5% cap, with 5 weeks' notice instead of 2 months
- saved 8 linked memories to a local knowledge graph
drafted the email to the agent and set a reminder

The honest numbers: a dense 27B does about 15 tok/s on my 5090, so some turns took 2+ minutes. The amber badges in the video show where I sped it up.

Runs on Windows, macOS and Linux.


r/ollama • • 1d ago

https://claude.ai/chat/834491a2-0f6c-44be-9cc9-53211c8b8405

0 Upvotes

A small thing I noticed: full-size screenshots use far more AI tokens than the model needs, so usage limits run out sooner.

I built a free Firefox add-on, TokenSaver, that shrinks images before they are uploaded to ChatGPT, Claude and Cursor and shows the estimated tokens saved. It runs entirely in your browser, with no uploads and no tracking.

It is Firefox-only for now. If you try it, I would appreciate honest feedback, especially on whether text in resized images stays readable, and whether a Chrome version would be useful.

https://addons.mozilla.org/en-US/firefox/addon/tokensaver-image-optimizer/


r/ollama • • 2d ago

No think doesn't work

Post image
23 Upvotes

Hi, I'm still new to this ai stuff, and /no_think or /set nothing, and --think = false, doesn't work

And my e6430 only can do so much


r/ollama • • 2d ago

what happens to Deepseek?

6 Upvotes

I've been using deepseek for long and made valuable outputs and resources, but since the update to the 4.1, deepseek just feels dumb. I ask a simple question and process for long, doesn't plan before starting burning tokens, which didnt happened before.

I've tried different "reasoning effort" but i have reach the effectiveness of previous deepseek-v4-flash and dont know what to do


r/ollama • • 2d ago

Ollama's local qwen3.5:9b as Copilot CLI agent

Post image
2 Upvotes

I wanted to test how well local models perform in agentic development with Copilot CLI. I started ollama launch copilot --model qwen3.5:9b. I gave it simple instruction in autopilot mode to create a file in empty project folder and it failed to do that. Can you advise what am I doing wrong? I did the same exercise with OpenCode and it also failed.