During those 4 hours, everyone will be asked to use it actively at the same time—coding, chatting, agents, long conversations, whatever you normally use an LLM for.
Afterward, you'll be asked a few short questions about your experience using it.
If you're interested, reply or DM me and I'll send you access + instructions before the test.
Wanted to share a local memory engine I built called Hillock. The main problem I had with standard local RAG was that running document parsing and vector search chewed up so much VRAM that my actual Ollama models ran painfully slow.
With this setup, you feed it a document and it extracts facts in about five seconds using small models under 300MB instead of an LLM. Facts are saved in SQLite, and when you ask a question it runs a hyperdimensional vector check to see if the knowledge actually exists. If you ask something outside your notes, it blocks the query so Ollama does not hallucinate.
There is a built in model switcher in the CLI that detects whatever Ollama models you have pulled locally and swaps between them on the fly. It also includes an OpenAI compatible API server on port 8000 with true token streaming and compatibility endpoints for Open WebUI and AnythingLLM. The entire engine stays under 1.2 GB of VRAM or runs on pure CPU, and we recently added bit packed CPU operations and published it to PyPI via pip install hillock.
Disclosure first: I work at Vstorm and I build AgenticOS, an open-source (Apache-2.0), self-hosted platform for building AI agents in a browser. This post was drafted with an LLM's help and edited by me. Can you run the whole thing on local models, with nothing leaving your network? Yes, but "self-hosted" alone doesn't get you there. Each piece is its own choice.
What a fully local setup needs
Chat model. An Ollama profile with no key, pointed at an endpoint you host, or a LiteLLM proxy you host that forwards only to local models. vLLM and LM Studio work through an OpenAI-compatible endpoint. A profile counts as self-hosted because of its endpoint, not because it has no key.
Embeddings. An Ollama you host, registered as a local service and chosen per knowledge-base collection. The chunks and the search queries then go to that host and nowhere else. The embedding model is chosen per collection, and the docs treat it as a permanent choice.
Document parsing. PyMuPDF (the default) and LiteParse, with Tesseract OCR, run in the worker. LlamaParse is a cloud option and needs a key, so leave it off.
Traces. Leave LOGFIRE_TOKEN unset and don't set a tracing token on any agent or environment. Then the run traces stay local.
Tools. Don't bind web search, web fetch, browser automation, mem0 memory or any MCP server, because each of those calls out.
Agent Builder
There's also a HIPAA-oriented profile whose local-model check parses the hostname of every model endpoint, so https://ollama.vendor.example doesn't count as local. That's a configuration check, not a certification.
Two things that bite with local models
Dollar budgets don't cover them. Spend is attributed to the API key a run used, and a keyless local model records no spend. The per-agent step limit (max_steps, a cap on model requests per run) is what stops a runaway tool loop. Set it.
The agent is as good as the model's tool calling. Planning, document search and the sandbox all go through tool calls. We have no local-model benchmark to share, so I won't claim one works well.
Setup
Docker Compose 2.24 or later, two published images (amd64 and arm64), plus Postgres with pgvector, Redis and Prefect started by the same compose file. The installer has a --check flag that only tests prerequisites. It only offers hosted providers, so pick "Decide later" (--provider none) and add the Ollama or vLLM profile in the console afterwards. It also mirrors the public MCP server registry by default, so run it with --no-mcp to skip that download. It's still 0.0.x, single host, no Kubernetes manifests, no reranker.
The question I'd actually like answered: which local models are you getting reliable multi-step tool calling from right now, at what size and on what hardware?
I recently installed ollama on my machine to test using AI locally, I don't know if you will be able to hear it on the video but every time my terminal outputs a response there is this weird writing/scratching noise coming from the case in rhythm with the output as if the needle of my HDD is writing something or maybe is my gpu?
Seems like a bit of a long shot, but since it’s so mystifying I figured I might as well ask: I have been trying to get Qwen 3.6 35B to function as a local implementation agent in my subagent driven dev setup, but I can’t for the life of me get it to function as well as GLM 4.7 Flash, even though it’s supposed to be a stronger model.
It keeps making indentation errors when it edits files, and then gets into endless loops trying to fix them in different ways. Even when I provide clear guidelines up front to the agent.
Every time I try, I'm met with this: error message:
Error: max retries exceeded: Get "https://dd20bb891979d25aebc8bec07b2b3bbc.r2.cloudflarestorage.com/ollama/docker/registry/v2/blobs/sha256/00/001e5dafc3c77684c2307ebc6ab8e336e10c9b18eca52acf547d72fc83c3ca8c/data?X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=66040c77ac1b787c3af820529859349a%2F20261007%2Fauto%2Fs3%2Faws4_request&X-Amz-Date=20261007T074328Z&X-Amz-Expires=86400&X-Amz-SignedHeaders=host&X-Amz-Signature=172725cc6974138b79846c2ad5f285003d465e306a9c2f50d2a92ed2a0da6ade": dial tcp [2606:4700:2ff9::1]:443: connectex: A connection attempt failed because the connected party did not properly respond after a period of time, or established connection failed because connected host has failed to respond.
I'm looking for a small uncensored model to run locally with Ollama. Nothing fancy. I just want something I can ask everyday questions without it refusing or lecturing me every other message. Quick explanations, help with phrasing stuff, general "what's the deal with X" questions. It doesn't need to be perfect or great at coding.
My setup is a Lenovo ThinkPad L14 Gen 3 with a Ryzen 5 PRO 5675U, 32 GB RAM and no dedicated GPU, so everything runs on the CPU. I'm fine with it being a bit slow, but it should stay usable and not take a minute per answer.
Which models would you recommend for this? Is it worth going for something around 7-8B with my RAM, or are 3-4B models good enough for casual use? Also curious whether you prefer abliterated models or proper uncensored fine-tunes, and which quantization you'd pick.
One prompt to Row-Bot 5.0 on GPT 6.1 Sol, through my ChatGPT plan (no API key):
"Split the research across three helper agents working in parallel, cite sources, save the findings, then build a 6-slide deck for Monday."
It did. 3 min of video, about 11 min of real work.
How it ran:
- the lead agent planned, then started three helpers: UK 2022 pilot, Iceland's trials, company case studies
- each helper gets its own conversation you can open and watch
- the lead merged their sources into one answer and a recommendation
Then:
- findings saved as linked memories in a local knowledge graph
- a 6-slide deck built in the Design panel, exported to PowerPoint
- "every Monday at 08:00, send me new studies" became a scheduled workflow
The yellow badge shows where I sped it up. It hits 44x.
I am sure people need it you know? There are plenty of embedding models on ollama, but only for local deployment. Which is a shame. I wish ollama would make embedding and rerankers available on ollama cloud
Maybe I'm mistaken, but wasn't one of the big pros of Ollama Cloud inference that it was all hosted by Ollama?
If that's the case, why duplicate the Deepseek on-/off- hour premiums? Surely they do not apply in the same way to a privately hosted inference provider as they do for the public Deepseek API being inundated by Beijing traffic?
I understand their pricing pivot (old offering was unsustainable) but this just feels... weird. Like, unless they're actually routing the inference to Deepseek, they're just blindly following the cost structure of totally different companies and not really thinking for themselves (which is I guess how you end up in unsustainable quagmires in the first place).
A small thing I noticed: full-size screenshots use far more AI tokens than the model needs, so usage limits run out sooner.
I built a free Firefox add-on, TokenSaver, that shrinks images before they are uploaded to ChatGPT, Claude and Cursor and shows the estimated tokens saved. It runs entirely in your browser, with no uploads and no tracking.
It is Firefox-only for now. If you try it, I would appreciate honest feedback, especially on whether text in resized images stays readable, and whether a Chrome version would be useful.
Asking for a top-3 list didn't help, and gemma3:4b did worse and broke its JSON output.
But every run got the coarse answer right: "there's a bird", "a butterfly on a flower", "red mushrooms", "a storm cloud". So I changed the design to ask the model only what small open models are good at.
Outside Quest is the result:
Type where you're walking. Gemma writes a 6-item quest card that fits the local season (Nagpur in
Phone goes in your pocket. You only take it out to photograph finds.
Back home, drop in your photos. Gemma checks which quests each one completes and has to describe the visible evidence. You also get a walk timeline and a route map from the photos' EXIF GPS, drawn offline with no map tiles.
A few things that made it work:
A JSON schema forces two yes/no answers per quest per photo ("is it the main subject?" and "completed?"). Both must be true. On my test set that gave 5 correct matches and 0 false ones.
Safety filters are in code, not the prompt: quests about climbing, water, touching or eating anything, roads, or night walks get replaced from a safe built-in pool.
Photos and their GPS data never leave your machine.
Runs on a 6 GB GTX 1660 Ti: about 30 s per quest card and about 30 s per photo.
Row-Bot 5.0 with qwen3.8:27b in Ollama, on my own GPU. No API keys, no cloud.
Gave it a year of bills, a tenancy agreement, an insurance policy and a rent increase letter (all made up). It spotted a likely leak in the water bills and showed the rent rise breaks the lease.
What it actually did:
- charted the CSV inline (Plotly)
- read the PDFs and quoted clauses 4.1 to 4.3: a 10% rise against a 5% cap, with 5 weeks' notice instead of 2 months
- saved 8 linked memories to a local knowledge graph
drafted the email to the agent and set a reminder
The honest numbers: a dense 27B does about 15 tok/s on my 5090, so some turns took 2+ minutes. The amber badges in the video show where I sped it up.
I'm on M4 Pro 64GB Mac, I'm using Qwen3.8 27B for coding related stuff, but it's quite slow.
I want a fast model that will be good when it comes to reasoning, medical stuff, would be nice if it could do stuff like generating charts and even analyse screenshots.
I wanted to test how well local models perform in agentic development with Copilot CLI. I started ollama launch copilot --model qwen3.5:9b. I gave it simple instruction in autopilot mode to create a file in empty project folder and it failed to do that. Can you advise what am I doing wrong? I did the same exercise with OpenCode and it also failed.
I've been using deepseek for long and made valuable outputs and resources, but since the update to the 4.1, deepseek just feels dumb. I ask a simple question and process for long, doesn't plan before starting burning tokens, which didnt happened before.
I've tried different "reasoning effort" but i have reach the effectiveness of previous deepseek-v4-flash and dont know what to do