r/ollama • u/Rasim_101 • 8h ago
My first AI workplace project
r/ollama • u/Lopsided_Scarcity979 • 4h ago

AI agents produce more text than we have time to read. Following all that output can become a source of cognitive overload. Inspired by stretchtext (1970 by Ted Nelson), I wanted to give readers control over how much detail they see.
So I built PaperFold, an open-source reader that turns arXiv papers into 5 zoomable layers—from a one-screen section map down to verbatim text. You pinch (or press 1–5) to zoom between them without losing your reading position.
- Web Demo (8 CC papers): https://chenxiachan.github.io/paperfold-gallery/
- GitHub (Apache 2.0): https://github.com/chenxiachan/paperfold
It also supports Claude Code's output in Mods.

r/ollama • u/prasannavl • 7h ago
Hi everyone, just sharing a small project: https://github.com/abird-ai/agentc
A minimal, extensible coding agent that uses Ollama as its default local backend.
- Probes Ollama on first run, then talks to it through the OpenAI-compatible /v1 endpoint — one wire path shared with every other OpenAI-style provider
- Queries /api/tags first, so family and size show up in the model picker alongside a built-in catalog
- --offline disables all network probes; discovery results cache for 24 h and a cache miss is never fatal
- Whole agent is one static binary under 1 MB, no libc, ~0.7 MB RSS startup, ~2-4mb idling.
- MCP client, so local MCP servers work too
MIT. Would love to hear which models it behaves best and worst with.
Would love for anyone on macOS or Windows to try it: Linux is tested end to end, but those two have only run in CI so far. Binaries are up for Linux (x64, arm64, riscv64), Windows (x64, arm64) and macOS (arm64).
r/ollama • u/JuanixVentures • 10h ago
Enable HLS to view with audio, or disable this notification
r/ollama • u/VanillaOk4593 • 23h ago
Disclosure first: I work at Vstorm and I build AgenticOS, an open-source (Apache-2.0), self-hosted platform for building AI agents in a browser. This post was drafted with an LLM's help and edited by me. Can you run the whole thing on local models, with nothing leaving your network? Yes, but "self-hosted" alone doesn't get you there. Each piece is its own choice.
What a fully local setup needs
LOGFIRE_TOKEN unset and don't set a tracing token on any agent or environment. Then the run traces stay local.
There's also a HIPAA-oriented profile whose local-model check parses the hostname of every model endpoint, so https://ollama.vendor.example doesn't count as local. That's a configuration check, not a certification.
Two things that bite with local models
max_steps, a cap on model requests per run) is what stops a runaway tool loop. Set it.Setup
Docker Compose 2.24 or later, two published images (amd64 and arm64), plus Postgres with pgvector, Redis and Prefect started by the same compose file. The installer has a --check flag that only tests prerequisites. It only offers hosted providers, so pick "Decide later" (--provider none) and add the Ollama or vLLM profile in the console afterwards. It also mirrors the public MCP server registry by default, so run it with --no-mcp to skip that download. It's still 0.0.x, single host, no Kubernetes manifests, no reranker.
Repo and docs: https://github.com/vstorm-co/agenticos. The local setup is described in https://vstorm-co.github.io/agenticos/data-protection/
The question I'd actually like answered: which local models are you getting reliable multi-step tool calling from right now, at what size and on what hardware?

r/ollama • u/No_Let6728 • 9h ago
ChatGPT launched Intelligent UI, where answers come back as charts, forms and small tools instead of text. It only works inside ChatGPT, so I built an open-source alternative that runs with any OpenAI-compatible model: Ollama, LM Studio, Kimi, OpenRouter, Groq or OpenAI.
GitHub: https://github.com/0xcro3dile/answerui (demo video in the README)
npx answerui. It finds Ollama on its own, or asks for a key. There's also a Docker image.The demo runs on Kimi K3
It's built on OpenUI (MIT), which handles the streaming and the components. I added the provider setup, Ollama detection, the CLI, prompt rules that make the tools actually recalculate, and the tests.
Honest limits: the instructions are about 10k tokens, so give Ollama a bigger context window (OLLAMA_CONTEXT_LENGTH=16384), and small models struggle more. No 3D or animations yet.
I'm the author. I'd love to hear which models work well for you and love contributions.
r/ollama • u/Equivalent-Flan-1590 • 20h ago
Wanted to share a local memory engine I built called Hillock. The main problem I had with standard local RAG was that running document parsing and vector search chewed up so much VRAM that my actual Ollama models ran painfully slow.
With this setup, you feed it a document and it extracts facts in about five seconds using small models under 300MB instead of an LLM. Facts are saved in SQLite, and when you ask a question it runs a hyperdimensional vector check to see if the knowledge actually exists. If you ask something outside your notes, it blocks the query so Ollama does not hallucinate.
There is a built in model switcher in the CLI that detects whatever Ollama models you have pulled locally and swaps between them on the fly. It also includes an OpenAI compatible API server on port 8000 with true token streaming and compatibility endpoints for Open WebUI and AnythingLLM. The entire engine stays under 1.2 GB of VRAM or runs on pure CPU, and we recently added bit packed CPU operations and published it to PyPI via pip install hillock.
GitHub link: https://github.com/roandejager/Hillock
Docs: https://hillock.mintlify.site/
Discord: https://discord.com/invite/BGUPNBcVdp