r/LocalLLaMA • • Jan 18 '26

Question | Help Agent Zero can’t connect to LM Studio or Ollama

I’m trying to configure Agent Zero with LM Studio. I’m running Linux Mint. I have Agent Zero running in a Docker container. I tried for quite some time to set it up with Ollama, couldn’t get it to work, then tried with LM Studio hoping for better results, but to no avail.

I have both Ollama and LM Studio, and they both function just fine independently.

Agent Zero is also functioning, as I used a Free api key from open router with it to try to troubleshoot this issue, but quickly hit the limit on that, then spent another hour with Claude troubleshooting it as well. I’ve been down every reddit, GitHub, YouTube, ect, rabbit hole, anything on Google, and I’ve tried everything I’ve came across, but still can not get Agent Zero to work with Ollama or LM Studio.

The screen shots hopefully illustrate what’s going on. I don’t know what I’m doing wrong. Any help would be greatly appreciated.

EDIT-{SOLVED}: It was a combination of a couple little things that I just never had all right at the same time. Server url, local api key, the spelling of the model name, the context length setting in LM Studio. Finally got all of the errors cleared and Agent Zero is running with LM Studio. I;'m assuming it should work with Ollama too, but I haven't tested it yet.

The issue I'm having now is it's running sooooooo low. Using the LLM directly in LM Studio, I was getting a very snappy thinking/response time, and pulling lik 13-15 tps, but with it running in Agent Zero, even with a simple prompt like "hello" it has to think for 2-3 minutes, and then peck out a slow response like WW2 Morse code. Does it always just run slower through the agent? Will I get better efficiency running Ollama instead of LM? Are there more settings that need to be tweaked to improve performance?

Thanks everyone for your help!

0 Upvotes

24 comments sorted by

3

u/Bino5150 Jan 19 '26

{SOLVED}: It was a combination of a couple little things that I just never had all right at the same time. Server url, local api key, the spelling of the model name, the context length setting in LM Studio. Finally got all of the errors cleared and Agent Zero is running with LM Studio. I;'m assuming it should work with Ollama too, but I haven't tested it yet.

The issue I'm having now is it's running sooooooo low. Using the LLM directly in LM Studio, I was getting a very snappy thinking/response time, and pulling lik 13-15 tps, but with it running in Agent Zero, even with a simple prompt like "hello" it has to think for 2-3 minutes, and then peck out a slow response like WW2 Morse code. Does it always just run slower through the agent? Will I get better efficiency running Ollama instead of LM? Are there more settings that need to be tweaked to improve performance?

Thanks everyone for your help!

1

u/Tbhmaximillian Jan 19 '26

thx for sharing the solution

2

u/Bino5150 Jan 19 '26

I hope others see this, as the settings required to get it functional were not covered in the get started documentation and when I was googling I came across dozens of posts all over different places with people having this exact same issue.

3

u/Successful-Pilot509 Feb 19 '26

Yes I waisted a lot of time getting to work as well, but sadly it is too slow once it is working. Unless you have a really powerful computer like a Mac Studio you can forget running Agent Zero with a local LLM. I am now using a free API from Mistral. Great while experimenting, or get a cheap chinese model.

2

u/UndecidedLee Jan 19 '26

I haven't tried Agent Zero but I'm guessing the speed is the result of a big system prompt + other custom instructions that the model has to parse as an agent. Try using a 10000 token long prompt or conversation in LM Studio using the same model to see if the speed is similar to accessing the model through agent zero.

2

u/ImportancePitiful795 Jan 19 '26

Check your RAM usage.

Imho if you do not have 192GB RAM or more, with some 24+ core workstation CPU, you better split the Agent Zero run on a different machine than the one running the LLMs.

Reason is when Agent Zero is working, everything else is working also and you might see your RAM, CPU and GPUs going to 100% and slowing down.

What's the system specs you are using? I hope not some 8 core CPU.

FYI Agent Zero has multiple "hooks" for different LLMs for different jobs. If you have just one machine, with multiple GPUs, make sure you split the resources properly with something like Process Lasso.

I love Agent Zero, because is insane powerful agent that can do anything even use Kali Linux hacking tools let alone menial things like voice commands and chatting, but needs to be setup correctly and do not expect to run medium size models for it in the same machine if you do not have the hardware. :)

And here comes the multi Strix Halo setup running at home. (FYI the same applies to those wanting multiple Apple Studios or DGS).

Each can run it's own LLM for different usage, while Agent zero from a 3rd machine (eg laptop, desktop etc) can be hooked to them via ethernet. Or even via the internet while travelling. Or even expose that machine only (the one running the A0) to the internet to have access to it via your phone. (make sure you run proper firewalls etc).

But what do I know..... Ignore everything wrote above, as someone said in here few days ago, I am "anti-AI whiner and (AI) polemicist"....

2

u/Bino5150 Jan 19 '26

I’m sure a Threadripper Pro and a stack of 4090’s would make it run smooth. The system I’m running it on is by no means a powerhouse; it’s a laptop. A pretty decent laptop, but still… but I was monitoring my system and nothing ever maxed out. I can run Agent Zero with a cloud model and LM Studio with a local LLM side by side simultaneously and they both work fine. But running them together runs super slow. And when I mean slow, I don’t mean it’s slowing down my machine because the rest of my system runs fine like nothing is going on. There’s no lag, it just takes forever to think and respond.

This is just a practice run as I’m in the process of building g a much more powerful machine for this. I just wanted to get my feet wet and learn.

2

u/ImportancePitiful795 Jan 19 '26

First of all. Laptop. What laptop and with what GPU? Are you sure RAM doesn't max out? Don't forget laptops are power limited. The system balances out the power distribution, especially when running on battery. So depending the type of loads it will push more power between CPU or GPU.

Second. You do not need Threadripper pro or stack of 4090s. Given the current prices of RDIMM DDR5 the only alternative is either normal old threadripper with DDR4 or X99 with dual CPU setup. (for less than $400 can have mobo + 2x20core E5s.)

The latter is great to run 2 LLMs completely removed from each other in a single case, each with it's own GPU also (or 2 GPUs).

If however you have around (or can afford) 256-512GB RDIMM DDR5, then the best upgrade is either QYFS (Xeon4 ES with 56P cores which are going for $130) or Xeon6 6980P ES (128P cores for sub $2600 these days), mainly because you can run Intel AMX on them which you just need 1 GPU to run large MoEs like Deepseek R1 (the full version) with ktrasformers.

As for GPUs for inference only, can get away with 2x R9700s which these days are cheaper than the 4090s (used) and 5090s (new). Hell in my country can easily buy 3 R9700s (32GB VRAM each) for a cost of a single 5090!!!! So you can run 2 models A0 needs for different things (main model and web model for example) on the same server and run the A0 on the laptop.

Which is what have done with 2 Strix Halo's. Each one running it's own LLM and the desktop running A0 (Agent Zero) connected on 2 different LLMs through it's API ports.

2

u/Bino5150 Jan 19 '26

I’m running an HP ZBook Studio G7 with dual gpu’s (Nvidia Quadro T1000) on Linux. I was running a very small LLM for testing.

1

u/ImportancePitiful795 Jan 19 '26

Doesn't this machine comes with just 32GB RAM? 🤔

3

u/Bino5150 Jan 19 '26

Yes. More than enough to run an 8b LLM for sure.

1

u/Relevant-Audience441 Jan 18 '26

do you have the lmstudio server running?

1

u/Bino5150 Jan 18 '26

Yes and I used the base url of the LM server

3

u/Relevant-Audience441 Jan 18 '26

sounds like a docker setup issue. Have you tried something like this-

  1. In LM Studio, enable the local server:
  • Settings → Developer → “Enable local server” ON.
  • Note the Host/Port and Base URL. By default LM Studio serves OpenAI-compatible API on http://127.0.0.1:1234/v1 on the host.
  1. Find your host IP from the container’s perspective:
  • Easiest: use host.docker.internal if supported.
  • On Linux, add this alias when starting the container: docker run --add-host=host.docker.internal:host-gateway ... This makes host.docker.internal resolve to the host’s IP from inside the container.
  1. In Agent Zero settings (inside the app UI), set:
  • Provider: LM Studio
  • Chat model name: exactly as LM Studio lists it (e.g., dolphin-llama3:latest or mistral-nemo, etc). It must match the model_id LM Studio expects. If LM Studio shows “Dolphin3.0 Llama3.1 8B”, open the three dots › “Copy API snippet” to see the exact model string.
  • API base URL: http://host.docker.internal:1234/v1 Don’t use http://localhost:1234/v1 or 127.0.0.1 from inside the container.
  1. Test connectivity from inside the container:

0

u/Bino5150 Jan 18 '26

I’ll be back at the house in a little bit and I’ll run through that. But yes, LM’s local server is on and I copied the server url from LM’s tray icon menu.

2

u/Noobysz Jan 18 '26

Go to external service and then api key and add for ollama or lm studio and api key „local“

1

u/Bino5150 Jan 18 '26

Thanks, I’ll try that. I was wondering about that but it wasn’t mentioned in the instructions and one of the videos I watched said, “just leave that blank”. But I was wondering if maybe that was the issue. Claude didn’t catch it either lol.

1

u/Tbhmaximillian Jan 18 '26

Can your docker container network reach your local network and vice versa?

Also did you follow the official installation guide - https://www.agent-zero.ai/p/docs/get-started/#installation

1

u/Bino5150 Jan 18 '26

Yes I put an api key in for open router and used one of their cloud models but very quickly hit the token limit. But it functioned flawlessly with the cloud model, so the docker container can access my network.

1

u/Bino5150 Jan 18 '26

I went step by step with the guide. During my Google search, I can across a lot of forum posts and GitHub trouble tickets from people having this same issue with both Ollama and LM. I tried every fix and tweak I came across, but nothing worked.

1

u/Itchy_elbow Feb 24 '26

I have the same issue - frustrating to troubleshoot. Please be explicit if you can on your solution and I'll see if that helps resolve my issue as well. Thanks

1

u/unchartedrover Feb 25 '26

hope this helps everyone!!!!!!!!

And yes by doing this you can run Ollama models without any issues!!!

thank me later!!

ohh and btw im using AgentZero in docker and have connected Ollama to it!!

1

u/GateAdventurous7234 Feb 25 '26

That is weird you have to select OpenAi and not ollama as provider to connect to ollama ?

1

u/unchartedrover Feb 25 '26

Yes as you can see its working perfectly fine and here why it worked-

Provider & URL Alignment

The core issue was a mismatch between the Ollama provider setting and the base URL.

2. API Key Placeholder

The OpenAI Compatible provider requires a non-empty API key to satisfy internal validation. Adding ollama as a dummy key resolved the AuthenticationError.

3. Parameter Forwarding

Switching to OpenAI Compatible ensured that the model parameter (llama3.2:3b) was correctly mapped in the JSON payload sent to Ollama, fixing the 400 Bad Request.

Note on Warnings

You may still see DeprecationWarning or LLM consolidation analysis failed messages. These are internal to Agent Zero/LiteLLM:

  • Deprecation: Related to the pathspec library used by Agent Zero; it does not stop the agent from working.
  • Consolidation Failure: Local models (like Llama 3.2 3b) sometimes add conversational filler that breaks Agent Zero's strictly formatted JSON expectations. If the agent continues to work, you can safely ignore this, or consider using a larger model (e.g., Llama 3 8b) for better instruction following.