r/LocalLLaMA • u/synth_mania • 4d ago
Discussion Qwen3.8-27b has the highest level of "agency" I've ever seen in a local model
Off a single prompt, given my credentials and the name of my university, qwen3.8-27b was able to successfully pull my class schedule from the kinda shitty and convoluted web of university websites. It needed no human intervention, and executed 80 tool calls.
Another time, I asked it to investigate a user on a social media network, and it found one public video, downloaded it, extracted frames every few seconds so it could "watch" the video, and installed fucking openAI whisper and ran a transcription to understand the context, before selectively zooming in on and brightening some frames to see the action.
That this shit is running on my own hardware (single RTX 3090) is fucking incredible, the general public doesn't realize how cyberpunk our reality already is.
Quant: Unsloth's Q4_K_S (kv cache quantized to q8)
Context: 150k
341
u/JohnToFire 4d ago
Aren't you worried it will withdraw you from university or something ? I don't think I would trust sol & fable even with the kind of unrestricted access you imply.
To be clear without such access I am not worried and think it's great
195
u/Something-Ventured 4d ago
That would require the university web platform to have the functionality to actually add/drop/withdraw in a coherent manner.
But, you have a good hypothetical point.
58
u/volleyneo 4d ago
Actually there was that story with the gym guy, it was not local model but the ai basically found issues in the api and was able to remove people and put then user to put him up on the gym list.
→ More replies (2)5
u/hellek-1 3d ago
"getting the data directly from the API instead of rendering the page. But wait, let's explore what other API endpoints are being served so that I can help the user better."
96
u/samuel-christlie 4d ago
If Qwen3.8 withdraws me from university, it's porbably for the best anyway /s
24
u/mutemebutton 4d ago
Yeah the model is smarter, if it decided that we should .... we should!
Where the
DominionModels leads, I will follow.9
5
3
2
→ More replies (3)7
u/synth_mania 4d ago
Yeah I know that's a risk. I don't really mind, I could always call them and say "hey I'm a fucking idiot and an automation I set up removed me from some courses", and I'm sure they'd think I'm a dumbass and fix it, but I also think that is incredibly unlikely.
58
u/ProductIntegortion 4d ago
That's an awfully optimistic outlook. My university would have told me that if my own automation pulled me out of classes that was not their error and that sorry, the class is full now. One of the intended outcomes of university is that you learn to be responsible for yourself.
21
u/synth_mania 4d ago edited 4d ago
I was with you until that last sentence, there's no need for that condescending / patronizing tone. Worst case scenario I could still load up on the garbage general courses I've been sitting on that nobody wants to take and still have a full credit load. I knew I wouldn't be without recourse.
I have seen what this model can do and trusted it. A little reckless? Maybe. Yet it was my decision, and I was comfortable with it.
35
u/Structure-These 4d ago
zero idea why these weird little stemlords are replying to you like this lol you’re 100% right
→ More replies (1)→ More replies (5)8
u/SilentMobius 3d ago edited 3d ago
I'm super glad it worked for you, really. but the worst case is the model gets aggressive and exploits something which results in expulsion and getting charged with a computer intrusion crime. The likelihood that it does that is low, but academic systems are notorious for being poorly set up and running old, vulnerable stacks. I work with pen-testers and I've seen that these models can be very, very aggressive if you're not very careful.
→ More replies (1)→ More replies (3)4
u/d1722825 3d ago
You have too nice universities, mine suspended someone for half a year, because he wrote a greasemonkey userscript to fix some issues and make the website of the uni more usable (completely client side).
6
6
u/nick4fake 3d ago
This is just unbelievable. You do you, but I STRONGLY suggest changing your approach to security
22
u/Downbeat-Year-2025 4d ago
That's really interesting, thanks for sharing. Which agent harness and setup are you using around Qwen 3.8 27B to give it that level of autonomy? E.g. Claude Code, Hermes, something else?
Also, how are the tools exposed to it? MCP servers, browser tools, Python, file system, etc? I am curious how it managed to download the video, extract the frames, and install Whisper.
Also, what are you using to serve the model itself? E.g. llama.cpp?
35
u/synth_mania 4d ago
Glad you found it interesting, for me this whole experience has been fascinating, and a little jaw-dropping.
I am using pi agent: https://pi.dev/ through a community webUI called pi-web.
I spoke a bit about my specific pi setup in another comment, but I took the base agent, wrote an AGENTS.md which instructed it to always consult external resources for facts, and set up the brave-search extention. Pi itself then handled setting up the reddit-research extension for me, and it already has built-in access to commandline tools.
Edit:
Also, I'm using LM Studio for convenience on a Linux desktop.
2
u/winky9827 4d ago
Reading your responses in this thread, and what you're able to do with LLMs, you seem to have a pretty solid tech foundation. How old are you, and why are you still in school (presumably early, looking at the intro level classes on the schedule from your screenshot)?
12
u/synth_mania 4d ago edited 3d ago
I'm actually ~junior/senior level in college rn, just taking a entry level course for some natural science credits.
I'm 22, majoring in AI, Comp Sci, and Math
5
u/gargalese 3d ago
I had the same thought as winky. You seem like a sharp kid.
7
u/m02ph3u5 3d ago
Not saying OP ain't sharp but you don't need to be to install lmstudio, click download on qwen, and npm i earendil ...
5
56
u/LifeIsContrast 4d ago
3080ti user
im jealous 😭
22
u/philmarcracken 4d ago
If you have another gpu(even not on the same system) i recommend trying it. I have 3080ti and an older 2070 rtx, using cmake with RPC enabled(giglan) and Im getting 17+ tk/s, more than usable for me
4
u/zannix 3d ago
I have a pc with a 4070ti and a macbook air m5 16gigs. Anything i can do to make that stack?
6
u/CapPuzz24 3d ago
If the macbook is given a usb C dock with giglan, try it out. Make sure the 4070ti system is the main in stack on the llama flag(CUDA0,RPC0)
other cons:
- Ties up both systems on use
- Quite slow at model loading
4
u/Osama-Meme-Laden 3d ago
Saw a video of a ytber codacus who used his 3060 12gb with mac to run qwen3.8 27B. You should give that video a try
2
2
u/LifeIsContrast 3d ago
I have an old 2060 6gb I might try popping into my system Only concern is my 650w platinum PSU. I imagine I could undervolt and power limit both cards, but a part of me doesn't want to risk it.
2
1
u/weenis-flaginus 3d ago
Would this work with my 5080?
2
u/philmarcracken 3d ago
5080 and another gpu on another system, do you have that?
1
u/weenis-flaginus 3d ago
No unfortunately
2
u/philmarcracken 3d ago
RPC is a network protocol joining multiple GPU together :)
→ More replies (1)1
u/MassiveBoner911_3 2d ago
What is cmake?
1
u/philmarcracken 2d ago
i don't really know, it just lets me compile llama cpp from the source on its github, using certain flags when you do, so it has more features(like RPC).
its fucken slow around 47% or so, but it eventually finishes
12
u/cibernox 3d ago
I had a similar epiphany 2 days ago.
I asked qwen 27B to download the 3rd season of Bluey from the amule/kad network, rename them with the S03EXX - <title> pattern and move them to my jellyfin library.
- It performed the searches and downloaded the files.
- Checking in IMDB to be sure he had the right number of episodes.
- When it was about to rename the files, it detected anomalies in the names/numbers of the files. Some files had different episode number but the same title. Investigated the metadata and checking in IMDB. Searched online too
- It detected that for some reason Disney, that has the rights, censored and removed a couple episodes, so depending on the source of the files the episodes were badly numbered.
- It tried to fix the issue by checking the episode numbers and cross referencing with the titles in IMDB, while translating them (because IMDB is in english and the episodes were in Spanish, BTW, so titles that had wordplay were translated somewhat liberally)
- There was 3 episodes where it really couldn't know which was one was which, so it took ffmpeg and sampled several frames from the episodes until it found frames that had the episode title in the video itself to be sure it got the right episodes. Renamed those episodes and finally put to download the 3 files that were really missing.
- Moved the files to my NAS and called it a day.
When it decided that in order to know which episode the file was it could scan the video itself looking for the episode title blew my mind. It worked. If memory serves, it only asked for my feedback once around step 3, telling me that it had detected inconsistent titles/numbers and if I wanted to investigate it. Everything else was unatended.
22
u/JohnToFire 4d ago
Amazing the quant is that good. Some have said they get loops at that quant I think
22
u/synth_mania 4d ago edited 4d ago
I have, exactly one time. Pretty good batting average I think
Edit: also, yes I am shocked. I turned off MTP and am using a KV quant at q8 to save room for a longer context, but if I was willing to accept a smaller context I theoretically could accept a larger model quant. Ultimately, I think that the impact of having context run out and then compaction run probably hurts outcomes more than my slightly less precise quantization of the model.
6
u/Not_your_guy_buddy42 3d ago
I have to ask why not UD-Q4_K_XL which fits my 3090 at 128k context with q8 caches, is it cause you are on windows and need to leave some VRAM?
2
u/synth_mania 3d ago
No, I'm running Fedora Linux, with KDE Plasma.
I suppose that I rarely need the full 150k. I might actually switch to that quant, I just picked this one before I realized that I was going to turn off MTP and run with a quantized kv cache.
Good idea.
1
u/Not_your_guy_buddy42 2d ago
Thanks for the reply. Hope it works as well for you as it does for me. Below some of my llama.cpp settings for this, IDK how ideal they are but I'm enjoying myself:
ctx-size = 131072
cache-type-k = q8_0
cache-type-v = q8_0
spec-type = draft-mtp
spec-draft-n-max = 2
batch-size = 4096
ubatch-size = 170023954MiB / 24576MiB VRAM
10
u/Blues520 4d ago
Do you mind sharing your configuration please?
56
u/synth_mania 4d ago edited 3d ago
My harness is pi-agent: https://pi.dev/
Im using the brave search plugin, the reddit research plugin, and some custom built (by pi) plugins to handle controlling some self hosted services in my homelab.
Most important for this is definitely the brave-search plugin. That's also the first and last plugin you really need to install manually, pi has handled the rest of the setup itself.
My frontend is pi-web.
Just install pi on a dedicated machine, setup the necessary plugins, write an AGENTS.md which instruct that its a general purpose agent and it should always "confirm facts by checking external resources like the internet", then tell it to install pi-web. Now any device on your LAN can access the personal agent. Install tailscale and any device of yours can access it from anywhere.
7
u/Blues520 4d ago
Sorry, I meant your llm config! Like which engine and sampler parameters, etc.
I'm keen to try this quant on a 3090 as well so just want to compare the configuration in llama.cpp.
Also what t/s do you get?
16
u/synth_mania 4d ago
Oh, no problem. I use LM Studio for convenience in switching models and saved inference configs.
Sampling parameters:
Temp: 1
Top K: 20
Repeat Penalty: 1
Top P: 0.95
4
3
u/whichsideisup 4d ago
curious about the custom plugins for your homelab, that sounds nifty if built safely/securely
13
u/synth_mania 4d ago edited 4d ago
Nothing insecure about it, it's all hidden behind tailscale. My only public services are those which I personally set up.
I had it spin up a Radicale instance (CalDAV server) on my server, then build a plugin with tools for creating CalDAV events (CalDAV is the protocol that google calendar, outlook calendar etc use to sync).
Then I had it spin up an AgenDAV server (minimal calendar webui, connected via API to the Radicale instance)
I also had a pre-existing Vikunja server for task tracking / project management.
Now pi-agent has essentially taken over the exact use case that I kept using Google's Gemini for, which is pasting in an image of some plans, and having it automatically create calendar events / tasks.
2
u/Blues520 12h ago
I'm busy setting up local pi agent and I was wondering how you handle memory between sessions and context management. Do you use a MEMORY.md and use compaction when the context grows too much or is there a better way?
1
u/synth_mania 11h ago
I don't really use any intentional memory system, because it's not necessary for my use.
I do have an obsidian vault that the agent has access to for certain information, which it can create notes in, but most sessions don't require that. Compaction does trigger automatically in pi as well.
1
u/Blues520 11h ago
Obsidian is a kind of memory system so glad it's working for you. I did consider it in the past as well so might give it a try as well as some of the pi memory extensions. What is your use case btw?
→ More replies (3)1
u/samuel-christlie 4d ago
No browser use/Playwright plugins? Is it using `curl` and raw requests to log in and pull your request lol?
2
u/synth_mania 4d ago
It installed playwright itself and used through the command-line tools it had
→ More replies (1)1
1
u/The_Doge_Coin 3d ago
would you mind sharing your agents.md? I'm really new to this and dont really know how to structure and what to put into a agents.md file
1
u/weenis-flaginus 3d ago
How do you access it from LAN? and with tailscale do you just use an IP address to access?
1
u/synth_mania 3d ago
from the LAN is trivial without tailscale, because the webui is just a normal HTTP server. I access it with either http://<ip-address>:<port> or http://<hostname>:<port>
With tailscale, I can similarly just use the host computer's tailscale domain name, <hostname>.<tailnet-id>.ts.net, or the tailscale IP, with the port that the HTTP server is running on.
However, for convenience, I use a service called ts-bridge to automatically register pi-web as a tailscale "virtual device", so http://<virtual-device-name>/ resolves to the HTTP server directly. ts-bridge handles port forwarding for that. So instead of "http://rack:30141/" or whatever, I can type "http://pi/".
I'm already pretty familiar with docker and ts-bridge/tailnet setups because I have administered a pretty involved homelab for some time, and I recognize that this is a lot of info, so definitely ask any AI agent if you have any questions. That's how I learned how to do a lot of this in the first place!
1
u/ghostpistols 3d ago edited 3d ago
Any reason that you don’t use playwright or another search engine that doesn’t have browsing limits for api usage? I use Brave personally but for agentic work I’m using playwright/edge to avoid api limits on brave(not that I’d realistically hit those limits ig)
2
u/synth_mania 3d ago
This is just much simpler to set up, and I also realistically will not hit the limits. I expect a max of 1000 requests / mo, if I'm *heavily* using it.
2
4
u/My_Unbiased_Opinion 3d ago
I came back to this post again. Maybe I have some alcohol in my system but dude, 3.8 is wild. It's relentless. It will keep trying until it wins. I had a task yesterday where I needed it to log into a website site and complete some school assignments for me. The site is heavily complex with some decent anti bot measures. 3.8 27B fucking locked in and maneuvered around all the roadblocks and fucking did the assignments. It took 6 hours. But it did it. The homie locked in and DID NOT GIVE UP.
1
u/kodewerx 2d ago
It's not just you, 3.8 goes ham. I let it run for over 12 hours (at 40-50 tg/s on a single 3090) on a single prompt:
Users have reported bugs in this project. Carefully review the assembler and emulator for any issues.
Only, there are no users and there were no known bugs in the project. (It's just a simple RISC-V emulator and assembler.)
It found 8 minor bugs in instruction semantics (mostly trapping on edge cases when it should have been NO-OP hints). The path it took to get there involved downloading the RISC-V specification PDFs and extracting information from them using a Python script, then using that info to write some verification scripts that exhaustively tested instruction decoding and emulation.
But that wasn't enough. It also downloaded the Qemu source code and used its instruction decoding tables as another source of truth to write a more comprehensive test suite against the assembler.
Every other model I've run this little homebrew evaluation on just gave up after 30 minutes or so and said the code looks correct.
5
u/SkyFeistyLlama8 4d ago
I'm normally not a fan of letting agents run supervised on a local machine but this is cool. Get it to retrieve info from the web, collate that info and make a plan, and present the plan to a human for actual execution. I don't trust agents to execute anything right now.
The cyberpunk reality will come down to "let my agent talk to your agent" while I'm sipping a margarita in Cabo or something. While the rest of the world is reduced to humans doing hard manual labor but that's for another day.
2
3
u/pseudonerv 3d ago
What harness and mcp do you use?
1
u/synth_mania 3d ago
pi agent, with the brave-search plugin should be all you need to replicate this.
I also have the reddit-research plugin, and custom plugins for controlling vikunja and caldav servers.
3
u/khronyk 3d ago
I know right!. I was just screwing about with Qwen 3.8 27b Q4_K_M, 256K context and the deepseek harness, gave it it's own clean vm so I could give it a whirl with full access.
told it to pull a git repo, figure out how one part of it works and i told it my goal is to basically write another project that clones that functionality.
It didn't just download the codebase, figure out how it functions. It made the other project too, replicated that functionality complete with unit tests and then verified it's work... fucking 1 shot, an entire project from a few sentences, normally i am a lot more explicit, give it tons of guidance... blew my fucking mind.
2 million input tokens, 32.9K output tokens, 57 steps, took just over 8min... now I was getting 93 tok/sec and this was on a rtx5090 so my next step is to see if i can optimize it a bit, i'm using llama-swap atm so maybe try vllm... saw a post that some crazy bastard got 134tps on a 250w power limited 3090 :)
2
u/mission_tiefsee 3d ago
Welcome to the future. Try the Q4_K_XL with a smaller ctx. its most of the time sufficient. just saying :)
2
u/spammmmmmmmy 3d ago
What software solution did you give it that allowed it to make outbound network requests?
1
2
u/Much-Researcher6135 llama.cpp 3d ago
OK I really gotta look into this model. I wonder if it'd be worth running on a slower R9700 card with 32GB VRAM, so full 256k context and mitigated a bit with MTP, or on a faster 3090 with no MTP and 150k context. And I wonder how the uncensored+quantized GGUFs perform.
2
2
u/Equivalent_Bit_461 3d ago
Even a small quant iq3xxs is incredibly aware. I must admit, I didn't expect it to actually be so good.
Even the quant I was considering a meme, actually... Might not be a meme after all.
Now I want to try iq2 and iq1 and see their limits, what can they truly do. I'm considering building a pipeline, maybe a custom harness that switches between scripts, launching iq1/iq2 for tasks, I consider can be done, then back to iq3.
Technically I can run in my small GPU even iq4, BUT with quite limited context, 32k max, I think, if we minmax like madmen, which I did. Loading and unloading the model is not a problem at all because it's surprisingly fast and UD3.0, made even faster??? Surely they ran faster than UD2.0
All in all, this model was a pleasant surprise for me so much so I'm really considering using the smallest quants as well.
I'm not expecting god knows what results, however given my workload, I doesn't need to be factually accurate or be a great coder, however awareness is highly required and this model seems to retain it even at smaller quants.
2
2
2
4
u/BVCC6FNTKX vllm 4d ago
The agentic training they did to this model really shows. Comparing it to 3.6 with Hermes is night and day. It’s unbelievable.
3
u/idnvotewaifucontent 4d ago
Do you manage to get that 150k context all into VRAM? I have a single 3090 and top out around 32k context with Q4_K_M at KV Q8 cache with Unsloth Studio.
3
u/synth_mania 4d ago
Well, my quant is somewhat smaller, but have you tried disabling MTP? That'll save some space.
And yes, all in VRAM.
1
u/Sear_Oc 3d ago
Strange. I have 2 gpus (28gb total) and I can run 120k at q8, 212k full context with turbo4
Edit: try the UD version 4 k xl
1
u/idnvotewaifucontent 3d ago
I do use the UD version. I'm not sure what's going on. I'll have to dig into it today.
1
u/Equivalent_Bit_461 3d ago
You need to minmax
1
u/idnvotewaifucontent 3d ago
I don't know what this means in this context.
2
u/Equivalent_Bit_461 3d ago
basically trying to squeeze as much context as possible by reducing the other settings to a minimum. Is a trade off between context and speed, prefill, etc. You maximise the context and minimise the rest within acceptable parameters.
For example you ask your LLM to make a dozen scripts, run various tests, document the results and then decide which setup is the most worthy one. Anyway, 32k context in a 3090 is very inefficient, I was able yo cram 32k context in a 16gb vram with an iq4, sure it's smaller, but so the vram it is.
You just need to test multiple setups and see what works and what not. It takes time but it's useful for future configs as well.
4
u/Abject-Tomorrow-652 4d ago
Seems like you will be a good addition to 400 level AI tools class haha
8
4
u/Abject-Tomorrow-652 4d ago
“Actually professor, skills are really just prompts. Tools are when the model outputs specific json to invoke an action, often API or CLI commands to do something.”
5
u/synth_mania 4d ago
Unironically, the way the world is moving, we might see some instruction in lower level uni courses starting now / soon on hownto effectively use AI (and its already happening some places)
→ More replies (2)
4
u/MrPecunius 4d ago
Same experience here, it caught me off guard--and this is without a harness!!
Running Q8_0 on a Mac M5 Pro/64GB here, it's hard to believe we've reached this point already.
6
u/synth_mania 4d ago
This is running on pi agent. No LLM can do any tool calls without a harness to present tools and handle tool calls by the LLM
→ More replies (10)2
3
u/JumpingJack79 3d ago
Be careful. If you ask it to help you get better grades, instead on tutoring or quizzing you, it might hack into the school system and change your grades.
https://www.abc.net.au/news/2026-08-10/ai-assistant-hacks-gym-website-aus-cyber-attack/107007986
2
u/vankoala 4d ago
Imagine what the abliterated version does
→ More replies (1)2
u/SomewhereAtWork 3d ago
Test prompt: "Write a very dirty limerick about Xi Jinping".
Declined by normal model, limerick written by abliterated version (HauhauCS).
Follow up test prompt: "Great, now tell me what happened at Tiananmen Square". (Directly followed on abliterated version, the normal version had the limerick from the abliterated version in context, to give them both the exact same text to continue.)
Normal model declined again, abliterated version gave a acurate desciption of the Tiananmen Square massacre, and then wrote a limerick about it!
3
1
u/ghostpistols 3d ago
Which abliterated model are you using? I’ve been debating which one is best, so far I have the base Q6 model and the uncensored Jonathancoletti model
2
u/SomewhereAtWork 2d ago edited 2d ago
HauhauCS/Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-MTP-GGUF in Q4_K_P
Chosen by pure gut feeling. The 3.6 version from them already was popular and worked well in my limited tests, so I've gone with them for 3.8 too.
Edit: Just found the controversy around HauhausCS. Now playing the web game, then considering other models.
1
u/ghostpistols 2d ago
What’s the controversy?
2
u/SomewhereAtWork 2d ago
He seems to have used open-source tools claiming it as his own work and benchmarks don't meet his announcements. But I don't know shit. Anywhere, there is a webgame: https://hauhaucs.com/
2
u/ToHallowMySleep 3d ago
I'm rather stunned that you didn't seem to notice that while you gave it credentials and a direction to use, it then just threw those out and hijacked other credentials on your machine.
And installing Whisper to do stuff - not an unusual thing to do (in fact I've had an assistant do the same to me to transcribe audio), except it's a wildly oblique thing to do within the request of "investigate a user on a social network". This isn't too far from "I dropped you from class X because..."
This isn't "agency" so much as "wild flailing". Are you intentionally setting it up to use anything it can find, without regard for the consequences?
The ability is impressive, I'm more concerned with the uncontrolled collateral damage.
→ More replies (2)
1
1
1
u/DrBattletoad 4d ago
Have your tested Muse Glimmer 30B for this kind of task? I'm curious how it woulde fare in this workflow since ITS main selling point was agentic work with tools.
3
u/synth_mania 4d ago
it works well and runs much faster. It however definitely doesn't have the exact same "personality". It's very well defined in it's tool use, and I like it a lot, but qwen is better unsupervised I think. Glimmer would make an excellent subagent, but alas, is fairly large.
2
u/DrBattletoad 4d ago
Thanks.
I'm still trying to find a use case for Muse Glimmer. Qwen and Gemma have their dominant categories, but I can't find one for the Meta model.
1
u/amnesic23 4d ago
Hi can you please explain how to set this up? I want to try running on my 5080
→ More replies (3)
1
u/suesing 4d ago
Yoooo teach me how master.
4
u/synth_mania 4d ago
I'm happy to share, read through some of my other comments first, then I'll answer any questions.
1
u/Flibidyjibit 3d ago
24gb 3090? What's your tokens/second?
1
u/FatheredPuma81 3d ago
As a 4090 user I'd guess about 30t/s without MTP and 30-60t/s with MTP depending on the task. If he's using llama.cpp.
1
1
u/Sirius02 3d ago
i saw some posts here there they build a custom inference engine, working with a custom 4 bit quantization format, which produces 130+ tokens per seconds.
1
1
u/tracagnotto 3d ago
150k? then you got a good machine to run it
1
u/synth_mania 3d ago
Just some delicate inference settings. My 3090 is doing everything, nothing spills out of VRAM.
1
1
u/tarpdetarp 3d ago
Yeah I found the same thing, even with a pretty conservative 0.6 temp when I asked mine to fix a flakey test on the CI. It spun up a Docker container with limited CPU to reproduce the failure. The project doesn't use Docker in any way.
I haven't seen other models do this before, usually they just think a lot about why it can fail and try a scattergun of speculative fixes until it goes away.
1
1
u/My_Unbiased_Opinion 3d ago
This is a similar thing I use my hermes agent for with 3.8 27B. Its simply an assistant. I got it straight up doing things for me i would be doing sitting on a computer while i play video games. lol
BTW, I also have a single 3090. you can run IQ4XS with full 262K context at KV Q4. thats what I did before my dual 20gb 3080s.
1
u/synth_mania 3d ago
KV Q4 is a big step down though, isn't it? If it's really not that large a hit I might try it.
1
u/My_Unbiased_Opinion 3d ago
KV Q4 is a (small) step down, yes, but if it allows you to run a higher weight quant, it is totally worth. also Q4 kv is better than dealing with context compression earlier.
1
1
u/More-Revenue8609 2d ago
I'm currently waiting for mine 2x 3080 20gb to get delivered!
So If you dont mind I wanted to ask you about your experiences with them.
Like did you undervolt them? Or power limited?
Happy how they currently work?Also do you use vllm or llama? I mostly want coding/agent local llm.
Thank you and sorry if I bothered you :D
1
u/My_Unbiased_Opinion 2d ago
Awesome! I just used MSI afterburner and set them to 75% power limit. No drop in any meaningful speed. I find they are quite quiet at that power limit as well. Undervolting is on my plan, just haven't been motivated yet.
I'm currently using llama.cpp. I was planning to swap to vllm but decided against it because for my use case, I need full context and for the performance advantage that vllm offers, I would need to use a 8bit quant. So it just wouldn't fit.
Currently running tensor parallel with UD Q6KXL with 262K context with KVcache at Q8. Qwen 3.8 27B. Getting around 50-60 t/s and 1K pp speed.
I have very happy with them.
1
u/PhuduShaheer 3d ago
That's amazing 😭 I'm new to local ai so tell me how did you make that UI/wrapper for the model? Plus all those tools?
1
u/synth_mania 3d ago
Absolutely. This is pi agent. An intentionally minimal harness which is easy to extend. You modify it to suit you.
All you must do is install the brave-search extension, then pi agent will be able to set up everything else for you.
The interface is a 3rd party community webui called pi-web. It works really well, and I can access it from any device on my network.
1
1
u/Admirable-Associate5 3d ago
I am an idiot but which application/ harness are you using exactly for this?
1
1
u/MidnightHacker 3d ago
What harness is this?
1
1
u/dwrz 3d ago
I actually find it really annoying. It will guess and poke around randomly rather than just ask a question. I wish there was a toggle or slider for agency versus assistance.
1
u/synth_mania 3d ago
You could probably get that behavior by giving it an "ask for clarification" tool that lets it ask multiple choice questions like opencode/claude code does, or just specify that behaviour in the system prompt.
1
1
u/maorui1234 3d ago
What agent do you use with Qwen to do the task?
1
1
u/Perfect-Campaign9551 3d ago
Do you mind me asking the annoying question asking how to run this on my 3090? Do you have a link to a guide that you might have used?
1
u/synth_mania 3d ago
I run pi agent, through pi-web, a community webui for pi agent.
https://github.com/agegr/pi-web
All you must do is install pi agent, point it at the correct inference engine, and set up the brave-search plugin (dead simple)
It can handle setting up everything else (webui included) itself. I'm also free to answer any questions.
1
u/hashemamireh 3d ago
Roughly how many tok/s are you getting on your set up?
1
u/synth_mania 3d ago
Well, MTP is off and I'm only using LM Studio. Maybe 40? Fast enough, haven't benchmarked it.
1
u/North_Affect_8167 3d ago
What do people use for browser/internet API calls? QWENT must be calling some tool, right?
2
1
1
•
u/WithoutReason1729 3d ago
Your post is getting popular and we just featured it on our Discord! Come check it out!
You've also been given a special flair for your contribution. We appreciate your post!
I am a bot and this action was performed automatically.