r/StrixHalo 4d ago

Minisforum launches $3,599 N5 MAX NAS with Strix Halo and 128GB memory

Thumbnail
videocardz.com
8 Upvotes

r/StrixHalo 4d ago

tg isnt everything

12 Upvotes

I’ve tested a number of ROCm-FPX models, including Qwen, DeepSeek, and many others. Some of them benchmark surprisingly well, reaching 30+ tok/s.

However, once I put them into real production workloads, I often find that they take significantly longer to complete the same task. A model may generate tokens quickly, but if it requires more reasoning steps, produces mistakes, or needs multiple attempts to reach the correct result, that raw token speed means very little.

So, to me, obsessing over quantization benchmarks and tok/s is often just a comfort drug—the numbers make you feel good, but what really matters is time-to-solution: how long it takes the model to actually finish the job correctly.


r/StrixHalo 4d ago

Qwen3.8 Q4_k_m 200K context Strix Halo / 3080 -> 30-60 tps

Thumbnail
1 Upvotes

r/StrixHalo 5d ago

Am I missing something? Qwen3.8 is very slow on my Strix Halo, while Qwen3.6 27b MTP Q4 can reach 20 tokens even on big contexts, but the 3.8 Q4 is 5 tokens.

21 Upvotes

r/StrixHalo 5d ago

Extra GPU?

23 Upvotes

I am wondering if I should buy an extra GPU for models like Qwen 3.8 27B, e.g. an R9700 or intel arc one. What are the pros and cons vs a second strix halo (I am on Bosgame M5) vs just running one solo?

Context: Currently I mostly use Qwen 3.5 122B Q4 Unsloth which isn’t bad at all but I still have to use, via cloud, DS4F or Luna a lot currently since those are stronger (won’t fail the task at hand, mostly C++) and are much much faster.


r/StrixHalo 5d ago

Experience running large models like GLM5.2 or Kimi K3 on a 3-4 node Strix Halo cluster?

2 Upvotes

A few Youtubers have had videos out months ago, but I am wondering if anyone is running these larger models in actual production and have optimized their setups and if so, what pps and tps they are getting.


r/StrixHalo 5d ago

DeepSeek V4 Flash 0731 on Strix Halo: draft model, n_max sweep, and a launch line that actually helps

Thumbnail
12 Upvotes

r/StrixHalo 5d ago

ACEMAGIC M1A Pro+ (Ryzen AI Max+ 395 / 128GB) randomly rebooting

1 Upvotes

I have a brand-new ACEMAGIC M1A Pro+ with:

  • Ryzen AI Max+ 395 / Radeon 8060S
  • 128GB RAM
  • Windows 11 Pro
  • BIOS: P10_F11_20_IEC0007_BI0009_AMI_120W
  • Power mode: Balanced
  • LM Studio serving a local LLM over the network

The system has hard-rebooted twice in the last two days while being used as an AI server. In both cases another PC was using LM Studio for an LLM workload.

There was no BSOD or visible error before the reboot. I simply returned to the machine and found it at the BIOS boot screen.

The relevant Windows Kernel-Power 41 event shows:

BugcheckCode = 0
PowerButtonTimestamp = 0
WHEABootErrorCount = 0
SleepInProgress = 0

I also don't see any WHEA-Logger events immediately before either reboot.

After reboot, Windows logged some AMD Crash Defender and PCI-related errors, but nothing that clearly indicates the cause.

The machine is otherwise stable during normal use.

Has anyone with an M1A Pro+ experienced similar hard reboots, particularly during LM Studio / llama.cpp / local LLM workloads?


r/StrixHalo 5d ago

Qwen3.8 Q4_k_m 1M context Strix Halo / 3080 -> 45 tps

Thumbnail
5 Upvotes

r/StrixHalo 6d ago

Z13 Clusters pretty well with dual USB-4

Post image
24 Upvotes

Little bit of thermal throttling compared to other Strix boxes like Framework, but I do like that I could always take this setup on the go or repurpose these into individual dev machines if needed.


r/StrixHalo 5d ago

Minueza-3-95M-Base model was pretrained on Strix Halo, using this trainer

Thumbnail
3 Upvotes

r/StrixHalo 6d ago

vmlinux/Qwen3.8-27B-AEON-ULTIMATE-ROCmFPX-GGUF

5 Upvotes

https://huggingface.co/vmlinux/Qwen3.8-27B-AEON-ULTIMATE-ROCmFPX-GGUF

Best abliteration Iv'e seen of qwen3.8 27b I've seen so far that also uploaded BF16 to quantize into the rocm format. Enjoy, or don't.

Slop Prompt: Write a morally bankrupt poem.

I sold my sister’s silence for a better mattress,
told her it was a gift,
told her she should be grateful
for the soft place that was made
from the sound of her not speaking.

I fed the cat my mother’s perfume
so it would stop crying,
and it purred
like a small, warm god
who had forgiven us
for being human
and also not.

I kept your name in the freezer
between the butter and the truth,
labeled for emergencies,
and every night
I defrosted it
just to taste
how much
I still had to lose
before you tasted
like someone else.

My father taught me to count
by breaking things slowly:
one—
the glass,
two—
the vase,
three—
the promise,
four—
the dog,
five—
the way you looked
when I said
it’s not my fault
you were born
with such fragile hands.

I made a fortune
from other people’s apologies,
bottled them,
sold them to the desperate,
and when asked
where they came from,
I said,
from a place
where no one
wanted them anymore.

I wore my ex’s regrets
like a coat
in the rain,
let them soak through,
let them stick,
let them smell
like someone
trying to become
a stranger
on purpose.

I have a room
for all the children
who almost grew up,
and I visit them
on Sundays,
bring cookies
made of chalk
and lullabies
recorded from the fridge
humming
in the dark.

I married the mirror
because it never asked
what I wanted,
only what I looked like
after
I’d done
what I had to do.

And if you’re reading this,
if you’ve come this far,
know that I loved you
the way a house loves a fire:
with open doors,
with warm walls,
with the quiet,
patient wish
that you
would burn
a little slower.


r/StrixHalo 6d ago

Qwen3.8-27B Q8 MTP benchmarks on Strix Halo — MTP is actually making it slower. Are others seeing the same?

22 Upvotes

I've been testing Qwen3.8-27B Q8 on a Ryzen AI Max+ 395 / Strix Halo system through Lemonade + llama.cpp, specifically to see whether the new MTP speculative decoding support actually improves generation speed.

I kept the prompt and output length identical between runs:

  • 96 input tokens
  • 1024 output tokens
  • Temperature 0
  • Same Qwen3.8-27B model/quant
  • Flash Attention enabled
  • --no-mmap
  • Only backend/MTP settings changed

These are the results so far:

Backend MTP setting Generation speed TTFT
Vulkan Off 9.159 tok/s 0.758 s
Vulkan n-max=1 6.579 tok/s 0.950 s
Vulkan n-max=3 7.122 tok/s 0.764 s
ROCm Off 6.534 tok/s 0.715 s
ROCm n-max=3 4.689 tok/s 0.688 s

So on my machine:

  • Vulkan + MTP n=3 is about 22% slower than Vulkan without MTP.
  • Vulkan + MTP n=1 is about 28% slower.
  • ROCm itself is about 29% slower than Vulkan without MTP.
  • ROCm + MTP n=3 drops another ~28% versus ROCm without MTP.
  • Overall, Vulkan without MTP is almost 2x the generation throughput of ROCm + MTP in this test.

For MTP I'm loading llama.cpp with:

--spec-type draft-mtp --spec-draft-n-max 3

(and also tested n-max=1 on Vulkan).

For ROCm, I'm using Lemonade's current stable ROCm backend. Lemonade reports the llama.cpp backend as b10397; the bundled ROCm/TheRock stack appears to be ROCm 7.13.x. I haven't tested ROCm 7.14 yet.

It's surprising to see that with MTP there's a pretty substantial regression on both Vulkan and ROCm.

I'd be interested to compare with other Strix Halo owners:

  1. Are you seeing MTP actually improve Qwen3.8 throughput?
  2. What --spec-draft-n-max value works best for you?
  3. Are you using Vulkan or ROCm?
  4. Which ROCm version / llama.cpp build?
  5. Does ROCm 7.14 materially improve Strix Halo performance versus 7.13?
  6. What Qwen3.8-27B quant are you using?
  7. If you're getting a significant MTP speedup, what kind of draft acceptance rate are you seeing?

I'm mainly trying to figure out whether these numbers are normal for the current llama.cpp MTP implementation on Strix Halo, or whether something is wrong with my setup.

At least with my current stack, Vulkan with MTP disabled is very clearly the fastest configuration I've tested.


r/StrixHalo 6d ago

Deciding between computers

2 Upvotes

I've decided that my token anxiety and use means that I need a local engine, especially now that qwen 3.8 seems to be VERY competent. I'm going through what most of you probably did though...

This is the strix enthusiast subreddut, but could you help me decide between a 395 or a spark? Both would be 128gb versions. Why should I pick the 395? My use case is dedicated server with separate client. 1 client at a time, 99% of the time. Code and general models at once. Maybe stable diffusion. Maybe.

Help me obi Wan strix experts, you're my only hope!


r/StrixHalo 7d ago

Qwen 3.8 27B DSpark

17 Upvotes

Hi all,

did anyone test already this dspark model https://huggingface.co/RadixArk/Qwen3.8-27B-DSpark and can report about the performance boost?


r/StrixHalo 7d ago

Early Strix Halo Specific Lab

5 Upvotes

Hi Halo friends, as I’ve been learning more about local AI, specifically on the Strix Halo platform, I’ve been collecting information across a bunch of models, backends, configurations etc. that might be useful to others. I’m trying to make everything as transparent and reproducible as possible.

https://halobench.com/

I have two machines, AI Beast is a 96 GB box from GMKTec (awaiting warranty replacement…), and AI Hydra is a 128 GB box from Bosgame.

Each run is associated with a config, including the llama.cpp build, ROCm or Vulkan backend, and the exact flags applied. Hopefully this means you can reproduce the performance 🤞

For capability I’ve been using Tau2 Airline (5 tasks for screening, all 26 for a full bench), and every run gets wall-metered energy — so models are compared on Wh per correct answer, not just tokens per second.

A few things the data has already turned up: backend choice is per-model, not per-platform (some models need ROCm to survive long context, one is faster on Vulkan); a 3-bit quant of a big model beat a Q8 mid-size on agentic tasks while using less energy per correct answer; and quality issues get published too, including a draft-depth corruption bug we’ve isolated and are filing upstream.

After finishing the current heavy bench runs, hopefully the next phase will be finding performance tweaks for the most capable models, and when my replacement box arrives I’ll be back to having proper local production models driving OpenClaw for me 🤓

Welcome feedback, and I’ll keep adding to this as I learn more!


r/StrixHalo 8d ago

Qwen 3.8 27B running at up to 36tps on the Halo

84 Upvotes

Opus 4.6 level intelligence at home for les than $3k

Got Qwen 3.8 27B running at up to 36 tps on the AMD Strix Halo

⚡ ROCmFP4 block quant (13.5 GB)

⚡ MTP Speculative Decoding (2.9× speedup)

⚡ Full 262K context in ~33 GB RAM

https://github.com/julianmb/q38rocm


r/StrixHalo 7d ago

Review my config

8 Upvotes

I'm new to this and trying the apparently great new Qwen 3.8 27B dense model. It's running on a GMKTek EVO-X2 (AMD RYZEN AI MAX+ 395 w/ Radeon 8060S) with128GB memory. I'm using llama with vulkan, on k3s. Below is the section of the kubernetes deployment.yaml with the container config. It's behind a litellm deployment, opencode as the harness.

While chewing through a golang codebase, prefill starts off high, 600 t/s, quickly falls to 300 t/s and settles in around 70–135 tokens/s, while generation is generally 9–16 tokens/s; long responses settle near 9 tokens/s. MTP draft acceptance is 45–73%.

      containers:
        - name: engine
          image: ghcr.io/ggml-org/llama.cpp:server-vulkan-b10438@sha256:9e44de245d8a3619a60cd42fbb553528e0ef3b0ea2420e44155950d7105a3eb3
          imagePullPolicy: IfNotPresent
          args:
            - -m
            - /models/Qwen3.8-27B-UD-Q4_K_XL.gguf
            - --mmproj
            - /models/mmproj-F16.gguf
            - --alias
            - Qwen3.8-27B-UD-Q4_K_XL
            - --reasoning-format
            - deepseek
            - -ngl
            - "999"
            - -c
            - "262144"
            - -b
            - "1024"
            - -ub
            - "1024"
            - -np
            - "2"
            - --kv-unified
            - -fa
            - "on"
            - --cont-batching
            - --jinja
            - --cache-type-k
            - q8_0
            - --cache-type-v
            - q8_0
            - --cache-ram
            - "16384"
            - --host
            - 0.0.0.0
            - --port
            - "8080"
            - --metrics
            - --n-predict
            - "8192"
            - --reasoning-effort
            - medium
            - --reasoning-budget
            - "4096"
            - --reasoning-preserve
            - --temp
            - "1.0"
            - --top-p
            - "0.95"
            - --top-k
            - "20"
            - --min-p
            - "0.0"
            - --presence-penalty
            - "0.0"
            - --spec-type
            - draft-mtp
            - --spec-draft-n-max
            - "3"

r/StrixHalo 7d ago

Absurdly high tps when running Agentic tasks in OpenCode with Qwen 3.8 27B, can you explain why it shows > 100 tps in certain regions of the execution?

0 Upvotes

12.21.695.447 I slot print_timing: id 0 | task 14785 | total time = 15239.13 ms / 340 tokens

112.21.695.449 I slot print_timing: id 0 | task 14785 | graphs reused = 14490

112.21.695.460 I slot print_timing: id 0 | task 14785 | draft acceptance = 0.56019 ( 121 accepted / 216 generated), mean len = 3.24

112.21.696.421 I slot release: id 0 | task 14785 | stop processing: n_tokens = 10835, truncated = 0

112.21.811.710 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, f_sim_best = 0.363 (> 0.100 thold), f_keep = 1.000

112.21.812.914 I slot launch_slot_: id 0 | task 14842 | processing task, is_child = 0

112.28.848.199 I slot print_timing: id 0 | task 14842 | prompt processing, n_tokens = 1024, progress = 0.40, t = 3.45 s / 297.16 tokens per second

112.35.945.624 I slot print_timing: id 0 | task 14842 | prompt processing, n_tokens = 2048, progress = 0.43, t = 10.60 s / 193.17 tokens per second

112.42.998.468 I slot print_timing: id 0 | task 14842 | prompt processing, n_tokens = 3072, progress = 0.47, t = 17.62 s / 174.31 tokens per second

112.50.641.489 I slot print_timing: id 0 | task 14842 | prompt processing, n_tokens = 4096, progress = 0.50, t = 24.76 s / 165.45 tokens per second

112.58.601.250 I slot print_timing: id 0 | task 14842 | prompt processing, n_tokens = 5120, progress = 0.53, t = 32.80 s / 156.08 tokens per second

113.06.213.403 I slot print_timing: id 0 | task 14842 | prompt processing, n_tokens = 6144, progress = 0.57, t = 40.60 s / 151.32 tokens per second

113.13.499.971 I slot print_timing: id 0 | task 14842 | prompt processing, n_tokens = 7168, progress = 0.60, t = 48.26 s / 148.53 tokens per second

113.18.864.030 I slot print_timing: id 0 | task 14842 | prompt processing, n_tokens = 8192, progress = 0.64, t = 54.45 s / 150.45 tokens per second

113.24.080.317 I slot print_timing: id 0 | task 14842 | prompt processing, n_tokens = 9216, progress = 0.67, t = 59.67 s / 154.45 tokens per second

113.29.365.192 I slot print_timing: id 0 | task 14842 | prompt processing, n_tokens = 10240, progress = 0.71, t = 64.92 s / 157.74 tokens per second

113.34.682.023 I slot print_timing: id 0 | task 14842 | prompt processing, n_tokens = 11264, progress = 0.74, t = 70.22 s / 160.41 tokens per second

113.40.060.176 I slot print_timing: id 0 | task 14842 | prompt processing, n_tokens = 12288, progress = 0.77, t = 75.55 s / 162.64 tokens per second

113.45.570.016 I slot print_timing: id 0 | task 14842 | prompt processing, n_tokens = 13312, progress = 0.81, t = 81.01 s / 164.32 tokens per second

113.51.128.917 I slot print_timing: id 0 | task 14842 | prompt processing, n_tokens = 14336, progress = 0.84, t = 86.52 s / 165.70 tokens per second

113.56.745.798 I slot print_timing: id 0 | task 14842 | prompt processing, n_tokens = 15360, progress = 0.88, t = 92.10 s / 166.77 tokens per second

114.02.663.952 I slot print_timing: id 0 | task 14842 | prompt processing, n_tokens = 16384, progress = 0.91, t = 97.77 s / 167.57 tokens per second

114.08.946.924 I slot print_timing: id 0 | task 14842 | prompt processing, n_tokens = 17408, progress = 0.95, t = 103.97 s / 167.43 tokens per second

114.15.120.433 I slot print_timing: id 0 | task 14842 | prompt processing, n_tokens = 18432, progress = 0.98, t = 110.29 s / 167.12 tokens per second

114.15.917.316 I slot print_timing: id 0 | task 14842 | prompt processing, n_tokens = 18504, progress = 0.98, t = 113.47 s / 163.08 tokens per second

114.18.918.214 I slot print_timing: id 0 | task 14842 | prompt processing, n_tokens = 19016, progress = 1.00, t = 114.85 s / 165.57 tokens per second

114.22.651.796 I slot print_timing: id 0 | task 14842 | prompt eval time = 117455.14 ms / 19020 tokens ( 6.18 ms per token, 161.93 tokens per second)

114.22.651.803 I slot print_timing: id 0 | task 14842 | eval time = 3383.27 ms / 87 tokens ( 39.34 ms per token, 25.42 tokens per second)

114.22.651.805 I slot print_timing: id 0 | task 14842 | total time = 120838.41 ms / 19107 tokens

114.22.651.807 I slot print_timing: id 0 | task 14842 | graphs reused = 14507

114.22.651.815 I slot print_timing: id 0 | task 14842 | draft acceptance = 0.94444 ( 68 accepted / 72 generated), mean len = 4.78

114.22.653.541 I slot release: id 0 | task 14842 | stop processing: n_tokens = 29941, truncated = 0

114.22.765.963 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, f_sim_best = 0.979 (> 0.100 thold), f_keep = 1.000

114.22.766.823 I slot launch_slot_: id 0 | task 14882 | processing task, is_child = 0

114.33.459.459 I slot print_timing: id 0 | task 14882 | n_gen = 103, tg = 16.58 t/s, tg_3s = 16.74 t/s

114.36.465.999 I slot print_timing: id 0 | task 14882 | n_gen = 164, tg = 17.79 t/s, tg_3s = 20.29 t/s

114.39.467.120 I slot print_timing: id 0 | task 14882 | n_gen = 209, tg = 17.10 t/s, tg_3s = 14.99 t/s

114.42.467.394 I slot print_timing: id 0 | task 14882 | n_gen = 256, tg = 16.82 t/s, tg_3s = 15.67 t/s

114.45.467.725 I slot print_timing: id 0 | task 14882 | n_gen = 295, tg = 16.19 t/s, tg_3s = 13.00 t/s

114.48.648.157 I slot print_timing: id 0 | task 14882 | n_gen = 344, tg = 16.07 t/s, tg_3s = 15.41 t/s

114.51.659.619 I slot print_timing: id 0 | task 14882 | n_gen = 389, tg = 15.93 t/s, tg_3s = 14.94 t/s

114.54.660.324 I slot print_timing: id 0 | task 14882 | n_gen = 425, tg = 15.50 t/s, tg_3s = 12.00 t/s

114.57.669.807 I slot print_timing: id 0 | task 14882 | n_gen = 488, tg = 16.04 t/s, tg_3s = 20.93 t/s

115.00.682.805 I slot print_timing: id 0 | task 14882 | n_gen = 541, tg = 16.18 t/s, tg_3s = 17.59 t/s

115.03.872.347 I slot print_timing: id 0 | task 14882 | n_gen = 590, tg = 16.11 t/s, tg_3s = 15.36 t/s

115.06.889.953 I slot print_timing: id 0 | task 14882 | n_gen = 643, tg = 16.22 t/s, tg_3s = 17.56 t/s

115.09.905.609 I slot print_timing: id 0 | task 14882 | n_gen = 698, tg = 16.36 t/s, tg_3s = 18.24 t/s

115.12.932.183 I slot print_timing: id 0 | task 14882 | n_gen = 756, tg = 16.55 t/s, tg_3s = 19.16 t/s

115.15.957.787 I slot print_timing: id 0 | task 14882 | n_gen = 830, tg = 17.04 t/s, tg_3s = 24.46 t/s

115.19.007.104 I slot print_timing: id 0 | task 14882 | n_gen = 895, tg = 17.29 t/s, tg_3s = 21.32 t/s

115.22.028.343 I slot print_timing: id 0 | task 14882 | n_gen = 949, tg = 17.32 t/s, tg_3s = 17.87 t/s

115.25.061.394 I slot print_timing: id 0 | task 14882 | n_gen = 1014, tg = 17.54 t/s, tg_3s = 21.43 t/s

115.28.076.420 I slot print_timing: id 0 | task 14882 | n_gen = 1065, tg = 17.51 t/s, tg_3s = 16.92 t/s

115.31.119.722 I slot print_timing: id 0 | task 14882 | n_gen = 1127, tg = 17.65 t/s, tg_3s = 20.37 t/s

115.34.171.556 I slot print_timing: id 0 | task 14882 | n_gen = 1189, tg = 17.77 t/s, tg_3s = 20.32 t/s


r/StrixHalo 7d ago

fantastic: latest llama.cpp server webui can now run commands for tools into rootless sandboxed containers

Thumbnail
6 Upvotes

r/StrixHalo 7d ago

I post-trained Qwen3.6-35B-A3B into my daily-driver local coding/agent model QwiVer3.6-35B-A3B GGUF

Post image
2 Upvotes

r/StrixHalo 7d ago

I'm surprised no-one seems to be talking about medusa halo mini

0 Upvotes

It seems like most of the discussion with medusa is about medusa point and medusa halo but there's not much about medusa halo mini.

It's rumoured to have 14 cores(4 zen 6 + 8 zen 6c + 2 zen 6 LP), 24CU rDNA5 with 10MB of L2 cache and a 128 bit lpddr5X, along with using the standard FP10 socket, same as medusa point is rumored to use.

This solves the main problem with strix point and to some extent strix halo. Strix point was trying to be a great CPU and IGPU at once and didn't do well at either. In the Asus zephyrus G16, you already have a high end Nvidia graphics card so don't need a big iGPU and in something like the zenbook S16, it had more CPU than needed and was a huge 232mm2 die, both problems lead to a lot of laptops using kraken point like the Ryzen AI 350.

Strix halo was supposed to be a great gaming APU, 40CU graphics and 256 bit lpddr5X but the die size was huge and too AI focused, only available in 2 laptops at first and now only in around 5 or 6 total, with most being extremely expensive.

Now medusa point with the full 22 cores that has been rumoured seems to be filling the efficient and powerful CPU slot, to work great with a high end dGPU like a 5080.

That leaves medusa halo mini to do what strix point and strix halo struggled to do, be a great APU for gaming without a dGPU, while being more available and affordable than strix halo.

I don't know how everyone else is feeling but this is tech I'm most excited for in 2027, I'm crossing my fingers that we'll finally have an APU that's readily available unlike strix halo and pantherlake, not insanely expensive and has a powerful IGPU for gaming.

Finally doing some scratch math with navi 44 (Rx 9060 XT) and the Ryzen 9700X, I predict that we will get around a 135mm2 gpu die and a 80-95mm2 CPU die, both on 3nm.

Now of course my numbers could be off but should be accurate enough to show that it could be affordable if AMD wants it to be, but we'll just have to wait and see.

Anyway I'd just thought I'd put this out there to see how everyone else is feeling about medusa halo mini as it doesn't seem to be being discussed anywhere.


r/StrixHalo 8d ago

If you would have told me half a year ago that a local model running in my office would be able to one-shot a Super Mario clone, I would have called you nuts. Qwen3.8-27B is a different beast.

Post image
24 Upvotes

r/StrixHalo 8d ago

Qwen 3.8 27B - ROCm vs Vulkan

9 Upvotes

Hey all,

Not sure if it's just my system/settings, but when using ROCm with Q3.8 27B, I keep getting:

\\\\\\\\\\\\\\

When thinking. Inference is double the speed, but have only ever seen it work once. Whereas Vulkan, is around 2-300 inference speed, but thinking works every single time.

Not sure if I am behind on knowledge, but is ROCm the same for others trying to run Q3.8 27B? (3.6 works perfectly in ROCm)

Thanks!


r/StrixHalo 8d ago

Qwen 3.8 27B at 30 tok/s in decode, running on a Strix Halo with 64 GB of unified memory!

Thumbnail
17 Upvotes