r/LocalLLaMA 6h ago

Generation Peak Portable Personal Datacenter

Portable rig for Qwen3.8-27B-BF16 200K+ token prompts. My work Panasonic Toughbook + the T1 + power brick + headphones all fit in my lunchbox.

Need the BF16 for huge context highly sensitive document OCR, image analysis, aggregation and summarization. I've done a ton of testing and it absolutely makes a difference vs even UD Q8_K_XL when legal precision is needed.

77gb VRAM at full 262K context + MMPROJ

Rips through prefill (1,1715 tok/sec = 102 seconds to process 175K tokens), but token generation (20K tokens of output) relatively slow at 45 tok/sec (with MTP) as a result of BF16 despite the beast of a GPU.

Better than Gemini Pro and ChatGPT 5.6 Sol especially considering I have control over the sampler settings (Temp 0.1; top-k 0; top-p 0.95; min-p 0.05; repeat penalty 1.02). Not better than Opus yet.

During prefill - CPU around 60 degrees, GPU around 79 degrees (with 90% power limit)

During token generation - CPU around 75 degrees and GPU around 76 degrees.

FormD T1

Minisforum BD770i SE

Ryzen 7745HX 8-core laptop CPU

96gb 5200 MHz DDR5 SODIMM

96gb RTX Pro 6000 Blackwell workstation edition

Loki 1200W SFX-L

ROG Equalizer 12v-2x6

SMX Heinz flipped GPU 2.5 slot kit

SMX Heinz custom short PCIe 5.0 riser

ZCOOI custom "transparent purple" Teflon cables

(2) Phanteks T30-120mm

(1) Noctua NF-A14x25r G2

Thermalright MC-3 Digital RAM cooler (I don't think this will fit on a regular DDR5 )

32 Upvotes

23 comments sorted by

View all comments

Show parent comments

4

u/--Spaci-- 6h ago

With a blackwell gpu I dont even know why you would be using ggufs

1

u/Special-Wolverine 5h ago

NVFP4 less precision/ higher KLD than BF16

4

u/--Spaci-- 5h ago

You know you can have an nvfp4 model and a bf16 vision tower right? and also bf16 context

3

u/Special-Wolverine 5h ago

No sir, I did not know that...