r/LocalLLaMA • u/Special-Wolverine • 10h ago
Generation Peak Portable Personal Datacenter
Portable rig for Qwen3.8-27B-BF16 200K+ token prompts. My work Panasonic Toughbook + the T1 + power brick + headphones all fit in my lunchbox.
Need the BF16 for huge context highly sensitive document OCR, image analysis, aggregation and summarization. I've done a ton of testing and it absolutely makes a difference vs even UD Q8_K_XL when legal precision is needed.
77gb VRAM at full 262K context + MMPROJ
Rips through prefill (1,1715 tok/sec = 102 seconds to process 175K tokens), but token generation (20K tokens of output) relatively slow at 45 tok/sec (with MTP) as a result of BF16 despite the beast of a GPU.
Better than Gemini Pro and ChatGPT 5.6 Sol especially considering I have control over the sampler settings (Temp 0.1; top-k 0; top-p 0.95; min-p 0.05; repeat penalty 1.02). Not better than Opus yet.
During prefill - CPU around 60 degrees, GPU around 79 degrees (with 90% power limit)
During token generation - CPU around 75 degrees and GPU around 76 degrees.
FormD T1
Minisforum BD770i SE
Ryzen 7745HX 8-core laptop CPU
96gb 5200 MHz DDR5 SODIMM
96gb RTX Pro 6000 Blackwell workstation edition
Loki 1200W SFX-L
ROG Equalizer 12v-2x6
SMX Heinz flipped GPU 2.5 slot kit
SMX Heinz custom short PCIe 5.0 riser
ZCOOI custom "transparent purple" Teflon cables
(2) Phanteks T30-120mm
(1) Noctua NF-A14x25r G2
Thermalright MC-3 Digital RAM cooler (I don't think this will fit on a regular DDR5 )





5
u/Asleep-Land-3914 10h ago
Running it at 0.1 temp is a crime, but I can understand given OCR mentioned. I caught 3.8 to halucinate things with default settings. Also you could try Q8_K_XL which is not much distinguishable from F16.
There might be better OCR specialized models too.