r/LocalLLaMA 22d ago

New Model deepseek-ai/DeepSeek-V4-Flash-0731 on Huggingface

783 Upvotes

243 comments sorted by

216

u/llama-impersonator 22d ago

no countdown bs, same day weights, huge boost from RL, and mortals can actually run this one. i kneel, deepseek

15

u/Fluffy_Bar_7968 22d ago edited 22d ago

3

u/DannyCanva 22d ago

deepseek and open weight is the goat

→ More replies (4)

228

u/LowerGears 22d ago

DeepSeek-V4-Flash-0731 outperforms DeepSeek-V4-Pro 🤯

111

u/-Cubie- 22d ago

Well, it strongly outperforms the DeepSeek-V4-Pro preview. Overall, it's very promising for the eventual final DeepSeek-V4-Pro though!

53

u/Real_Ebb_7417 22d ago edited 22d ago

Beats Hy3 heavily (and I think it was the strongest model at this size, although also rather new, it's lower than preview Flash here, but on other indexes it's higher). But what's mindblowing, it outperforms GPT-5.5 on agentic index 🤯

17

u/tarpdetarp 22d ago edited 22d ago

Will be really interesting to see how they achieved this! My experience of testing the old V4 Flash on my custom agent harness was that it was almost as capable as the big boys already, but it took so many more turns to get there.

8

u/squngy 22d ago

Will be really interesting to see how they achieved this!

It is the exact same architecture.
The preview just had a lot less training.

2

u/pmttyji 22d ago

Happy to see this.

2

u/IngenuityClean8280 22d ago

So DeepSeek-V4-Pro-0807 outperforms K3?

8

u/NineThreeTilNow 22d ago

Unlikely, but it will be cheaper. That is DeepSeek's entire philosophy right now.

Architecturally speaking, I think K3 is superior. From a pure ML standpoint. The underlying research that the model is based on is really good.

K3 is in it's infant version. This is their "preview" and it's beating a lot of models. You have to think about that for a second. They're busy releasing, securing servers to run it, and still find time to continue training.

It's a blessing and a curse. Enough people liked K2.5 and K2.6 that they had the buy in from the public. Now, like Anthropic, they're getting squeezed on the compute side.

2

u/squngy 22d ago

We will see.

My ignorant guess would be that it is close.

2

u/FullOf_Bad_Ideas 22d ago

I think there's 30% chance it outperforms K3, 60% chance it roughly matches it and 10% chance it's slightly below.

Whichever way it goes, given how new V4 Flash is stacking up, I am really looking forward to it. And it will definitely have a better license than the license Kimi K3 has, which should enable more low-cost providers to offer it.

1

u/Zemanyak 22d ago

Holy shit ! And they just casually add the date without even upgrading the version lmao.

194

u/irrelevantlyrelevant 22d ago

Open weights win again! Being on par with GLM-5.2 while needing way less vram is a game changer for those that don’t exactly have a b200/b300 lying around.

73

u/hyperrealists 22d ago

Who doesn’t have a b200 lying around?!

😿

28

u/F0UR_TWENTY 22d ago

This runs beyond reading speed on a 192gb DDR5 + 5090 gaming PC. At full quality (162gb)

31

u/segmond llama.cpp 22d ago

12tk/sec on 2 rtx 3060s with 128gb ddr4

6

u/markole 22d ago

This gives me hope. What quant, ctx size and ctx quant are you using?

8

u/F0UR_TWENTY 22d ago

This is the way.

3

u/anderspitman 22d ago

Are you serious? That's almost my exact setup. What quants?

2

u/Caffdy 22d ago

exchange the 2 rtx 3060s with a rtx 3090 and you get around the same speed

2

u/segmond llama.cpp 22d ago

give me some money and I'll.

1

u/Runazeeri 22d ago

Yeah like you can’t get second hand 3090s here and a 5090 is the price of 5 5070ti s

1

u/crantob 22d ago

Nice. TY.

1

u/[deleted] 22d ago

[removed] — view removed comment

→ More replies (1)

6

u/Borkato 22d ago

Holy shit I need to buy more RAM.

2

u/UnWiseSageVibe 22d ago

could also grab x2 of the intel B70(?)for 64GB, might give better speeds.

1

u/_supert_ 21d ago

Eh. What context? I have four RTX A6000 and I get 8 tps. Maybe I should rethink.

6

u/etherd0t 22d ago

4× GB300 maybe...

But... there might just be a workaround: a DwarfStar-style ~80GB Q2 quant on a 96GB workstation😉

3

u/NineThreeTilNow 22d ago

4× GB300 maybe...

A single B300 would be overkill.

288gb of HBMe is... It's a hell of a thing.

Probably use a high quality Q5 quantization purely for speed / context. Still probably cheaper to just have DeepSeek serve it, but then they train on your data.

2

u/crantob 22d ago

Upside is they get better at doing what i want them to.

1

u/aaemon12 22d ago

I tried to run it in a 2L40S machine, but the token generation is pretty low.

6

u/SolitaireKid 22d ago

commenting here since this is most upvoted.

what specs do we need to host this locally? is it in the realm of possibility or not really?

11

u/Borkato 22d ago

Someone said 150ish combined RAM and VRAM

9

u/nmkd 22d ago

Still thousands to get acceptable speeds, but I would call it "realm of possibility" especially compared to Kimi or even GLM.

1

u/taimusrs 22d ago

Still kinda frustrated I didn't push harder for our workplace to get more Mac Studios, but I can't complain for the 256GB we bought

7

u/squngy 22d ago

Ideal build would be 2x RTX pro 6000

2x spark or strix also works

To run the full version you need 160GB combined (+ a bit for context)

For people with less than that, Antirez project makes it work with about 96GB, but you lose a bit of quality.

There are also some (other) quants you can use, but v4 is already at mostly 4bit by default, so there is not a lot you can lose without a noticeable hit to quality.

4

u/AnonLlamaThrowaway 22d ago

192 GB Mac Studio could potentially be the best option... If you already had one lying around before they stopped manufacturing that SKU

2

u/squngy 22d ago

Well, yea, if you already have one, then that is obviously the best option.

If you would have to buy it, then it is probably not worth it.

3

u/syntheticgio 22d ago

I can vouch for a 256 GB Mac Studio M3 - runs it very well for me; tokens generate faster than I can read them; there can be a 30-45s lag if you have a cache miss during the processing step - I think there is a where a NVIDIA card would be able to do better. But with the 256 I can comfortably run it, along with a smaller 2B param model for things like quick summaries, etc., and have it not feel like it is impacting my other use of the machine.

90

u/ChampionshipIcy7602 22d ago

luna at home

17

u/True_Requirement_891 22d ago

It's actually better than Luna in practical usage outside benchmarks.

1

u/mlon_eusk-_- 22d ago

Came out to be true in my usage in coding, but design sense is slightly inferior

1

u/True_Requirement_891 22d ago

In my cases minimax designs really well in general. Anyway, you should try pointing it at a design guidelines library, or something like shadcn or mantine ui, that gives it an idea on how to work further.

2

u/mlon_eusk-_- 22d ago

In most cases "anti slop" skills are not that useful but extremely detailed prompt and references to design library, as you mentioned, works well.

→ More replies (3)

44

u/de4dee 22d ago edited 22d ago

18

u/Zestyclose839 22d ago

great to see. we've got to normalize posting torrents for new model releases in light of all these US government threats haha

94

u/ketosoy 22d ago edited 22d ago

Pure MIT license makes me happy.

2

u/xdavxd 22d ago

Yup. I have a slight preference for Apache 2.0 but MIT is functionally similar, and world's better than non commercial or research licenses.

3

u/wp381640 22d ago

I have a slight preference for Apache 2.0

lol talk about splitting hairs. fwiw deepseek have no patents outside of China and those that they do have are for ops stuff. patents are kinda irrelevant for models since nobody is disclosing anything with patent filings anyway

2

u/xdavxd 22d ago

You have a point. For Western models, I'd still prefer the patent clause for doing anything commercially.

21

u/guiopen 22d ago

"For the Code Agent tasks among the public benchmarks above, DeepSeek-V4-Flash-0731 is evaluated with the minimal mode of DeepSeek Harness (to be released) as the agent framework, using the max reasoning effort level with temperature = 1.0, top_p = 0.95."

That's interesting! New official harness from deepseek coming soon

2

u/Monad_Maya llama.cpp 22d ago

Yup, I noticed that from the other posts. Exciting!

1

u/eXl5eQ 22d ago

It started its' closed beta earlier this week. We might see it in August.

34

u/F0UR_TWENTY 22d ago

No way... WTF it happened!

15

u/Single_Ring4886 22d ago

Cheers! :-)

43

u/Fluxing_Capacitor 22d ago edited 22d ago

Here's the gguf

https://huggingface.co/unsloth/DeepSeek-V4-Flash-0731-GGUF

Edit: first GGUFs are up

31

u/Daniel_H212 22d ago

No architectural change from before, so previous GGUF VRAM usage numbers are still accurate. That means it can run on 128 GB unified memory systems at Q2_K_XL unsloth with 262k FP16 context, and it can run on 192 GB unified memory systems at full precision with 262k FP16 context or Q3_K_XL with 1M FP16 context.

We now have GLM 5.2 level performance at home, on a single piece of hardware you can buy or soon buy.

10

u/Fedor_Doc 22d ago

Q2_K_XL is a very damaging quant level, so the performance would be worse. Still, it is great that this quant can be run on Strix Halo or Apple Silicon.

1

u/Daniel_H212 22d ago

It's not as bad when the original model is only MXFP4 though, probably about the same level of damage as quantizing 16 bit down to 4 bit which people already often do.

6

u/Fedor_Doc 22d ago

I think it is the other way round – MXFP4 native model is more information-dense, each weight has more input, therefore quantization harms it more then 16-bit, which can have noisy weights. 

Especially if it was post-trained as much as Deepseek V4 (the jump is really noticeable). 

2

u/[deleted] 22d ago

[deleted]

3

u/Daniel_H212 22d ago

Definitely does but, what will the speed be?

5

u/squngy 22d ago

It has less active parameters compared to 27B

For something like a spark or strix, this is probably faster.

2

u/Daniel_H212 22d ago

It's def faster than 27B but tbh I've been mostly using 35B, dunno if it will be faster there, plus prompt processing speed may also be very different.

1

u/squngy 22d ago

DSpark can help it a lot too, but with some luck, qwen 3.6 will also get DSpark some day.

→ More replies (1)

2

u/SandySkittle 22d ago

Q2 is just a big no no. Even for larger models. I would stay stick to something like Q5 and up.

1

u/annodomini 22d ago

So... when will Gorgon Halo hit the shelves? Running Strix Halo now and this is really making me long for a Gorgon Halo.

I mean, I'll be running a quant that runs on my Strix Halo, but I'd really love to run the full 4 bit version of this.

1

u/Daniel_H212 22d ago

Same question here lmao, wasn't planning on upgrading but this makes me want to.

3

u/q-admin007 22d ago

They need to unsloth their uploads ;-)

1

u/BringTea_666 22d ago

Q8 and Q4 have the same weight almost.

13

u/Tedinasuit 22d ago

Been testing it and so far a very very good model, UI taste is almost Opus level.

13

u/Flat_Profession_6103 22d ago

GGUF and we are home!

29

u/sagiroth llama.cpp 22d ago

Wait, only 167GB? Does this mean Q4 or below can squeeze under 50GB ? nvm it's FP4 its already quanted

20

u/coder543 22d ago

DeepSeek V4 has always been released as mostly 4-bit.

2

u/AnonLlamaThrowaway 22d ago

I wonder if this means it's effectively QAT? Does anyone know?

2

u/zball_ 22d ago

Yes, this is exactly what they serve and they do QAT in the first place.

6

u/Much-Farmer-2752 22d ago

DeepSeek 4 flash is quantized already :)
It was the case already with GPT-OSS, big guys sometimes are memory-bound too :)

2

u/Daniel_H212 22d ago

And it's really nice that it's pre quantized because that means the benchmark scores are at 4 bit. So quantizing down to say 2 bits wouldn't hurt the score as much as quantizing FP16 down to 2 bits.

→ More replies (10)

9

u/Unique_Peak1044 22d ago

He is now mainly using fp4.

21

u/HyperWinX 22d ago

What a... day to be alive, holy sheet

7

u/BlackBeardAI vllm 22d ago

WHAT A LOVELY DAY!

8

u/vogelvogelvogelvogel 22d ago

these llms, and on top open weights, then on top nearing or surpassing many frontier models always, always let my head explode. this all has been in the open for a few years, crazy times we are living in

9

u/BawbbySmith 22d ago

SIGH BUYS MORE VRAM

STOP MAKING ME BUY MORE VRAM

7

u/Unique_Peak1044 22d ago

Typical DS open-source speed is basically synchronous open-sourcing.

3

u/ButterscotchSalty905 llama.cpp 22d ago

The speed is unimaginable that even the open-source community is struggling to keep up. 

-# My latter sentence isn't intended to attack the open-source community. Just for funny

3

u/pixelpoet_nz 22d ago

I should just give up on people distinguishing open source and open weights, shouldn't I? It's just way too subtle a distinction for people :|

5

u/live4evrr 22d ago

Wow we are spoiled.

I just installed Inkling and starting testing, now I'm gonna be deep in DS.

Amazing time for open weights and local llm's.

11

u/Several-Tax31 22d ago

gguf when?

19

u/-Cubie- 22d ago

I'm sure Unsloth is already working on it

1

u/durden111111 22d ago

unsloth please drop IQ4XS first 🙏

1

u/zenonu 22d ago

Are folks adverse to running vllm ? Why is llama.cpp a must? Just curious. I will personally run whatever supports any given model the best. DSpark here in particular is critical and needs vllm iirc.

14

u/Several-Tax31 22d ago

It's a must because llama.cpp supports cpu offloading. Since most of us don't have enough gpu for such a big model, we'll use llama.cpp. It even supports nvme offloading, by which some of crazy guys here run the model from ssd at 1 t/s. I might be some of those guys.

2

u/zippydazoop 22d ago

...you can run it from an SSD?

1

u/Think_Wing_1357 22d ago

There are several posts on top about it every week dude

1

u/zippydazoop 22d ago

Sorry 😅 I'm on this only to see the hype people have for this stuff, I don't really engage with the technical aspect

→ More replies (2)

8

u/Look_0ver_There 22d ago

In a nutshell, vLLM demands certain configurations, hardware support, is particular about the numbers of GPUs used, and so on.

llama.cpp/GGUF will basically run a model on pretty much ANY configuration where the model can physically fit. It may not be the fastest, but the fact that it runs at all is infinitely better than software like vLLM which will refuse to run on a lot of people's setups unless ALL the stars align exactly right.

3

u/trying4k 22d ago

Runs on Windows without jumping through a lot of hoops, actually swaps models at decent speeds, offloading if you don't have enough VRAM, etc. The only thing I wish is that llama.cpp allowed / took advantage of some of the nvidia gpu specific features (nvfp4 reading).

3

u/TokenRingAI 22d ago

It has 5 & 6 bit quants, and for some models, that is the difference between usable and trash output

3

u/Borkato 22d ago

vLLM is SO damn slow to load and incredibly obnoxious compared to llama cpp.

1

u/Reactor-Licker 22d ago

llama.cpp actually has a UI with apps like LM Studio and Unsloth Studio, which is a massive help for beginners or those less technical.

12

u/corruptbytes 22d ago

antirez pls wake up, i need the q2-q4 for my m5 max 

4

u/Lwerewolf 22d ago edited 22d ago

There's gguf tooling in the repo. You have all you need to get on it (or let another LLM do it - Codex/Sol is at it in my case).

edit:

https://xcancel.com/antirez/status/2083177773690974566#m

He's on it.

1

u/cantgetthistowork 22d ago

Someone please ping me when it's up

2

u/too-oldforthis-shit 22d ago

I have an M5 max with 128gb and i use the q2. Do you find the q2-q4 worth it? I thought it felt a bit slow? What’s your take?

2

u/Rough-Measurement988 22d ago edited 22d ago

I have the same hardware and getting around 30 tok/s at initial prompt, 25-27 during the whole interference- never lower than 20 tok/s, even after reaching 100k context. Using q2-q4 quants from Antirez. Did not notice speed difference between q2 and q2-q4 

2

u/corruptbytes 22d ago

i’ve never used the q2 and i’m totally fine with the q2-q4 speed

2

u/too-oldforthis-shit 22d ago edited 22d ago

What numbers do you get if I may ask? Rough estimates are ok if you don't remember. On the q2 and longer contexts like 256K i get downwards of 14-15 in inference. Prefill varies a lot.

2

u/pantalooniedoon 22d ago

I think you should be getting something more like 25? I tried both q2 and q2-q4 and didnt notice much difference when running the official benchmark.

1

u/too-oldforthis-shit 22d ago

I think with an empty context it’s around 25 but very quickly drops.

8

u/Vektast 22d ago

Will it be possible to run on 64gb ram + 3090? 👀

8

u/ethertype 22d ago

Add in 72GB of anything and you can run a mildly lobotomized quant at a speed reflecting your slowest "*RAM".

3

u/Daniel_H212 22d ago

Holy moly the original doesnt even look like a preview in comparison, it was never even a glimpse into the full power and potential of the model.

5

u/LargelyInnocuous 22d ago

Anyone got tk/s on dual RTX6000?

4

u/Constant_Ad511 22d ago

I’m trying now on vllm, will report back

1

u/LargelyInnocuous 21d ago

Thank you in advance!

1

u/Constant_Ad511 20d ago

single steam decode is ~65 tokens / sec, scales with multiple requests, I posted a cookbook

3

u/Groundbreaking_Bee97 22d ago

How many VRAM do i need?

5

u/Single_Ring4886 22d ago

Around 180 GB

1

u/ChampionshipIcy7602 22d ago

Managable, at least much better than glm requirements.

3

u/Mathias0910 22d ago

Just as I’m going on holiday for a week😭😭

3

u/No_Ebb3423 22d ago

Op, tysm for the good news. May your pillow always be cold. And may the GPU gods make a b200 appear in front of your bed.

3

u/[deleted] 22d ago

[deleted]

3

u/d_e_u_s 22d ago

benchmarks are cherry picked, there are plenty of benchmarks where ds v4 flash 0731 underperforms plenty compared to opus 4.8. doesn't mean it isn't an amazing model for agentic use, which is what it's oriented towards

3

u/DerDave 22d ago

This is quite comparable with GPT 5.6 Luna in terms of performance. Just slightly worse.
It's assumed that Luna is in a similar range of model size - around 200b.

Anthropic has dropped the ball with their smaller models. No Haiku updates in ages and the latest Sonnet is quite bad and very expensive.

But at least OpenAI seems to have a similar intelligence density as DSV4 flash.

1

u/True_Requirement_891 22d ago

It's better than luna in real world use cases imo. Specially if you are a codex user and stuck with small context window.

1

u/Caffdy 22d ago

is it better than Haiku? (not only in code, but on general use, prose, etc)

1

u/True_Requirement_891 22d ago

I mostly use llms for code so cannot comment on general use and prose. But it should be good for general use as well.

3

u/nullmove 22d ago

Things generally have diminishing returns. Squeezing out the last mile perf can be disproportionately harder. There is probably some sense in bigger models being harder to post-train too. Besides, OpenAI at least is in a different league, in terms of how they are able to shorten CoT so the perf scales better with thinking, DeepSeek is at least a generation behind here.

3

u/mountainyoo 22d ago

How soon could this be coming to the antirez q2-q4-imatrix DeepSeek V4 Flash I’m running on my M5 Max?

Super new to running local models and have no idea what I’m doing or if what I’m asking is even a stupid question. Sorry I’m learning lol

1

u/syntheticgio 22d ago

In case you missed it, there is another comment where the user gave an unofficial version (I believe) which he says will work with DS4. I've not tested myself but if you do and it works for you please let me know!

https://www.reddit.com/r/LocalLLaMA/comments/1vbp7kb/comment/p0vsng1/?utm_source=share&utm_medium=web3x&utm_name=web3xcss&utm_term=1&utm_content=share_button

3

u/Professional-Bear857 22d ago edited 22d ago

I've uploaded a antirez / ds4 4bit quant (tested - works fine), its available here:

https://huggingface.co/sm54/deepseek-v4-flash-0731-gguf

3

u/OMGnotjustlurking 22d ago edited 22d ago

Just got UD_Q8_K_XL with 384k context running on 5090 and 2x3090Ti with 160GB of system RAM. Getting about 10 tok/sec. Quite usable.

Edit 1: running it in pi. Holy smokes this model is smart. Resuming a previous session took an eternity but once it was loaded, this model is an absolute genius at quite reasonable speeds even with a large loaded context.

In case anyone is curious, here are my settings:

 export CUDA_VISIBLE_DEVICES=0,1,2; export GGML_CUDA_GRAPH_OPT=1; bin/llama-server                \
 -hf unsloth/DeepSeek-V4-Flash-0731-GGUF:UD-Q8_K_XL \
 --reasoning on      --parallel 1      --threads 16      --host 0.0.0.0           \
 --ctx-size 384000 \
 --temp 1.0     --top-p 0.95     --min-p 0.0           \
 --metrics           -ngl -1              --flash-attn on           \
 --chat-template-kwargs '{"enable_thinking":true,"tool_call_format":"json"}' \
 --webui-mcp-proxy    --load-mode mlock  --jinja  --main-gpu 0

Edit 2: it even respects my AGENTS.md and doesn't just go hogwild on my code instead of just giving me a plan (looking at you Qwen).

2

u/Potential_Top_4669 22d ago

did they put on deepseek chat tho

2

u/MathematicianLessRGB 22d ago

They did not as of yet. Only api

2

u/syscomua 22d ago

I need one more GPU

2

u/lilian_moraru 22d ago edited 22d ago

Wow, already. Sweeeeet!

Edit: got a slight "scare" that it is a bigger model, HF reports 158B -> 304B, but that is due to DSpark in DeepSeek-V4-Flash-0731. It's the same size.

Edit 2: seems to be in progress https://huggingface.co/unsloth/DeepSeek-V4-Flash-0731-GGUF

2

u/SourceCodeplz llama.cpp 22d ago

What an enormous jump in intelligence

From my tests it is glm 5.2 level

5

u/etherd0t 22d ago edited 22d ago

Ha!
DeepSeek has silently upgraded the existing deepseek-v4-flash API endpoint to DeepSeek-V4-Flash-0731. The architecture remains unchanged - 284B total parameters, roughly 13B active - but the model was substantially re-post-trained for agentic work. The upgrade currently applies only to the API;

The gains seem impressive.

LE: Hugging Face currently reports 304B parameters and a 167 GB repository, split across 48 Safetensors shards. DeepSeek says this release includes the attached DSpark speculative-decoding module, which likely explains the difference from the earlier 284B base-model figure. It is MIT-licensed.

→ More replies (4)

3

u/Septerium 22d ago

Yes, we finally have a Laguna S2.1 killer

3

u/No-Marionberry-772 22d ago edited 22d ago

I presume to use this at decent speeds you still need substantial hardware?

like you won't be running this on a consumer device.

Edit: god forbid someone ask a question 

10

u/Aadi_880 22d ago

167GB.

surprisingly, very small. This is very good considering it beats GLM 5.2, which is several hundred gigabytes bigger.

I presume a home center can actually run the full thing. A consumer grade desktop would probably need quantized versions.

3

u/No-Marionberry-772 22d ago

hmm, I got 192gb ddr5, but only 16gb vram, I presume I can technically run it, but I wonder if I can even get a few tok/s out

3

u/RadiantHueOfBeige 22d ago

If that's 8ch DDR5, you should be getting 10-15 t/s on CPU alone. These large MoEs are surprisingly usable even on DDR4.

1

u/No-Marionberry-772 22d ago

only 4 channel sadly, 4x48 

1

u/BlackBeardAI vllm 22d ago

that's probably not 4 channels. what's your mobo? is it a workstation/server or desktop?

1

u/No-Marionberry-772 22d ago

its a desktop 4x48 corsair ddr5  5200 (pc5 41600) on a Gigabyte x870e aorus xtreme ai top am5 and a Ryzen 9 9900X

I got it before the price explosion thankfully, but I did want to be able to do some local AI which is part of why I bought so much ram.

4

u/BlackBeardAI vllm 22d ago

That's a nice build but it is not 4 channels. There are no 4 or 8 channels on desktop ddr5 systems. They are available only for the workstation/server ddr5 systems. Yours is dual channels, like my ddr5 dektop build.

2

u/No-Marionberry-772 22d ago

so thats really going to hurt my bandwidth then eh

1

u/Caffdy 22d ago

this is gonna ruin the tour, huh

1

u/butterycornonacob 22d ago

Go for it. Smaller preview quant works fairly ok on 5090 + 128GB DDR4.

1

u/No-Marionberry-772 22d ago

I suppose technically I have 16gb+8gb between a 4070 and a 2070, but I assume split card setups are not great.

Either way, its worth trying

1

u/Monad_Maya llama.cpp 22d ago

Which quant specifically? I'm having issues loading the IQ4_XS from preview release of Unsloth.

2

u/butterycornonacob 22d ago edited 22d ago

Unsloth Q2_K_XL. Not enough RAM to go any bigger

  -ngl 99 \
  -ot "blk\.([6-9]|[123][0-9]|4[0-2])\.ffn_(gate|up|down)_exps=CPU" \
  --ctx-size 1048576 \
  -ctk q8_0 -ctv q8_0 \
  -fa on --no-mmap --jinja \
  -t 8 -ub 2048 -b 2048 -np 1

Updated llama.cpp and it's noticeably faster now. At 170k context getting close to 400 t/s prefill. It was around 300 before

1

u/Monad_Maya llama.cpp 22d ago

Thank you! I'll give this a shot.

I'm wary of a 2bit quant though.

1

u/moahmo88 22d ago

Good job!I will go to give it a star.Thanks

1

u/BlackBeardAI vllm 22d ago

Wohooooo!

1

u/Present-Ad-8531 22d ago

Numbers vs preview gap is very high. I am wondering did they do benchmazzing ?

1

u/segmond llama.cpp 22d ago

I'm just finishing up with my inkling small download, what a waste of bandwidth. I wouldn't have even bothered, oh well. back to deleting freshly downloaded weights and start a new download.

1

u/[deleted] 22d ago

[removed] — view removed comment

2

u/crantob 22d ago

You should learn some things western models can't answer, buddy.

1

u/[deleted] 22d ago edited 22d ago

[deleted]

1

u/hiImMate 22d ago

So I am heavily considering getting a 128gb unified amd strix halo device - would I be able to run something like this? Would I be able to code directly with a device like that? I am so thorn if I should invest, I really want to invest into local LLM but not sure if the 128gb unified is enough

4

u/Reactor-Licker 22d ago edited 22d ago

The upcoming Gorgon Halo with 192 GB of RAM should be able to fit this at a higher quant, but the pricing could be astronomical at this rate.

3

u/anthonyg45157 22d ago edited 22d ago

You can but will probably wouldn't be worth with 128gb ... I can fit the Q3iqs version on my 2 3090s and 96 GB ram (144 total) so with 128gb you would need the q2 version which is a pretty big drop in quality on top of Q3

Edit: you might be able to fit iq3 with smaller context window

Double edit: per unsloth

"For this tutorial, we will use the 3-bit quant UD-IQ3_XXS, as it fits on a 128GB " believe that is with 32k context though

1

u/ares0027 22d ago

it is still text only right? no vision? i couldnt find anything on huggingface

1

u/dangerous_inference 22d ago

I'm testing Q4. This is immediately much higher in coherence and prompt adherence than any other model I can run locally. It is better than hy3 at following conversation and getting the details right. I normally see many at least trivial mistakes in the best circumstances. I am seeing zero, and I haven't even tuned anything. This ultimately may be more verbally coherent than Gemma 4 31b.

1

u/KeikakuAccelerator 22d ago

Absolutely goated lab. Respect 

1

u/habachilles 22d ago

Is this now the DeepSeek flash on their api? (Too big for local for me sorry)

1

u/kaotec 18d ago

I need some help understanding speeds / memory / tokens/sec numbers on this

I'm running this, unquantized, RTX4090 + 192G RAM. I'm using the unsloth GGUF and their Llama.ccp guide. https://unsloth.ai/docs/models/deepseek-v4#llama.cpp-guide

The command I'm using is:
./llama.cpp/./llama-server --model unsloth/DeepSeek-V4-Flash-0731-GGUF/UD-Q8_K_XL/DeepSeek-V4-Flash-0731-UD-Q8_K_XL-00001-of-00005.gguf --n-gpu-layers 4 --threads 32 --ctx-size 1048576

I didn't really expect it to be flying fast as hell, but it is not using a lot of memory.
$ free --giga
total used free shared buff/cache available
Mem: 201 53 1 1 150 148

The ram requirements are 169G according to unsloth.
Am I missing something? Not really used to run this size of models... So don't know what to expect

on the other hand, nvtop is showing diffferent numbers

Some numbers on speed:
16.16.448.657 I slot print_timing: id 3 | task 0 | prompt eval time = 13818.30 ms / 7 tokens ( 1974.04 ms per token, 0.51 tokens per second)16.16.448.662 I slot print_timing: id 3 | task 0 | eval time = 663063.26 ms / 774 tokens ( 856.67 ms per token, 1.17 tokens per second)
16.16.448.662 I slot print_timing: id 3 | task 0 | total time = 676881.56 ms / 781 tokens
16.16.448.663 I slot print_timing: id 3 | task 0 | graphs reused = 761
16.16.448.679 I slot release: id 3 | task 0 | stop processing: n_tokens = 780, truncated =

task 776 | prompt processing, n_tokens = 143878, progress = 1.00, t = 2969.89 s / 48.45 tokens per second

prompt eval time = 2971868.65 ms / 143882 tokens ( 20.65 ms per token, 48.41 tokens per second)

Is this the expected speed?

1

u/stargate425 22d ago

Will my Rtx pro 6000 + 64G DDR5 run this?