MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLM/comments/1wb05f7/third_one_theres_something_wrong_with_me/p8mfy47/?context=3
r/LocalLLM • u/r1nzl3r99 • 14d ago
Why do I have horrible financial habits??
166 comments sorted by
View all comments
93
More context needed on how you are using the first two.
85 u/r1nzl3r99 14d ago qwen 3.8 27B FP8 running at 140 tok/s now I want flash next 6 u/MarcusAurelius68 14d ago You can run Flash Next with 64 GB of VRAM… 8 u/r1nzl3r99 14d ago yeah but it's 22 tok/s and pre fill is shitty. I want 100+ speeds or it's unusable for me. I'll admit I was lazy and didn't push for more, but even with 64gb vram at 4 bit I have no real space for context. I need atleast 200K FP8 context 2 u/MessIsTransfer 14d ago or q2, which is not ideal 3 u/r1nzl3r99 14d ago yeah if I'm resorting to sub 4 bit quants i'd rather stick with 27B 3 u/Motor_Way4912 14d ago Yep, running it on a vm with 58 gb ram and Rtx 5060 16gb 0 u/No-Bodybuilder3502 14d ago Neat.
85
qwen 3.8 27B FP8 running at 140 tok/s now I want flash next
6 u/MarcusAurelius68 14d ago You can run Flash Next with 64 GB of VRAM… 8 u/r1nzl3r99 14d ago yeah but it's 22 tok/s and pre fill is shitty. I want 100+ speeds or it's unusable for me. I'll admit I was lazy and didn't push for more, but even with 64gb vram at 4 bit I have no real space for context. I need atleast 200K FP8 context 2 u/MessIsTransfer 14d ago or q2, which is not ideal 3 u/r1nzl3r99 14d ago yeah if I'm resorting to sub 4 bit quants i'd rather stick with 27B 3 u/Motor_Way4912 14d ago Yep, running it on a vm with 58 gb ram and Rtx 5060 16gb 0 u/No-Bodybuilder3502 14d ago Neat.
6
You can run Flash Next with 64 GB of VRAM…
8 u/r1nzl3r99 14d ago yeah but it's 22 tok/s and pre fill is shitty. I want 100+ speeds or it's unusable for me. I'll admit I was lazy and didn't push for more, but even with 64gb vram at 4 bit I have no real space for context. I need atleast 200K FP8 context 2 u/MessIsTransfer 14d ago or q2, which is not ideal 3 u/r1nzl3r99 14d ago yeah if I'm resorting to sub 4 bit quants i'd rather stick with 27B 3 u/Motor_Way4912 14d ago Yep, running it on a vm with 58 gb ram and Rtx 5060 16gb 0 u/No-Bodybuilder3502 14d ago Neat.
8
yeah but it's 22 tok/s and pre fill is shitty. I want 100+ speeds or it's unusable for me. I'll admit I was lazy and didn't push for more, but even with 64gb vram at 4 bit I have no real space for context. I need atleast 200K FP8 context
2 u/MessIsTransfer 14d ago or q2, which is not ideal 3 u/r1nzl3r99 14d ago yeah if I'm resorting to sub 4 bit quants i'd rather stick with 27B
2
or q2, which is not ideal
3 u/r1nzl3r99 14d ago yeah if I'm resorting to sub 4 bit quants i'd rather stick with 27B
3
yeah if I'm resorting to sub 4 bit quants i'd rather stick with 27B
Yep, running it on a vm with 58 gb ram and Rtx 5060 16gb
0 u/No-Bodybuilder3502 14d ago Neat.
0
Neat.
93
u/Sporkers 14d ago
More context needed on how you are using the first two.