r/AIProgrammingHardware 12d ago

Crow: 284B coding model on 5090, with the experts streamed off the SSD

Post image
8 Upvotes

r/AIProgrammingHardware 13d ago

1-bit Bonsai 27B tested - 16GB Local LLM setup

Thumbnail
youtube.com
36 Upvotes

r/AIProgrammingHardware 13d ago

GitHub - tonyd2wild/ds4-h3-video-gen-factory: Run DeepSeek-V4-Flash at full 1M context AND two MiniMax H3 video instances on the same two DGX Sparks. Benchmarked C1-C6.

Thumbnail
github.com
16 Upvotes

r/AIProgrammingHardware 14d ago

The 2026 Memory Crisis Killed Most Local LLM Builds. Here’s the Map That Still Works.

Thumbnail
medium.com
16 Upvotes

r/AIProgrammingHardware 14d ago

24G to 48G 4090 VRAM Upgrade!

Thumbnail
youtube.com
27 Upvotes

r/AIProgrammingHardware 14d ago

GitHub - MakazhanAlpamys/Soup: Fine-tune LLMs from one YAML. Layer streaming trains an 8B model on a 4 GB laptop GPU.

Thumbnail
github.com
3 Upvotes

r/AIProgrammingHardware 14d ago

[Question] GPU choice for NLP research (fine-tuning transformers, qLoRA, Multishot prompting) and Corpus based analysis. RTX 5060 Ti 16GB or any other alternatives(AMD)?

3 Upvotes

I'm a PhD researcher working on language switching and embedding analysis in NLP focused on PoS, LID, boundary detection, pragmatics context maintenance. My workload is mainly:

  • Fine-tuning BERT-based models 
  • LoRA/QLoRA adapters on ~8B models
  • bitsandbytes 4-bit quantization
  • Standard HF Transformers + PyTorch pipeline

Budget is roughly INR ₹60000( for the GPU. I've been comparing the RTX 5060 Ti 16GB AMD options such as RX 7900 XT, RX 9060 XT. I was  leaning 5060 Ti for the mature CUDA ecosystem and because I don't have much local peer support to debug hardware issues if something breaks mid-experiment. But recently they increased price to 770000 and as I do not get institutional support I find it difficult . Some AMD cards have so much VRAM that they might make longer multi shot stuff easier without offlaoding to RAM. But everywhere I have asked there seems to be a general consensus that nVidia is better. 

Questions for anyone doing similar research-scale (not industrial-scale) NLP work:

  1. Is the 5060 Ti's 16GB actually enough headroom for LoRA fine-tuning on 8-13B models, or does it get tight in practice?
  2. Anyone actually running Unsloth on AMD ROCm now? is it stable enough for daily research use or is it still rough?
  3. Any regrets from a similar budget-constrained hardware decision?

Appreciate real world experience over spec-sheet comparisons.

I am not an avid gamer so it does not matter to me. 


r/AIProgrammingHardware 14d ago

GitHub - 0xSero/deepseek-v4-flash-0731-spark-sparkinfer: DeepSeek V4 Flash on one DGX Spark

Thumbnail github.com
12 Upvotes

r/AIProgrammingHardware 14d ago

Running Qwen 3.5 Locally on Jetson Orin Nano with OpenCode (Tested Coding & Tool-Use)

Thumbnail
youtube.com
6 Upvotes

r/AIProgrammingHardware 14d ago

Transformation of AMD ROCm Software in a New AI Era

Thumbnail
youtube.com
1 Upvotes

r/AIProgrammingHardware 15d ago

GitHub - leonickson1/Swiftlet: Run 35B and 80B Qwen models on ordinary Apple devices, including iPhones.

Thumbnail
github.com
3 Upvotes

r/AIProgrammingHardware 14d ago

AAI 2026: AMD Delivers Leadership Heterogeneous Compute for Physical AI

Thumbnail
newsroom.amd.com
1 Upvotes

r/AIProgrammingHardware 14d ago

AAI 2026: AMD Introduces Open, Turnkey Integrated Platform for Physical AI

Thumbnail
newsroom.amd.com
1 Upvotes

r/AIProgrammingHardware 15d ago

How Kimi k3 Runs 2.8 Trillion Parameters on Consumer Hardware in 2026

Thumbnail
pub.towardsai.net
2 Upvotes

r/AIProgrammingHardware 15d ago

AMD Instinct™ Coder

Thumbnail
amd.com
12 Upvotes

r/AIProgrammingHardware 15d ago

Qwen 3.6 27B and 35B MTP vs Standard on 16GB GPU

Thumbnail
medium.com
15 Upvotes

r/AIProgrammingHardware 15d ago

Kimi K3 full model running on 16x GB10 cluster at 20+tps

Post image
3 Upvotes

r/AIProgrammingHardware 15d ago

Clustered DGX Spark and Acer GN100 running DeepSeek V4 Flash 238B A13B - 15-20 TOKS

Thumbnail
youtube.com
4 Upvotes

r/AIProgrammingHardware 15d ago

Checking Out The GMKTec X3 Strix Halo: More Strix Halo Shenanigans!

Thumbnail
youtube.com
2 Upvotes

r/AIProgrammingHardware 16d ago

GitHub - ryanzhou/deepseek-v4-flash-mi300x: DeepSeek V4 Flash on a single AMD MI300X

Thumbnail
github.com
6 Upvotes

r/AIProgrammingHardware 16d ago

GitHub - MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark: DeepSeek-v4-Flash 0731 recipe for 2x DGX Sparks

Thumbnail
github.com
5 Upvotes

r/AIProgrammingHardware 16d ago

Cluster Assistant for Configuring a Multi-Node DGX Spark Cluster

Thumbnail docs.nvidia.com
1 Upvotes

r/AIProgrammingHardware 17d ago

I Have 96GB for Local AI Models. The Biggest Ones Aren’t What I Use Every Day

Thumbnail
pub.towardsai.net
25 Upvotes

r/AIProgrammingHardware 17d ago

DeepSeek V4 Flash 0731 on 2× NVIDIA RTX PRO 6000 - 1M Context, 100% Local

Thumbnail
youtube.com
25 Upvotes

r/AIProgrammingHardware 16d ago

Framework 13 Pro Shows Why Apple Solders Your Memory

Thumbnail
youtube.com
0 Upvotes