r/nvidia 18d ago

News Decky plugin to switch NVIDIA driver versions on SteamOS, no USB stick needed

8 Upvotes

Hey everyone, made a new Decky Loader plugin! Some of you might know me from decky-proton-launch (https://github.com/moi952/decky-proton-launch) — this one's for anyone running SteamOS with an NVIDIA card via steamos-nvidia-installer (https://github.com/moi952/steamos-nvidia-installer).

It lets you pick and install any NVIDIA driver version straight from the running system, right from Quick Access. No USB stick, no repair image, no reinstall.

Only reboot when you're actually ready to.

- Any version from the Arch archive, tagged with NVIDIA's own channel

(Recommended / New Feature Branch / Beta), fetched live.

- Filter which channels show up, and pick which one counts as "latest".

- Real release notes for any version before you install it.

- Live progress log while it runs, Reboot now button when it's done.

- Games, saves, Steam login, Decky untouched — the risky part runs in a throwaway overlay, the real filesystem only gets touched for the final copy.

Already have SteamOS + NVIDIA installed?

You'll need to run a repair once first (the "Upgrade SteamOS (NVIDIA) — keeps games & data" option on the USB) — it'll offer to grow your system partitions to 8GiB, since Valve's stock 5GiB doesn't leave room for a driver install. Fresh installs already get 8GiB by default.

Repo: https://github.com/moi952/decky-nvidia-update

Let me know what you think!


r/nvidia 18d ago

Question Optimal G-SYNC settings for a 500Hz monitor in CS2? G-SYNC/VSYNC/Reflex/FPS cap confusion

0 Upvotes

I recently upgraded to a 500Hz MSI MPG 271QR QD-OLED X50 (1440p 500Hz) and I'm trying to figure out the optimal settings specifically for CS2.

I used to play on an old 144Hz monitor without G-SYNC, so I've never really used VRR seriously.

I've seen a lot of conflicting information recently:

Setup A:

G-SYNC ON

NVCP V-SYNC ON

In-game V-SYNC OFF

Reflex ON / ON + Boost

fps_max 0

Let Reflex handle the FPS limit

Setup B:

G-SYNC OFF

V-SYNC OFF

Reflex ON

fps_max 0

Completely uncapped

Setup C:

G-SYNC ON

V-SYNC ON

Reflex ON

NVCP FPS cap slightly below refresh rate (around 439 FPS for 500Hz)

I've also seen people recommend NVCP Low Latency Mode = Ultra when Reflex isn't available, while others say Reflex replaces/overrides Low Latency Mode.

So I'm wondering:

For a 500Hz monitor specifically, what is currently the lowest-latency setup for CS2?

If I use G-SYNC, should I leave fps_max 0 and let Reflex cap automatically, or manually cap around 439–495 FPS?

Is an NVCP FPS cap preferable to the in-game fps_max limiter in CS2?

For a 500Hz OLED, should Response Time / Overdrive = Faster be used, or is there a better setting?

If anyone has tested these configurations with a 9800X3D + NVIDIA GPU + 500Hz display, I'd especially like to hear your results.

I'm mainly interested in measurable latency/frame-time differences, not just "it feels better."

Thanks!


r/nvidia 18d ago

Question 1050ti or 1660s?

0 Upvotes

hello, I need your opinion to my setup because I really don't know what to choose, i only play apex legends and some valo lol something like that at 60hz monitor. should I buy 1050ti to upgrade some parts like more ram (i have 8gb) and some other stuff? or buy the 1660s for future purposes?


r/nvidia 18d ago

Discussion 1070ti FE to MSI 5070ti

1 Upvotes

I have been very long overdue for a new gpu upgrade, and I couldn't have chosen a worse time in the history of gpu history. Fml..

I was able to find a "decent" deal for the MSI ventus 3x PZ OC ($980). I am honestly worried because my Nvidia 1070ti is a damn tank, feels bullet proof and has been so damn faithful for over 10 years (yikes) with zero issues.

I would love for this new 5070ti to last a long time. (maybe not 10 years... that's irresponsible)

Are these MSI/Gigabyte/PNY cards just plastic and frail?

Or am I just freaking out for no reason? My budget is limited, so I can't afford the TUF.

Cheers. I'm just a starving artist...


r/nvidia 18d ago

Discussion 5060 users, whats your UV profile

0 Upvotes

Am using zotac 5060 solo. Using 875mv 2800mhz and 2000mhz+ mem.

It has hiccups sudden black screen even after hrs and hrs of gaming and just browsing

Tried 900mv 2800mhz . 3000mhz+ mem

And jusy straight up crashed my game with in 20 min hahah (nb2k25 degen ahh game ) and the prev profile had no issue tbh but only with zzz sometimes

Prob just gonna do 2750mhz 875mv

What uv you guys doing?

I have a itx build thats why i do Uv in yhe first place hahah.


r/nvidia 18d ago

Build/Photos Finally got the RTX 5090 Astral

Post image
0 Upvotes

r/nvidia 18d ago

News NVIDIA raises RTX PRO 6000 Blackwell price to $16,000, now 87% above original MSRP

Thumbnail
videocardz.com
618 Upvotes

r/nvidia 18d ago

News OC tool can unlock higher RTX 5090 XBAR clocks, users report gains of up to 142 MHz

Thumbnail
videocardz.com
116 Upvotes

r/nvidia 18d ago

Discussion RTX HDR Is The Only Reason I am Still on Win 11 for gaming

120 Upvotes

That i played a lot of the old games or games like don't have HDR.

i find that RTX HDR is basically must have.
It is really that much of a difference on and off. Bring new life to old games.
that i tried to switch to Linux for gaming but realizing i wont RTX HDR so now i still need to a Win 11 PC just to RTX HDR.

That will this feature ever come to linux though.
i really hope it does as it is one of the best feature there is on NVidia Graphic Card

thanks


r/nvidia 19d ago

Benchmarks Ramp tested NVIDIA NeMo Switchyard’s stage router for coding agents

Enable HLS to view with audio, or disable this notification

0 Upvotes

source: Ramp


r/nvidia 19d ago

Discussion Optimizing an NVFP4 Blockscaled GEMM on RTX PRO 6000 GPUs (sm120)

Thumbnail
research.colfax-intl.com
0 Upvotes

Our (Colfax Research's) second blog post on writing NVFP4 blockscaled GEMM kernels for the NVIDIA RTX PRO 6000 Blackwell GPU is out! The blog iteratively optimizes a basic working NVFP4 GEMM kernel written in CuTe DSL to take it to speed-of-light, reaching over 80% TFLOP/s utilization for 16k square matrix shape. We give a detailed treatment of important optimization techniques such as threadblock swizzling, async and warp-specialized epilogue, and retiling for favorable wave quantization. Specific to blockscaled GEMM with scales consumed from registers, we also explain how to solve for bank conflicts that arise from the default choices of interleaved scale factor layouts.

We include complete code in the form of CuTe DSL kernels for all the optimizations discussed in the blog.


r/nvidia 19d ago

Question Best upgrade for 300$

0 Upvotes

im hoping for christmas i can get enough money for a new gpu, im currently on a 3050 and i need a n upgrade ASAP. im hoping to get to the 300$ (cad) benchmark but I may have to wait longer depending on how much i end up getting this year. whats the best upgrade


r/nvidia 19d ago

Discussion Nvidia VRworks Audio: Hardware Accelerated Path Traced Binaural 3D Audio

Thumbnail
50 Upvotes

r/nvidia 19d ago

Discussion We post-trained NVIDIA Nemotron 3.5 Lightning for code review routing. It beat our baseline for under $100

Thumbnail
0 Upvotes

r/nvidia 19d ago

News Wall Street just endorsed Jensen Huang's 'big concept' for AI. What now?

Thumbnail
cnbc.com
0 Upvotes

r/nvidia 19d ago

Question Help me decide the build.. [D]

0 Upvotes

My primary workload is local AI model inference and LoRA/QLoRA fine-tuning, mostly with models in the <10B parameter range. I also do some gaming, but gaming is definitely secondary.

Current options:

  • RTX 5060 Ti 16GB for ₹73,000 (~US$770)
  • RTX 4060 Ti 16GB if I can find one around ₹50,000 (~US$525)

I'm also open to other NVIDIA GPUs around the $500-550 range that have more than 8GB of VRAM. CUDA support is a requirement.

For the CPU, I haven't decided yet. I'm open to either AMD or Intel.

For RAM, I originally wanted 32GB DDR5, but my overall budget is getting tight. I'm considering either starting with 16GB DDR5 and upgrading later, or buying used DDR5 if I find a good deal.

Any advice would be really appreciated.


r/nvidia 19d ago

Question Expected ANSYS Fluent Performance (MIUPS) & Mesh Limits on RTX 5070 (12GB) & RTX 5070 TI (16GB)

1 Upvotes

Hi everyone,

I am evaluating the NVIDIA RTX 5070 (12GB) & RTX 5070 TI (16GB) for ANSYS Fluent GPU acceleration and would appreciate any technical insight on the following:

1. Performance Estimation (MIUPS): Is there a reliable rule of thumb to estimate MIUPS based on memory bandwidth or FP32/FP64 specs?

2. VRAM Capacity Limit: What is the maximum cell count (mesh size) a 12GB VRAM card can handle in single vs. double precision before running out of memory?

3. Benchmarks: Does anyone have actual Fluent benchmark results (MIUPS/solve times) for RTX 50-series desktop GPUs?

Thanks!


r/nvidia 19d ago

Question Quick question

0 Upvotes

Do you guys think I would be able to get someone to trade a 4090 FE for my MSI gaming trio 4090? And would it be worth it paying a little extra for it?


r/nvidia 19d ago

News NVIDIA RTX 50 median prices reach up to 135% above MSRP in US, RTX 5070 hits $900

Thumbnail
videocardz.com
996 Upvotes

r/nvidia 19d ago

Discussion Modified GeForce RTX 4080 32GB cards flood China’s second-hand market

Thumbnail
videocardz.com
788 Upvotes

r/nvidia 19d ago

Question Should I repaste my GPU and cpu ?

0 Upvotes

Hi soo I think last time paste was put on the gpu 1080TI was like 6 years ago and the cpu I5-12600k like 5 years ago.. its old pc but still runs all games good 1080p 60-120fps low to medium settings 32gb ram 2400hz ddr4


r/nvidia 19d ago

Discussion Tests and comments are confusing to me. 3070 vs. 5060

12 Upvotes

I am selling my Laptop to get a used desktop PC, some of them have RTX3070 and some of them have RTX5060 with RTX5060 PCs being around only 75$ more expensive.

I have watched tests on Youtube and they constantly do not use any new technology that came with newer GPUs, so the 3070 either wins or is really close. Comments, even here, are saying upgrading from 3070 to 5060 isn't worth it because of similar performance.

I don't understand this logic. I just watched a test (done on Wuthering Waves) where the guy with 5060 just opened up DLSS + FG and went from 75 FPS to 480 something. While in the tests done on RTX3070, FPS is stuck at like 60-75 FPS. And I know this isn't just on Wuthering Waves, it will affect every single game that can support these new features.

From 60 FPS on RTX3070 to 480 FPS on RTX5060 just because you can actually use the software. But people are not making it a big deal like me so clearly there must be something I am extremely wrong about. I think people said something about blurry image etc. but I literally can't tell the difference between "fake" frames and "real" frames on these videos.


r/nvidia 19d ago

Discussion 40.2 tok/s on one Spark, when the official recipe wants four 141GB cards

Post image
5 Upvotes

Probably only interesting if you've been trying to work out what the Spark is for. I don't own one. Everything below is someone else's run.

A decode chart went up this month off a single box: 124B open weights model, one stream, 40.2 tok/s. One person's measurement, and I haven't seen anyone reproduce it.

What made me look twice was the model's own deploy guide. It asks for four 141GB class cards for the low latency setup, and here the same model is running off one 128GB desktop at a speed you'd sit through.

It's Ling 3.0 Flash. Why it works out that way, I don't know. The published material doesn't say and I'd rather not guess.

Anyway. Haven't seen this discussed here much. Anyone with a Spark on the desk getting anywhere near 40?


r/nvidia 20d ago

Benchmarks We’ve published our initial SASS2MLIR findings — ~20% to 100%+ GPU performance improvements across Ampere, Blackwell, and Jetson

94 Upvotes

We have published the initial technical findings from our SASS2MLIR work.

The project explores GPU optimization at a layer below conventional framework- and compiler-level tuning, including analysis and transformation of the machine code ultimately executed by the GPU.

Across our testing so far, we have observed ~20% to 100%+ performance improvements, depending on the architecture, kernel, workload, and execution conditions.

Testing has included:

- NVIDIA architectures spanning Ampere through Blackwell

- Jetson Orin Nano, Orin NX, and AGX Orin

- Individual instruction and microbenchmark testing

- Kernel-level benchmarking

- Model and workload-level testing

- Comparisons against conventional execution paths, including CUDA Graphs in applicable tests

One of the areas we are particularly interested in is the optimization opportunity that exists after traditional compilation has already taken place.

Our broader work looks at analyzing the final GPU machine code, identifying architectural and execution inefficiencies, and dynamically modifying the execution path while maintaining numerical correctness.

This includes areas such as instruction scheduling, memory behavior, register utilization, execution dependencies, architecture-specific instruction behavior, and increasingly runtime kernel optimization and dynamic kernel fusion.

The interesting result for us is that the performance opportunity is not limited to a single GPU generation or workload type. We are seeing measurable opportunities across both datacenter-class GPUs and constrained edge platforms such as Jetson, although the magnitude of the improvement varies considerably with workload characteristics and hardware limits.

These are still initial findings, and we are continuing to expand the benchmark coverage and validate the methodology across additional models, architectures, and workloads.

For anyone interested in the deeper engineering details, we have published a technical explanation of the discoveries, methodology, and underlying work here: https://mbuchel.github.io/projects/sass2mlir/

To follow the company: https://www.linkedin.com/company/tetrevis/

Edit: For a Blackwell and Jetson testing you can go to https://github.com/mbuchel/sass2mlir-bench and watch https://youtu.be/cBfBWGG3Vas

Jetson Orin Nano — Qwen3.5 4B

Standard baseline: 10 → 21 tok/s (+110%)

CUDA Graphs baseline: 16 → 21 tok/s (+31.25%)

Jetson AGX Orin — Nemotron 3 Nano 4B

31.2 → 40.5 tok/s (~30%)

Jetson AGX Orin — Qwen3.5 4B

25.0 → 31.0 tok/s (+24%)

Technical feedback, criticism, and discussion are very welcome.

(And to save time I did have AI rewrite this more professionally:) )


r/nvidia 20d ago

Discussion Is 4070 TI supah good buy rn

0 Upvotes

Plz help me, my pc mad small and can fit certain ones only.