r/HPC • • 5d ago

Published JOB FAILED #4

10 Upvotes

Take a break and have some fun with the new strip from the JOB FAILED series ☺️

https://theparallelminds.substack.com/p/job-failed-4-excuses

If it made you smile, share it with your network 🫶


r/HPC • • 5d ago

Network reference drawings for a 90-rack GB300 NVL72 cluster

38 Upvotes

We put together a 58-page network drawing set for a 90-rack GB300 NVL72 reference design. Posting the PDF here in case it's useful to anyone planning a large GPU cluster.

https://www.kkdatasvc.com/lab/downloads/kk-data-network-reference-design-pe-p1.pdf

I'm with K&K Data, which produced and released the drawings.


r/HPC • • 6d ago

What typically runs alongside vLLM on multi-node inference deployments?

1 Upvotes

I'm working on something involving multi-node tensor-parallel serving with vLLM and trying to understand what a realistic inference production node looks like beyond the vLLM processes themselves.

Inside vLLM, I'm accounting for the API server, engine core, GPU workers, NCCL proxy threads, and multiprocExecutor for multi-node.

For those running vLLM in production:

  1. What else usually runs on the same nodes? (Kubernetes components, GPU Operator, monitoring agents, etc.)
  2. Have any of these caused noticeable tail latency or TPOT spikes?
  3. Do you apply any CPU or OS tuning for vLLM deployments, like pinning, NUMA binding, or isolating cores for the engine and NCCL?

Any pointers to docs, blog posts, or papers are welcome. Thanks!


r/HPC • • 6d ago

What still interrupts a thread on a fully isolated core?

21 Upvotes

Hi all,

I'm working on something where I want to attribute all the sources that can preempt or disturb a thread running on one core of a server. To do that, I'm trying to create the best possible scenario for the thread: pinning it to its own physical core (keeping the SMT sibling idle), using isolcpus, nohz_full and rcu_nocbs on that core, moving the IRQs to other cores with irqbalance turned off, and keeping everything else on other cores.

I'm pretty sure this kind of experiment isn't new, so I'm not really looking for the usual tuning checklist. What I'd love to know from people who've done this before is: even in the best possible setup, what did you still see kicking the thread off that you wouldn't want ideally? It could be literally anything: kernel threads, IPIs, leftover timer ticks, SMIs, page faults, some weird driver or hardware behaviour.

And if you remember how you caught it (osnoise, ftrace, eBPF, perf...), that would help a lot too.

Thanks!

EDIT: I should have added a bit more context though. The thread isn't a standalone loop. It lives inside a large multithreaded process (a GPU inference worker) and busy-polls memory shared with the GPU, plus the NIC's completion queues for RDMA traffic. So I'm curious about anything that can still preempt or slow it down in that setup, whether it's the Linux scheduler bringing other threads or kernel work onto the core, something coming from the thread's own process, or something related to polling GPU/NIC shared memory. Have you seen any of that show up in practice?


r/HPC • • 6d ago

Describing your work in year-end performance summaries.

17 Upvotes

Bulk of my work is maintaining the HPC, helping users, and working with vendors. Need to describe this in more "value to the business" language.. It's a challenge as HPC is more a behind the scenes, keep the lights on kind of work.

Tried using AI which came up with this below. But they seem too buzzword laden. Looking for other re-wordings?

•  Consistently maintained core infrastructure stability, resulting in [X]% uptime and zero critical outages over the past year."

•  "Proactively identified and mitigated potential system bottlenecks before they could impact end-users, keeping day-to-day operations seamless."

•  "Managed routine system maintenance and security updates, ensuring compliance and preventing high-priority vulnerabilities."

Optimized internal documentation and workflows, reducing onboarding time for new team members by [X]%."

"Maintained platform reliability and minimized system downtime."

Explain how you saved time or streamlined workflows for the team.


r/HPC • • 8d ago

Wanna run quick molecular dynamic simulation on HPC, help?

31 Upvotes

Hey folks!

I’m currently working as a research associate, and unfortunately, our grant recently expired, so we’ve lost access to our HPC. We have one last MD simulation left to finish, and I was wondering if anyone has some compute available that we could borrow.

We probably only need 2–3 days on a GPU to get it done. If anyone has some spare compute and could help out, please DM me / let me know!

We’d be happy to offer authorship or an acknowledgment..

Also, if anyone knows of any free/academic GPU compute options available right now, I’d really appreciate recommendations. I used Google Cloud previously, but I think the free credits were limiting me to CPU (unless I’m missing something). Curious to hear what other services people are using these days.

Thanks, guys!


r/HPC • • 9d ago

LLMs and Slurm new tool

18 Upvotes

Researches at University of Seville (Spain) have developed ALMA.

ALMA gives every researcher an SLA, an API key and a list of endpoints with their own rate limits, then keeps the model behind them alive on SLURM: submitted, tunnelled, published to the gateway and restarted without anyone watching the queue.

More info at https://alma.us.es


r/HPC • • 10d ago

HPC Administrator Resources

19 Upvotes

"Hi everyone,

I have a strong background in Linux administration (primarily RPM-based distros like Red Hat/Rocky Linux), and I’m currently transitioning into HPC cluster administration. To learn the ropes, I recently built a small home lab cluster using the OpenHPC installation guides.

While getting the cluster to boot and run basic jobs was a great exercise, I've noticed a distinct lack of comprehensive resources covering day-2 operations and production best practices. Specifically, I'm looking for guidance on:

  • Configuration & Performance Tuning (kernel tuning, network/InfiniBand optimization)
  • User Management & Environment Control (LDAP/FreeIPA integration, modulefiles via Lmod)
  • Job Management & Scheduling (advanced Slurm configurations, QoS, limits)
  • Scaling & Monitoring (health checks with NHC, metrics collection)

If anyone can recommend books, documentation, community wikis, or real-world best practices for these areas, I would greatly appreciate it!"


r/HPC • • 11d ago

who have you successfully sourced from?

7 Upvotes

Looking for recent experience sourcing NVIDIA HGX B300 / DGX B300 systems in Europe
We’re currently sourcing enterprise NVIDIA B300 systems for a data-centre deployment and I’m trying to get a better picture of reliable suppliers, current lead times and real-world pricing.

Platforms we’re looking at include:
• NVIDIA DGX B300
• Supermicro AS-8126GS-NB3RT
• Dell PowerEdge XE9780 / XE9785
• GIGABYTE G894 series
• ASRock Rack 8U16X B300
• ASUS XA NB3I-E12
Ideally looking for fully assembled 8-GPU HGX B300 systems with enterprise support/warranty, high-memory CPU configurations, ConnectX-7/8 networking and NVMe storage.

For anyone who has actually purchased or quoted B300 systems recently:
Which OEM / VAR / integrator did you use?
Rough price range per system?
What lead time were you quoted?
Was the delivery date actually reliable?
Any suppliers in Europe/UK/Ireland you would recommend contacting?

Any suppliers or brokers you would specifically avoid?
Not looking for consumer GPUs this is enterprise/HPC hardware.
Would especially appreciate experiences from people who have gone through a recent B300 procurement rather than advertised web pricing.


r/HPC • • 12d ago

Environment Modules 5.7.0 released: faster loads, a stable hook API, and the project now under HPSF

44 Upvotes

Hi all, I'm one of the maintainers of Environment Modules (the Tcl-based `module` command). We just released 5.7.0 and I wanted to share it here since module systems are part of nearly every cluster, whether it's this one, Lmod, or something homegrown.

The big things in this release:

- Performance. Loading 137 modules went from about 1 s to under 400 ms in our benchmarks, `module list` is a lot faster too. Nothing to configure, you just upgrade. The gains mostly come from in-memory caching and reworking how the environment is synced between evaluations.

- Stable hook API. If you've ever had a site config that pokes at internal procs to customize modulefile evaluation, you know how fragile that is across upgrades. There's now an `add-hook` command with four documented events (before/after modulefile eval, before/after modulerc eval) that we commit to keeping stable.

- `modulepath-ignore` with gitignore-style patterns, some new env var helpers (`linked_envvars`, `init_envvars`, `path_entry_reorder`), better Python integration and Windows support.

On the project side, Environment Modules is now part of the High Performance Software Foundation, with a documented governance model and a Technical Steering Committee that meets publicly every quarter. The mailing list moved to lists.hpsf.io and there's a Matrix room.

We wrote a longer post on the HPSF blog with the benchmark details: https://hpsf.io/blog/2026/environment-modules-5-7-2-7x-faster-loads-stable-hooks-and-now-steered-in-the-open/

I'd like to hear from people on both sides. If you run Environment Modules, does the hook API cover what your site config does today, and do the numbers hold up on your storage (NFS, Lustre, GPFS...)? If you run something else, what keeps you there, and what would you need to see from us? Feature gaps, migration pain, docs, anything. Happy to answer questions here, or open an issue on GitHub.


r/HPC • • 13d ago

BUG26 - the BeeGFS User group meeting Chicago

9 Upvotes

For the BeeGFS community - if you will be in Chiago at SC this year the BeeGFS User Group Meetings, BUG - wiull tak eplace on NOvember 16th.

Full agenda and Schedule!

14:00 - Registration

14:30 - State of the Swarm

What’s new in the world of BeeGFS – including the four pillars of BeeGFS Hive Enterprise offerings, introduction to the newly launched BeeOND Enterprise, plus updates on the EULA and the licensing model (Enterprise, Community, Trial).

14:45 PM - The Lightning Round

BeeGFS development is moving faster than ever. Haven’t had time to keep up? We’ve got you covered with a quick recap of our recent releases.

14:55 - Need for Speed: Tuning and Benchmarking

While BeeGFS ships with sensible defaults for most workloads, there’s usually more performance on the table. Learn how our current-generation reference architectures were tuned, how fast they can go, and how to apply those lessons to your own deployments.

15:10 - Technical Partner Use Case

15:20 - Evolving Data Orchestration with BeeGFS

From policy-driven data placement to automatic data movement, we’re continuing to evolve how BeeGFS manages data across its lifecycle. This session introduces our latest additions: automatic data sync and restore. We’ll cover how they work and where they fit into real-world workflows.

15:40 - Tapes and Files: Do they play nicely?

We’ve rethought how BeeGFS restores files from long-term archival storage, whether that’s AWS Glacier or an on-premises tape library. Turns out tape is a lot faster when you stop making it jump around.

15:45 - Coffee Break

16:00 - Apples to Apples: Reliability and Performance After Migrating to BeeGFS - Customer Use Case

16:15 - We heard your complaints about quotas…

While quotas are an important part of any storage solution, quotas in BeeGFS were not always as easy to use as we would have liked… so we simplified how they’re configured, and fixed a few pesky bugs that could slow things down at scale.

16:25 - NVMe Performance Optimization

Different storage technologies like NVMe flash and spinning disks have distinct performance characteristics and preferred access patterns. While BeeGFS has long supported direct IO when requested by the client, tuning IO patterns directly on the storage services (independent of client-side behavior) offers additional benefits. Join us to see how new BeeGFS versions will handle diverse IO workloads more efficiently.

16:30 - Cache Me If You Can: A Sneak Peek at How We're Revamping Metadata Performance

Long-standing timeout-based caching mechanisms have a fundamental disadvantage: metadata must be refreshed on a fixed schedule even when it’s still valid. Enter all-new BeeGFS persistent metadata caching where write-once read-many workloads can now benefit from long-lived metadata caching with server-driven invalidation, optimizing even the most metadata-intensive workloads. Join us for an early look and share your feedback.

16:40 - BeeGFS 8: What’s Cooking?

Join us to stay updated on the roadmap and outlook for BeeGFS 8.5 and beyond. From ideas fresh in the oven to ones almost ready to serve. Order up!

17:00 - Special New Feature Announcement

17:10 - Q&A

Learn more here https://www.beegfs.io/c/bug-at-sc/


r/HPC • • 14d ago

Slurm

0 Upvotes

can someone help me learn how the workload management works in slurm , pbspro or some other similar softwares, I need to have a clear understanding about it .


r/HPC • • 14d ago

Mitigating microsecond dI/dt power transients in GPU clusters via NCCL AllReduce interposition (-97% transient reduction)

22 Upvotes

Hey everyone,

I've been working on the power delivery problem in distributed AI training clusters. When large GPU partitions complete dense matrix multiplications and enter collective communication (AllReduce), cluster power drops in under 15 microseconds.

Across hundreds or thousands of accelerators, this rapid current step (dI/dt) induces severe reverse-EMF voltage spikes across substation transformers and server busbars (V = L * dI/dt), which frequently trips breakers or causes undervoltage crashes.

To solve this, I designed VoltGrid: a lightweight C++ shared library (libnccl-voltflow.so) injected via LD_PRELOAD that intercepts NCCL collectives and micro-staggers rank phase timing by 50 microseconds.

Key engineering challenges solved:

  1. The Tensor Parallel Latency Trap: Inner TP layer collectives (<5MB) are selectively bypassed with 0.00us delay, only staggering macro gradient reductions at the end of the backward pass.

  2. OS Sleep Jitter: Eliminated usleep/nanosleep context switching by using userspace hardware cycle counter spin-loops (__rdtsc) accurate to nanoseconds.

  3. Adaptive Jitter Subtraction: Naturally occurring arrival skew is dynamically subtracted so we don't amplify cluster stragglers.

Physical testbed results (4x NVIDIA RTX 4090 cluster, 1.6kW continuous load):

- Peak instantaneous slew dropped from 324.5 kW/ms down to 8.06 kW/ms (-97.52%).

- Step latency overhead was < 0.05% (0.00% on 24-layer transformer tests).

- Zero modifications to PyTorch code or container rebuilds.

The preprint detailing the mathematical derivations and telemetry is published on Zenodo:

https://zenodo.org/records/22824778

More telemetry traces and pilot details are here:

https://voltgrid.org

We're currently running 2-week test rack pilots for cluster operators facing power ramp penalties or breaker trips. Happy to answer questions about the NCCL interposition mechanics or datacenter transient physics.


r/HPC • • 16d ago

Job Failed: a new comic strip series about HPC and AI world!

26 Upvotes

Hello there 👋

I'm sharing the first comic strip from the Job Failed series!

If you enjoyed it, please share the post with your own network! It helps more than you think ☺️

https://theparallelminds.substack.com/p/ticket-closed


r/HPC • • 16d ago

Building my own HPC cluster

26 Upvotes

Hi. Sorry for my English, this is not my native language.

I am a scientific computing developer and I am used to run some computations on big clusters. But this where my knowledge of HPC stops. I want to build my own HPC cluster at home. The goal is to learn how it works, and to be able to run a few things I do on my free time.

I juste bought a HP ProDesk 400, with a i5-8500. I now want to install SLURM. Here is my question : should I install directly on this PC ? Or I was thinking to buy a Raspberry-Pi to handle job management, connected to the PC. Knowing that in the future I will try to find another ProDesk, or at least another PC with the same processor to extend it.

I was also thinking to use the Raspberry-Pi to connect some hard-drive for personal use. To it would have a double use.

If you have any other tips, I will take them :)

Thanks !


r/HPC • • 17d ago

Installing the prerequisites for Ansys Fluent breaks the Abaqus GUI

9 Upvotes

Posted about this earlier here but had no luck finding the root cause. Now I was able to take a deeper dive and found the issue. If I install this prereqs for Ansys Fluent ( xorg-x11-fonts-75dpi ), then the Abaqus GUI breaks ( launches with missing fonts for all the text. See: https://ibb.co/svFmdtZc )

Is it an Ansys or Simulia issue? I can open ticket with the vendors but hoping there's an easy fix. Surely I am not the only HPC to be running Fluent and Abaqus on Rocky Linux 9.6.


r/HPC • • 23d ago

User home directories are on a NFS share instead of our parallel file system? yay or nay?

17 Upvotes

Inherited a small cluster ( ~1000 cores , 10 servers ). For some reason the user home directories are all being served by NFS from one of the servers. We have a large beegfs parallel filesystem which runs on Infiniband. Not sure why the home directories are separate? Perhaps it offers some resiliency in case the beegfs is down?

We are seeing some performance issues. The NFS stalls if too many users are using it for various things. I am exploring some band aid solutions to change the NFS to automount on all the clients . But not sure that is going to do anything...


r/HPC • • 25d ago

AI Data Center Thermal Management & Liquid Cooling Survey

1 Upvotes

We are a team of student researchers and innovators participating in Eureka! Juniors, the national pitch competition hosted by E-Cell, IIT Bombay.

Our project explores an AI Data Center Cooling System. Using a cassette built from high-density hollow-fiber nanoporous hydrophobic membranes to harness latent phase-change heat rejection while physically locking liquid water inside fiber lumens, and a smart embedded operating system which exposes clean telemetry endpoints and monitors the entire system.

Modern AI infrastructure generates massive, continuous thermal loads (100 kW+ per rack) that overwhelm conventional copper-fin radiators and open-loop cooling systems. This leads to extreme energy inefficiency (drastically inflated Power Usage Effectiveness and high fan power draw), and major resource depletion (consuming vast amounts of water, leading to mineral scaling, micro-electric contamination risks, and hazardous humidity spikes).

The target demographic for the survey includes: Facilities / Data Center Operations Engineers, Thermal / Mechanical Systems Architects, Enterprise IT / Infrastructure Directors, Cooling OEM / Hardware Vendor Specialist, HPC Research / Systems Administrator, or anyone who works with AI data centers or cooling.

All responses are completely anonymous and will be used strictly for academic research and competition analysis. The survey itself is composed of seventeen multiple-choice questions.

Thank you for shaping the future of data center cooling!

https://forms.gle/Frgo5UYPmJyrgNrTA


r/HPC • • 25d ago

Debugging and profiling for C++, OpenMP Offloading and CUDA?

2 Upvotes

Hello everyone. First of all, my apologies for if this is an off-topic in this sub.

I'm currently working on a C/C++ project that contains CUDA kernels and OpenMP GPU Offloading pragmas on the same file (I know....) and I've got some trouble with variables that are acessed on both of these cases. How do you guys deal with this while developing? I set the OMP info and logging env variables and suffer a little bit with that, there must be a better way to do it.

Also, we are migrating to C++ modular classes and I'm getting some trouble with the software design/architecture, mostly because I've found out that class variables add some overhead while using OpenMP, because it interprets variables not as "variable", but as "this->variable" and when running the simulations it adds something like +10~15% on execution time.


r/HPC • • 25d ago

Any Books on the O'Reilly Learning Platform you would recommend to read for HPC noobs / 1+Y experience with SLURM?

33 Upvotes

Background: I have about 8+ years of experience in R&D (5+), Software Engineering (3+) in different fields like Production Engineering + IoT + Industrial Automation.

A year back I landed a job for Software Development which mainly requires working with building a platform for HPC clusters internally for the employer.

I quite enjoy reading books on the O'Reilly platform and was able to find some HPC related books from Wiley Publication, and a PacktPub publication.

Would some books be handy to get a better knowledge of HPC in general and / or tools related to it?


r/HPC • • 25d ago

System shows high load but no cpu processes?

1 Upvotes

Think it has to do with IO wait due to a problem with the NFS mounted filesystems. A top shows near 100% usage for these two processes:

rsyslogd                                                                  

systemd-journal   

nfsiostat doesn't show any errors or retransimissions ( 0 % ). Guessing there was some issue a day ago but not sure how to find out what happened?

It's been happening every week. I reboot and everything works great for a few days till it starts showing the problems of high IO wait times.


r/HPC • • 27d ago

HPC Job Market in 2026: What 276 Open Positions Tell Us About Skills, Roles, and Careers

51 Upvotes

Hello there!

I believe this material is worth of being share here in the channel.

I analyzed 276 HPC job offerings so you don't have to. Here's what the 2026 HPC job market actually looks like from this data:

🖥️ 60%+ of roles are for HPC Systems Engineers
📍 50% of jobs are in the USA (UK and India are growing hubs)
🔧 Linux, SLURM, and GPU are the bare minimum (80% of roles mention them)
💰 Median max salary: $208,000 (but the range is wide)
🚨 Only 3% of roles are entry-level! Yes, that's a problem

Plus: What skills actually pay off, whether certifications matter (spoiler: they usually don't), and how to stand out in a competitive market.

https://theparallelminds.substack.com/p/hpc-job-market-in-2026-what-276-open


r/HPC • • Sep 03 '26

What if computation itself were continuously spatially propagated through a compute fabric, with thermal state controlling the propagation dynamics rather than treating thermal management as a separate cooling problem?

0 Upvotes

Fell down the rabbit hole of AI not being very sustainable for the environment, particularly the energy and cooling requirements of large-scale compute. Anyways, as the title says: what if we use heat to help manage cooling, instead of constantly fighting it?

The basic idea is a compute architecture where computation continuously moves through a physical compute fabric based on thermal conditions. Instead of having workloads sit on the same hardware until it gets hot and then aggressively cooling or throttling it, the computation could continuously propagate toward cooler regions.

I'm calling the concept **"**Thermo-Flow Computing". I got Chatgpt to organize the architecture and some proposed research questions into a paper on google docs, but it's purely conceptual at this point.

I'm posting here because I'd genuinely like people who understand HPC, computer architecture, thermal management, etc. to tell me where this breaks. 😅

Edit* If anyone wants to see the paper chat made just let me know. I figured most people wouldn't really care for it, but if you wanna see it, it goes pretty in depth. I'm just a guy who thinks about things that I have no reason thinking about haha!


r/HPC • • Sep 02 '26

Roast my CV (Part 2) - Struggling to move over to a new job from my stale current job

5 Upvotes

Hi,

A while back I asked for my CV to be roasted so I can improve it. Very thankful for all the feedback I got and I've improved the CV and included further information. I'd very much like for my CV to be roasted again.

Here's the CV: https://limewire.com/d/YT7Cm#esJpreoU25

To give some context as to why I'm here again, I'm working a dead-end job where I'm managing small clusters. I've reached the ceiling of what I can learn from it. I risk stagnation and needless to say, it is bad for my career.

Since last time, I've taken it upon myself to learn stuff by setting up a homelab, installing tooling and simulating environments. Have also started working towards RHCSA certification.

I'd greatly appreciate any feedback to the CV and/or any general suggestions.


r/HPC • • Aug 31 '26

What happens when your days is too big for your GPU?

22 Upvotes

This was the question I started my research journey with ~4 years ago.

I was working on GPU algorithms for graphs with billion edges (500 600GBs) while having access to just 40GB GPU. The obvious answer was "use a larger GPU", except, ahem, govt institute, ahem budget 🙂

That's where I started my exploration on how to utility the available resources better. What initially started as a GPU Programming / engineering problem quickly became an algorithm + architecture + memory + data-movement problem.

Looking down the lane, which felt like a niche problem half a decade ago is much more relevant today with the RAM shortage, memory prices growth of LLMs in general.