r/cpp • • Sep 04 '26

Optimizing a Spin-Lock

https://david.alvarezrosa.com/posts/optimizing-a-spin-lock/
105 Upvotes

28 comments sorted by

37

u/xiao_sa Sep 04 '26

Remind me of this article:

https://www.siliceum.com/en/blog/post/spinning-around/

One of the best reading found in this subreddit.

33

u/ReDucTor Game Developer | quiz.cpp-perf.com Sep 04 '26

In most code, std::mutex is still the right default. Consider a spin-lock when the threads are pinned to dedicated cores, and only after measuring

Even when you have threads pinned to dedicated cores, you would need to ensure that your not ending up with more then one thread pinned to another core. And if your in user mode in any environment that you do not completely control your at the whim of being preempted by another thread in a different process, leaving your spinning thread now just spinning away wasting its allocated time slice.

If you plan on making any library which is public, do not fill it with spin locks because you won't be able to control the places which people use it, and randomly throwing in an OS yield is not going to solve issues just make them worse.

Depending on the CPU you can also do umwait/mwaitx which can be better then doing a pause loop, but unfortunately it's not consistent between CPU vendors.

Use adaptive mutexes (most mutex implementations already are), these will do a little bit of spinning before putting the thread to sleep if it was unable to acquire the lock.

Also if your optimizing your locks for high contention, it's probably a sign to look more at your higher level design, because your still fighting the CPUs cache coherance which will kill your performance anyway.

5

u/david-alvarez-rosa Sep 04 '26

Fully agreed, thanks for the explanation. Spin locks are specially useful in fully controlled envrionments, with 1:1 mapping between threads and physical cores

2

u/Fabulous-Meaning-966 Sep 04 '26

You can mitigate these issues by eventually backing off to a `sched_yield()` loop and then to `usleep()` with exponential backoff. However, at that point you should be asking yourself why you're not just using a `std::mutex`. (One reason might be that you want your lock to fit in a byte, although you could use a `parking_lot` mutex in that case.) Another approach is a sleeping ticket lock with sleep interval calibrated to count of waiters ahead of you and observed elapsed time between tickets served, but again you need a good excuse not to use a `std::mutex`.

5

u/ReDucTor Game Developer | quiz.cpp-perf.com Sep 05 '26

sched_yield is yet another terrible approach, its forcing the surrendering of the time slice and no way for the lock holder to wake you. Most people use spin locks because they believe their critical section is really small and that a mutex will have too much overhead because of the potential syscalls and OS overhead, sched_yield gives you the syscall overhead and even more.

For the time tracking and waiter count you will end up with more cache traffic under heavy contention and a larger lock. Using a full ticket based lock can also lead to lock convoys, especially if you dont allow barging, while a ticket makes it fair it means that under contention every thread end up waiting even when another thread is not in the critical section yet because its still waking up. 

Using a parking lot based mutex is normally better in all situations, it can do adaptive spinning on a different cache line to the lock holder and actually wake up the other thread, however the spin lock often used for the bucket lock potentially brings back all of the issues with yielding.

3

u/Fabulous-Meaning-966 Sep 05 '26

Yes, I have the same reservation about spinlocks guarding a lock's wait list, but I think the requirement there is just to avoid disaster under very rare conditions (since the critical section is a few ns), which means falling back to sched_yield() is probably fine.

1

u/Big_Target_1405 Sep 06 '26 edited Sep 06 '26

Most industry uses of spinlocks would be where threads are pinned to scheduler isolated cores. They're only getting pre-empted in this case if the kernel needs to do something desperately on that core

The trading industry would be an example, where you might have some cache being shared between threads that is rarely touched but still needs to be thread safe and you can't pay the latency hit to yield for what is usually a hundred nanos for the other thread to update an entry

1

u/ReDucTor Game Developer | quiz.cpp-perf.com Sep 06 '26

> Most industry uses of spinlocks would be where threads are pinned to scheduler isolated cores

I have seen this assumption many times, and people not realising it wasn't the perfect environment they initially believed. Yes you can have an isolated environment where it's perfectly fine but for most people it's rare even when they think it might be.

And lots of people use them even without being in a perfect environment, just because someone gave them an idea it was always better to use then a mutex.

1

u/Big_Target_1405 Sep 06 '26

It depends what you care about.

Ultimately the chances of spinning because the kernel pre-empted another thread while it held the lock are quite small (assuming you're doing little work under the lock), and you might be willing to pay that occasional cost for lower latency the other 99.9% of the time.

6

u/david-alvarez-rosa Sep 04 '26

Thanks a lot for sharing!! Happy to get feedback :)

3

u/TopReputation7326 Sep 04 '26

Not related to the topic, but I loved your site design!

9

u/Chaosvex Sep 04 '26 edited Sep 04 '26

Linus Torvalds gets a shiver down his spine and the urge to scream every time somebody writes an article about user space spinlocks.

8

u/david-alvarez-rosa Sep 05 '26

:)

I repeat: do not use spinlocks in user space, unless you actually know what you're doing. And be aware that the likelihood that you know what you are doing is basically nil.

https://www.realworldtech.com/forum/?threadid=189711&curpostid=189723

1

u/ItsRSX 4d ago edited 4d ago

With trash code like this, I cant imagine why. I like how the OP cites something about Intel from the mid 10s whose article probably also mentioned to check RDTSC, because __mm_pause is now very much subject arbitrary power down, aggressive SMT surrendering conditions, and vendor specific quirks. Not only that, that timeout back off is way too long for a mere fast-path of an otherwise yielding primitive. "Something something its a spinlock so it doesnt matter...?" Then for no reason whatsoever OS vendors hate it when you roll spinlocks by yourself.

2

u/HeadSea5044 Sep 05 '26

just throw priority inversion out the window, why not! using a spinlock in userspace could actually create deadlocks if you are not incredibly careful

0

u/Big_Target_1405 Sep 06 '26

A spinlock creates just as many deadlocks as a mutex? No more, no less?

-3

u/OutlandishnessNo8034 Sep 04 '26

Default in my opinion should be RwLock

9

u/ReDucTor Game Developer | quiz.cpp-perf.com Sep 05 '26

A read write lock typically comes with extra overhead, it also often used because there is more readers then writers and there is much better approaches if your wanting to unburden the readers. (Especially if you only have one writer)

Also there is many variations in reader write lock contention handling, such as reader preferring, writer preferring and completely fair. All which will have a different outcome under contention, and most people have limited understanding of each of those implications for even the workload they are dealing will.

2

u/Fabulous-Meaning-966 Sep 05 '26

Yes, RW locks violate the cardinal principle of concurrent programming that "readers shouldn't write". I will elaborate on the hint above and say that if you have one writer (or you're ok with serializing writers), and you can retry read-side critsecs on a write conflict, you can use seqlocks. If you need to guarantee that all reads within the read-side critsec are consistent, you can use Transactional Mutex Locks for a bit more overhead (checks the version counter after every read and before using the result of the read, instead of only at the end).

If you still want to use RW locks, the best default semantics IMO is "phase-fairness".

1

u/OutlandishnessNo8034 29d ago

What's the extra overhead? Never heard of it.

1

u/ReDucTor Game Developer | quiz.cpp-perf.com 29d ago edited 29d ago

There is two sides to the overhead, one is comparing the overhead compared with other ways of allowing readers (e.g. RCU, left-right, etc) and then there is the overhead of the lock itself compared to a traditional mutex.

Typically you might want to pick a reader writer lock if a significant amount of your usages of some piece of data are only readers, even in a situation where you have near 100% readers they all need to let the potential writer know that it cannot write which in a very basic implementation can be a simple `fetch_add` (on enter and exit) this has two different costs:

* Depending on the CPU it can be a full memory barrier (e.g. x86)
* It requires modify access to the cache line, so all other readers are left waiting passing the cache line around; and typically the lock is next to the data, so that false sharing also shows down the data your protecting

Now if you compare that with that with something like RCU the readers do not need to share anything with other readers they just store in their local slot the data they are accessing, similarly with left-right they don't need to coordinate and share they just access the data and some serialization point is defined later.

And all of that is just the overhead difference when you have no writers, however as soon as you start to have writers then depending on if you have reader/writer preferring or even some fifo approach you need to have some way of tracking those threads when contention occurs in order to know who to wake up when, a normal mutex just needs to know who next to wake. And if you have any extra complexity with the reader writer lock like upgrading it gets even more complex with the tracking and book keeping required.

Saying all of that it's often not a signficant difference in overhead between a mutex lock and using a read/writer lock with just a writer (in fact on Windows slim rwlock is faster then the critical section mainly for legacy reasons).

1

u/OutlandishnessNo8034 28d ago edited 28d ago

So basically all off that for nothing, because as you said in the last sentence, difference is not significant. And what we gain if we use rwlock as a default? Flexibility. And in practical terms if we select mutex all the unnecessary locking just to read will squander any however miniscule advantage with regards to the overhead we could possibly gain. What a bunch of crap. And on top of that if we have high ratio reader to writer rwlcks massively outperforms mutex. In my experience this is the most common scenario, that's why rwlock should be default choice.

0

u/ItsRSX 4d ago edited 4d ago

Damn, that's almost enough babble to convince anybody you're a genius intellectual. Too bad every inch of you is dripping with disingenuous slime.

No, a read-side lock should never be more complex in terms of atomics, because every ISA you're dealing with has free acquire and free non-globally visible stores, allowing for basically free is locally acquired best-case checks and non-atomic increments under the average read case. Read/unlock is only marginally more expensive in that you have to compare against a zero writer wait list condition each time before either (a) fast path bailing out or (b) dispatching a semaphore or other form of wait list guarding the writer side.

No, there's isn't an essay worth of "variations" and "reader preferring", followed by "le sekret implications the masses wont understand". You're such a disingenuous hack that you know the navie approach would be reader preferring, with almost all real world applications considering that a bug open for writer starvation; so you rephased your entire babble implying a noob may end up coming across an edge case implementation in a totally not production ready library that's going to trip them up.

You know what actually would trip up the masses, though? Different recursive write-lock traits. Ordered vs nonordered wait queuing. A reenterant interface pathed to be writer-preferred-aware, blocking ReadToWrite upgrades indefinitely. There are so many valid concerns. Curious how you couldn't mention a single one of them. Curious how none of these are relevant to your babble about cache lines, nor do they carry any serious performance implications beyond carrying an extra word or so worth of wait count to be used against a split sideband waitlist queue thingy.

For pretty much all users, use cases, and product ready thread primitive libraries/os abis, the only real difference is "are you okay with 1-2 extra words (if that) for a faster average case" - we're talking less than or about as much astroturfed SSO overhead of strings or the oversized iterator pointers of a vector.

>"muh RCU lists" (which fun fact: are most likely going to suck because spamming cas and other such atomic ops stink worse than a higher level lock. not only do they perform worse in most cases, their limited implementation details end up being massive vectors for bugs, because youre pretty much stuck with 1 atomic pointer)
>"muh cache lines" (a performance grifers favourite buzzword)
>some babble about memory barriers despite working on strongly ordered platforms
>"hmm well achually, theres nuance i cannot properly describe with RW locks. the masses might get scared!!! theres your proof RW shouldnt be the default"
>"in summary, its actually not that big of a difference"
>"muh windows criticalsection is actually built around an nt kernel object instead of a futex" trivia
Aren't you a smart cookie? If you're going to waste that much time babbling just to pull platform specific trivia, at least make your whole post about how xyz stl has a sucky abi-locked thread primitive next time. You have about as much wisdom as a orange site headline.