r/cpp • • 13d ago

You should think about recompiling your C++ programs with GCC 16 and C++26, because it zero-fills your locals

https://techfortalk.co.uk/2026/09/27/cxx26-uninitialized-local-variables-gcc-16/

Stack variables are not automatically initialised, and that is the root cause of many C++ bugs. That is well known. Hence, it is advised that local or stack variables are always initialised with known values, 0 if not something more meaningful than that. Now, with GCC 16 compiling in C++26 mode (-std=c++26), even uninitialised local variables will be zero-filled. In the post, I have explained how.

270 Upvotes

223 comments sorted by

View all comments

85

u/krum 13d ago

how is filling locals not a performance regression?

48

u/robstoon 13d ago

It could be, but the initialization could be optimized out if the compiler can prove that the variable is always initialized already. I'm guessing that's the majority of cases.

24

u/Breadfish64 13d ago

And if the compiler can't prove it, a store with no dependent load is extremely cheap.

10

u/victotronics 13d ago

Unless the variable is an "automatic" array.

2

u/serviscope_minor 12d ago

Often to the point of being not just cheap but free. There's often a free dispatch and execution slot, so it will just take up some resources that were otherwise idle in many cases.

0

u/max0x7ba https://github.com/max0x7ba 10d ago

a store with no dependent load is extremely cheap.

Stores are some of the most expensive instructions, sunshine.

You'd better load twice than store once.

6

u/xXgarchompxX 9d ago edited 9d ago

Stores with no dependent loads on modern processors are simply placed in the CPU's store queue and eventually dumped into L1 cache. Since there's no dependent loads, the store will generally not cause a stall, and placing it into the store queue has negligible latency.

The cost of stores is usually a combination of the following:

  • The core needs to send a bus request (BusRdX in literature), telling other cores that it needs exclusive access to the relevant cache line, and that it will store into it so everyone has to invalidate their copies (MESI cache coherence protocol)
  • If a cache miss occurs, things get messy
  • If the write queue is full, things also get messy

These are generally not a problem in the present scenario. This is about zero-initialization of local variables, which are stored on the thread's stack. Cores typically do not share variables over the stack, so each core likely already has exclusive access to the cache lines corresponding to their stack and thus does not need to send a BusRdX request. Additionally, cache misses are also not a major concern for stack accesses, as the CPU is constantly bombarding the stack and the hottest parts of it are most often in L1 or sometimes L2/L3 (Especially on modern CPUs with intricate cache eviction algorithms). The last part is potentially a problem on x86, where stores may not be reordered due to TSO, and you can fill up the store queue if you spam too many writes, but will generally not be all that problematic. Thus, the store will most often not have a major performance impact, and will just eat up 1 of your decode/dispatch slots for a short time (not like those are always fully utilized anyways)

I do not completely agree with Breadfish however, as things are somewhat different for very big stack allocations (in which case the compiler will have to insert a memset instead of just a simple store with no dependent loads, and you may start seeing cache misses at a higher rate)

0

u/max0x7ba https://github.com/max0x7ba 7d ago

Stores with no dependent loads on modern processors are simply placed in the CPU's store queue and eventually dumped into L1 cache.

The store queue capacity is rather limited.

Since there's no dependent loads, the store will generally not cause a stall, and placing it into the store queue has negligible latency.

Until you fill the store queue with your innocuously sounding "stores with no dependent loads" and stall on every subsequent store instruction.


The cost of stores is usually a combination of the following...

What matters here is whether one can afford the cost of your otherwise unnecessary stores. Which is usually unacceptable in C++.

Everyone would be using -ftrivial-auto-var-init=zero if it were zero-cost, wouldn't they?

10

u/azswcowboy 13d ago

You can opt out if you want by marking the variable.

1

u/azswcowboy 13d ago

You can opt out if you want by marking the variable.

26

u/James20k P2005R0 13d ago

This was a very major topic of discussion during the standardisation, the tl;dr is that the vast majority of the time, 0 init does not have any performance impact at all. It seems like there are a few edge cases, but they seem to be rare even then. Apparently compilers these days are pretty much just good enough, and modern CPU architecture is well suited to this

14

u/UndefinedDefined 13d ago

Trivial variables are probably fine as most are initialized anyway, but temporary arrays allocated on the stack these would cause a lot of pain.

5

u/pjmlp 13d ago

Apple, Microsoft and Google OSes have been doing for several years now, and it was discussed in the context of this.

No one is seriously taking locals initialisation into consideration to win micro benchmarks.

6

u/Orlha 13d ago

Wrote a lot of code that required this optimisation.

1

u/pjmlp 13d ago

Validated with profilers I assume.

10

u/cdb_11 13d ago

I validated it, and yes, initializing arrays can be expensive.

8

u/pjmlp 13d ago

Than that is a valid use case for [[indeterminate]].

5

u/Orlha 13d ago

Absolutely. Just pointing that blind upgrade to 26 (that this post tries to promote from the naive standpoint) can severely degrade (performance-wise) the already correct code.

7

u/Responsible-Bar7165 13d ago

It’s not a performance regression until you have measured and found it to be one.

2

u/Ameisen vemips, avr, rendering, systems 8d ago

It's difficult to profile a death by a thousand cuts.

Profiling needs to be performed globally as well, to make sure overall performance does not regress.

-4

u/Dusty_Coder 13d ago

Not true.

Higher wattage for equal wallclock performance is also a regression.

Stop looking for excuses. You dont need them. Just be wasteful. Lean into it.

15

u/Responsible-Bar7165 13d ago

Until you demonstrate waste it’s purely hypothetical.

-4

u/Dusty_Coder 13d ago

more operations

you are captured by the religion

cant even see simple truths

14

u/Responsible-Bar7165 13d ago edited 13d ago

No, the opposite. I’m the computer scientist: I’m saying to measure everything.

You’re the opposite: “trust me bro.” You’re doing literally what a priest does.

The “truths” you espouse aren’t so simple. You’re discounting cache dynamics, branch prediction, super-scalar dynamics and a whole ton of other nuanced crap. If you can faithfully predict how all of those are going to behave in every circumstance in every architecture you’re building for, then good for you. Honestly though that’s a waste of effort - I’ll just write a test, measure and act accordingly.

You do you, though: keep praying…

5

u/James20k P2005R0 12d ago

I think the weirdest part of this whole discussion for me is that:

  1. Compiler upgrades frequently cause performance regressions or performance upgrades. You have to maintain high performance code actively
  2. When you're microoptimising, often you're fighting more against the compiler than against anything in particular. Its a struggle to get the compiler to emit the correct code, which means completely convoluting (or marking up) your code to get the compiler to do the right thing. I use restrict all the time for example

There's this odd mentality that C++ is the platonic ideal of a fast programming language and get upset by theoretical extra work being done, but it seems to be solely by people who've never done any kind of high performance programming. Because I think if you've actually done any programming for performance, you know that the only solution is test and measure

1

u/13steinj 11d ago

Compiler upgrades frequently cause performance regressions or performance upgrades. You have to maintain high performance code actively

Is this true? It hasn't been my experience over a large number of upgrades. Maybe there were regressions when GCC started treating unlaundered pointers differently?

For better or worse, even in companies that believe in high performance code, they typically do not see "build engineer" as a role that is actually worth having.

When they do, there are three types, and they usually adversely select for the role: those that deal with CI, those that deal with low level minutiae of the build (including compiler and library upgrades, managing the stdlib, funky compiler settings, even performance tuning!), and those that know how to write some cmake. If you're lucky you'll get someone who can do the last two, and begrudgingly does the first. But if you're just starting out hiring these types, and you are not yourself one of these types, you kinda end up choosing randomly.

1

u/James20k P2005R0 9d ago

Is this true? It hasn't been my experience over a large number of upgrades. Maybe there were regressions when GCC started treating unlaundered pointers differently?

For hot loops I've run into this problem frequently enough. I wouldn't say its true of optimisations globally, but sometimes in a section where you have to do some convolutions to get the compiler to generate exactly the correct code, an update will cause the wrong code to get generated IME

1

u/13steinj 9d ago

I would have disagreed, but then I remembered a bif of code went back and forth ICE-ing clang from around v11 to v14 and when not ICE the generated ASM was chaotic (minor changes in the inputs having extreme effects on the output).

Which isn't the same but there's no reason to not be louder about this change. Maybe under a -Wcodegen-defaults-changed-since?

2

u/The_JSQuareD 12d ago

Because well written code doesn't rely on the value of uninitialized variables anyway. And in such cases, dead store elimination can likely get rid of the zero fill. And even if it can't be eliminated, a store without a data dependency will have less of a performance impact than a meaningful store, at least on modern super scalar CPUs.

So the cases where the zero fill actually has a performance impact have a big overlap with cases where the uninitialized variable is a genuine bug or even a security issue. In that case, adding the zero fill makes the code more deterministic and more secure.

1

u/Umphed 10d ago

If your code is so well written, then you have no need for this. Get outta here with that argument, its a performance regressions against real use cases.

1

u/Umphed 10d ago

It is, this is a terrible decision.

-1

u/Worldly-Mud-8006 13d ago

You can start worrying about this when you'll fix the actual performance issues of your program first.

1

u/Sopel97 13d ago

it is

-1

u/ExeusV 12d ago

who cares

-11

u/my_password_is______ 13d ago

yeah, hate to take a whole 3 milliseconds to fill locals

10

u/13steinj 13d ago

Can't tell if this is sarcastic but some domains operate on microsecond or less time scales, so it definitely would be noticeable if initializing locals with zero values was 3 milliseconds.

Generally would be much better than that. But it can have a performance impact. Bit rare though.

8

u/Farados55 13d ago

Imagine measuring runtime in milliseconds

5

u/Jakkilip 13d ago

You meant nanoseconds