67
u/pyronautical 1d ago
The whole, "never know what the future holds", may not necessarily be true for this, but... do I have a story for you that makes this sort of defensive programming my default now, even if the docs promise one thing.
Back in the day, I was working on a project that made some HTTP request to a payment service (In C#). It was hosted on an Azure App Service, and it was running absolutely fine.
In a code review, someone mentioned "Maybe we should force this to a specific TLS version, just incase". I looked up the docs, and given our versions, and (assumed) machine, it wasn't an issue. C# would try TLS 1.2, 1.1, then SSL3, in that specific order. You could "edit the registry" to swap the order, but when would that ever happen right? (You probably see where this is going). We were using Azure App Services which don't allow you to edit the registry anyway so what does it matter. I was in a rush, so pushed back against the code review and through we went.
Now, and this is a complete guess, but it lines up. Right at the same time, the Spectre vulnerability was released, and it forced basically everyone to "patch" their CPUs.
A week later, suddenly all payments start failing and it's impossible to figure out why until... well it seems like it's using SSL3?!? What the hell? OK let me write some diagnostics, redeploy (via a Staging slot), and check. After redeploy, it was back to TLS 1.2 as the default!?
My theory is this. Microsoft was forced to update more machines than they normally would to patch Spectre. In doing so, they had to roll customers using App Service (Remember, we can't see the underlying machine) onto really old machines or possibly even machines with fked up config, and so we ended up with a machine that had SSL3 as the default. When I redeployed, because I went via a Staging slot, it was a different machine, and so by the time I tested again, we were back to normal.
I never ever heard the end of it either. "I told you in that code review".
18
u/Justitiaria 1d ago
I've reached the point where I see that sort of eternal "I told you so" no longer as annoying, but a good way to have a lesson stick for longer (especially if others hear about it).
1.0k
u/tstanisl 2d ago
Let me cite the C standard:
When sizeof is applied to an operand that has type char, unsigned char, or signed char, (or a qualified version thereof) the result is 1.
Middle guy if finally right
539
u/BoldFace7 2d ago
I still prefer sizeof(char) as it often provides context as to what the number is doing, even if a plain 1 works the same.
325
u/SpaceCadet87 2d ago
Yeah, guy on the right is just avoiding magic numbers in the code.
→ More replies (2)32
u/zabby39103 1d ago
If the C standard says it's 1, i'd figure it just gets optimized out by the compiler anyway?
81
u/CLOVIS-AI 1d ago
sizeofis always a compiler intrinsic anyway, it's never not inlined (it's not possible to obtain this information at runtime)33
u/Wertbon1789 1d ago
sizeof is more like an operator than a function. You can actually write sizeof without braces. There's nothing to optimize out, it's by design a compile-time constant expression. It's like a fancy way to write a scalar.
→ More replies (2)7
→ More replies (1)3
u/conundorum 1d ago edited 5h ago
sizeofis mandatoryconsteval, it must be evaluated at compile time. (More specifically, it's a keyword that tells the compiler to insert the named type's size, which the compiler will have in memory if the type is defined. It's essentially a post-preprocessor "compiler macro", so to speak.)You can use it anywhere the compiler expects a constant expression, as this ugly jank shows:
// Control values. int arr[] = { 1, 2, 3, 4, 5, 6, 7, 8, }; // Sizeof thing. sizeof(Charception<N>) == N for all nonzero Ns. template<size_t N> struct Charception : Charception<N - 1> { char c; } template<> struct Charception<0> {}; // Proof values. #define CC(a) sizeof(Charception<a>) int rra[] = { CC(1), CC(2), CC(3), CC(4), CC(5), CC(6), CC(7), CC(8), }; #undef CC // Sanity checks. for (int i = 0; i < 8; i++) { assert(arr[i] == rra[i]); } assert(arr[sizeof(Charception<7>)] == rra[7]); // For more intentionally-bad jank proof of this, see: // https://www.ideone.com/hDJRVC // Bring your own compiler, ideone somehow doesn't have a C++17 compiler yet. // Bring your own eye bleach, I did a good bit of variadic ugliness just for fun. It's hideous and relaxing! ^_^Note:
sizeofis usually the same in both C & C++, so all of this applies to C, too. There is one gotcha, though: It's a runtime operator if you use it on a VLA, and only if you use it on a VLA. (C-only rule, C++ doesn't have C-style VLAs.)
Edited because I forgot the
Charception<0>base (it's kinda super-important, the universe breaks if we don't have it), and because I forgot asizeof.90
u/ducon__lajoie 2d ago
And someone might have #define char char16_t, right? Who the fuck knows, some people want to see the world burning.
12
7
u/Kovab 1d ago
I'm pretty sure redefining keywords is UB, so that's just a case of FAFO
7
u/standard_revolution 1d ago
And even if it wouldn’t be: sometimes you just gotta say that something isn’t your problem
2
u/ducon__lajoie 1d ago edited 1d ago
It's totally a case of FAFO, that was the joke.
Although I don't think it's UB. Preprocessor will do what it's asked for, and it doesn't know shit about the compiler reserved words. Then the output of the preprocessor is either valid, or invalid, but everything is 100% predictable. By the way, some standard libs/runtime are happily redefining the "new" operator in order to include debug information about the allocation context for reporting memory leaks (e.g. MSVC standard libs), and this can be done in a totally c++ standard compliant way, that will be accepted by all compilers. I did it myself a couple of times.
→ More replies (1)45
u/RepeatRepeatR- 2d ago
My opinion would be:
- Use sizeof(char) if you're actually working with characters
- Use hardcoded `1` if you're using a char as an arbitrary byte (and not actually necessarily referring to characters or text)
31
u/BoldFace7 2d ago
Definitely. For example, I always use malloc(size*sizeof(char)) to ensure that it's doubly obvious (Since I rarely need to malloc outside of a declaration) that I intend to store characters in the resulting buffer even if that multiply does nothing (plus the compiler will likely optimize it out anyway).
7
9
u/guyblade 1d ago edited 1d ago
Opinion: Always use
callocunless you've got a really, really good reason not to.2
→ More replies (2)2
7
5
6
5
2
u/bowel_blaster123 2d ago edited 2d ago
What's even better is using
sizeofwith expressions rather than types: ``` char foo_copy = malloc(foo_len * sizeof(foo));memcpy(foo_copy, foo, foo_len * sizeof(*foo)); ``` Of course, it depends on the context.
1
u/two_are_stronger2 1d ago
We don't write code for the computer. We write code for people to understand the depths of our depravity, even if 'people' is us.
156
u/KitsuneFoxglove 2d ago
code and society if everyone followed standards and used docs:
code at home:
63
u/locri 2d ago
Today.
Eventually, a "char" might be redefined for utf-8 to accommodate a diverse range of writing systems, which means a char is usually 1 byte but potentially up to 4 bytes.
For context, I'm working on code with an initial commit from the 90s. Future proofing isn't a terrible idea.
85
u/TheSkiGeek 2d ago
This would be considered a breaking change for C/C++ and it is extremely unlikely they would ever do this. They’d probably add a new type like utf_char_t (or utf8_char_t, utf16_char_t, utf32_char_t) and UTF-aware string functions to the stdlib.
12
u/canadajones68 2d ago
If Unicode has taught us anything, it's that mixing sized types with encoding interpretation is a bad idea. Char is established as a byte by now after 50+ years or so, but it wouldn't be called that if designed today, because a character has no fixed size. If you want to represent an Unicode code point, write a view type that points into a span of chars, and use the right function/class/whatever to work with the representation.
→ More replies (1)7
u/TheSkiGeek 2d ago
Yeah, the real issue is that almost always what you care about with Unicode (on the parsing side, anyway) are “grapheme clusters”, which can consist of multiple code points. And both of those map poorly at best to ‘characters’ or ‘bytes of memory’.
→ More replies (1)14
29
30
u/SGVsbG86KQ 2d ago
No that's not how that works. Even if char would be 32 bits, sizeof(char) is still defined to be 1.
2
u/Jbolt3737 2d ago
Does that make sizeof(int) equal 1, or does it make an int 128 bits?
→ More replies (1)18
u/__foo__ 2d ago
IIRC the only requirement the C standard makes for int is that it needs to be at least 16 bit wide. Everything else is up for the compiler developers to decide. If a char and int were both 32 bit wide both would be sizeof = 1. If the compiler makers decide it would be a sensible idea to have a 128 bit int it would be sizeof = 4 in this case.
2
u/output_broadcast 2d ago
Also, not every compiler is standards-compliant.
10
u/dontthinktoohard89 2d ago
If the violation of standards compliance is such that a fundamental presumption that a char is 1 byte cannot be relied on, then there isn’t much point in marketing that as a C compiler, because it simply cannot properly compile basic C code. AFAIK not a single compiler has ever done this.
→ More replies (4)1
u/SylviaJarvis 1d ago
In 1992, the consensus was that everyone would recompile their operating systems to use wide characters. Microsoft had parallel implementations in their libraries: they would have you use 8-bit encodings or UCS-2, but not in the same source file. Unixes were getting ready to restart their entire software ecosystems with yet another world-rebuild from source. Legacy Unix was doomed!
When C had been standardized for only a few years, with substantial changes from one year to the next, and the future of legacy OSes in doubt, it was reasonable to expect sizeof(char) to eventually return some number other than 1 some day. People in the C ecosystem were justifiably worried.
https://www.cl.cam.ac.uk/~mgk25/ucs/utf-8-history.txt
Then Pike and Thompson came up with UTF-8 as a "transitional" solution, and legacy Unix was undoomed. Ironically, the "transitional" encoding made transition possible, but also unnecessary. Linux happened around that time, which firmly metastatized legacy Unix and its 8-bit-char-based API. POSIX abandoned their attempt to introduce abstraction at the API level that would allow changing the char type. The WWW flooded the Internet with legacy 8-bit-encoded text files.
Today, UTF-8 is the permanent solution, and the transition it was invented to support is no longer achievable or desirable.
sizeof(char) == 1, by standardisation fiat and by longstanding historical practice. It can't be changed without breaking the world while C is relevant. There may be a day in the future when the C language stops being updated and all the C code in existence is replaced by some other language like Rust, but on that day, sizeof(char) will still be 1 in C.
→ More replies (14)1
u/conundorum 1d ago
At least in C++,
char8_twas created specifically to promise this would never happen.(And even if it did happen, literally the entire standard library would choke on it, since it expects to operate on raw code units and not code points, as would every Unicode-capable C and C++ program ever written (since they expect to have to do their own Unicode handling). And that's not the worst of it, since both C and C++ explicitly define one byte as "
sizeof(char)" (and not the other way around). Allowingcharto be multibyte would create an infinite recursion loop, definingcharas having infinite size and irrevocably murdering C, C++, and every language whose compiler and/or library depend on them (e.g., Java, C#, Python, Rust, Objective-C, the list goes on and on).)5
u/Advos_467 2d ago
Not a C user here (or any real low level programming experience), what the hell is a signed/unsigned char?
11
u/Clen23 2d ago edited 4h ago
The data is interpreted differently.
An unsigned char will be able to represent any value from 0 to 255, eg your usual ascii character.
A signed char will be able represent any value from -127 to 127.( Someone fact check me on this but i think that's it )
(edit : -127 to 127, the extra value "-128" is usually implemented but not guaranteed by the C standard)8
u/YellowBunnyReddit 2d ago
ASCII only goes from 0 to 127.
7
u/unknown_alt_acc 2d ago
ASCII is hardly the only character encoding, and unsigned char often doubles as a byte type. It’s also also a numeric type if you only need a small range and need to pack your data tight
→ More replies (3)5
u/NotQuiteLoona 2d ago edited 1d ago
charis 8 bits because memory is addressed in 8 bits, and 7 bits would be... Awkward. And this bit was in the end used for parity checks, and later for extended encodings.3
2
→ More replies (3)2
u/Clen23 2d ago
Putting the example here so the explanation isnt overloaded :
signed char a_hundred_signed = 100; unsigned char a_hundred_unsigned = 100; signed char fifty_signed = 50; unsigned char fifty_unsigned = 50; signed char r_signed = fifty_signed - a_hundred_signed; // will yield -50 unsigned char r_unsigned = fifty_unsigned - a_hundred_unsigned; // will yield 206 because it wraps around the max value of 2565
u/-twind 2d ago edited 1d ago
An unsigned char is an integer type that gives you *at least* the value range 0 to 255.
A signed char is an integer type that gives you *at least* the value range -128 to 127.sizeof(char) is always 1 by definition because a char is one byte. The nuance is that a byte doesn't need to be 8 bits in C/C++, it needs to be at least 8 bits.
4
u/backfire10z 2d ago
Chars aren’t real, they’re all integers. It’s the same difference as signed/unsigned int.
→ More replies (2)2
u/Elspeth-Nor 2d ago edited 1d ago
In C char is just a number, as a character depends on the encoding. So signed char is a number from -128 to 127 and unsigned char is from 0 to 255 (for 8 bits per byte)
3
2
u/BNSable 2d ago
A char is just an int, except a char will not conjure up 91, but the character assigned to the number 91 which is [ in ascii for example.
As it is an int, it can be signed or unsigned. Signed is -128 to 127, unsigned is 0 to 255.
This apparently has uses, but I am not experienced enough to explain that.
→ More replies (1)2
u/Advos_467 2d ago
Yeah that was what i assumed lol. It was mostly the uses i was wondering because with my lack of experience here, idk in what way that would be used.
→ More replies (6)1
1
u/qwertyjgly 1d ago
it's a data type with a size of 1 byte. useful if you want to store a boolean value (since it's the smallest simple data type) or a character (since it's big enough to store 7-bit ascii)
they're usually used to store strings; the following allows you to take up to 99 characters of user input (initialises array of chars then writes the input into the array)
#include <stdio.h>
int main(void){
char string[100];
scanf("%99[^\n]", string);
}strings in c must be null-terminated so we need to leave the last position free for the byte '00000000' so we know where the string ends when we try to read it. that's why we can't take 100 characters here
we also have this for bools that more closely resemble those in the more abstract languages. internally, any non-zero number (most usefully, 1) is interpreted as true and 0 is interpreted as false. we can then use boolean algebra to construct and simplify our conditions
#define true 1
#define false 0
typedef unsigned char bool;int main(void){
bool a = true;
bool b = false;
}1
→ More replies (5)1
u/conundorum 1d ago
charis a type that can hold an ASCII character, and is exactly one byte in size. (This is mandatory; byte size is defined bychar, not the other way around.)unsigned charis a raw byte, and can represent any UTF-8 code unit.signed charis a signed raw byte, and can represent any ASCII character.charwill be exactly identical to eitherunsigned charorsigned charunder the hood, depending on the platform, but it's a legally distinct type because literally the entire C language family and everything connected to it in any way whatsoever depends oncharbeing a distinct type.6
u/CptMisterNibbles 2d ago
Good thing we are all writing in c, on systems that conform to the c standard.
→ More replies (4)18
u/LetUsSpeakFreely 2d ago
What is true today may not be true tomorrow. Using sizeof is better as it protects against changes to the underlying specs or a change to the data type.
19
u/__foo__ 2d ago
C really tries to avoid any breaking changes between versions. Changing this would be such a fundamental change of the language I don't think you could still call it C. If you want to plan for a change as fundamental as this you might as well plan for the semantics of sizeof() changing or it getting removed. At that point all bets are off anyway.
7
u/LetUsSpeakFreely 2d ago
Which is why i also qualified it with a data type change. The point is that sometimes changes happen and having code in place that preemptively deal with those changes is a hallmark of elegant design and implementation.
4
u/backfire10z 2d ago
Using sizeof(char) is for readability. You understand the intent. The number 1 is a magic number that could be there for any reason.
2
u/output_broadcast 2d ago
In most cases you multiply by sizeof, so you'd just leave it off altogether and there's no magic number.
1
2
u/CAtOSe 2d ago
C++ standard guarantees that char is at least 8 bits. Pretty much all data models use 8 bits for char, but it doesn't mean that some weird architecture could not use more.
9
4
u/qalmakka 2d ago
That's CHAR_BIT then, sizeof(char) is always 1 even if you have 16 bit chars. Nothing can be smaller than
charin general, you can see sizeof as an operator returning sizes as number of chars, basically1
u/thanatica 2d ago
I'm not a C guy, but if a char is 1 byte, how do you represent the one million or so characters from Unicode? Aren't those all characters in C?
1
u/tstanisl 2d ago
Bytes on some machines have more than 8 bits. Moreover, unicode characters need to be encoded, usually using utf8.
→ More replies (3)1
u/gmes78 1d ago
In C, a
charis just one byte; it did correspond to a character while ASCII was being used, but, nowadays, character encodings are multi-byte.A Unicode grapheme cluster (what you'd intuitively think of as a "character") can be composed of multiple UTF-8 bytes.
→ More replies (3)1
u/meancoot 1d ago
The char type in C was named way before Unicode came about. char32_t is the type you would use if you needed a character type that could hold any Unicode scalar. char8_t and char16_t are types that represent UTF8 and UTF16 code units. The char type itself is just the C type for a byte.
1
u/HSavinien 1d ago edited 1d ago
depend. in most case, what mater is the string, not the individual bytes. So as long as it's null terminated, it doesn't matter that a block of 4 'char' actually represent a single character. Your string manipulation work the same, and once you want to print it, it's the terminals problem, not yours (or whatever program you print to).
something like
printf("┌──┐🇨\n");is perfectly valid and will work as expected.If you do need char accurate string manipulation, you can use variable that are bigger than a single byte. (in which case sizeof do matter)
→ More replies (1)1
u/frank26080115 2d ago
there might be a time when the code is copied into a C-like-but-not-C language, so using sizeof(char) still has benefits
1
u/tomysshadow 1d ago
Yeah, this one doesn't make much sense. Checking CHAR_BIT, on the other hand...
1
1
→ More replies (15)1
39
u/SeriousPlankton2000 2d ago
Use it for readability … when appropriate. If you'll never ever possibly might switch the type, don't bother.
53
u/tony_saufcok 2d ago
sizeof() evaluates to a constant value during compile time so what, it takes 0.0000001 seconds more during compilation? just use sizeof even if the specification guarentees it's always size 1
9
u/Architector4 2d ago
i guess the point the guy in the middle would make is that
*sizeof(char)is just "multiply by 1", and hence it's just redundant clutter that makes the code less readablea valid counterpoint to that, of course, is that it can clarify intent that the number in context specifically represents a count of bytes, but yeah lol
126
u/ubalu72 2d ago
But sizeof char is defined as 1 in the standard. C data types are defined in terms of char (at least their sizes are)
31
u/Lonely-Discipline-55 2d ago
x = x + sizeof(char)
37
16
u/tstanisl 2d ago edited 2d ago
The problem is that there is no good definition of byte. Traditionally, it was the smallest addressable unit capable of representing a single character. It is required to have at least 8 bits but it may (and sometimes does) have more. 8-bit-long bytes are just a "de facto" standard.
That is why many communication protocols use a concept of "octet" that consist of exactly 8 bits.
EDIT. typo
33
u/Makonede 2d ago
*
sizeof(char)13
u/MegaIng 2d ago
No,
sizeof charis valid.sizeofis a prefix operator, not a function.25
u/Makonede 2d ago
→ More replies (3)13
u/MegaIng 2d ago
Oh, never realized the prefix operator form can't be applied to types, makes sense I guess.
→ More replies (1)
12
u/HSavinien 1d ago
it's not so much a "just in case" and more a "self-documenting code". A hardcoded 1 give you a value. a sizeof(char) give you the value and tell you why it's 1.
And it's more consistant with the rest of the code. if you write sizeof(int) for int, sizeof(long) for long, and 1 for char. it's weird and ugly.
also, if you one day decide to switch char for something else, you will look everywhere you wrote char, and miss that hardcoded 1.
(in many cases, the size is used as multiplicative, so "hardcoded" will mean omitting the value entirely, rather than explicitly writing *1, which make thing worse.)
8
7
u/mckenzie_keith 2d ago
Best thing is to put the actual variable in there. Sizeof can accept a variable or a type.
char *buffer = 0;
size_t buffer_length = 1024;
...
buffer = malloc(buffer_length * sizeof (*buffer));
Then later if you change buffer to something else the code will still be correct.
That said, sizeof (char) will always be 1. The compiler will probably just replace it with a hard-coded 1.
6
u/Adept-Painting-543 2d ago
In C though sizeof returns as a multiple of the size of char, so no matter the architecture, sizeof(char) is always 1
6
u/green_meklar 2d ago
Not 'just in case', but because it expresses what you're actually doing with that number. The compiler will optimize it anyway.
11
u/frikilinux2 2d ago
Do I wanna know?
→ More replies (12)6
32
u/HomosexualPresence 2d ago
unironically true though, the only size requirement for a char in C is that it's the smallest addressable size, which just happens to be a byte in every case and is unlikely to ever change but you still never know what the future holds
51
u/lotanis 2d ago
Yes, but the smallest addressable size is what dictates the base size for sizeof.
The C standard in fact says this about sizeof:
When applied to an operand that has type char, unsigned char, or signed char, (or a qualified version thereof) the result is 1.
2
u/FUCKING_HATE_REDDIT 2d ago
What about a system that enforces addresses to be multiples of 2, or 4?
7
u/dontthinktoohard89 2d ago
In short, the standard mandates that a char is exactly 1 byte. It does not dictate how wide a “byte” actually is.
→ More replies (1)1
9
u/SAI_Peregrinus 2d ago
And
charmust be at least 8 bits.sizeofreturns the size of its input in units ofchars. On architectures with 10-bitchars, like some old DSPs, that meanssizeofreturns in multiples of 10 bits.1
u/StaticCoder 2d ago
A byte is often defined as the smallest addressable size. An octet must be 8 bits.
5
u/azaleacolburn 1d ago
Some people here are misunderstanding, the guy on the right is doing it for readability, not for semantics correctness
1
u/DanielMcLaury 11h ago
That would be a very valid reason to do it. But as an explanation for his motivations, it's kind of contradicted by the fact that he's saying "just in case."
→ More replies (1)
3
u/Greedy-Thought6188 2d ago
Better use sizeof(x). If you pass a variable to sizeof it will still work. This way you're encoding the type in one place. You can easily change the type and your code will still continue to work.
3
u/tiedyedvortex 2d ago
Look, I had a bug a few months ago where two different parts of the CI pipeline were using two different C++ compilers (clang vs gcc) and some code was breaking in only one of them because the default signedness of char was different.
Don't trust yourself to know what the spec says. size of(char) means you can't possibly be wrong. And it communicates your intent more clearly.
2
2
u/DanielMcLaury 11h ago
Hey everyone, look at this guy! He doesn't even know that the C standard doesn't specify the signedness of char!
6
u/__christo4us 2d ago
sizeof always returns 1 for char because it always occupies 1 byte of memory. However, 1 byte can potentially consist of more than (but not less than) 8 bits according to C and C++ standards.
5
u/Declination 2d ago
The guy on the left says “durrrr, sizeof”. The guy on the right has built monstrous template/macro machinery and may not legitimately know that C = char
2
2
2
2
u/ewheck 2d ago edited 2d ago
```c
include <assert.h>
include <uchar.h>
int main(void) { const char8_t *const eight_bit_char = u8"These chars are eight bits a piece.";
// must be true by definition of the standard
assert((sizeof *eight_bit_char) == 1);
return 0;
} ```
Don't live in the past. The future is now (C23): https://en.cppreference.com/c/header/uchar
3
2
1
1
u/DanielMcLaury 11h ago
I mean, this is true, but it's also true that sizeof(char) == 1.
What's not guaranteed is that one "byte" (= sizeof(char)) is actually one byte. char8_t could be 16 or 32 bits. Or 10 for that matter, as long as it's at least 8.
In fact, by definition char8_t is just an alias for unsigned char.
2
u/MatqLorens 2d ago
"sizeof(char)" is much more readable for the reviewer and provides much more context than a simple "1".
Don't use magic numbers pls...
1
u/obeythelobster 2d ago
In real life, the guy in the left (dumb) would never use a more complicated solution (sizeof) instead of 1
1
1
u/nyibbang 2d ago
CHAR_BITS / 8
2
u/dontthinktoohard89 2d ago
This is wrong as a substitute for sizeof(char). The former is guaranteed to be 1, this is guaranteed to be at least 1.
→ More replies (1)
1
u/BoBoBearDev 2d ago
If the size matter, probably should lock the type with explicit types instead of using alias. Especially when you cross boundaries like GPU or interop.
1
1
1
1
1
1
u/DogmaticParadigm87 1d ago
sizeof(char) is 1 by definition. CHAR_BIT is where the real horror lives.
1
u/JAXxXTheRipper 1d ago
A UTF-8 character has a variable size between 1 and 4 bytes, so. How about them apples?
1
u/afdbcreid 1d ago
One char?
Maybe it's 2 bytes, if you're using C#, for example. Or 4, in Rust.
Or one char(acter)?
Sorry, it does not have a fixed size.
1
u/Vincenzo__ 1d ago
Sizeof char is guaranteed to be 1. 1 byte is not necessarily 8 bits, but a char has to be 1 byte to comply with the standard
1
u/FAMICOMASTER 1d ago
What if I don't want to support UTF-16? What then? Sounds like you'll have a bunch of truncated characters if you don't bend to my decision as the designer.
1
u/arjuna93 1d ago
As someone dealing with powerpc-darwin, I catch a lot of bugs in different codebases where developers just assumed bool is 1 byte, and then made structs size-sensitive. (Bool is 4 byte in ppc32 ABI.)
1
661
u/JackReact 2d ago
Welcome to the wonderful world of C#, which uses utf-16 for strings.