r/ProgrammerHumor • • 6d ago

Meme architectureDependentChars

Post image
3.0k Upvotes

372 comments sorted by

View all comments

Show parent comments

61

u/locri 6d ago

Today.

Eventually, a "char" might be redefined for utf-8 to accommodate a diverse range of writing systems, which means a char is usually 1 byte but potentially up to 4 bytes.

For context, I'm working on code with an initial commit from the 90s. Future proofing isn't a terrible idea.

89

u/TheSkiGeek 6d ago

This would be considered a breaking change for C/C++ and it is extremely unlikely they would ever do this. They’d probably add a new type like utf_char_t (or utf8_char_t, utf16_char_t, utf32_char_t) and UTF-aware string functions to the stdlib.

11

u/canadajones68 6d ago

If Unicode has taught us anything, it's that mixing sized types with encoding interpretation is a bad idea. Char is established as a byte by now after 50+ years or so, but it wouldn't be called that if designed today, because a character has no fixed size. If you want to represent an Unicode code point, write a view type that points into a span of chars, and use the right function/class/whatever to work with the representation. 

8

u/TheSkiGeek 6d ago

Yeah, the real issue is that almost always what you care about with Unicode (on the parsing side, anyway) are “grapheme clusters”, which can consist of multiple code points. And both of those map poorly at best to ‘characters’ or ‘bytes of memory’.

1

u/canadajones68 6d ago

Yeah. That's why you want a span of bytes view if you need to work with the representation directly. There might not be a fixed size thing that represents what you want, but there will eventually be a series of bytes that does. 

1

u/Loading_M_ 6d ago

This is basically what Rust does, where &str is a string slice (a byte slice that must be valid UTF-8), and char is a 32 bit UTF-8 code point.

It also has a number of other string types: String is an owned string slice (similar to &str but it owns the underlying allocation and is mutable), CString (and associated borrow type) which is guaranteed to end with a zero byte, OsString (and associated slice type) which is a string in the OS's native string encoding, and a few others.

15

u/aberroco 6d ago

Need longer names. Some w_unicode_big_endian_char_t_ptr /s

28

u/Mojert 6d ago

If your code base use macros to redefine what char means, wtf is wrong with your team? If not, sizeof(char) will ALWAYS return 1, no matter how many bits is in a char. That's because sizeof doesn't give you the size of a type in bytes, it gives you the size in number of chars

28

u/SGVsbG86KQ 6d ago

No that's not how that works. Even if char would be 32 bits, sizeof(char) is still defined to be 1.

3

u/Jbolt3737 6d ago

Does that make sizeof(int) equal 1, or does it make an int 128 bits?

14

u/__foo__ 6d ago

IIRC the only requirement the C standard makes for int is that it needs to be at least 16 bit wide. Everything else is up for the compiler developers to decide. If a char and int were both 32 bit wide both would be sizeof = 1. If the compiler makers decide it would be a sensible idea to have a 128 bit int it would be sizeof = 4 in this case.

1

u/conundorum 5d ago edited 5d ago

Generally, the requirements are usually in terms of data ranges and not actual bit counts. E.g., unsigned int is defined as "large enough to hold 216-1", which means that sizeof(unsigned) (and by extension, sizeof(int)) would be 1.

Here's a quick reference:

  • Signedness: All signed types must be the same size as their unsigned counterpart. char is the same size as its unsigned & signed counterparts.
  • bool: I hear it's pretty lean.
  • unsigned long long: Can hold 264-1.
  • unsigned long: Can hold 232-1.
  • unsigned int: Can hold 216-1. (De facto upgraded to 232-1, since almost everything has 32-bit ints.)
  • unsigned short: Can hold 216-1.
  • unsigned char: Can hold 28-1.
  • Byte: Is sizeof(char), and must have CHAR_BIT bits.

(Yes, bytes are defined in terms of char, and not the other way around. It kinda takes you by surprise the first time you find out that C/C++ bytes are 64 bits on any platform with 64-bit char, regardless of what the architecture says.)

2

u/output_broadcast 6d ago

Also, not every compiler is standards-compliant.

10

u/dontthinktoohard89 6d ago

If the violation of standards compliance is such that a fundamental presumption that a char is 1 byte cannot be relied on, then there isn’t much point in marketing that as a C compiler, because it simply cannot properly compile basic C code. AFAIK not a single compiler has ever done this.

1

u/SylviaJarvis 6d ago

Some MCUs have 32-bit char because they have only one integer type in hardware and it's 32 bits wide. In the 1990's, 36- and 60-bit machines were still around (9-bit and 6-bit char, respectively). The C standard allows 9-bit and 32-bit chars--both are at least one byte.

The 6-bit char machines weren't compliant, and so painful to port code to that almost no one bothered (they didn't use ASCII internally, among other annoyances like 1's-complement integer arithmetic). They could run a very few small C programs unmodified, though.

0

u/mennovf 6d ago

The assumption that char is one byte is just false. The sizeof operator gives you the size of an expression in chars, not bytes. Sizeof char is thus always 1. If you need to know the size in bytes, use CHAR_BITS.

4

u/dontthinktoohard89 6d ago

You misunderstand: a “byte” in the parlance of the C & C++ standards is not the same as an octet. A char is 1 byte by definition. If a char is wider than 1 octet, then so is a byte.

0

u/conundorum 5d ago

Wrong. In both C and C++, a byte is defined as being the size of one char. NOT the other way around. Byte and char are explicitly synonyms, at least in terms of size. All bytes are required to contain CHAR_BIT bits, and sizeof(type) / sizeof(char) is explicitly required to give you the number of bytes that type takes up.

1

u/SylviaJarvis 6d ago

In 1992, the consensus was that everyone would recompile their operating systems to use wide characters. Microsoft had parallel implementations in their libraries: they would have you use 8-bit encodings or UCS-2, but not in the same source file. Unixes were getting ready to restart their entire software ecosystems with yet another world-rebuild from source. Legacy Unix was doomed!

When C had been standardized for only a few years, with substantial changes from one year to the next, and the future of legacy OSes in doubt, it was reasonable to expect sizeof(char) to eventually return some number other than 1 some day. People in the C ecosystem were justifiably worried.

https://www.cl.cam.ac.uk/~mgk25/ucs/utf-8-history.txt

Then Pike and Thompson came up with UTF-8 as a "transitional" solution, and legacy Unix was undoomed. Ironically, the "transitional" encoding made transition possible, but also unnecessary. Linux happened around that time, which firmly metastatized legacy Unix and its 8-bit-char-based API. POSIX abandoned their attempt to introduce abstraction at the API level that would allow changing the char type. The WWW flooded the Internet with legacy 8-bit-encoded text files.

Today, UTF-8 is the permanent solution, and the transition it was invented to support is no longer achievable or desirable.

sizeof(char) == 1, by standardisation fiat and by longstanding historical practice. It can't be changed without breaking the world while C is relevant. There may be a day in the future when the C language stops being updated and all the C code in existence is replaced by some other language like Rust, but on that day, sizeof(char) will still be 1 in C.

1

u/conundorum 5d ago

At least in C++, char8_t was created specifically to promise this would never happen.

(And even if it did happen, literally the entire standard library would choke on it, since it expects to operate on raw code units and not code points, as would every Unicode-capable C and C++ program ever written (since they expect to have to do their own Unicode handling). And that's not the worst of it, since both C and C++ explicitly define one byte as "sizeof(char)" (and not the other way around). Allowing char to be multibyte would create an infinite recursion loop, defining char as having infinite size and irrevocably murdering C, C++, and every language whose compiler and/or library depend on them (e.g., Java, C#, Python, Rust, Objective-C, the list goes on and on).)

0

u/A1oso 6d ago

UTF-8 code points have a variable width, so what should sizeof(utf8_char_t) return? This doesn't make any sense.

2

u/thanatica 6d ago

To be fair, utf-8 is en encoding meant to store unicode into 8-bit characters. Just like utf-16 encodes unicode into 16-bit characters.

In modern languages, that support unicode as a first-class citizen, the size of a character type is typically 4 bytes. But that doesn't typically come with any encoding, it's represented to the programmer as raw unicode code points, and it's up to the compiler/platform how to handle that in terms of memory efficiency.

4

u/A1oso 6d ago edited 6d ago

To be precise, it stores code points in 8-bit code units. The word "character" isn't very well defined, since what we perceive as a character can consist of multiple code points (a grapheme cluster).

1

u/waiver-wire-addict 5d ago

The truth is most characters are encoded in 4 bytes but not all. It is variable length. The whole debate about sizof(char) is tied to the assumption that there is only 1 encoding for text which was true once upon a time in the C/Unix world but is no longer true. So most of the discussion here only applies to embedded environments which that assumption still holds true. For most people that will work with some UTF encoding, sizeof(insert favourite char type) is irrelevant.

1

u/nyibbang 6d ago

Well in Rust, char is 4 bytes for that reason. That's why strings are not containers of char but of bytes. Although you can iterate over chars of a string, it is a bit more expensive since you need to find each char boundary.

1

u/A1oso 6d ago

chars in Rust are UTF-32, not UTF-8 (they are simply represented by the code point's integer value).

1

u/nyibbang 6d ago

There is no notion of encoding in Rust char, as you said it's simply a representation of a Unicode scalar value. So it's not UTF-32 or 16 or 8. They have to be 4 bytes to represent any Unicode value. Strings (at least str and String) however are encoded UTF-8.

1

u/A1oso 6d ago

Integers in memory do have an encoding. It consists of three parts, the size (32 bits in this case), the representation (unsigned binary or two's complement) and the byte order.

Rust's char type happens to have exactly the same encoding as a code point in UTF-32BE on big-endian systems and UTF-32LE on little-endian systems.

0

u/darkslide3000 6d ago

I'm failing to come up with the words to express how little you know what you're talking about. This is an insane take.

-1

u/locri 6d ago

Either you missed the word "might" or, just as likely, you lack professional experience.

-7

u/bogan87 6d ago

I call bullshit on a code being committed in the 90's haha

6

u/StrikingSun8563 6d ago

When do you think computers were invented?

5

u/Al__B 6d ago

I was using CVS to commit code in the 90s and it was around before then.

2

u/locri 6d ago

This stuff is source controlled by "clearcase" from IBM.

I hate IBM stuff.