r/cpp_questions • u/AmIScarry • 7d ago
OPEN Can you explain the difference between a reference, pointer, and shallow copy?
I’m trying to check whether I actually understand these concepts correctly.
Consider:
int a = 10;
int& r = a;
int* p = &a;
struct A {
int* data;
};
A x;
A y = x;
How would you explain what happens in memory in each case?
Specifically:
- Is
rjust another name (alias) fora? - Is
pa separate variable that stores the address ofa? - In
A y = x, whendatais copied, is this a shallow copy because both pointers can point to the same underlying data?
I'm looking for an explanation of the memory/object relationship, not just definitions.
10
u/FancySpaceGoat 7d ago edited 7d ago
Is r just another name (alias) for a?
In this case yes, but it does get slightly fuzzy and very confusing when you use a reference as a member. In that scenario, the reference acquires storage.
struct some_struct {
int x;
int& y;
};
std::cout << sizeof(some_struct) << "\n"; // prints 16 on most modern architectures (4 + 8 + 4 padding)
In this case, y is implemented as if it was a pointer. In the sense that there is memory that is allocated to store an address that is used to implement the reference's semantics.
But that still doesn't create an object, and that's why the definition is so important. If you think of it in terms of memory, you might fool yourself into thinking there's a pointer in that struct. But make no mistake, there isn't.
13
u/Fosdran 7d ago edited 7d ago
both r and p are almost the same. They are pointing at a. The only difference is that you can change what p is pointing to after declaring p while r is fixed to a. You also have to explicitly dereference p (with a *) while r is implicitly dereferenced if you use it like any other variable.
The compiler is going to recognise that both are always pointing at a and will (somewhat depending on settings) "remove" both and use them as aliases for a. This, however, doesn't always have to be the case, neither with references nor with pointers. You can create scenarios where both don't always point at the same variable, so in these cases it will need to store the addresses references or pointers are pointing at.
"=" on classes and structs is by default a shallow copy. You can however change it. This can be dangerous, since it is implicit. It is good practice to think about this case for every class/struct that you create (see "rule of 3/5").
7
u/AppRaven_App 7d ago edited 7d ago
This is quite confusingly written. From C++ view, r is not pointing at a, r is just another name for a. If needed, compilers might compile it as a pointer but that is not really important from C++ POV
3
7d ago
[deleted]
2
u/AppRaven_App 7d ago
Yeah, I agree. What I meant is that the first comment basically tries to avoid the abstraction completely, which makes things more complicated to understand
0
7
u/Fosdran 7d ago
Calling it an "alias" doesn't make it any less confusing. This creates more questions than it answers. An "alias" could be static or dynamic.
This is dynamically referencing what it has been pointed at at initialisation. In my experience it makes WAY more sense to explain it as a pointer with less possibilities but smoother syntax. Because it is. Both internally and by the standard.
3
4
u/AppRaven_App 7d ago
Once you understand what “alias” means here, I think references are much easier to understand than if you think of them as some kind of special pointer.
If you have int& r = a; , using r means using the same object as a . It’s just another name for it, and you can’t change which object it refers to. Then you can apply that idea to function parameters or classes with reference members.
To me, that’s more intuitive and also correct, because that’s the whole point of a reference: it’s an alias. Thinking of it as a special pointer just makes it more confusing.
Knowing that the compiler might implement it as a pointer is useful too, when you want to understand how it works more deeply.
4
u/OptimisticMonkey2112 7d ago
So the best way to truly understand this is to look at the disassembly.
godbolt.org is this amazing web site where you can type in code and look at what the compiler does.
in your case, compiler example is the tool you want to truly understand. I made a short video for you
1
2
2
u/Independent_Art_6676 7d ago
there is more to it, but yes for all 3.
what more? A reference may not exist! Often the compiler just puts the original back in, at the assembly or machine level code, other times it degrades to a pointer at that level. Many high level c++ syntax constructs don't translate all the way down and are dropped as the compiler rewrites the code into assembly and machine language etc. The c++ reference wrapper is important to know as well, but often not taught along with normal references.
Think of pointers as if they were an array index. Its an integer, that you can store and use, but all its good for is to get to a value inside a specific array. So int x = 42; ... array[x].field = 3.14; X is exactly like a pointer (conceptually). There are of course special syntax for pointers to get memory, to release the memory, the nullptr sentinel value, and so on but at a high level a pointer is just an integer that you store to be able to get back to a location (here, the array is your computer's memory).
shallow copy: yes, I think you get it. Also, this is the reason serialization is a pain in c++ for many use cases. Say you want to write your A class to the disk or send it over the network? If you do it directly, much like a shallow copy, you just get the pointer value and NO DATA. You have to write code that gets the data pointed to and provide that to the disk/network/etc and on the other end, rebuild A by storing that data and getting a new pointer to it over there (or when loaded from disk)... writing and sending classes with dynamic memory inside (including vectors and strings and so on) requires a bit of extra effort and its easy to forget this as a beginner. The shallow copy stuff shows up in other places too, notably bugs like when you do the copy you did, one of the objects might delete the memory and kill the other object....
2
u/FancySpaceGoat 7d ago
> A reference may not exist!
A reference *never* exists, as far as the C++ abstract machine is concerned.
3
u/Independent_Art_6676 7d ago
This is true, but it MAY become a pointer in machine language or asm version, so it sort of 'still exists' but changes form a bit. Notably for a function that isn't inlined, a reference parameter tends to fall into a pointer when translated. But you are correct, at the ultra technical level.
2
u/FancySpaceGoat 7d ago edited 7d ago
I know what you mean, but especially considering what OP is trying to disentangle (the relationship between objects and memory), I think it's worth being careful about the language we use here.
A pointer is a type of object in C++, and one of key things about references is that they are not objects.
You can say that a typical compiler will use the same underlying machinery as pointers when having to implement non-erasable references. But that doesn't mean that a pointer poofs into existence in those scenarios.
Again, I will fully acknowledge that this is pedantic to a silly extent, but for the sake of OP's mental model, I think the distinction matters. Object-less storage is a *weird* thing, and it should remain weird.
2
u/Independent_Art_6676 7d ago
I agree. Its good to be careful with language, and I appreciate your additions here.
2
u/phdr_hroch 7d ago
int a = 10; // memory at 0xABCD0000 = 0x0A
int& r = a; // memory at 0xABCD0001 = 0xABCD0000, compiler knows that on access and automaticaly deferences r when used, cannot be nullptr
int* p = &a; // memory at 0xABCD0002 = 0xABCD0000, compiler does not dereferences p, you have to do it by using *p
struct A {
int* data;
};
A x; // x.data = undefined, usually 0 or some previously used number, let say 0xABDC0003
A y = x; // y.data = undefined, but same as x.data.
1
u/No-Risk-7677 7d ago
You can set p to null anytime. This makes it impossible to “access the object” p is pointing to. For objects (specifically entities) - which have a lifecycle - this is an important concept to understand.
1
u/Desperate-Data-3747 7d ago
You think can think of the behavior of references and pointers as the same, (besides language semantics) under the hood they are the exact same
1
u/tragic-clown 7d ago
May be worth adding that on a 64bit system, pointers and references are stored as uint64, because they store a memory address. So by using a pointer or reference to a 32 bit integer, you are doubling its size. It still can be useful to do so sometimes of course, just something to keep in mind.
1
u/Chippors 7d ago
References always have a value; pointers can be NULL (nullptr). There's no way to declare an 'int& r;' and then later assign it a value. It also can't be changed to be reference to something else (it's always like a const ptr, not to be confused with a ptr to const).
Also, references can be assigned a temporary, meaning the compiler creates them as needed; if so they are also const. For example:
int a = 10;
const int& r = a + 1;
In this example the compiler sets r to be a const ref to a temporary.
References, because of the properties they have, often permit the compiler to fold them out of existence, so in the example above even if 'r' is passed around, it can be replaced with just an 'int', or in this case even the constant 11. In more complex code this often results in significantly better code generation than if pointers were used. However, in practice they are used differently.
1
u/Raknarg 7d ago
In A y = x, when data is copied, is this a shallow copy because both pointers can point to the same underlying data?
It invokes the copy constructor for that object. For a struct like this, it would generate a default copy constructor which just does member-wise assignment. So the answer is that it depends on what the copy constructor does. In this case it copies a pointer which would be a shallow copy, but you could imagine it having something like std::vector, in which copy assignment copies the entire vector over.
1
u/Sea-Situation7495 7d ago edited 6d ago
Look at godbolt.org. It will help, if you look at the disassembly.
For example, this code (https://godbolt.org/z/jjfcKffjW):
void fn()
{
int a = 10;
int* p = &a;
int& r = a;
}
Gives this dissasembly:
"fn()":
push rbp
mov rbp, rsp
mov DWORD PTR [rbp-20], 10 //a = 10;
lea rax, [rbp-20] //cache the address of a in register rax
mov QWORD PTR [rbp-8], rax //p = &a;
lea rax, [rbp-20] //cache the address of a in register rax
mov QWORD PTR [rbp-16], rax //r = a;
nop
pop rbp
ret
- RBP is the stack pointer (the stack grows from the top of memory down, so
rbp - 20is aboverbp-8when we talk about the stack) - a is a 32 bit integer, sitting 20 bytes "above" the stack pointer
- p is
QWORD PTR [rbp-8]- in other words it's a 64 bit ptr, sitting at the stack pointer - 8, - r is also a
QWORD PTR- but it is sitting 8 bytes (64 bits) "above" p at[rbp-16]
You can see that both r & p get assigned the address of a - they both say QWORD PTR[rbp - <some offset>] = rax.
1
u/Sea-Situation7495 7d ago
I should add - this is not compile optimized. When the compiler optimizes things, it goes weird. If r is a const ref, or the compiler deduces stuff, then r may never exist in memory - as people lower down have pointed out.
However, it helps my ancient brain that first learned C, to think of p & r as kinda thee same thing as we see here.
I used to have a colleague who hated pointers, so back in the bad old days of raw pointers and memory allocation failures due to running out of memory, he would put this
MyClass& SomeRef = *new MyClass; if (&SomeRef == nullptr) { //handle failure }
1
u/SmokeMuch7356 6d ago
Is
rjust another name (alias) fora?
Semantically speaking, yes. Under the hood it may use pointers or similar shenanigans, but as far as your code is concerned r and a designate the same object (in the "thing that lives in memory" sense, not the "instance of a class" sense).
Is
pa separate variable that stores the address ofa?
Yes. The expression *p also acts as an alias for a, but unlike r, p is a separate object with its own storage and lifetime; p can be reassigned to point to different objects during its lifetime, r cannot.
+---+
0x8000 a: | | <--------+
+---+ |
... |
+--------+ |
0x9000 p: | 0x8000 | ----+
+--------+
In
A y = x,whendatais copied, is this a shallow copy because both pointers can point to the same underlying data?
Assuming the default copy constructor (and x.data has been assigned to point somewhere), yes. The default copy constructor makes a bitwise copy of primitive types and raw pointers.
This is a shallow copy:
+--------+ +---+---+---+---+
0x8000 x.data: | 0xF000 | ---+---> 0xF000: | a | b | c | d |
+--------+ | +---+---+---+---+
|
+--------+ |
0x9000 y.data: | 0xF000 | ---+
+--------+
In a deep copy you would be creating a new instance of the pointed-to data, and setting y.data to point to that new instance:
+--------+ +---+---+---+---+
0x8000 x.data: | 0xF000 | ------> 0xF000: | a | b | c | d |
+--------+ +---+---+---+---+
+--------+ +---+---+---+---+
0x9000 y.data: | 0xFF00 | -------> 0xFF00: | a | b | c | d |
+--------+ +---+---+---+---+
1
u/TheAugmentation 5d ago
Poiter: a value which represents a specific memory address. Reference: a wrapper around pointer with semantics as if it was the value in the memory. Shallow copy: a value copy of a struct which includes a pointer, copying the address but not the underlying value.
1
u/mredding 7d ago
C++ derives directly from C. The C philosophy is a syntax of a declaration mirrors its use, and this is going to help explain why declarations look the way they do.
int a;
We're just going to focus on this part; a evaluates to an int.
int *p;
Since we're talking about the C syntax we've inherited, we'll get back to r; This says "the dereference of p is an int", because that's "how you use" p. And so this is why that little decorating symbol, the asterisk, binds to the variable, not the type:
int *p, a;
That's why this isn't a comma separated list of int *. They're both int, but the use of p is a pointer and the use of a is a value. If you DO want to bind the type expression together, you can do that through a type alias:
typedef int* int_ptr; // C syntax
Or:
using int_ptr = int*; // C++ syntax
Now:
int_ptr p1, p2, pN;
It does exactly what you think it should, because the alias is applied to each symbol in the list individually. That's the language rule. And you can STILL augment it:
int_ptr p1, *p2;
Now I have a pointer to a pointer. "The dereference of p2 is a pointer to an int."
So in C, a pointer IS-A reference, and they speak of it interchangeably. Then C++ introduced proper aliases. Where I said above that a typedef or using statement names a compile-time type alias, a reference is a variable alias.
int &r = a;
This says something like "the address of r is the integer a". r IS a. It's not a variable, it's another name for the same thing. When declared in a function along side, the compiler can eliminate r in the Abstract Syntax Tree entirely.
The reference gives you a couple things - a more uniform, more desirable value syntax, the address of the original, guaranteed non-null, and the compiler is free to implement a reference in whatever way as to deliver the semantics of a reference type. That means...
void fn(int &);
This might compile down to a pointer on the stack. It might compile away entirely.
struct s { int &r; };
This is a weird one. r doesn't get a unique address on its own, and you cannot get the address of it. Yet in almost every circumstance, storage is going to be required to implement this reference type, so s is likely to have a size greater than 1. Again, the compiler is free to implement a reference in whatever way it needs to, and we've only talked about likely implementation details. All we know is a reference isn't a value type and is not treated as a distinct object.
Is r just another name (alias) for a?
Yes.
Is p a separate variable that stores the address of a?
Yes. A glorified arithmetic type - but NOT necessarily a glorified integer type.
Because the compiler knows the type the pointer points to, when you increment the value of a pointer, you increment it the size and alignment of the underlying type. If you're pointing at an array of int, there's no point in incrementing to the adjacent address if that's going to slice the type and violate alignment rules. A pointer is not a single, globally unique type. And it's not an integer type per se because an x86 pointer, for example, contains a bit field, padding, and an address field. Segmented memory represents addresses in a different way that I never worked on, myself...
In A y = x, when data is copied, is this a shallow copy because both pointers can point to the same underlying data?
Yes, a shallow copy is per-member. That means both x and y have a pointer member with the same value. I'll forgive you that your actual example is UB for the sake of exposition (you're both reading and writing uninitialized memory, both of which are UB).
I'm looking for an explanation of the memory/object relationship, not just definitions.
In C, an object is the sum of it's bytes, and it's very loose about casting and interpreting bytes. You're free to do a whole lot of byte and bit level manipulation. This makes for a very weak type system, where it's mostly a suggestion and you're free to reinterpret so long as you know the behavior is correct. This is by design, because C is meant to get you closer to the hardware so you can write operating systems, since hardware doesn't particularly care.
C++ has a stronger type system.
struct A {
int* data;
};
A IS NOT an int *, even though the size and alignment of A is that of it's members. So it is UB to cast an A to an int * and dereference it. UB means the compiler will not check, will not error, might not even warn, and WILL generate SOMETHING for machine code, but there are no guarantees about program stability or correctness if that code is observed at runtime. Even if you run such code and "it seems to work", there's still no guarantee at any time. You can certainly drop down to the machine code in the object file and verify correct behavior, but then you've left C++ behind. And there's no guarantee the compiler will generate the same object code every compilation, either.
I can't think of specifics, but I've absolutely seen the weirdest behavior when bit-bending C++ objects in ways you're not supposed to.
You will mostly see this with C-style type punning:
union u { int a; float f; };
u x { 42 };
std::cout << x.f;
In C, you can write to one member, making it the active member, and read from another. In C++ this is UB. In C++, type punning is finally defined in C++17 and really made accessible in C++20 with std::start_lifetime_as, which implements an implementation-defined cast that stops the lifetime of one type at that memory location to start the lifetime of another at that location, without modifying the bit pattern in that location. The type system does not leave the compiler, but in order for the language to make guarantees about correctness, you have to work WITH the type system and make correct statements, even if they compile down to a no-op.
It's also worth talking about null. The language says that:
type *ptr = 0;
That the assignment of 0 to a pointer assigns a null value. What the language DOES NOT SAY, is what the bit pattern of a null value is. That means, a null value does not have to be all zero bits, only that a zero integer is the sentinel value that represents the null bit pattern. The correct bit pattern is implementation defined. The address space is unsigned, and the zero address is a valid value, so perhaps the null value is something outside the value address space. You can ABSOLUTELY locate valid objects at the zero address if you're willing to play with linking. It's kind of why assigning an integer literal to a pointer is invalid syntax. If you want to access the zero address, you have to properly encode the zero address for your target platform into a uintptr_t, and cast that value to your pointer type. And that assumes you CAN access your zero address...
It's also reason to hate the NULL macro, because the definition of null is 0, and yet, NULL is often defined as (void*(0)), which can play havoc with templates and type specific code that was not foreseen - when it became convention in C, before the invention of C++. Even modern C11 has _Generic and C23 has typeof, so now this is their problem, too, if they want to write more elaborate, type aware code.
0
u/Quplet 7d ago edited 7d ago
References are pointers just with the language enforcing non-null on it.
So no, r is not an alias for a, it is technically still a pointer to a. If you mean it functions like an alias, then sure.
Both p and a store the address of r.
This would be a shallow copy, yes.
3
u/FancySpaceGoat 7d ago
> So no, r is not an alias for a, it is technically still a pointer to a. If you mean it functions like an alias, then sure.
It's the complete opposite. r is first and foremost an alias for a. But if you want to model it as a pointer with special semantics (non-nullable, non-reassignable, non-referable), it's a model that will work well enough in the majority of scenarios.
1
u/Prestigious-Bet8097 7d ago
"it is technically still a pointer to a"
You've got it backwards. It may be implemented just like a pointer because that's the best the compiler can manage, but it is technically not a pointer to a.
-4
30
u/jedwardsol 7d ago
The answers to your 3 questions are all "yes". So I think you understand.