Sections7
  1. Where Is Data Stored?
  2. Scope
  3. Setup
  4. Line by Line
  5. What Happens When a Scope Ends?
  6. Another Scope
  7. Quick Conclusion

Binary

From this point onward, I will assume that you are comfortable working with binary. If you are not, feel free to pause here and do some exercises.

The main binary skills you will need are:

  • Converting binary to base 10 and vice versa.
  • Reading and looking up ASCII values.
  • Understanding how a sequence of bits can be interpreted as different types of data.

Where Is Data Stored?§

Before we write any code, there is already memory available to our program.

From the program's perspective, memory is simply a large collection of bytes that it can use to store data. The operating system manages the memory available to each program, and when our program runs, it gives the program access to a region of memory.

For the purpose of this chapter, let's imagine that the memory currently looks like this:

00000010 10010000 00010100 00000010 10101010 10101010
00000000 10010100 00010101 00100101 01010101

Let's assume these are leftovers from previous programs that used this memory.

Now, just for fun:

  • Can you find an integer in this memory?
  • Can you find a character?
  • If you think you found one, can you explain why you believe those bits represent an integer or a character rather than some other type of data?

The answer is:

No. You cannot determine what these bytes represent just by looking at them.

For example, the sequence

01100001

could be:

  • the integer 97,
  • the character 'A' in ASCII,
  • part of a larger integer,
  • part of a floating-point number,
  • or simply a meaningless collection of bits.

The bits themselves do not tell us what they mean.

We need context.

That context comes from the program itself: what type of data the program expects to find at a particular memory location, and how the generated machine code accesses that memory.

This is one of the most important ideas in understanding memory:

Memory stores bits. The program gives those bits meaning.

Scope§

Setup§

Now let's see what happens when we actually write and run some code.

For simplicity, we will imagine the program's stack as a sequence of bytes and watch what happens to it line by line.

We will also assume that each integer is stored in big-endian byte order ( we will explain later ) so that the diagrams are easier to read. Real machines may use a different byte order.

Consider this code:

char alpha = 'A'; // 1 byte
int six = 7;      // 4 bytes

if (six != 6)
{
    char beta = 'D';
    alpha = beta;
    int D = 100;  // 4 bytes
    six = 6;
}

{
    int scope1 = 123;
    {
        short scope2 = 456; // 2 bytes
    }
    scope1 = 0;
}

Let's execute it step by step.

Line by Line§

char alpha = 'A'; // 65 in ASCI; binary: 01000001

So conceptually, our memory now looks like:

01000001 10010000 00010100 00000010 10101010 10101010

The important part is that alpha is associated with that particular memory location.

int six = 7; // binary: 00000000 00000000 00000000 00000111

So our conceptual memory becomes:

01000001 00000000 00000000 00000000 00000111 10101010 10101010

Notice that the new variable occupies the next 4 bytes in our simplified stack model.

The condition

if (six != 6)

is true because six is currently (7).

Therefore, we enter the scope.

Now:

char beta = 'D'; 68 in ASCI; binary: 01000100

01000001 00000000 00000000 00000000 00000111 01000100 00000000 10010100 00010101 00100101 01010101

So far, we can see the pattern as new local variables are created, the stack allocates more space for them. And the variable name itself is not stored next to the data.

Instead, the compiler determines where each variable will be stored and generates machine instructions that access those locations.

By the time C++ becomes machine code, the names alpha, six, and beta generally no longer matter to the CPU.

The compiler has essentially transformed something like:

alpha = beta;

into instructions equivalent to:

"Read the value from this memory location and write it to that memory location."

The CPU does not need to maintain a dictionary mapping variable names to memory addresses.

This is one reason programming in assembly directly becomes much more difficult: the higher-level compiler is doing much of this bookkeeping for us.

Now we are not creating a new variable.

We are simply changing the value already stored for alpha.

beta contains 01000100

so we copy that value into alpha's existing memory location:

01000100 00000000 00000000 00000000 00000111 01000100 00000000 10010100 00010101 00100101 01010101

Notice that no new memory was required for alpha.

We simply overwrote the previous value: 01000001 with 01000100

The variable alpha still occupies the same location. Only its value changed.

Now we create another local variable:

int D = 100; // binary 00000000 00000000 00000000 01100100

01100100 00000000 00000000 00000000 00000111 01000100 00000000 00000000 00000000 01100100 01010101

Again, the compiler knows where this variable belongs and generates the appropriate memory accesses.

You do not have to manually remember which byte belongs to D.

Finally:

six = 6; // binary: 00000000 00000000 00000000 00000110

Again, we are not creating a new variable.

We simply overwrite the value already stored at six's memory location.

So:

01000100 00000000 00000000 00000000 00000110 01000100 00000000 00000000 00000000 01100100 01010101

What Happens When a Scope Ends?§

Now we reach the closing brace:

}

The scope ends.

This means that the local variables created inside that scope have reached the end of their lifetime.

In particular:

char beta = 'D';
int D = 100;

are no longer alive.

A useful simplified mental model is that the stack space used by these variables is now available for reuse.

It is important to understand that this does not necessarily mean that the bytes are immediately erased.

The old bits may still physically remain in memory:

01000100 00000000 00000000 00000000 00000110 01000100 00000000 00000000 00000000 01100100 ...

But the objects beta and D no longer exist as C++ objects.

Later, another variable may reuse the same memory.

This is the key connection between scope and memory lifetime:

A local variable normally exists only during the lifetime of its scope. Once that lifetime ends, the storage can be reused for something else.

Another Scope§

We now enter another scope:

{
    int scope1 = 123;
    {
        short scope2 = 456;
    }
    scope1 = 0;
}

We are now creating:

int scope1 = 123; // binary: 00000000 00000000 00000000 01111011

So the stack can reuse the space that became available when the previous scope ended:

01000100 00000000 00000000 00000000 00000110 00000000 00000000 00000000 01111011 01010101

The exact placement of variables in a real program is more complicated than this simplified diagram. The important idea is that storage whose previous lifetime has ended can be reused.

Now we enter another nested scope and create:

short scope2 = 456; // binary: 00000001 11001000

So:

01000100 00000000 00000000 00000000 00000110 00000000 00000000 00000000 01111011 00000001 11001000

Now we reach the closing brace of the inner scope:

}

scope2 has reached the end of its lifetime.

Its old bits may still be sitting in memory: 00000001 11001000

but scope2 is no longer a valid C++ object.

The memory can now be reused by another object.

This distinction is extremely important:

A variable going out of scope does not necessarily erase its bytes. It means that the variable's lifetime has ended.

Trying to access the old object after its lifetime has ended can therefore lead to undefined behavior, even if the old bits happen to still be sitting there.

We will return to this much later when we discuss pointers, references, and object lifetime.

Finally:

scope1 = 0;

We overwrite the 4 bytes belonging to scope1:

01000100 00000000 00000000 00000000 00000110 00000000 00000000 00000000 00000000 00000001 11001000

Nothing automatically cleared them.

The important difference is that scope2 is no longer alive, while scope1 is still alive and has simply had its value changed.

Quick Conclusion§

What we have just been looking at is the basic idea behind stack storage.

A stack is useful because its allocations can be managed in a LIFO — Last In, First Out — manner.

Conceptually:

Create A
    ↓
Create B
    ↓
Create C
    ↓
Destroy C
    ↓
Destroy B
    ↓
Destroy A

This makes stack allocation extremely cheap and predictable.

Local variables are commonly stored in stack frames, and when a scope ends, the storage associated with variables whose lifetimes ended can be reused.

However, remember that our diagram is a simplified model.

A real compiler may:

  • keep variables entirely in CPU registers,
  • reorder or optimize variables,
  • reuse stack slots,
  • eliminate variables completely,
  • arrange memory differently,
  • or generate no memory access at all for a particular variable.

The purpose of this model is not to describe every detail of a real CPU.

It is to establish the fundamental idea:

Memory is just storage for bits. Variables are our program's way of giving those bits names, types, lifetimes, and meaning.

Discussion

no comments
Commenting as a guest — sign in to comment as yourself.

No comments yet — yours could open the discussion.