Search

What Actually Happens When You Run a Program

The short answer

You double-click an icon or type a command, and a moment later a program is running. In between, the operating system does a surprising amount of work.

Quick answer: When you run a program, the operating system creates a new process, reads the executable file from disk, and sets up a private virtual address space for it. It maps the program's code and data into that space, loads any shared libraries the program needs, prepares a stack with the command-line arguments, and finally jumps to the program's entry point, which eventually calls main(). From then on, the CPU scheduler gives the process slices of CPU time alongside everything else that is running.

Step 1: A program is just a file

A program on disk is not "running" in any sense. It is a file in a specific format that the operating system knows how to read:

Operating systemExecutable format
LinuxELF (Executable and Linkable Format)
WindowsPE (Portable Executable), the .exe and .dll files
macOSMach-O

The file contains machine code, the program's initial data, and a header that acts like a table of contents: which bytes are code, which are data, where each part should be placed in memory, and the address of the first instruction to run. A compiler and linker produced this file from source code long before you ran it.

Scripts are slightly different. A Python file is not machine code, so the OS actually runs the Python interpreter and passes your script to it. On Unix-like systems, the #!/usr/bin/env python3 line at the top tells the OS which interpreter to launch.

Step 2: The shell asks the OS to start it

Something has to request the launch. In a terminal, that is the shell; on a desktop, it is the file manager or launcher. Either way, the request is a system call.

On Linux and macOS this traditionally happens in two steps:

  1. fork() creates a copy of the current process. The copy is cheap, because memory is shared until one side writes to it (a trick called copy-on-write).
  2. execve() replaces the copy's contents with the new program. The execve manual page describes exactly what is kept and what is discarded.

Windows does it in one step with CreateProcess, which builds the new process directly.

Step 3: The kernel creates a process

The kernel sets up a data structure describing the new process. It records:

  • A process ID (PID).
  • The owner and permissions, which decide what files it may touch.
  • A table of open files, usually starting with standard input, output and error.
  • A fresh virtual address space.

That last item matters most. Every process believes it has a huge range of memory to itself. The kernel and the CPU translate those private addresses into real RAM behind the scenes. This is covered in depth in how virtual memory works.

Step 4: The loader maps the file into memory

The kernel reads the executable's header and sets up regions of the address space:

RegionContainsNotes
TextMachine codeRead-only and executable
DataGlobal variables with initial valuesRead-write
BSSGlobal variables that start at zeroNo space needed on disk
HeapMemory requested at run time (malloc, new)Grows as needed
StackFunction calls and local variablesGrows downward

A key detail: the kernel does not copy the whole file into RAM. It maps it, noting "these addresses correspond to that part of the file." Pages are loaded from disk only when the program first touches them, which is called demand paging. That is why a huge application can start quickly while only reading a fraction of its file.

For security, modern systems also place these regions at random addresses on every run (ASLR, address space layout randomisation), which makes many attacks harder. The difference between the last two regions is explained in stack vs heap.

Step 5: Shared libraries are linked

Most programs do not contain all the code they use. Printing text, opening files and drawing windows live in shared libraries (.so files on Linux, .dll on Windows, .dylib on macOS).

If the executable is dynamically linked, the kernel first starts a small helper called the dynamic linker. It:

  1. Reads the list of libraries the program needs.
  2. Finds them on disk and maps them into the address space.
  3. Fixes up references so that a call to printf reaches the real printf.

Because many programs map the same library file, its code is stored in RAM once and shared between all of them. A statically linked program skips this step, because the library code was copied into the executable at build time.

Step 6: The stack is prepared and execution begins

Before any of your code runs, the kernel places the command-line arguments and environment variables on the new stack. Then it sets the CPU's instruction pointer to the entry point and switches to user mode.

The entry point is usually not main(). It is a small piece of runtime start-up code (often called _start) that initialises the language runtime, runs global constructors, and then calls main(argc, argv). In languages with a managed runtime, such as Java, Go or C#, this stage also starts the garbage collector and other background machinery.

Step 7: Running, waiting and exiting

Your program is now one of many. The CPU scheduler switches between runnable processes many times per second, so each appears to run continuously. Whenever the program needs something outside its own memory, such as reading a file or sending data on the network, it makes a system call and the kernel does the work.

When main() returns or the program calls exit():

  1. Exit handlers run and buffered output is flushed.
  2. The kernel closes open files and frees all the process's memory.
  3. The exit code is saved for the parent process to collect. By convention, 0 means success.

Frequently asked questions

What is the difference between a program and a process?

A program is a file on disk. A process is a running instance of that program, with its own memory, open files and state. You can run the same program many times and get many separate processes. See processes vs threads.

Why do some programs take longer to start than others?

Start-up time depends on how many libraries must be loaded, how much initialisation the runtime does, and whether the needed files are already cached in RAM. A second launch is usually faster because the operating system keeps recently read files in memory.

What does "segmentation fault" mean?

The program tried to use a memory address it has no right to access, such as an unmapped address or a read-only region. The CPU raises a fault and the kernel stops the process.

Is the whole program loaded into RAM?

No. Only the pages the program actually touches are loaded, and they are loaded on demand. Unused parts of a large executable may never leave the disk.

Conclusion

Running a program is a hand-off between several pieces of software: the shell asks, the kernel creates a process and maps the executable, the dynamic linker connects libraries, and start-up code finally calls main(). Knowing these steps explains everyday mysteries, from slow first launches to "library not found" errors and segmentation faults.

Related articles

Sources and further reading

Usama Muneer

Usama Muneer

Coder, Blogger, Tech Speaker & Web Technologies Enthusiast. Passionate about working on open-source Programming languages & Tools while utilizing my Product Development skills.

Your experience on this site will be improved by allowing cookies Cookie Policy