The short answer
Quick answer: A compiler translates source code into machine instructions in a series of stages. A lexer splits the text into tokens. A parser arranges the tokens into a tree that reflects the program's structure. Semantic analysis checks types and names. The compiler then converts the tree into a simpler intermediate representation, runs optimisations on it, and finally generates machine code for a specific processor. A linker combines the results with libraries into an executable file.
Why a compiler is needed
A CPU understands only simple numeric instructions: load this value, add those two registers, jump to that address. People write in languages with variables, loops, functions and types. A compiler bridges that gap, and a good one also makes the result run fast.
We will follow one line through the pipeline:
int area = width * (height + 2);
Stage 1: Lexing
The lexer (or scanner) reads raw characters and groups them into tokens, discarding whitespace and comments:
KEYWORD(int) IDENT(area) EQUALS IDENT(width) STAR
LPAREN IDENT(height) PLUS NUMBER(2) RPAREN SEMICOLON
Lexers are typically built from the same theory as regular expressions: each token type is a pattern.
Stage 2: Parsing
The parser checks the tokens against the language's grammar and builds an abstract syntax tree (AST):
VarDecl(int, area)
└── Multiply
├── Var(width)
└── Add
├── Var(height)
└── Number(2)
The tree captures structure that the flat text only implies. The parentheses are gone, because the tree's shape already says the addition happens first. A missing semicolon or unbalanced bracket is caught here as a syntax error.
Stage 3: Semantic analysis
A program can be grammatically correct and still meaningless. This stage checks meaning:
- Name resolution. Are
widthandheightdeclared? Which declaration does each name refer to? - Type checking. Can these types be multiplied? Does the result fit the declared type?
- Other rules. Is a
returnmissing? Is a constant being modified?
The compiler records what it learns in a symbol table. Errors like "undefined variable" or "cannot add a string to an integer" come from this stage.
Stage 4: Intermediate representation
Most compilers do not go straight from the tree to machine code. They first lower the program into an intermediate representation (IR): a simple, machine-independent form that looks like idealised assembly.
t1 = height + 2
t2 = width * t1
area = t2
Many compilers use static single-assignment form, where every variable is assigned exactly once, because it makes optimisation much easier to reason about.
The IR is also what makes compilers reusable. A toolkit such as LLVM splits the job in three:
| Part | Job | Example |
|---|---|---|
| Front end | Source language to IR | Clang (C, C++), rustc (Rust), Swift |
| Middle end | Optimise the IR | Shared by all languages |
| Back end | IR to machine code | x86-64, ARM64, RISC-V, WebAssembly |
A new language needs only a front end to gain every supported processor. A new processor needs only a back end to gain every language.
Stage 5: Optimisation
The optimiser rewrites the IR into something that does the same thing faster or smaller. Common transformations:
| Optimisation | What it does |
|---|---|
| Constant folding | Computes 60 * 60 * 24 at compile time |
| Dead code elimination | Removes code whose result is never used |
| Inlining | Replaces a call with the function's body |
| Common subexpression elimination | Computes a repeated expression once |
| Loop-invariant code motion | Moves unchanging work out of a loop |
| Vectorisation | Processes several array elements per instruction |
Optimisation levels such as -O0, -O2 and -O3 control how much of this happens. -O0 compiles quickly and keeps the code easy to debug; higher levels take longer and produce faster code.
Optimisers assume your program follows the language's rules. In C and C++, code with undefined behaviour, such as signed integer overflow or reading past an array, can be transformed in surprising ways because the compiler is allowed to assume it never happens.
Stage 6: Code generation
The back end turns IR into real instructions for the target CPU. It involves three main tasks:
- Instruction selection. Choose machine instructions for each IR operation.
- Register allocation. A CPU has only a handful of registers. The compiler decides which values live in registers and which are kept on the stack. Doing this well has a large effect on speed.
- Instruction scheduling. Order instructions to suit the processor's pipeline.
Our line might become something like:
mov eax, [height]
add eax, 2
imul eax, [width]
mov [area], eax
Stage 7: Assembling and linking
The output so far is one object file per source file. Each contains machine code plus a list of symbols it defines and symbols it needs from elsewhere.
The linker combines object files and libraries, resolves those references, and produces a final executable in the operating system's format. With static linking, library code is copied in. With dynamic linking, the reference is left for the system to resolve at start-up, as described in what happens when you run a program.
"Undefined reference" errors come from the linker, not the compiler: every file compiled fine, but a function someone called was never provided.
Not every compiler targets machine code
- Java and C# compile to bytecode for a virtual machine, which later compiles hot code at run time.
- TypeScript compiles to JavaScript; this is sometimes called transpiling.
- Many languages compile to WebAssembly to run in browsers.
These variations, and how interpreters and JIT compilers differ, are covered in compiled vs interpreted vs JIT.
Frequently asked questions
What is the difference between a compiler and an interpreter?
A compiler translates the whole program ahead of time into another form. An interpreter executes the program directly, step by step. Many real systems mix both.
What is an abstract syntax tree?
A tree-shaped representation of the program's structure, produced by the parser. Linters, formatters and code editors use ASTs too, not only compilers.
Why is compiling large projects slow?
Mostly optimisation and, in some languages, heavy features such as templates or generics that are processed repeatedly. Incremental builds and caching avoid recompiling unchanged files.
How can I see what my compiler produces?
Use your compiler's assembly output flag (for example gcc -S), or an online tool such as Compiler Explorer, which shows source and generated assembly side by side.
Conclusion
A compiler is a pipeline of small, well-defined translations: text to tokens, tokens to a tree, a tree to checked meaning, meaning to IR, IR to better IR, and finally to instructions a processor can run. If you want to build one yourself, Robert Nystrom's free book Crafting Interpreters walks through every stage in working code.
Related articles
- Compiled vs Interpreted vs JIT: Why Languages Run Differently
- What Actually Happens When You Run a Program
- How Regular Expressions Actually Match Text
- Stack vs Heap: Where Your Variables Actually Live
