The short answer
Quick answer: A storage device is just a long row of numbered blocks. A file system is the software layer that organises those blocks into named files and folders. For each file it keeps a metadata record (an inode on Unix-like systems) holding the file's size, owner, timestamps and the locations of its data blocks. Directories are special files that map names to those records. To survive crashes, most modern file systems write changes to a journal first, or never overwrite data in place.
The raw material: blocks
A disk or SSD offers a very simple interface: read block number N, write block number N. Blocks are usually 4 KB. The device has no idea what a "file" is.
The file system adds everything else:
- Names and a folder hierarchy.
- A record of which blocks belong to which file.
- A record of which blocks are free.
- Permissions, timestamps and other metadata.
- Protection against corruption when the power fails mid-write.
Inodes: the file's identity card
On Linux and macOS file systems, each file has an inode (index node). It stores:
| Stored in the inode | Not stored in the inode |
|---|---|
| Size | The file's name |
| Owner and permissions | The file's contents |
| Timestamps (modified, changed, accessed) | |
| Link count | |
| Pointers to the data blocks |
That the name is missing surprises many people. Names live in directories, which is what makes hard links possible: two different names in two folders can point to the same inode. The file's data is only freed when the last name is removed and no program still has it open. Windows' NTFS uses a similar structure called the Master File Table.
Finding the data blocks
Older designs stored a list of block numbers, with extra "indirect" blocks for large files. Modern file systems such as ext4 use extents: compact ranges like "blocks 5,000 to 5,999". One extent can describe a large contiguous file in a few bytes.
Directories: names to inodes
A directory is itself a file. Its content is a table:
name inode
notes.txt 1432
photos 2210
report.pdf 1433
Opening /home/sam/notes.txt means walking the path one piece at a time:
- Read the root directory and find
home. - Read
home's directory and findsam. - Read
sam's directory and findnotes.txt, giving inode 1432. - Read inode 1432 to check permissions and find the data blocks.
Each step is a disk lookup, so the operating system caches directory entries and inodes in memory. Your program triggers all of this with a single open system call.
Tracking free space
When a file grows, the file system needs to find unused blocks quickly. It keeps a bitmap or tree marking each block as free or used, and tries to place a file's blocks close together.
On spinning hard drives, files split into many scattered pieces (fragmentation) are slow to read, because the head must move between them. On SSDs, physical position does not affect speed, which is why defragmenting an SSD is unnecessary. See why SSDs are faster than HDDs.
Surviving a crash
Appending to a file changes at least three things: the data block, the inode (new size and block pointer) and the free-space map. If power fails after one or two of those are written, the file system is inconsistent. Blocks may be marked used but belong to nothing, or worse, belong to two files.
There are two main solutions.
Journaling
Before making changes in place, the file system writes a description of them to a dedicated log area called the journal. After a crash, it replays any complete entries and discards incomplete ones. The file system returns to a consistent state in seconds, without scanning the whole disk.
ext4, NTFS and XFS all work this way. By default, most journal only metadata, which keeps the structure consistent but does not guarantee the contents of a file being written at the moment of the crash. It is the same idea databases use; see how write-ahead logs prevent data loss.
Copy-on-write
File systems such as ZFS, Btrfs and Apple's APFS never overwrite live data. They write the new version to free space, then switch a pointer in one atomic step. The old version stays intact until the switch happens. A useful side effect is cheap snapshots: keeping the old pointers preserves the previous state of the whole file system.
What "delete" really does
Deleting a file removes its directory entry, reduces the inode's link count, and marks the blocks as free. The data itself is usually not erased. That has two consequences:
- Deleting a 50 GB file is nearly instant.
- Recovery tools can often find the data until something else overwrites it.
On SSDs, the operating system also sends a TRIM command telling the drive the blocks are no longer needed. The drive may then erase them in the background, which makes recovery much less likely.
The page cache and why you "eject" drives
For speed, writes do not go straight to disk. They land in memory first (the page cache), and the operating system writes them out a little later. Reads are cached the same way.
This is why you should eject a USB drive before unplugging it: ejecting forces pending writes to be flushed. Programs that must be sure data is on disk, such as databases, call fsync to force it.
Common file systems
| File system | Typically used on | Notable traits |
|---|---|---|
| ext4 | Linux | Journaling, extents, mature |
| XFS | Linux servers | Scales well for large files |
| Btrfs, ZFS | Linux, storage servers | Copy-on-write, snapshots, checksums |
| NTFS | Windows | Journaling, permissions, compression |
| APFS | macOS, iOS | Copy-on-write, snapshots, encryption |
| exFAT, FAT32 | USB sticks, SD cards | Simple, works everywhere; FAT32 has a 4 GB file size limit |
Frequently asked questions
What is an inode?
A record that stores everything about a file except its name and contents: size, owner, permissions, timestamps and where its data blocks are.
Why can I run out of space when the disk is not full?
Some file systems create a fixed number of inodes. Millions of tiny files can use them all up while free blocks remain. df -i on Linux shows inode usage.
Is a deleted file really gone?
Usually not immediately. The space is marked reusable, but the data stays until overwritten. On SSDs with TRIM, it is often erased soon after.
What is the difference between a hard link and a symbolic link?
A hard link is another name for the same inode. A symbolic link is a small file containing a path to another file, and it breaks if the target is moved.
Conclusion
A file system is an elaborate bookkeeping system laid over a row of numbered blocks. Inodes describe files, directories give them names, free-space maps find room, and journals or copy-on-write keep everything consistent when things go wrong. Once you know the pieces, behaviour such as instant deletes, hard links and "eject before unplugging" stops being mysterious.
Related articles
- Why SSDs Are Faster Than HDDs, Explained From the Hardware Up
- How Write-Ahead Logs Prevent Data Loss During Crashes
- What Is a System Call and Why Is It Expensive?
- Why Do We Need Both RAM and Storage?
