Search

How File Systems Store Your Files on Disk

The short answer

Quick answer: A storage device is just a long row of numbered blocks. A file system is the software layer that organises those blocks into named files and folders. For each file it keeps a metadata record (an inode on Unix-like systems) holding the file's size, owner, timestamps and the locations of its data blocks. Directories are special files that map names to those records. To survive crashes, most modern file systems write changes to a journal first, or never overwrite data in place.

The raw material: blocks

A disk or SSD offers a very simple interface: read block number N, write block number N. Blocks are usually 4 KB. The device has no idea what a "file" is.

The file system adds everything else:

  • Names and a folder hierarchy.
  • A record of which blocks belong to which file.
  • A record of which blocks are free.
  • Permissions, timestamps and other metadata.
  • Protection against corruption when the power fails mid-write.

Inodes: the file's identity card

On Linux and macOS file systems, each file has an inode (index node). It stores:

Stored in the inodeNot stored in the inode
SizeThe file's name
Owner and permissionsThe file's contents
Timestamps (modified, changed, accessed)
Link count
Pointers to the data blocks

That the name is missing surprises many people. Names live in directories, which is what makes hard links possible: two different names in two folders can point to the same inode. The file's data is only freed when the last name is removed and no program still has it open. Windows' NTFS uses a similar structure called the Master File Table.

Finding the data blocks

Older designs stored a list of block numbers, with extra "indirect" blocks for large files. Modern file systems such as ext4 use extents: compact ranges like "blocks 5,000 to 5,999". One extent can describe a large contiguous file in a few bytes.

Directories: names to inodes

A directory is itself a file. Its content is a table:

name            inode
notes.txt       1432
photos          2210
report.pdf      1433

Opening /home/sam/notes.txt means walking the path one piece at a time:

  1. Read the root directory and find home.
  2. Read home's directory and find sam.
  3. Read sam's directory and find notes.txt, giving inode 1432.
  4. Read inode 1432 to check permissions and find the data blocks.

Each step is a disk lookup, so the operating system caches directory entries and inodes in memory. Your program triggers all of this with a single open system call.

Tracking free space

When a file grows, the file system needs to find unused blocks quickly. It keeps a bitmap or tree marking each block as free or used, and tries to place a file's blocks close together.

On spinning hard drives, files split into many scattered pieces (fragmentation) are slow to read, because the head must move between them. On SSDs, physical position does not affect speed, which is why defragmenting an SSD is unnecessary. See why SSDs are faster than HDDs.

Surviving a crash

Appending to a file changes at least three things: the data block, the inode (new size and block pointer) and the free-space map. If power fails after one or two of those are written, the file system is inconsistent. Blocks may be marked used but belong to nothing, or worse, belong to two files.

There are two main solutions.

Journaling

Before making changes in place, the file system writes a description of them to a dedicated log area called the journal. After a crash, it replays any complete entries and discards incomplete ones. The file system returns to a consistent state in seconds, without scanning the whole disk.

ext4, NTFS and XFS all work this way. By default, most journal only metadata, which keeps the structure consistent but does not guarantee the contents of a file being written at the moment of the crash. It is the same idea databases use; see how write-ahead logs prevent data loss.

Copy-on-write

File systems such as ZFS, Btrfs and Apple's APFS never overwrite live data. They write the new version to free space, then switch a pointer in one atomic step. The old version stays intact until the switch happens. A useful side effect is cheap snapshots: keeping the old pointers preserves the previous state of the whole file system.

What "delete" really does

Deleting a file removes its directory entry, reduces the inode's link count, and marks the blocks as free. The data itself is usually not erased. That has two consequences:

  • Deleting a 50 GB file is nearly instant.
  • Recovery tools can often find the data until something else overwrites it.

On SSDs, the operating system also sends a TRIM command telling the drive the blocks are no longer needed. The drive may then erase them in the background, which makes recovery much less likely.

The page cache and why you "eject" drives

For speed, writes do not go straight to disk. They land in memory first (the page cache), and the operating system writes them out a little later. Reads are cached the same way.

This is why you should eject a USB drive before unplugging it: ejecting forces pending writes to be flushed. Programs that must be sure data is on disk, such as databases, call fsync to force it.

Common file systems

File systemTypically used onNotable traits
ext4LinuxJournaling, extents, mature
XFSLinux serversScales well for large files
Btrfs, ZFSLinux, storage serversCopy-on-write, snapshots, checksums
NTFSWindowsJournaling, permissions, compression
APFSmacOS, iOSCopy-on-write, snapshots, encryption
exFAT, FAT32USB sticks, SD cardsSimple, works everywhere; FAT32 has a 4 GB file size limit

Frequently asked questions

What is an inode?

A record that stores everything about a file except its name and contents: size, owner, permissions, timestamps and where its data blocks are.

Why can I run out of space when the disk is not full?

Some file systems create a fixed number of inodes. Millions of tiny files can use them all up while free blocks remain. df -i on Linux shows inode usage.

Is a deleted file really gone?

Usually not immediately. The space is marked reusable, but the data stays until overwritten. On SSDs with TRIM, it is often erased soon after.

What is the difference between a hard link and a symbolic link?

A hard link is another name for the same inode. A symbolic link is a small file containing a path to another file, and it breaks if the target is moved.

Conclusion

A file system is an elaborate bookkeeping system laid over a row of numbered blocks. Inodes describe files, directories give them names, free-space maps find room, and journals or copy-on-write keep everything consistent when things go wrong. Once you know the pieces, behaviour such as instant deletes, hard links and "eject before unplugging" stops being mysterious.

Related articles

Sources and further reading

Usama Muneer

Usama Muneer

Coder, Blogger, Tech Speaker & Web Technologies Enthusiast. Passionate about working on open-source Programming languages & Tools while utilizing my Product Development skills.

Your experience on this site will be improved by allowing cookies Cookie Policy