Search

What Is a System Call and Why Is It Expensive?

The short answer

Quick answer: A system call (syscall) is how a program asks the operating system's kernel to do something it is not allowed to do itself, such as read a file, send network data, allocate memory or start another process. The program places a request number and arguments in CPU registers and executes a special instruction that switches the CPU into privileged kernel mode. The kernel checks the request, does the work, and switches back. That round trip costs far more than an ordinary function call, so fast programs try to make fewer, larger system calls.

Why programs cannot do everything themselves

CPUs run code at different privilege levels:

  • User mode is where applications run. Code here cannot touch hardware directly, cannot access other processes' memory, and cannot change memory mappings.
  • Kernel mode is where the core of the operating system runs, with full access to the machine.

This split is what keeps a buggy or malicious program from corrupting the disk or reading another program's data. But programs still need files, networks and screens. System calls are the controlled doorway between the two worlds.

What happens during a system call

Take read(fd, buffer, 4096), which reads up to 4,096 bytes from a file:

  1. A library wrapper sets up the call. Your code calls a normal function in the C library. It puts the system call number and arguments into specific CPU registers.
  2. A special instruction traps into the kernel. On x86-64 this is syscall; on ARM64 it is svc. The CPU switches to kernel mode and jumps to a fixed entry point set up by the kernel.
  3. The kernel saves state and dispatches. It saves the user program's registers, switches to a kernel stack, and looks up the handler for that call number.
  4. The kernel validates everything. Is the file descriptor valid? Does the process have permission? Does the buffer address really belong to this process?
  5. The work is done. Data is copied from the kernel's file cache into the program's buffer. If the data is not available yet, the process may be put to sleep and another process scheduled.
  6. Return to user mode. The result goes into a register, state is restored, and the program continues after the call.

Linux has several hundred system calls, listed in the syscalls manual page. Most programs use only a few dozen.

Common system calls

CategoryExamplesWhat they do
Filesopen, read, write, closeWork with files and devices
Processesfork, execve, exit, waitCreate and manage processes
Memorymmap, brkRequest or map memory
Networksocket, connect, send, recvNetwork communication
Waiting for eventsepoll_wait, select, pollWait on many files or sockets at once

High-level operations break down into these. Printing a line in Python eventually becomes a write call. Starting a program involves fork and execve, as described in what happens when you run a program.

Why it is expensive

An ordinary function call takes around a nanosecond. A simple system call typically takes from around a hundred nanoseconds to a microsecond or more, and often longer once the real work is counted. The reasons:

  • Mode switch. Entering and leaving kernel mode means saving and restoring registers and switching stacks.
  • Security mitigations. Protections added after the Spectre and Meltdown CPU vulnerabilities made the boundary crossing noticeably more costly on many processors.
  • Cache disruption. Kernel code and data push the program's own data out of the CPU caches, so the program runs slower for a moment after returning.
  • Validation and copying. The kernel cannot trust anything passed in, and data usually has to be copied between kernel and user memory.
  • Possible context switch. If the call has to wait, the kernel runs another process, which costs more again.

None of this matters for a call made occasionally. It matters a great deal for a call made millions of times.

How programs reduce the cost

Buffering

Writing one byte at a time with a system call per byte would be disastrously slow. Standard libraries collect output in a buffer in user memory and issue one write for thousands of bytes. This is why output sometimes does not appear until you "flush" it.

# Slow: one system call per line when unbuffered
for line in lines:
    os.write(fd, line)

# Fast: one system call for everything
os.write(fd, b"".join(lines))

Batching and vectored I/O

Calls such as readv, writev and sendmmsg handle several buffers or messages in one trip.

Avoiding the kernel entirely

Some frequently used calls, such as getting the current time, do not need privileged access at all. Linux exposes them through the vDSO, a small piece of kernel-provided code mapped into every process, so they run as ordinary function calls.

Memory mapping

With mmap, a file appears as a region of memory. After the mapping is set up, reading it needs no further system calls; the virtual memory system loads pages as they are touched.

Asynchronous interfaces

Linux's io_uring lets a program queue many I/O requests in a shared memory ring and submit them with one call, or sometimes none. High-performance servers use event-driven designs for the same reason; see what is the event loop.

Seeing system calls for yourself

On Linux, strace shows every system call a program makes:

strace -c ls        # summary: count and time per syscall
strace -e trace=openat cat /etc/hostname

macOS has dtruss, and Windows has Process Monitor. Watching a familiar command this way is one of the quickest ways to understand what software really does.

Frequently asked questions

Is a system call the same as a function call?

No. A function call stays inside your program. A system call crosses into the kernel, with a change of privilege level and extra checks.

Is a system call the same as a context switch?

Not necessarily. A system call switches mode within the same process. A context switch changes which process or thread is running. A system call that blocks will usually cause one.

What is the difference between an API and a system call?

A library API such as fopen or printf is ordinary code that may or may not make system calls underneath. One library call can trigger zero, one or many system calls.

Do containers and virtual machines change this?

Containers share the host kernel, so their system calls work normally. Virtual machines run their own kernel; some operations cost more because they also involve the hypervisor.

Conclusion

System calls are the boundary where your program's world ends and the operating system's begins. The boundary exists for safety, and crossing it has a real cost. The practical lesson is simple: do I/O in large chunks, let libraries buffer for you, and when performance matters, count your system calls before optimising anything else.

Related articles

Sources and further reading

Usama Muneer

Usama Muneer

Coder, Blogger, Tech Speaker & Web Technologies Enthusiast. Passionate about working on open-source Programming languages & Tools while utilizing my Product Development skills.

Your experience on this site will be improved by allowing cookies Cookie Policy