The short answer
Quick answer: Streaming services do not send you one big video file. After upload, each video is transcoded into many versions at different resolutions and bitrates, and each version is cut into segments a few seconds long. Your player downloads a manifest listing the available versions, then fetches segments one after another over ordinary HTTP from a nearby CDN server. After each segment it measures your connection speed and picks the quality for the next one. This is adaptive bitrate streaming: when your network slows down, the picture gets softer instead of stopping.
This article explains how large video platforms work in general. The techniques are industry standards; the specifics of YouTube's own systems are only partly public.
Why video is hard
Raw video is enormous. A single uncompressed 1080p frame is about 6 MB, and at 30 frames per second that is well over a gigabit per second. Viewers have connections ranging from fibre to a weak mobile signal, on screens from a phone to an 85-inch television. A service must deliver something watchable to all of them, start instantly, and never pause.
Stage 1: Upload and transcoding
When a creator uploads a file, a processing pipeline takes over.
- Upload. Large files are uploaded in chunks so an interrupted upload can resume.
- Split. The video is divided into pieces that many machines can process in parallel.
- Transcode. Each piece is re-encoded into a "ladder" of versions.
- Package. The outputs are cut into segments and manifests are generated.
- Other jobs. Thumbnails, captions, audio tracks, and automated copyright and policy checks.
A typical ladder looks something like this (bitrates vary by codec and content):
| Resolution | Rough video bitrate |
|---|---|
| 144p | About 0.1 Mbps |
| 360p | About 0.5 Mbps |
| 720p | About 2 to 3 Mbps |
| 1080p | About 4 to 6 Mbps |
| 4K (2160p) | About 15 to 25 Mbps |
Transcoding is very CPU-intensive, so the work is queued and spread over large pools of machines; see how message queues work. This is why a new upload is often available in low resolution first, with higher qualities appearing later.
Stage 2: Compression
A codec shrinks video by removing redundancy.
- Within a frame: neighbouring pixels are similar, as in a JPEG.
- Between frames: most of the picture does not change from one frame to the next. Instead of storing each frame, the encoder stores keyframes (complete pictures) occasionally and, in between, only the differences and motion.
| Codec | Notes |
|---|---|
| H.264 (AVC) | Plays on almost everything; the universal fallback |
| VP9 | Open and royalty-free; widely used by YouTube |
| H.265 (HEVC) | More efficient than H.264; licensing is complicated |
| AV1 | Newest of the four, most efficient, slow to encode; adoption growing as hardware support spreads |
Better codecs give the same quality at a lower bitrate, but need more computing power to encode and devices that can decode them. Platforms keep several codecs for each video and serve whichever the device supports best.
Stage 3: Segments and manifests
Each version of the video is cut into segments of roughly 2 to 10 seconds. Every segment begins with a keyframe, so it can be decoded on its own. That is what makes it possible to switch quality at a segment boundary.
A manifest file lists the versions and where their segments are. There are two dominant standards:
| HLS | MPEG-DASH | |
|---|---|---|
| Created by | Apple | An international standard (MPEG) |
| Manifest | .m3u8 playlist | .mpd XML file |
| Native support | Apple devices and Safari | Most other platforms via JavaScript players |
Apple documents HLS in its streaming developer pages; the DASH standard works on the same principles.
Crucially, both are just files served over HTTP. No special streaming server or protocol is needed, which means the whole web infrastructure of caches and CDNs can be reused.
Stage 4: Adaptive bitrate playback
The intelligence lives in the player.
- Fetch the manifest.
- Start with a low or medium quality so playback begins quickly.
- Download a segment and measure how long it took.
- Estimate available bandwidth and check how many seconds are in the buffer.
- Choose the quality of the next segment: step up if there is headroom, step down if the buffer is draining.
- Repeat for every segment.
In browsers, players feed downloaded segments to the video element using the Media Source Extensions API.
The player keeps a buffer of upcoming video, typically tens of seconds, to ride out short network dips. Buffering (the spinner) happens when that buffer runs empty before the next segment arrives. Good algorithms avoid it by lowering quality early, and avoid switching too often, which is distracting.
Stage 5: Delivery through CDNs
Video is the heaviest traffic on the internet, so where it is served from matters enormously. Segments are stored on CDN servers spread around the world, and viewers are directed to one nearby.
- Popular videos are cached at many edge locations.
- Rarely watched videos are fetched from larger regional or origin stores on demand.
- The largest platforms, including Google and Netflix, run their own CDNs and place caching servers inside internet providers' networks, so video travels only a short distance to the viewer.
Segments are immutable files, which makes them perfect for caching.
Why TCP and not UDP?
On-demand video is delivered over HTTP on TCP, or over QUIC with HTTP/3. A lost packet is retransmitted, and the buffer hides the small delay. Correct pictures matter more than instant ones.
Live video calls make the opposite choice: they use UDP-based protocols such as WebRTC, because a half-second delay ruins a conversation. See TCP vs UDP.
Live streaming
Live streams use the same segment-based approach, with the pipeline running continuously: the broadcaster's stream is ingested, transcoded in real time, and published as new segments every few seconds while the manifest updates.
The delay behind real time comes from encoding, segment length and the player's buffer, and is often 10 to 30 seconds with standard settings. Low-latency variants of HLS and DASH use shorter chunks to bring that down to a few seconds.
Everything else
A video platform is much more than playback. Around it sit a metadata database for titles and channels, a search index, view counting, comments, and the recommendation system that decides what to show next.
Frequently asked questions
What is adaptive bitrate streaming?
A technique in which the video is stored at several quality levels in short segments, and the player picks the level for each segment based on current network speed.
Why does video quality drop suddenly?
The player detected that your connection slowed or its buffer was shrinking, so it switched to a lower bitrate to avoid pausing.
What causes buffering?
The player's buffer emptied because segments were arriving more slowly than they were being played, usually due to a slow or unstable connection.
What is the difference between HLS and DASH?
They are two similar standards for segmenting video and describing it in a manifest. HLS comes from Apple and is native on its devices; DASH is an open international standard.
Conclusion
Smooth streaming is a combination of preparation and adaptation. The platform does heavy work up front: compressing each video into many qualities and slicing it into small cacheable files. Your player then makes a fresh decision every few seconds about which slice to fetch from a nearby server. The result feels like a continuous stream, built entirely out of ordinary HTTP downloads.
Related articles
- How CDNs Make Websites Load Faster Worldwide
- Why HTTP/2 and HTTP/3 Exist: The Problems They Solve
- TCP vs UDP: Why the Internet Needs Both
- How Recommendation Systems Know What You Want to Watch
