Search

How to Design a URL Shortener Like Bitly

The short answer

Quick answer: A URL shortener stores a mapping from a short code to a long URL and redirects visitors. To create a link, the service generates a unique short code, commonly by taking a unique number and encoding it in Base62, and saves the code with the long URL in a database. To follow a link, it looks up the code, usually in a cache, and replies with an HTTP redirect. The system is extremely read-heavy, so the design centres on fast lookups, caching and a collision-free way to generate codes.

This is one of the most popular system design interview questions, because it is small enough to finish and still touches storage, caching, scaling and trade-offs.

Step 1: Requirements

Functional

  • Given a long URL, return a short one.
  • Visiting the short URL redirects to the long one.
  • Optional: custom aliases, expiry dates, click analytics.

Non-functional

  • Redirects must be fast and highly available. A broken shortener breaks every link that uses it.
  • Short codes should not be guessable in sequence.
  • Reads vastly outnumber writes.

Step 2: Rough numbers

Estimates guide the design. Suppose:

QuantityAssumptionResult
New links100 million per monthAbout 40 writes per second
Read to write ratio100 to 1About 4,000 redirects per second on average
Retention5 years6 billion links
Size per recordAbout 500 bytesRoughly 3 TB

Two conclusions: writes are easy, reads are the challenge, and the data is large enough to plan for partitioning but not exotic.

How long must the code be? Base62 uses a-z, A-Z and 0-9.

  • 6 characters: 62^6 ≈ 57 billion combinations.
  • 7 characters: 62^7 ≈ 3.5 trillion combinations.

Seven characters comfortably cover 6 billion links with room to spare.

Step 3: The API

POST /api/links
{ "url": "https://example.com/a/very/long/path?with=params", "alias": "optional" }

201 Created
{ "short_url": "https://sho.rt/aZ3kP9x" }
GET /aZ3kP9x

301 Moved Permanently
Location: https://example.com/a/very/long/path?with=params

Step 4: Generating the short code

This is the heart of the design. Three approaches:

Option A: Hash the URL

Hash the long URL (for example with SHA-256) and take the first 7 characters of its Base62 form.

  • Simple, and the same URL always gives the same code.
  • Collisions are inevitable when truncating, so every insert must check for one and retry with a salt.

Option B: Random code

Generate 7 random Base62 characters and check the database for a clash.

  • Unpredictable codes.
  • Each write needs a uniqueness check, and clashes become more likely as the space fills.

Option C: Counter plus Base62

Give each new link a unique integer ID and convert it to Base62. ID 125 becomes 21; ID 1,000,000,000 becomes 15ftgG.

  • No collisions, by construction.
  • Needs a source of unique IDs across many servers.

Ways to produce the IDs:

  • A database sequence (simple, but a single point of contention).
  • Range allocation: each application server reserves a block of, say, 1,000 IDs from a coordinator and hands them out locally.
  • Snowflake-style IDs: combine a timestamp, a machine ID and a per-machine counter.

One drawback: sequential IDs make codes guessable. Anyone can walk through ...aZ3, ...aZ4, ...aZ5 and discover other people's links. Fix it by scrambling the number with a reversible function before encoding, or by adding random bits.

ApproachCollisionsPredictableComplexity
Hash and truncatePossible; must checkNoLow
RandomPossible; must checkNoLow
Counter + Base62NoneYes, unless scrambledMedium

A fourth option, a key generation service that pre-generates random unused codes and hands them out, combines unpredictability with no runtime collisions.

Custom aliases are simply user-chosen keys: insert with a uniqueness constraint and return an error if it is taken.

Step 5: Storage

The data model is one table:

ColumnNotes
codePrimary key
long_url
user_idOptional
created_at, expires_at

Access is almost entirely "get by key", with no joins. That is an ideal fit for a key-value or wide-column store such as DynamoDB or Cassandra, which scale horizontally by key. A relational database also works well; at this size you would shard it by a hash of the code. The decision follows the reasoning in SQL vs NoSQL.

Step 6: The read path

Redirects must be fast, and popular links are requested far more than others.

  1. The request reaches a load balancer, then a stateless application server.
  2. The server checks a cache such as Redis for the code.
  3. On a hit, it redirects immediately.
  4. On a miss, it reads the database, stores the result in the cache, and redirects.

Because mappings almost never change, they cache extremely well, with long expiry times and least-recently-used eviction. This is the cache-aside pattern from caching strategies. Caching "not found" results briefly also stops repeated lookups for codes that do not exist.

Step 7: 301 or 302?

301 Moved Permanently302 Found (temporary)
Browser caches the redirectYesGenerally no
Later clicks reach your serversOften notYes
Server loadLowerHigher
Click analyticsIncompleteAccurate

A 301 tells browsers the move is permanent, so they may skip your service next time. That saves load but hides clicks. Services that sell analytics tend to use temporary redirects, or a 301 with caching disabled, so that every click is counted.

Step 8: Analytics

Do not write to the analytics database during the redirect; it would slow the hot path. Instead, publish a small click event to a message queue and return immediately. Separate workers consume the events and aggregate counts by time, country, referrer and device.

Step 9: Safety and abuse

Shorteners hide destinations, which makes them attractive for spam and phishing.

  • Rate limit link creation per user and per IP. See how rate limiters work.
  • Validate URLs and block redirect loops back to the shortener itself.
  • Check destinations against malware and phishing blocklists, and let users report links.
  • Expire unused links if your policy allows, with a background clean-up job.

Step 10: Scaling and availability

  • Application servers hold no state, so add as many as needed behind the load balancer.
  • Replicate the database and cache across availability zones.
  • Serve redirects from multiple regions, or from a CDN edge, to cut latency.
  • The ID generator must not be a single point of failure; range allocation lets servers keep working if the coordinator is briefly unavailable.

Frequently asked questions

Why Base62 and not Base64?

Base64 includes + and /, which have special meanings in URLs. Base62 uses only letters and digits.

How do you avoid two links getting the same code?

Either derive the code from a guaranteed-unique ID, or generate a candidate and enforce a uniqueness constraint in the database, retrying on conflict.

Should the same long URL always return the same short URL?

It is a product decision. Reusing codes saves space; issuing a new code each time lets different users track their own clicks separately.

What database is best for a URL shortener?

Any store with fast key lookups. Key-value databases fit naturally; a relational database with a cache in front is perfectly adequate for most scales.

Conclusion

A URL shortener is a key-value lookup with a redirect on top. The interesting decisions are how to generate unique, non-guessable codes, how to keep the read path fast with caching, and how to count clicks without slowing redirects. Work through the requirements and estimates first, and the architecture largely follows from them.

Related articles

Sources and further reading

Usama Muneer

Usama Muneer

Coder, Blogger, Tech Speaker & Web Technologies Enthusiast. Passionate about working on open-source Programming languages & Tools while utilizing my Product Development skills.

Your experience on this site will be improved by allowing cookies Cookie Policy