Vibe Engines
YouTube
System Design

Design a URL Shortener

Step 1 / 9

Learn system design by building a URL shortener like Bitly or TinyURL step by step.

The numbers to beat62symbols7chars3.5Tunique codes

The whole design, in writing

Learn system design by building a URL shortener like Bitly or TinyURL step by step. An interactive guide covering the client–server skeleton, base62 key generation, caching the read-heavy redirect path, splitting read/write services, sharding billions of mappings, async click analytics, and the unhappy paths.

Every step of the build above, written out: the problem each piece solves, the option that was taken and the ones that were not, the numbers, and how it fails in production.

The big idea

What is a URL shortener?

Strip away the brand and a URL shortener does one tiny thing: take a long, ugly link and hand back a short one — then, when anyone taps the short one, instantly send them to the original.

Shorten UIVisitor
New in this step: Shorten UI, Visitor. · swipe to pan the diagram

Sounds trivial. But "instantly", "anyone", and "never collide" turn it into a real systems problem: how do we mint a unique code for every link, and resolve billions of clicks in single-digit milliseconds without ever sending someone to the wrong page?

What the new pieces do

Shorten UIclient
Someone with a long, ugly link. Sends it to the service and gets back a tiny code to share.
Visitorclient
Anyone who taps a short link. Expects to land on the original page instantly, with no idea a lookup happened.

Step 1 · The skeleton

Two jobs, one server

There are really only two operations: shorten (give me a long URL, return a short code) and redirect (given a code, send me to the long URL). Both need somewhere to keep the mapping — but where?

App ServerURL Store
New in this step: App Server, URL Store. · swipe to pan the diagram

Two jobs: shorten (long URL → code) and redirect (code → long URL). What’s the minimal structure?

  1. A client can’t share a mapping it created with everyone else who taps the link, and can’t be trusted to keep it. The mapping must live on a server, in durable storage.

  2. To shorten, the server saves code → long_url; to redirect, it looks the code up and returns a 301/302. The classic client → server → database spine: client asks, server decides, database never forgets.

  3. Packing a full URL into a few characters is impossible (URLs are far longer), and the code would no longer be short. You store the mapping and hand back a tiny key to it.

Put a server in the middle and a database behind it. To shorten, the server saves a row code → long_url. To redirect, it looks the code up and answers with an HTTP 301/302 to the original. This is the client → server → database spine under every app.

Why this piece earns its place

The quiet commitment in this step isn’t the server, it’s permanence. The moment the mapping lives on your box, every code you mint is a liability with no end date: a short link gets printed on a poster, pasted into a PDF, embedded in an email somebody opens four years from now. If the store loses that row the link doesn’t degrade, it 404s — on someone else’s page — and nothing on the internet can reconstruct it from the code alone. So I’d rank durability here above everything that comes later. The cache, the key service and the whole analytics side can each be down and the product still half-works; lose code → url and the product is gone retroactively. Two rules fall straight out of that. Never recycle a code: when a dead link is cleaned up you reclaim the bytes, not the key, or an old poster quietly starts sending people somewhere new. And keep the mapping exportable, because the only honest answer to “what happens if you shut down” is handing the table to whoever takes over.

What the new pieces do

App Serverbackend
The single front door for every request. Later it becomes an API Gateway sitting behind a load balancer.
URL Storestore
Durable storage for the one thing you must never lose: which short code maps to which long URL.

Step 2 · The short code

How do we mint the code?

Every link needs a unique code like /aZ9k2x. Hashing the URL gives collisions and runs long; pure random risks clashes and forces an "is this taken?" check on every single write. How do you guarantee uniqueness cheaply?

App ServerKey Gen ServiceURL Store
New in this step: Key Gen Service. · swipe to pan the diagram

Every link needs a unique short code like /aZ9k2x. How do you mint it?

  1. Truncated hashes collide (different URLs → same prefix), forcing a "taken?" check and retries, and identical URLs map to the same code. You want uniqueness without a lookup.

  2. Random codes force a collision check on every write, and clashes rise as the space fills. Counting guarantees uniqueness with no check at all.

  3. Counter 125 → "cb"; the KGS ensures two servers never mint the same number, so codes are unique by construction — no collision check. 62⁷ ≈ 3.5 trillion codes.

Treat it as a number. Keep a global counter and encode it in base62 (a–z A–Z 0–9): counter 125 becomes "cb". A dedicated Key Generation Service hands out ids — so codes are unique by construction, no collision check needed.

Why this piece earns its place

A counter is unique by construction, and it is just as surely predictable by construction. If 125 encodes to “cb”, then the code a second later is the one issued to whoever shortened next — so anybody can start at the short end of the keyspace and walk it, harvesting every link the service has ever minted. That matters because people treat a short link as if it were a secret: an unlisted draft, a document whose only lock is a URL nobody guesses, a one-off invite. None of that is protected by the original link being long and ugly, and a sequential shortener has just published an index of all of it. The fix keeps the property I actually wanted: run the counter through a reversible keyed permutation over the id space before base62-encoding it, so ids stay dense and collision-free while the emitted codes look nothing like a sequence. It costs one fixed function call, on the rare path. The leak has a commercial edge too — a sequence you can walk is a sequence you can difference, which hands a competitor your creation rate.

  • 62symbols
  • 7chars
  • 3.5Tunique codes

What the new pieces do

Key Gen Serviceservice
Hands out unique numbers (or pre-made codes) so two servers never mint the same short code.

Back of the envelope

62 symbols × 7 slots
62⁷ ≈ 3.5 trillion codes
counter ⇒ unique by construction
no collision check on write
KGS hands out ids
two servers never mint the same code

Step 3 · The hot path

Reads dwarf writes

People create a link once but click it thousands of times. The redirect path runs maybe 100× more than the write path — and every redirect that hits the disk database adds latency to someone’s tap. How do you keep it sub-millisecond?

App ServerCacheURL Store
New in this step: Cache. · swipe to pan the diagram

A link is created once but clicked thousands of times. How do you keep redirects sub-millisecond?

  1. Replicas add throughput but every redirect still does a full DB lookup for an unchanging mapping. When the same code is hit repeatedly, serve it from memory instead.

  2. A hit returns in <1ms; a miss falls through to the store and warms the cache. Because a mapping never changes, it’s perfectly cacheable — no invalidation, pure upside.

  3. You can’t rely on browser caching, and a permanent (301) redirect browsers cache means you never see the click — bad for analytics. Server-side caching keeps you fast AND in the loop.

Put a cache (Redis) in front of the database. A redirect checks the cache first; on a hit it returns in under a millisecond. On a miss it falls through to the database, then warms the cache. Because a mapping never changes, it’s perfectly cacheable.

Why this piece earns its place

The cache introduces the one number on this page you have to choose rather than derive: how much memory to buy. Clicks are brutally front-loaded — a link is shared, it runs hot for a few days, and then it is dead more or less forever — so nearly all of the hit rate comes from recency, and plain LRU over a modest budget does most of the work. Too small and the working set thrashes: every evicted-then-reclicked code pays a sharded store lookup, and the chaos run prices that shape at +9ms at p99. Too large and you are renting RAM to hold links nobody will open again. What I’d really size against, though, is the restart. A cold cache drops the entire read fleet through to the store at once, so the store has to be able to carry the full redirect rate unaided — that same slower-but-correct mode — or the cache has stopped being a latency layer and become a dependency you didn’t agree to.

  • 100 : 1reads : writes
  • ~4k/sredirects
  • <10msp99 lookup

What the new pieces do

Cachecache
In-memory store of the hottest code→url mappings. A redirect that hits here returns in under a millisecond.

Back of the envelope

~100:1 reads : writes
created once, clicked thousands of times
hit ≈ <1ms RAM
vs a sharded KV lookup
immutable mapping ⇒ no invalidation
caching is pure upside

Step 4 · Don’t fall over

Now serve the whole internet

One server was fine for a demo. Under real traffic it melts, and if it dies the whole service goes dark. Reads and writes also have wildly different shapes — mixing them on one box wastes resources. How do you scale?

Shorten ServiceCacheRedirect Service
New in this step: Shorten Service, Redirect Service. · swipe to pan the diagram

One server melts under real traffic, and reads outnumber writes ~100:1. How do you scale?

  1. Vertical scaling hits a ceiling and is still one machine — when it dies, the whole service goes dark. And it can’t shape resources to the 100:1 read/write split.

  2. Sharding the DB (step 5) helps storage, but the single app server is still the bottleneck and single point of failure. You also want to scale reads and writes independently.

  3. Run many small stateless copies behind a load balancer so one dying box never takes the system down, and split into a Shorten (write) and Redirect (read) service so the read fleet — 99% of traffic — scales on its own.

Add a load balancer and run many copies (horizontal scaling). Split the work into a Shorten Service (write) and a Redirect Service (read) so each scales — and fails — on its own. The read fleet, carrying 99% of traffic, can grow independently.

Why this piece earns its place

Splitting the fleets buys independent scaling; it does not by itself buy the failure isolation the split implies, because both services still land on the same store and the same shards. A burst of writes — a bulk import, someone scripting the API — can eat the store’s connection budget, and then redirects, 99% of the traffic, get slower because of the 1%. So the two services get separate connection pools and separate quotas against the store, and I’d settle the shedding order before the incident rather than during it: under pressure you refuse to shorten, never to redirect. A failed POST is a retry the creator sees and repeats; a failed redirect is a broken link on a stranger’s page. The split also changes what “healthy” has to mean to the load balancer. An instance that answers a static health endpoint but has lost its store connection will keep being routed to and fail every redirect it gets, so the check has to do a real lookup — otherwise scaling out just multiplies the number of machines confidently serving nothing.

What the new pieces do

Shorten Serviceservice
Handles the rare write: takes a long URL, gets a fresh code, and saves the mapping. Scales on its own.
Redirect Serviceservice
Handles the overwhelming majority of traffic: look up a code, return a redirect. The fleet you scale the most.

Step 5 · A mountain of links

Where do billions of rows live?

At ~100M new links a month you cross billions of rows within a few years. No single database holds that comfortably, and one disk can’t serve the read rate on its own. Where do they live?

Shorten UIVisitorAPI GatewayShorten ServiceCacheRedirect ServiceKey Gen ServiceURL Store
The system as it stands at this step. · swipe to pan the diagram

You cross billions of code→url rows in a few years. Where do they live?

  1. A single table of billions of rows outgrows one machine’s disk and read capacity, and you never need joins or queries — only exact lookups by code. A simpler, shardable store fits better.

  2. The access pattern is a pure lookup by code, so a KV store is ideal; sharding by code-hash spreads billions of rows across machines, each shard small and fast. Random-looking codes spread load evenly — no hotspot.

  3. RAM can’t durably hold billions of mappings, and a cache is for the hot subset, not the system of record. You need a durable sharded store behind the cache.

Use a key-value store — the access pattern is a pure lookup by code — and shard it: split the keyspace across many machines, routing each code to its shard by a hash of the code. Each shard stays small and fast; add more as you grow.

Why this piece earns its place

The arithmetic lands on six to eight shards; the decision that actually matters is the function that maps a code to one of them. Hash the code and take it modulo the shard count and growing from six to seven rehomes nearly every key — a migration you get to run underneath live redirect traffic, with a window where a code exists in two places and the wrong one can answer. So fix the mapping now: hash into a large, constant number of virtual buckets and map buckets onto machines. Adding capacity then moves whole buckets, one at a time, and a bucket in flight is a thing you can reason about and roll back. I’d also be precise about what “random codes spread load evenly” does and doesn’t cover. It is true of bytes and broadly true of traffic in aggregate, but one link going viral is a single key on a single shard, and no hash function splits a key. That case is answered by replicating the shard that holds it, not by choosing a cleverer shard key.

  • 100Mnew links / mo
  • 12B+rows in 10 yrs
  • ~6 TBof mappings

Back of the envelope

100M/mo × 12 × 10 yrs = 12B rows
the row count is the easy part — the bytes are what decide the architecture
12B × ~500 B/row ≈ 6 TB of mappings
too big for one box, which is what forces the next line
6 TB ÷ ~1 TB/shard ⇒ ~6–8 shards to start
sized from the arithmetic, with room to split as it grows
shard by hash(code)
random codes spread load evenly
KV lookup, no joins
each shard stays small and fast

Step 6 · Count the clicks

Who clicked, from where?

Owners want stats — clicks, countries, referrers, devices. But writing an analytics row on every redirect would slow the one thing that must stay fast: the redirect itself. How do you have both?

Redirect Serviceread pathAnalyticsclick statsClick EventsKafka
New in this step: Analytics, Click Events.

Owners want click stats, but writing a row on every redirect would slow it. How?

  1. Blocking the redirect on an analytics write couples the must-be-fast path to a slower one, and an analytics outage would break redirects. Counting must never gate the click.

  2. Owners genuinely want stats, and you can have both — you don’t have to choose. Decouple counting from serving instead of dropping it.

  3. The Redirect Service emits an event to Kafka and returns immediately; downstream consumers fold events into the analytics store at their own pace. If analytics lags or crashes, redirects keep flying.

Make analytics asynchronous. The Redirect Service fires a lightweight event onto a stream (Kafka) and immediately returns the redirect. Downstream consumers fold those events into an analytics store at their own pace — the click is never blocked on counting it.

Why this piece earns its place

Fire-and-forget is a delivery guarantee, and the guarantee is at most once. Clicks will be lost: sitting in a producer buffer when an instance is recycled mid-deploy, and across any window where the stream is unreachable — which is exactly the outage this page lets you trigger. That is the right trade, but it has to be a stated one. The owner-facing figure is an estimate, and if you label it “clicks” with no qualifier, the first customer who diffs it against their own site analytics files a bug you can never close. Call it approximate and reconcile toward the destination’s numbers instead of promising to match them. The implementation detail that decides whether the trade holds at all is what the producer does when its buffer fills. Most clients default to blocking the caller until there is room, which silently turns the decoupled path back into a synchronous one — the redirect now waits on the analytics system, the single thing this step exists to prevent. Configure it to drop, and count the drops, because that counter is your only evidence the loss happened.

What the new pieces do

Analyticsstore
Aggregates clicks, countries, referrers and devices — built up asynchronously so it never slows a redirect.
Click Eventsbus
A firehose of click events. Lets redirects fire-and-forget while analytics catches up at its own pace.

Step 7 · The sharp edges

Aliases, expiry & abuse

Real users want custom aliases (/launch), links that expire, and — less welcome — spammers who shorten malicious URLs or hammer your API.

Shorten UIVisitorAPI GatewayShorten ServiceCacheRedirect ServiceKey Gen ServiceURL StoreAnalyticsClick Events
The system as it stands at this step. · swipe to pan the diagram

Let writers request a custom code (check it’s free first). Store an optional TTL so expired links return 404 and get cleaned up. Add rate limiting at the gateway and a safe-browsing check on new URLs, so one bad actor can’t ruin the service for everyone.

Why this piece earns its place

Screening a URL on write quietly assumes a URL is a fixed thing. It isn’t — the destination is someone else’s server, and the profitable play is to shorten something harmless, pass the check, seed the link, then swap the page for a phishing form a week later once it has spread. So screening can’t only be a gate at creation. It has to re-check links as they get popular, which the click stream from the previous step already tells you about for free, alongside a report path and a way to kill a code on the spot. Killing it is where this design needs the one thing it has so far been able to avoid. The mapping is write-once, which is precisely why there is no invalidation anywhere in the caching story — but a takedown is a deletion, and deletion is the write the cache never hears about. Remove the row and a flagged link keeps redirecting from memory until the entry happens to age out. The kill path has to reach the cache explicitly. It is also the sharpest argument for the 302: a 301 you handed out months ago lives in browsers you cannot reach, and can never be recalled.

You did it

You just designed a URL shortener.

Shorten UIVisitorAPI GatewayShorten ServiceCacheRedirect ServiceKey Gen ServiceURL StoreAnalyticsClick Events
The finished design, end to end. · swipe to pan the diagram

Everything you assembled, in order

  • Client → server → database — the spine shared by shorten and redirect.
  • Base62 codes from a Key Gen Service — unique by construction.
  • A Redis cache makes the read-heavy redirect path sub-millisecond.
  • Load balancer + split read/write services to scale horizontally.
  • A key-value store sharded by code holds billions of mappings.
  • Async click events via Kafka — analytics never blocks a redirect.
  • Custom aliases, TTL expiry, 404s and rate limiting for the real world.

Where an interviewer pokes next

Getting the boxes right is the easy half. These are the questions that separate a candidate who drew the diagram from one who has run the thing. Answer each one out loud before you open it.

  1. 301 vs 302 — which redirect should you return?

    A 301 (permanent) lets browsers and proxies cache the redirect, so later clicks skip your server — great for load, but you lose click tracking. A 302 (temporary) routes every click through you, so you can count and re-target it, at the cost of more traffic. If analytics matters, use 302; if pure redirect performance matters and the target is fixed, 301.

  2. How do custom aliases coexist with generated codes?

    Custom aliases (/launch) are user-chosen, so unlike counter-based codes they CAN collide — you must check availability and reserve atomically (a unique constraint on the code). They live in the same code→url store; the only difference is generated codes are unique by construction while custom ones need a taken-check. Reserve a namespace or length so they can’t clash with the generator’s output.

  3. How do you expire links and handle the unhappy paths?

    Store an optional TTL; a redirect past expiry returns 410 (gone) and a background sweep (or the store’s native TTL) reclaims it. Unknown code → 404, too many requests → 429 (rate limiting at the gateway). Designing these explicit failure responses is what separates a demo from a service people trust.

  4. How do you stop abuse — spam and malicious URLs?

    Rate-limit creation per IP/account at the gateway, run new target URLs through a safe-browsing/reputation check before (or shortly after) accepting them, and support takedown of flagged links. Shorteners are attractive to phishers precisely because they hide the destination, so abuse prevention is core, not optional.

  5. Why a KGS instead of each server using its own counter?

    Independent per-server counters would collide (two servers both mint code 125). The KGS centralizes issuance so codes are globally unique; to keep it from being a bottleneck or SPOF, it hands out ranges/blocks of ids to each app server ("you own 1–1000"), which they consume locally and refill occasionally. Global uniqueness with almost no per-request coordination.

Check yourself — the answers, and why

Eight steps in, these are the calls you should be able to make cold. Pick one, then read why.

  1. The short code is generated by…

    Counting up then base62-encoding never collides; the KGS keeps it globally unique.

  2. Redirects are sub-millisecond because of…

    Mappings never change, so caching is pure upside — a hit returns from memory.

  3. Read and write are split into separate services because…

    The read fleet carries ~99% of traffic and can grow (and fail) on its own.

  4. Billions of mappings live in…

    Pure lookup-by-code fits a sharded KV store; random codes spread load with no hotspot.

  5. Click analytics is asynchronous so that…

    Fire an event to a stream and return; consumers aggregate later — redirects never wait.

How you’d open this design in an interview

Before any boxes: agree what it must do, pin the qualities that shape everything, then build — naming each trade-off as you make it. The walkthrough above is that exact order.

What it must do

Agree on these before drawing a single box.

  • Shorten: given a long URL, return a unique short code.
  • Redirect: GET /code sends the visitor to the original URL — the hot path.
  • Custom aliases: vanity codes like /launch, coexisting with generated ones.
  • Expiry: optional TTL — links can die on schedule.
  • Click analytics: owners see clicks, referrers, countries, devices.

The qualities that shape everything

Each one names the mechanism that buys it.

Redirects survive anything (availability first)
Stateless read fleet behind a load balancer; a code→url mapping is immutable, so serving slightly stale data is harmless — choose availability over consistency on the read path.
Sub-millisecond redirect latency
Cache-aside Redis holding the hot mappings in front of the store; a miss falls through and re-warms — the cache is a latency layer, never the source of truth.
Read-heavy scale (~100 reads per write)
Split the shorten (write) and redirect (read) services so the read fleet scales independently.
Billions of mappings
Key-value store sharded by a hash of the code — the access pattern is a pure key lookup, and random-looking codes spread load with no hotspots.
No two links ever collide
A counter encoded in base62, handed out by a Key Generation Service — unique by construction, no “is this taken?” check on the write path.
Analytics never slows a redirect
Fire-and-forget click events onto a stream (Kafka); analytics is best-effort and decoupled, redirects stay on the fast path.
Abuse doesn’t take you down
Rate limiting at the gateway, malicious-URL screening on writes, 404s for dead codes handled cheaply.

The trade-offs you say out loud

Senior signal isn’t the boxes — it’s naming what you gave up and why it was the right price.

302 redirects over 301 permanent

A 301 lets browsers cache the hop — fewer hits, but you go blind: no analytics, no expiry, no kill-switch on abuse. A 302 routes every click through you; you pay servers for control. Most real shorteners pay.

Counter + base62 (KGS) over hashing the URL

Hashes collide and need retry loops; the same URL hashes identically for different owners. A dealt-out counter is unique by construction — the KGS becomes a coordination point, so it deals codes in pre-allocated batches.

Sharded key-value store over one relational database

The query is always “code → url”, never a join. A KV store sharded by code-hash scales to billions of rows; you give up transactions you never needed.

Availability over strict consistency

Mappings are write-once. A redirect served from a stale cache is still correct, so the read path can favor being up over being perfectly in sync — the cheapest CAP call you’ll ever make.

The answer, out loud

What a strong answer to “Design a URL Shortener” sounds like, first question to last trade-off. It is about 7 minutes of talking; the whiteboard and the interviewer fill the rest of the 45. Read it aloud once, then close the page and give it yourself.

  1. 0–4 min

    Scope the two operations

    Let me scope it first. There are two operations and they could hardly be more different. Shortening is a write — a long URL in, a code back — and it happens once in a link’s life. Redirecting is a read, and it happens every time anyone taps that link — here, around a hundred times per link. So it is one system with a rare write path and a very hot read path, and most of what follows falls out of that ratio. I’d also confirm whether aliases, expiry and click stats are in scope, since each lands on the write path.

  2. 4–9 min

    The skeleton, and which redirect

    The skeleton is the plain one: a client, a server, a store. To shorten, the server writes a row mapping the code to the long URL; to redirect, it looks the code up and answers with a redirect. The decision worth arguing here is which redirect. A 301, permanent, lets browsers remember the destination and skip me next time — lovely for my servers, and it blinds me: no click stats, no expiry, no way to pull a link down. A 302 routes every click back through me. I pay for servers and keep control, and every real shortener I know of pays.

    Built in step 1: Two jobs, one server
  3. 9–16 min

    Minting the code

    Then the code itself. The obvious move is to hash the long URL and take the front of it, and I’d talk myself out of that: truncated hashes collide, so I’m back to checking whether a code is taken and retrying, and two people shortening the same URL land on the same code. Random is the same problem with worse odds as the space fills. So I treat it as a number. Keep a counter, encode it in base62, and hand it out from a small key generation service so two servers never mint the same one. Seven characters gives about three and a half trillion codes, and uniqueness costs no lookup at all.

    Built in step 2: How do we mint the code?
  4. 16–23 min

    The path that carries the traffic

    Now the part that actually carries load. A redirect that goes to disk is a redirect somebody is waiting on, so I put a cache in front — Redis, holding the hot mappings. A hit comes back in well under a millisecond; a miss falls through, answers, and warms the cache for the next person. What makes this unusually easy is that a mapping never changes once written. No invalidation, nothing to reconcile. It is the rare cache that is pure upside, which is worth saying out loud, because caching is normally where the bugs come from.

    Built in step 3: Reads dwarf writes
  5. 23–29 min

    Scale out and split the fleets

    One box is fine until it dies and takes everything dark with it, so I run many stateless copies behind a load balancer. And because reads outnumber writes so heavily, I’d split them into a shorten service and a redirect service. That isn’t tidiness — the read fleet carries essentially all the traffic, and I want to scale it, deploy it and lose instances of it without touching writes. It also lets me make a clean availability call: a mapping is written once and never edited, so a slightly stale one is still the right answer. On the read path I take availability over consistency.

    Built in step 4: Now serve the whole internet
  6. 29–35 min

    Where the rows live

    Storage. A hundred million new links a month is twelve billion rows over ten years, and at roughly five hundred bytes a row that is about six terabytes. That does not sit on one machine, so it is sharded — which raises what kind of store. I never join, never filter, never sort; every query is one exact lookup by code. So a key-value store, sharded by a hash of the code, giving up transactions I was never going to use. Codes look random, so rows spread evenly, and at about a terabyte per shard I’d start with six to eight.

    Built in step 5: Where do billions of rows live?
  7. 35–40 min

    Counting clicks without slowing them

    Click stats are the last piece, and the property I care about is that counting a click never delays one. So the redirect service fires a small event onto a stream — Kafka — and returns immediately, waiting on nobody. Consumers fold those events into an analytics store on their own schedule and build what the owner sees: clicks, referrers, countries, devices. If analytics falls behind, or falls over, a redirect never notices.

    Built in step 6: Who clicked, from where?
  8. 40–43 min

    Aliases, expiry, abuse

    Then the edges. Custom aliases are user-chosen, so unlike generated codes they genuinely can collide — that needs an availability check and an atomic reservation on the write. Expiry is an optional TTL on the row. And because a shortener hides where a link goes, it is a gift to phishers, so I rate limit creation at the gateway and screen new URLs against a reputation service. An unknown code is a 404, an expired one a 410, too many requests a 429.

    Built in step 7: Aliases, expiry & abuse
  9. 43–45 min

    Close on the trade

    To close: client, server, store; base62 codes from a key service; a cache that works because mappings are immutable; split read and write fleets behind a load balancer; a sharded key-value store underneath; analytics on a stream. The trade running through it is that I chose the 302 and paid for servers to keep control of the link, and chose availability over consistency because a stale mapping is still the correct one. With more time I’d push the read fleet geographically closer to visitors, since a redirect is one key lookup, and work through branded custom domains resolving against the same store.

What this teaches

Learn system design by building a URL shortener like Bitly or TinyURL step by step. An interactive guide covering the client–server skeleton, base62 key generation, caching the read-heavy redirect path, splitting read/write services, sharding billions of mappings, async click analytics, and the unhappy paths.

Key takeaways

  • Client → server → database — the spine shared by shorten and redirect.
  • Base62 codes from a Key Gen Service — unique by construction.
  • A Redis cache makes the read-heavy redirect path sub-millisecond.
  • Load balancer + split read/write services to scale horizontally.
  • A key-value store sharded by code holds billions of mappings.
  • Async click events via Kafka — analytics never blocks a redirect.
  • Custom aliases, TTL expiry, 404s and rate limiting for the real world.

Concepts covered

  • What is a URL shortener?
  • Two jobs, one server
  • How do we mint the code?
  • Reads dwarf writes
  • Now serve the whole internet
  • Where do billions of rows live?
  • Who clicked, from where?
  • Aliases, expiry & abuse
RUN IT YOURSELF

Base-62 short codes, in Python & TypeScript

A URL shortener turns a numeric id into a short slug with base-62 encoding — and back. Here it is in both languages, running live. Switch tabs, read the comments, edit, and hit Run.

HOW TO READ THE CODE — 4 IDEAS
  1. Each new URL gets an auto-increment id (a number); the slug is that number in base 62.
  2. Encode by repeatedly taking n % 62 and dividing, prepending each digit (steps 1–2).
  3. Decode is the reverse — Horner's method walks the slug back to a number (step 3).
  4. It is a bijection: every id has one slug and vice-versa, so no collisions.
CPython · WebAssembly
built to be redirected, not memorized — make the calls, drop the cache, run the gauntlet.
Finished this one? 0 / 65 System Designs done

Explore the topic

See this alongside everything else on the same subject — handbooks, system designs, challenges and tools, in one place.

More System Designs