A URL Shortener Is Easy, Until You Scale It

2026-09-26

When I was reading the URL shortener chapter from System Design Interview by Alex Xu, something clicked. On the surface, a URL shortener looks like the simplest app you can build: take a long URL, hand back a short one, redirect on click. What is the big deal?

But the longer I sat with it, the more I realized the engineering hiding behind that simplicity is genuinely nice to look at. A URL shortener is a CRUD app, right until it is not. The moment you add real traffic, real money, and real constraints, it turns into a system design problem: caching, replication lag, capacity planning, expiration cleanup, the whole thing.

So this post is my attempt to reason from the constraints to an architecture, instead of jumping straight to "frontend + backend + Postgres." It is about taking your thinking out of "write some backend, write some frontend, connect them with an API call" and into "engineer a system that actually holds up at scale."

I have also implemented this design in code so you can see how these ideas translate into a working service: URL SHORTNER implementation on GitHub.

The Constraints

These are the same constraints from the book:

  • Traffic: 100 million URLs generated per day
  • Short key: as short as possible, and if possible the user can provide a custom one (pick any hashing/scheme based on your requirements)
  • Operations: create URL, delete URL, edit URL, get URL
  • Availability: the shortened URL should be usable the moment it is generated
  • Expiration: every link has a validity window: 2 hours, 1 day, 1 month, 1 year, or custom
  • Read : write ratio: around 10:1

That last one is worth putting numbers on, because it drives the entire design.

text
writes/day  = 100,000,000
writes/sec  = 100,000,000 / 86,400 = ~1,160
reads/sec   = 1,160 * 10         = ~11,600  (10:1)
OperationPer dayPer second
Writes100M~1,160
Reads1B~11,600

Reads outnumber writes by an order of magnitude. Keep that in your head, because almost every decision below is really about serving reads without killing the database.

The Naive Design

This is how most of us think on day one: the client talks to the backend, the backend does reads and writes directly against Postgres, done.

Naive design: client to backend to a single Postgres instance

It is good enough on your local machine. In production, it is a disaster waiting to happen. At the scale we are targeting, a single database handling both the read and write load will fall over under the simultaneous request pressure. You can try to throw more hardware at it (vertical scaling), but that only buys you time, and it costs a ton of money on a cloud provider like AWS. And when something eventually does go wrong, the user experience gets sacrificed too: a database that crashes under high volume takes your redirects down with it.

A redirect that 500s is the worst kind of failure for this product, because a redirect is the one thing users expect to be instant and invisible.

Add an API Gateway (a.k.a. nginx as a Reverse Proxy)

Next, let us route every request through a reverse proxy / API Gateway before it reaches the backend, so the client is directed to the correct endpoint.

Adding a reverse proxy / API gateway in front of the backend

And honestly? Right now, this does nothing. It has no business logic. It just routes the request to the right destination. I am adding it because it becomes useful in the next step: it is the single place where I can decide where a request should go, based on what kind of request it is.

Read Replicas

Our constraint says reads dominate writes roughly 10:1, which makes complete sense for a URL shortener. So how do we take the read pressure off the primary database? There are two classic answers:

  1. Create a read replica of the primary DB and send all read requests there.
  2. Use caching to keep the most recently and frequently accessed links close at hand.

Let us start with the read replica. Using the gateway, you can send reads to the replica and writes to the primary.

Read replica: writes to primary, reads to replica

There is a catch, though. For a small amount of data, a user can read back what they just wrote almost instantly, and the replica looks perfectly in sync. But as data accumulates, replication lag grows. The replica falls further behind the primary, and that delay directly costs you the "instant availability" the URL is supposed to have. A freshly created link might 404 for a few seconds, which is exactly the thing our availability constraint forbids.

Replicas also cost real money on a cloud provider. So replication alone does not solve this: it just moves the problem around.

Caching

Now let us see what happens if we introduce a cache.

The client hits the load balancer (nginx), which picks an available server and routes the request based on its type (get_url, create_url, edit_url, delete_url). A read request checks the cache first: on a hit, we return the value immediately; on a miss, we fall through to the database and populate the cache.

Cache in front of the database

For writes, the pattern matters, so let us be precise about the naming because these two get mixed up all the time:

  • Cache-aside (lazy loading): the application reads the DB itself. On a write, you update the DB and then invalidate (delete) the cache entry, so the next read reloads fresh data.
  • Write-through: you write to the DB and the cache together, synchronously, before returning.

For this use case, a good default is cache-aside on reads and write-through (or DB-write-then-invalidate) on writes. What you do not want is to update the DB and leave a stale cache entry behind.

To be honest, for our use case this architecture actually fits really well if the cache size is chosen wisely. From my understanding, here is how I would size the RAM.

Assuming an average long URL of ~150 bytes, each entry also carries:

FieldSize
row_id16 bytes
long_url150 bytes
short_key7 bytes (Base62)
timestamp8 bytes

Pure data size per row:

text
data     = 16 + 150 + 7 + 8 = 181 bytes/row
overhead = headers, pointers, fragmentation, etc.
in Redis = ~280 bytes/entry (safe estimate)

I do not think we need to keep every link in hot RAM. Most traffic hits a small set of active or popular short URLs, so a modest cache can absorb the vast majority of reads. If I keep around 100 million active links in cache:

text
total_mem = 100,000,000 * 280
        = 28,000,000,000 bytes
        = 28 GB ~= 26.1 GiB

That is a much more realistic number to start with than "cache everything." Add proper TTL plus an LRU/LFU eviction policy, and let the cold keys fall back to the database. Then monitor real traffic and pick the right Redis plan (or run it on-prem) instead of over-provisioning a huge amount of hot RAM on day one.

A Quick Word on Generating the Short Key

One thing I would change from the usual "just hash it" approach: a hash function is deterministic, so two long URLs can collide in the first 7 Base62 characters, and now you need to detect that and regenerate. A cleaner approach is to generate a unique numeric ID (a counter, or something like a Snowflake ID), then encode that ID in Base62. You get guaranteed uniqueness with no collision handling, and the key is still as short as possible. Custom aliases then become a simple uniqueness check against that same key space.

Expiration and Cleanup

For expiration, the straightforward approach is to store an expires_at timestamp column alongside each link. A background job (a daily or hourly sweeper) checks that column, and when the window for a link has passed, it notifies the owner, the person who created the link, to renew it. If they decline (or do not respond), the row gets deleted so the database does not fill up with stale data.

A couple of details worth getting right:

  • The person clicking a redirect is anonymous, so "ask the user to renew" has to target the creator/owner, not the visitor. The visitor just gets an expired page.
  • On the cache side you do not need to do any of this manually: Redis TTL handles it. The DB is the source of truth, so the sweeper is only responsible for the durable rows.
  • If you want to avoid a sweeper entirely, you can delete lazily: check expires_at on read and treat an expired row as a miss. A sweeper is usually simpler at this scale, though, because it keeps the table from bloating.

Sharding and Periodic Cleanup

Techniques like sharding can also be introduced to improve performance and increase cost. I am joking. It totally depends on your use case. For a lot of setups, you can start with monthly or weekly database cleanups to remove stale data and only reach for sharding when you genuinely outgrow a single primary.

The takeaway I keep coming back to: the constraints decide the architecture, not the other way around. The moment you write down "100 million writes a day" and "reads are 10x writes," the naive design disqualifies itself, and every good decision after that (gateway, cache, TTL, eviction, cleanup) falls out of the numbers instead of out of habit.

Full Implementation

If you want to see how this looks in actual code, the complete working implementation is available here: URL SHORTNER on GitHub. Use it to follow along with the design above and see how each part fits together in a running service.

← Back to home