developer

How to Build Your Own URL Shortener (And Why You Probably Shouldn't)

A practical guide to building a URL shortener from scratch: architecture, base62 encoding, database and cache design, plus an honest look at when a managed platform saves you months of work.

Team U2L • 17 min read

Building a URL shortener from scratch is a great weekend project and a bad long-term product. The core is small: a base62 encoder, a key-value store, an HTTP redirect handler, and a cache. The hidden cost is everything around it: abuse detection, edge distribution, analytics pipelines, custom domains, SSL provisioning, uptime, and the on-call rotation that keeps a service that thousands of people bookmark from ever going down.

Every backend engineer eventually writes a URL shortener. It's the "hello world" of system design interviews, the weekend hack that teaches you Redis, and the side project behind countless abandoned GitHub repos. And honestly, it's a great learning exercise. You get to touch database design, caching strategy, HTTP redirects, and (once you actually deploy it publicly) a masterclass in dealing with spam.

But there's a wide gap between "I built a URL shortener that works on my laptop" and "I run a URL shortener that thousands of people trust with their branded links." This guide walks through both. We'll design a production-shaped system from scratch, look at what the code and infrastructure actually cost to run, then talk honestly about when doing this yourself is a good use of a couple of weekends and when it just quietly steals six months of your engineering time.

Table of Contents


What a URL Shortener Actually Does

A URL shortener is a service that maps a short, opaque token to a longer destination URL and issues an HTTP redirect when the short token is requested. The user-facing product is one line of storage (short key to long URL) and one HTTP handler (fetch the mapping and issue a 301 or 302). Everything else, and there is a lot of everything else, is around that core.

The apparent simplicity is exactly why it's a rite of passage. The real product has three axes of complexity: read throughput (redirects are hot and must be fast), write reliability (a lost link is a broken link forever), and abuse resistance (public URL shorteners are magnets for spam, malware, and phishing). Miss any one and you don't have a shortener, you have a support ticket queue.

On the SEO side specifically, Google's documentation on redirects is worth reading before you decide anything about 301 vs 302 handling. It's short, clear, and will save you from a class of bad architectural choices.

If you've read our what is URL shortening explainer and thought "I could build that in a weekend," you're right. You could also build a hobby car in a garage. Whether it's roadworthy is a different question.


The Minimum Viable Architecture

At its simplest, a working URL shortener needs six pieces:

  1. A client-facing API to accept long URLs and return short ones.
  2. A key generator that produces unique short tokens.
  3. A storage layer that persists the token-to-URL mapping.
  4. A redirect handler that reads the mapping and issues an HTTP redirect.
  5. A cache so hot links don't hammer the database.
  6. A safety layer so you're not helping phishers.

Here's a rough diagram of the request flow:

Client → API Gateway → Application Layer → Cache (Redis)
                                             ↓ (miss)
                                          Database → Cache (write) → Redirect

The write path (create a short link) is low volume, latency-tolerant, and straightforward. The read path (redirect) is the hot path. It has to be sub-100ms in almost all cases, or you'll produce a noticeable pause every time someone clicks a link. Most of the interesting design decisions live in the read path.

For a hobby version you can collapse this into a single Node.js or Go binary with an embedded SQLite database, and that'll happily handle a few thousand redirects per day on a $5/month VPS. That's absolutely a legitimate way to learn. Just don't confuse it with a service you can hand to marketing to use for the Q4 campaign.


Generating the Short Key

The short key is the alphanumeric token that goes after your domain. Something like u2l.ai/aB3xK9p. Two mainstream approaches exist, and both have honest tradeoffs.

Approach 1: Auto-increment ID plus base62 encoding. Insert into the database, get back an integer ID, base62-encode it into a short string. Base62 uses the 62 alphanumeric characters (0-9, a-z, A-Z), which gives you a 7-character token that can represent about 3.5 trillion unique URLs (62^7). It's collision-free by construction because IDs are unique, and it's easy to reason about.

const ALPHABET = "0123456789abcdefghijklmnopqrstuvwxyzABCDEFGHIJKLMNOPQRSTUVWXYZ";

function encode(id) {
  let short = "";
  while (id > 0) {
    short = ALPHABET[id % 62] + short;
    id = Math.floor(id / 62);
  }
  return short.padStart(7, "0");
}

The catch is that auto-increment IDs are sequential, meaning your short URLs are trivially guessable. Someone can iterate through and enumerate every link on your platform. That's fine for a hobby project. It's not fine if any of your customers care about link privacy.

Approach 2: Random token with collision check. Generate a random 7-character base62 string, check if it exists in the database, retry on collision. Any single insert almost never collides, but the birthday paradox means your first collision is likely once you pass roughly two million links in a 7-character space, so the check is not optional.

function random7() {
  return Array.from({ length: 7 }, () =>
    ALPHABET[Math.floor(Math.random() * 62)]
  ).join("");
}

Random tokens are non-enumerable, which is a real privacy win. The downside is you now have a "check, maybe retry" step on the write path, and at scale you need to think about whether your randomness source is good enough (Math.random is not; use crypto.randomBytes in Node or your language's equivalent).

Most production systems we've seen use random tokens for user-created links and a separate ID-based scheme for internal system generation. If you're building for real, start with random and don't look back.

Custom aliases are just token overrides: the user provides the short slug, you validate it doesn't collide with an existing one or a reserved word, and you skip the generator. Not hard, but the list of reserved words gets long fast (admin, api, login, help, about, on and on).


Database Choice and Schema

The workload is 95% reads, 5% writes, and each row is a tiny key-value pair. That points squarely at a key-value store, but "just use Postgres" is a completely reasonable answer for the first few years of any real product.

Here's a minimal schema:

CREATE TABLE links (
  short_code   VARCHAR(16) PRIMARY KEY,
  destination  TEXT NOT NULL,
  created_by   BIGINT REFERENCES users(id),
  created_at   TIMESTAMPTZ NOT NULL DEFAULT NOW(),
  expires_at   TIMESTAMPTZ,
  status       SMALLINT NOT NULL DEFAULT 0,
  click_count  BIGINT NOT NULL DEFAULT 0
);

CREATE INDEX idx_links_created_by ON links(created_by);
CREATE INDEX idx_links_expires_at ON links(expires_at) WHERE expires_at IS NOT NULL;

Six columns cover the core. Once you start adding branded domains, UTM presets, password protection, deep linking, and A/B testing, this schema grows fast. Real production tables at this shape end up with 20-40 columns; don't be surprised. If you're wondering what all those columns do, our writeup on what happens when you click a short link walks through the redirect-path features one by one.

Which database? For under 10 million links and a few hundred writes per second, Postgres or MySQL are the boring right answer. Both are battle-tested, both have great tooling, and you almost certainly already have expertise. Above that scale, or if you need geographic distribution, look at DynamoDB, Cassandra, or Cloudflare's KV / D1 offerings.

A common trap: reaching for MongoDB because "URLs are documents." They're not. A URL mapping is a single tiny KV pair, so a document model buys you nothing here. Use a KV store or a relational database you already know how to run.


The Redirect Path (Where Latency Lives)

Every redirect adds latency between a user clicking a link and reaching their destination. The design goal is to make this invisible: under 50ms end-to-end in most regions, ideally under 20ms in your hot regions. That number matters more than any other in the system.

The path looks like this:

GET /aB3xK9p
  → Edge (POP nearest user)
    → Cache lookup (Redis or edge KV)
      → HIT: return 301/302 immediately
      → MISS: read from database
              write to cache
              return 301/302

Cache hit ratio is everything. With a 95% cache hit ratio and Redis on the same machine as your app, cache reads take under a millisecond. Miss to your primary database costs 10-50ms depending on region and load. That difference is why a well-designed shortener feels instantaneous and a poorly designed one feels sluggish.

Cache TTL is a judgment call. Long TTLs (hours) minimize database load but make link edits slow to propagate. Short TTLs (seconds) make edits fast but load up the database. A common middle ground is 5-15 minutes with a manual invalidation path when a link is updated.

301 vs 302 is a real product decision, not just a technical one. A 301 is permanent and browser-cacheable, which means faster subsequent redirects and a stronger signal to search engines that the destination is the canonical URL, but it also means you can't change the destination without cache-busting. A 302 is temporary and never cached, so destination edits take effect immediately at the cost of more traffic hitting your servers. Our writeup on 301 vs 302 redirects covers the tradeoffs in more depth. Most managed shorteners let users pick per-link.

Edge deployment matters. A monolithic app in one region will always be slow for users in the opposite hemisphere. If you're serious about latency, you deploy the redirect handler to edge locations (Cloudflare Workers, Fastly Compute, AWS Lambda@Edge) and only fall back to a central database for cache misses. This is the biggest single architectural change between a hobby shortener and a production one.


Analytics Without Blocking the Redirect

Every real URL shortener logs analytics on redirect: timestamp, IP-derived geo, device, browser, OS, referrer. The mistake is doing this synchronously and slowing the redirect down.

The right pattern is fire-and-forget queue write:

  1. Redirect handler reads the destination from cache.
  2. Redirect handler enqueues an analytics event (Kafka, SQS, Redis Streams, Cloudflare Queues, whatever) without waiting for the write to complete.
  3. Redirect handler returns the HTTP 301/302 to the user.
  4. A separate consumer batches events and writes them to your analytics store.

That "without waiting" step is the discipline. A blocking analytics write adds 5-30ms to every redirect, and users notice.

The analytics store is a separate database decision. Time-series databases (ClickHouse, TimescaleDB) or purpose-built analytics engines are common. A cheap starting move: log to your primary database in a clicks table and switch to a real analytics engine only when volume forces you.

There's also a privacy question. If you're logging raw IPs and you have EU users, truncate them (drop the last octet) or hash them with a secret salt you rotate regularly. A plain SHA-256 of an IPv4 address is not anonymization: there are only about four billion of them, so the hash can be reversed by brute force in minutes. GDPR Recital 30 names IP addresses as online identifiers that can identify a person, which is the reference point most legal teams start from. Ignoring this is a GDPR problem waiting to bite you.


The Hidden Costs Nobody Talks About

This is where the "just build it yourself" honeymoon ends. The core system is a weekend. The care and feeding is a full-time job. In no particular order:

Custom domains and SSL. Users want their own short domains (yourbrand.co/promo). That means DNS validation, dynamic SSL certificate provisioning (Let's Encrypt, ACME protocol), certificate renewal automation, and edge routing that resolves the right rule for the right host. Building this properly takes weeks and breaks often.

Abuse and spam. The day you go public, you'll get abuse. Phishing kits, malware links, spam runs. You need URL scanning (the Google Safe Browsing Lookup API is the standard baseline for non-commercial projects, and Google points commercial products to its Web Risk API; either way you'll want additional layers), rate limiting per IP and per account, slug blocklisting, and a takedown process. Miss any of this and your domain lands on browser warning lists, at which point every link on your platform stops working. This has ended real projects.

Uptime. Once anyone bookmarks a link, it must work forever. That's a promise. Building the SRE muscle to hit four nines of uptime (52 minutes of downtime per year) is expensive: monitoring, alerting, on-call, incident response, chaos testing, capacity planning. Every one of these takes real engineering time you're not spending on your actual product.

Analytics depth. "Show me clicks by country" is easy. "Show me unique visitors filtered to mobile devices from Tuesday's Instagram post, grouped by referrer, with a time series overlay" is a full-blown analytics product. Users expect the latter. Every added dimension is another schema decision, another index, another query optimization pass.

Feature creep. Custom aliases. Password protection. Link expiration. Cloaking. Deep links for mobile apps. Bio pages. QR codes. UTM builders. A/B testing. Multi-user teams. API access. Webhooks. Every one is a real feature users will ask for the moment they compare you to any of the mainstream tools. Our developer walkthrough on the URL shortener API guide covers what a full API surface looks like.

Legal and compliance. Privacy policy, terms of service, DMCA response, GDPR data requests, data retention policy. Not glamorous. Also not optional.

Add it all up and a "simple URL shortener" is a multi-quarter engineering commitment for a small team, or a hobby project you regret every time it goes down at 2 AM.


When Building Makes Sense (And When It Really Doesn't)

Being honest: sometimes building is the right call. The situations where we'd tell someone to build their own are narrow but real.

Build your own if:

  • You're learning. There's genuine value in writing the redirect handler yourself and understanding what happens under the hood.
  • You have a genuinely unique routing rule that no managed platform offers, and it's core to your product (a game that routes based on live match state, for example).
  • You're operating at a scale where per-click pricing makes sense to internalize, and you already have edge infrastructure and SRE capacity.
  • You have a strict compliance requirement (data residency in a specific region, air-gapped deployment) that no vendor meets.

Use a managed platform if:

  • You want branded short links for marketing, sales, or product use.
  • You need analytics on how links perform.
  • You need QR codes tied to the same short links (change destination without reprinting).
  • You need a link-in-bio page for social profiles.
  • You need custom domains without becoming an SSL expert.
  • You need to be up and running this quarter.

This is where U2L AI fits. It's an all-in-one platform with the pieces we described above already built and battle-tested: global edge network with 330+ locations, detailed click analytics, dynamic QR codes, bio pages, deep links for popular apps, and custom domain support with automatic SSL. You skip the multi-quarter engineering commitment and get the marketer-facing product on day one. See u2l.ai/features for the current feature set.

Free plan and no-login shortening are also on the table if you just want to see what a fully built system looks like from the outside. Sometimes the fastest way to learn what to build is to use the finished version first.

Our take: build your own to learn, use a managed platform for anything anyone else depends on. That distinction has saved more engineering teams than any framework choice.


Frequently Asked Questions

How hard is it to build a URL shortener from scratch?

The core (generate a short key, store the mapping, redirect on lookup) is a weekend project for a competent backend engineer. The production version, with edge distribution, custom domains, SSL provisioning, abuse detection, analytics, and 99.9% uptime, is a multi-quarter effort for a small team.

What database should I use for a URL shortener?

For under 10 million links and a few hundred writes per second, PostgreSQL or MySQL are the boring right answer. Above that scale or if you need multi-region writes, look at DynamoDB, Cassandra, or Cloudflare KV. A Redis cache in front of any of these is essential for redirect performance.

What is base62 encoding and why do URL shorteners use it?

Base62 encoding maps a number to a string using the 62 alphanumeric characters (0-9, a-z, A-Z). It's URL-safe, case-sensitive for higher density, and a 7-character token can represent about 3.5 trillion unique URLs. Shorter than base16 (hex) and denser than base36.

How do I generate a unique short URL code?

Two mainstream approaches: auto-increment an integer ID and base62-encode it, or generate a random 7-character token and check for collision. Random tokens are non-enumerable (better for privacy) but need a lookup on write. Auto-increment is fully collision-free but exposes sequential IDs.

What are the biggest challenges of running a public URL shortener?

Abuse handling (phishing, malware, spam), custom domain and SSL provisioning, sub-50ms redirect latency worldwide, analytics pipeline design, and staying off browser warning lists. Plenty of hobby shorteners die because of abuse, not because of technical limits.

Should I use 301 or 302 redirects for short URLs?

301 (permanent) is browser-cacheable and passes SEO signal well, but destination edits propagate slowly. 302 (temporary) is not cached and lets you change destinations instantly, at the cost of every click hitting your server. Managed platforms typically let users choose per link.

How do I track click analytics without slowing down redirects?

Enqueue analytics events fire-and-forget to a queue (Kafka, SQS, Cloudflare Queues, Redis Streams) and process them in a separate consumer. Never write analytics synchronously on the redirect path. A batched consumer writing to a time-series database or analytics engine means click logging adds essentially nothing to redirect latency.

When should I build a URL shortener vs use an existing one?

Build if you're learning or have a truly unique routing requirement no vendor supports. Use a managed platform if you need branded links, analytics, QR codes, custom domains, or link-in-bio pages for real work. The gap between "works" and "works reliably at scale" is usually several engineer-quarters.

Build to Learn, Buy to Ship

Building a URL shortener is a genuinely great learning exercise. It teaches you HTTP, database design, caching, queues, and edge computing all in one project. It's also, for anyone who needs to actually rely on the service, almost always the wrong build-vs-buy decision. The hidden costs (abuse, uptime, custom domains, analytics, feature parity) dwarf the fun part.

If you want the finished version to compare notes against, start a free U2L AI account and see how the pieces fit together in production: unlimited free links without login, dynamic QR codes, bio pages, and analytics dashboards, all on a global edge network. No credit card required to poke at it.

Ready to try U2L AI?

Free forever plan. No credit card required.