Building a URL shortener sounds almost too simple.
Take a long URL:
https://example.com/articles/how-to-build-a-python-project
Turn it into:
http://localhost:5000/aB91x
When someone visits /aB91x, redirect them to the original URL.
That is the basic idea.
But while building a simple URL shortener with Python, Flask, and SQLite, I discovered that the redirect itself was probably the easiest part of the project.
The interesting problems appeared around it.
What happens if two URLs get the same short code?
What happens if someone submits an invalid URL?
What happens if the database contains thousands of links?
What happens if someone tries to use the service for malicious redirects?
And what happens when a project that worked perfectly on my laptop meets real users?
This is what I learned while building it.
What I Wanted to Build
I wanted a small application with four basic features:
- Accept a long URL
- Generate a short code
- Store the URL in a database
- Redirect users when they visit the short URL
The architecture looked like this:
User
|
v
Flask Application
|
+---- Generate Short Code
|
+---- Store URL
|
v
SQLite Database
|
v
Short URL
|
v
Redirect to Original URL
I deliberately kept the first version simple.
The goal was not to build the next Bitly.
The goal was to understand what actually happens inside a URL shortener.
The First Version
I started with Flask because it provides everything needed for a small HTTP application without adding unnecessary complexity.
The project structure looked like this:
url-shortener/
│
├── app.py
├── database.db
└── templates/
└── index.html
I installed Flask:
pip install flask
Then created the application.
from flask import Flask, request, redirect, render_template
import sqlite3
import string
import random
app = Flask(__name__)
def get_db():
connection = sqlite3.connect("database.db")
connection.row_factory = sqlite3.Row
return connection
def generate_code(length=6):
characters = string.ascii_letters + string.digits
return "".join(random.choice(characters) for _ in range(length))
@app.route("/", methods=["GET", "POST"])
def index():
if request.method == "POST":
url = request.form["url"]
code = generate_code()
db = get_db()
db.execute(
"INSERT INTO urls (code, url) VALUES (?, ?)",
(code, url)
)
db.commit()
db.close()
return render_template(
"index.html",
short_url=f"http://localhost:5000/{code}"
)
return render_template("index.html")
@app.route("/<code>")
def shorten_redirect(code):
db = get_db()
result = db.execute(
"SELECT url FROM urls WHERE code = ?",
(code,)
).fetchone()
db.close()
if result:
return redirect(result["url"])
return "URL not found", 404
if __name__ == "__main__":
app.run(debug=True)
It worked.
I entered a URL.
The application generated a short code.
I clicked the short URL.
The browser redirected me to the original website.
For about five minutes, everything looked perfect.
Then I started asking what could go wrong.
Problem 1: Short Codes Can Collide
My first implementation generated a random six-character code.
For example:
a91Bc2
But random does not mean unique.
Eventually, the application could generate the same code twice.
Imagine the database contains:
abc123 -> https://example.com
Then another user creates a URL and the application generates:
abc123
Now what?
The application cannot safely use the same code for two different URLs.
The first solution is to check whether the code already exists.
def generate_unique_code(db):
while True:
code = generate_code()
existing = db.execute(
"SELECT id FROM urls WHERE code = ?",
(code,)
).fetchone()
if not existing:
return code
Then:
code = generate_unique_code(db)
This is much better.
But there is still another problem.
Two requests could theoretically check the database at almost the same time and both discover that the code is available.
That is why the database should also enforce uniqueness.
CREATE TABLE urls (
id INTEGER PRIMARY KEY AUTOINCREMENT,
code TEXT UNIQUE NOT NULL,
url TEXT NOT NULL
);
The application checks.
The database enforces.
I learned an important lesson here:
Application logic should not be the only thing protecting data integrity.
Problem 2: I Forgot URL Validation
My first version basically trusted whatever the user submitted.
That means someone could enter:
hello
instead of:
https://example.com
Or they could enter a completely different URL scheme.
A basic validation function is better:
from urllib.parse import urlparse
def is_valid_url(url):
try:
parsed = urlparse(url)
return parsed.scheme in ("http", "https") and bool(parsed.netloc)
except Exception:
return False
Then:
if not is_valid_url(url):
return "Invalid URL", 400
This is not a complete security system, but it prevents many obviously invalid inputs.
It also makes the application behave more predictably.
Problem 3: I Started Thinking About Security
A URL shortener looks harmless.
But URL shorteners can be abused.
Someone could create a short link pointing toward a phishing page, malware download, or other malicious website.
That means a production URL shortener needs to think about abuse.
Depending on the application, additional protections could include:
- Rate limiting
- Abuse reporting
- Link scanning
- Domain reputation checks
- CAPTCHA for suspicious activity
- Authentication
- Expiration dates
- Link ownership
- Logging
- Blocking known malicious destinations
For a small learning project, I did not need to implement all of these.
But realizing that a technically functional application could still be abused changed how I thought about "finished."
Working is not the same thing as production-ready.
Problem 4: The Database Was Too Basic
Initially, my table only contained two useful fields:
code
url
That was enough for a demo.
But then I wanted to answer basic questions.
When was the link created?
How many times was it visited?
Who created it?
Should it expire?
The schema could evolve into something like:
CREATE TABLE urls (
id INTEGER PRIMARY KEY AUTOINCREMENT,
code TEXT UNIQUE NOT NULL,
url TEXT NOT NULL,
created_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP,
clicks INTEGER DEFAULT 0,
expires_at TIMESTAMP
);
Now the application can support additional features later.
For example:
Short URL
|
+-- Original URL
+-- Created date
+-- Number of clicks
+-- Expiration date
This made me realize something important about database design.
You do not need to design everything perfectly before writing your first line of code.
But you should think about what information the application will probably need later.
Problem 5: Counting Clicks Was Not as Simple as I Expected
I wanted to count how many times each link was opened.
The obvious solution was:
db.execute(
"UPDATE urls SET clicks = clicks + 1 WHERE code = ?",
(code,)
)
Then redirect the user.
Something like:
@app.route("/<code>")
def shorten_redirect(code):
db = get_db()
result = db.execute(
"SELECT url FROM urls WHERE code = ?",
(code,)
).fetchone()
if not result:
db.close()
return "URL not found", 404
db.execute(
"UPDATE urls SET clicks = clicks + 1 WHERE code = ?",
(code,)
)
db.commit()
db.close()
return redirect(result["url"])
For a small project, this is fine.
But with a large number of requests, constantly updating the database for every redirect can become a performance consideration.
A production system might need a different architecture involving caching, queues, analytics systems, or other infrastructure.
Again, the simple version worked.
But scaling changes the problem.
Problem 6: SQLite Was Great Until I Thought About Scale
SQLite was perfect for learning.
There was no database server to configure.
The entire database lived in one file:
database.db
That made development extremely easy.
But a public URL-shortening service could eventually need:
- More concurrent writes
- Replication
- Backups
- Monitoring
- High availability
- Better scaling characteristics
At that point, a production application might use a database such as PostgreSQL or another system designed around the application's workload.
This taught me another useful lesson:
The best technology for a prototype is not always the best technology for production.
That does not mean SQLite was a bad choice.
It was the right level of complexity for what I was building.
Problem 7: My URL Codes Were Not Really "Short"
I initially thought that six characters was automatically the correct answer.
It isn't.
The length of the code depends on the number of available characters and the number of URLs you need to represent.
If the code uses:
a-z
A-Z
0-9
there are 62 possible characters.
With six characters, there are:
62^6
possible combinations.
That gives a large address space for a small project.
But if the application becomes extremely large, code length and collision probability become important design considerations.
This was one of those moments where a simple programming project suddenly became a small lesson in probability and system design.
Problem 8: My Error Handling Was Terrible
The first version assumed everything would work.
Real applications do not get that luxury.
What if:
- The database is unavailable?
- The submitted URL is empty?
- The generated code already exists?
- The user requests a nonexistent code?
- The database insert fails?
- The application receives malformed input?
Instead of letting unexpected errors reach the user, the application should handle expected failures.
For example:
try:
db.execute(
"INSERT INTO urls (code, url) VALUES (?, ?)",
(code, url)
)
db.commit()
except sqlite3.IntegrityError:
db.rollback()
return "Could not create short URL", 500
finally:
db.close()
Error handling is not the most exciting part of programming.
But it is one of the things that separates a demo from a reliable application.
What the Final Flow Looked Like
After fixing the main problems, the application flow became:
User submits URL
|
v
Validate URL
|
v
Generate short code
|
v
Check uniqueness
|
v
Store in database
|
v
Return short URL
|
v
User opens short URL
|
v
Find code in database
|
v
Record click
|
v
Redirect to destination
It was still a small application.
But it was now a much better representation of how a real system needs to behave.
What I Would Add Next
If I continued developing the project, I would add:
Authentication
Users could manage the links they created.
Expiration
Links could automatically stop working after a specific date.
Analytics
Users could see:
Total clicks
Clicks per day
Referrer
Device type
Country
Analytics would need to be designed carefully because collecting additional information also introduces privacy considerations.
Custom aliases
Instead of:
aB91x
users could choose:
my-blog
Rate limiting
This would help prevent automated abuse.
Caching
Frequently accessed URLs could potentially be served without querying the primary database every time.
Production database
If the application grew significantly, I would consider moving beyond SQLite depending on the workload and deployment architecture.
The Biggest Lesson
The biggest lesson was not how to generate a six-character string.
It was learning how quickly a simple idea becomes a system.
At first, the project looked like this:
Long URL
↓
Short Code
↓
Redirect
After thinking about real usage, it became:
Input validation
↓
Code generation
↓
Uniqueness
↓
Database constraints
↓
Security
↓
Analytics
↓
Error handling
↓
Performance
↓
Abuse prevention
↓
Monitoring
That is what I like about small projects.
You can start with something that looks almost trivial and slowly discover the engineering decisions hiding underneath.
Final Thoughts
Building a URL shortener is one of those projects that looks like a beginner exercise until you start asking production-level questions.
The redirect itself took very little code.
The difficult part was everything around it.
I learned about database constraints, URL validation, collision handling, security, analytics, error handling, and scalability from one relatively small project.
That is why I think developers should build small applications instead of only following tutorials.
A tutorial can show you the happy path.
A real project shows you what breaks.
And those broken parts are often where the actual learning happens.
Top comments (0)