Introduction
Building a portfolio project often starts with a single script and a dream. That was exactly how Job Copilot began. I needed a tool to help me navigate the chaotic job search process, optimize my resume for different roles, and track where I had applied.
Job Copilot is an AI-powered career platform designed to make job hunting less painful. It features a dashboard for application tracking, profile management, and experimental AI tools powered by the Google Gemini API for resume tailoring and interview prep. It was a fun project that solved a personal need, and I wanted to use it to showcase my backend engineering skills to potential employers.
But there was a problem. Once the initial excitement faded and the features were "working," I took a hard look at the codebase. It was an unmaintainable mess. I realized that if a hiring manager looked at this repository, it wouldn't demonstrate production readiness. It demonstrated how to hack a prototype together over a weekend. That realization was the turning point. I decided to pause feature development and refactor Job Copilot into a platform with a clean, professional architecture.
Note: This isn't a tutorial. It's a case study in refactoring decisions, trade-offs, and lessons learned. Code snippets are illustrative, not copy-paste ready.
The Problems I Found
When I audited my own codebase, the reality was sobering. Here are the specific issues that were preventing this project from being production-ready:
- Zero Automated Tests: I was relying entirely on manual clicking to verify features. I found critical bugs simply by navigating the UI in unexpected ways. Without tests, any new feature risked breaking existing functionality.
-
Hardcoded Secrets in
.env.example: In my rush to build, I had accidentally committed a live Gemini API key into the example environment file. This is a classic security blunder that exposes credentials to the public internet. - Mixed Concerns: My FastAPI routes were a tangled web. The endpoint functions were parsing requests, querying the SQLite database directly, and applying business logic all in one place. There was absolutely no separation of concerns.
-
No Database Migration System: I was using SQLAlchemy's
Base.metadata.create_all(bind=engine)at import time. Every time I altered a model schema, it broke the existing database. My "migration strategy" involved manually dropping tables or writing ad-hoc SQL scripts to patch the schema. -
Security Vulnerabilities: The backend was configured with CORS
allow_origins=["*"], meaning any malicious website could read my local data. Furthermore, the server was binding to0.0.0.0instead of127.0.0.1, exposing it to the entire local network without any authentication mechanism. -
Cluttered Repository Root: My root directory contained over 40 ad-hoc scripts. Files like
patch_db.py,test_api.py, anddummy_data.pywere mixed in with the actual application code, making it impossible for a newcomer to figure out where the real application lived.
The Refactor
1. Repository Structure Cleanup
The first step was aggressively cleaning up the workspace. The root directory was a graveyard of deploy artifacts, legacy scripts, and test outputs.
I moved the core application into a clean backend/ directory, structured by domain features like app_tracker/. I created a docs/ folder for architectural documentation and API references. All legacy compatibility scripts were safely archived outside the repository.
Finally, I extended the .gitignore to strictly exclude node_modules/, static_*, .pytest_cache/, and .coverage. The repository instantly became navigable.
2. Layered Architecture
The biggest architectural change was splitting the tangled FastAPI routes into a strict layered architecture. I wanted the codebase to reflect enterprise patterns.
- Routes: Now strictly handle HTTP requests and responses. They contain zero business logic.
- Services: This layer owns the business rules, ensures valid state transitions, and handles idempotency.
- Repositories: Responsible exclusively for data access and ensuring queries are properly scoped to the owner.
- Models: Pure SQLAlchemy ORM entities representing the database schema.
Here is how the Application Tracker module is structured today:
backend/app_tracker/
├── domain.py # State machine, allowed transitions
├── schemas.py # Pydantic request/response models
├── repository.py # Owner-scoped queries
├── service.py # Business rules, idempotency
└── router.py # FastAPI routes
This separation ensures data flows in one direction and components remain highly testable. Below is the Mermaid architecture diagram showing how the data flows from the UI down to the SQLite database.
flowchart TD
subgraph API["API layer - FastAPI"]
V1["/api/v1/applications"]
end
subgraph Services["Service Layer"]
SVC["ApplicationService"]
DOMAIN["Domain Rules"]
end
subgraph Data["Data Layer"]
REPO["ApplicationRepository"]
ORM["SQLAlchemy models"]
DB[("SQLite/PostgreSQL")]
end
V1 --> SVC
SVC --> DOMAIN
SVC --> REPO
REPO --> ORM
ORM --> DB
3. Testing Strategy
I went from zero tests to a robust suite of 156 automated tests, covering both unit and integration boundaries.
By mocking external dependencies, I ensured tests run blazingly fast in isolated, temporary SQLite instances. They never touch the real database. We also added legacy compatibility tests to ensure that refactoring didn't break backward compatibility for older clients.
Here is an example of how we test state transitions in the Application Tracker to ensure valid domain logic:
def test_valid_status_transition(db_session):
# Setup
repo = ApplicationRepository(db_session)
service = ApplicationService(repo)
# Create application
app = service.create_application(
owner_id=1,
company="Tech Innovators",
role="Backend Engineer",
status="Applied"
)
# Transition status
updated_app = service.update_status(
app_id=app.id,
new_status="Interviewing"
)
# Assert successful transition
assert updated_app.status == "Interviewing"
assert len(updated_app.history) == 2
4. Database Migrations
Relying on create_all() was a recipe for disaster. To fix this, I integrated Alembic for proper database schema versioning.
Now, every schema change generates a migration file with clear upgrade and downgrade paths. For example, Migration 0002 safely added a users table, an owner_id foreign key, and a new application_status_history table without dropping existing user data.
To prevent accidental corruption, I implemented a strict Schema Guard. If a user tries to start the server with an outdated schema, the guard intercepts the startup, refuses to run, and prompts the user to explicitly run the migration command.
5. Security Fixes
Security in a local-first application is often overlooked, but it is critical. My first action was revoking and rotating the accidentally exposed Gemini API key.
Next, I addressed the network vulnerabilities. I replaced the dangerous CORS wildcard (*) with an explicit allowlist restricted strictly to the frontend dev server and the browser extension ID. I also changed the uvicorn binding from 0.0.0.0 to 127.0.0.1 so the server is only accessible from the local machine.
Finally, I scrubbed all hardcoded Personally Identifiable Information (PII) from the DEFAULT_PROFILE constants, replacing them with empty placeholders, and updated the .env.example to only show dummy values.
6. Dependency Management
A project is only as stable as its dependencies. The original requirements.txt was just a loose list of package names.
I added missing dependencies that were previously installed globally (python-docx, fpdf2, alembic, pydantic-settings). More importantly, I established a strict dependency management policy:
-
Upper Bounds: Every top-level dependency in
requirements.inis now pinned with a strict upper version bound to prevent unexpected breaking changes. -
Hash-Locked Lockfiles: I used
uv pip compile --generate-hashesto create a fully reproducible, hash-verifiedrequirements.txt. -
CI Enforcement: The GitHub Actions CI pipeline now installs packages using the
--require-hashesflag. This guarantees protection against supply-chain attacks.
Here is what the locked requirements process looks like:
# requirements.in
fastapi>=0.142.2,<0.143
uvicorn>=0.54.0,<0.55
sqlalchemy>=2.1.1,<2.2
# Generate secure lockfile
uv pip compile --generate-hashes --allow-unsafe -o requirements.txt requirements.in
Honest moment: I almost gave up halfway through the refactor. The legacy code was so tangled that fixing one thing broke three others. Tests made it possible to proceed with confidence.
Lessons Learned
Refactoring Job Copilot was a massive learning experience. Here are the core takeaways that will influence every project I build from now on.
1. Tests Are a Design Tool, Not Just Verification
Writing tests forced me to think critically about my interfaces and dependencies. When a route was hard to test, it usually meant the route was doing too much. The 156 tests I wrote made the massive refactor safe, immediately catching regressions. I am now fully converted to a TDD-ish approach: writing the test interface before implementing the feature.2. Hash-Locked Dependencies Prevent "Works on My Machine" Issues
The hash-lockedrequirements.txtensures reproducible builds across any machine. During the refactor,ruffupdated its ruleset, which caused my local environment to format code differently than the CI server. The lockfile caught this discrepancy immediately, proving its worth.3. Honest Documentation Builds Trust
It is tempting to make a portfolio project sound like an enterprise SaaS platform. Instead, I made the README explicitly honest: "local-only, single-user, no authentication yet." I added aSECURITY.mdfile admitting current limitations. Transparency over marketing speak builds immediate credibility with other engineers reviewing the code.4. Data Safety Over Convenience
The schema guard completely disabled the "convenience" of auto-migrating databases on startup. Instead, it forces an explicit backup and migrate step. It is always better to fail safely and loudly than to quietly corrupt user data.
What's Next
Job Copilot is in a much better place, but software is never truly finished. Here is the roadmap for the future.
Short-term (this month)
- User Authentication: Implementing robust JWT access and refresh tokens to move away from the single-user local model.
- Production Database: Integrating PostgreSQL via Docker Compose for production deployments, phasing out SQLite.
- Community: Creating "good first issues" and improving the contribution guidelines to welcome open-source collaborators.
Long-term (if traction grows)
- Multi-tenancy Support: Architecting the database and services to securely handle multiple distinct users on a hosted platform.
- Cloud Hosting Option: Providing a one-click deployment option to a cloud provider.
- Advanced AI Features: Adding unlimited cover letter generation and interactive voice-based interview mock sessions.
Try It Yourself
If you are a backend developer looking for a clean FastAPI reference architecture, or a job seeker who wants to run their own local AI career assistant, I invite you to try Job Copilot!
Quick start (Local SQLite):
git clone https://github.com/ManoharVit/job-copilot.git
cd job-copilot
pip install -r requirements.txt
./start.sh
With Docker (PostgreSQL + zero setup):
docker-compose up
Run the test suite:
PYTHONPATH=backend pytest backend/tests/ -v
Let's Connect
If you're building something similar or have questions about the refactor:
- GitHub: https://github.com/ManoharVit/job-copilot (open an issue!)
- LinkedIn: https://www.linkedin.com/in/manohar511/
I'm always happy to discuss architecture, testing strategies, or career advice.
Top comments (0)