DEV Community

Cover image for From Job Scraper to Production AI System: Building ApplyLens AI
Sriram Hariharan Neelakantan
Sriram Hariharan Neelakantan

Posted on AI-assisted

From Job Scraper to Production AI System: Building ApplyLens AI

Job searching started feeling like a second job.

I was jumping between different job sites and ATS platforms, seeing duplicate or stale postings, comparing multiple resume versions, and still manually trying to answer the questions that actually mattered:

Is this role really worth applying to? Which resume should I use? What am I missing? Should I tailor first โ€” or move on?

There are already strong products that solve valuable parts of this workflow: resume analysis, job tracking, autofill, matching, and application organization.

I wasn't trying to claim that those problems had never been solved.

The engineering question that interested me was different:

What would happen if I treated the entire workflow as one AI systems problem instead of a collection of separate tools?

So I started with something much simpler: a job scraper.

Somewhere between ATS integrations, deduplication, ranking, LLM evaluation, resume matching, provider routing, retries, observability, agentic workflows, persistence, and production deployment... things got slightly out of hand. ๐Ÿ˜…

That project became ApplyLens AI โ€” a production-deployed AI job intelligence and application planning platform.

Live deployment: applylensjobs.com (access-controlled)

Source code: GitHub โ€” ApplyLens AI


What ApplyLens does

At a high level, ApplyLens combines job acquisition, deterministic filtering, AI evaluation, resume intelligence, application prioritization, evidence-grounded tailoring, retrieval, and human review within one production system.

Three numbers give a useful sense of the current scope:

  • 11 acquisition adapters across ATS platforms and job sources
  • a 16-stage observable processing pipeline
  • shared production job acquisition running automatically every 6 hours

The acquisition layer currently includes:

Workday, Greenhouse, Lever, Ashby, Workable, Jobvite, Recruitee, SmartRecruiters, Built In, USAJobs, and Himalayas.

But the number of sources is not the part of the architecture I find most interesting.

The important part is what happens after the jobs arrive.


The architecture in one view

The system is built around a simple separation: shared acquisition, personalized intelligence, bounded AI reasoning, and human-controlled action.

ApplyLens AI system architecture

ApplyLens AI architecture: jobs are acquired into a shared corpus, then processed through user-specific intelligence, bounded AI evaluation, scoring, planning, and human review.

There are several design decisions behind this that became much more important than the original scraping problem.


1. Collect once. Personalize independently.

Shared acquisition with isolated personalization in ApplyLens AI

Jobs are acquired once into a shared corpus, while preferences, resumes, seen-state, ranking, AI settings, and application planning remain isolated per user.

One of the first architectural decisions was separating global acquisition from user-specific intelligence.

A naive multi-user design could effectively repeat acquisition work for every user.

I didn't want that.

Instead, ApplyLens maintains a shared job corpus and then projects that data through an authenticated user's own workflow.

The shared layer handles job acquisition.

The personalized layer owns things such as:

  • role and seniority preferences
  • location preferences
  • excluded terms
  • preferred skills
  • seen-job state
  • resume variants
  • ranking context
  • AI credentials
  • provider/model routing
  • application planning state

In other words:

Collect once. Personalize independently.

This also creates a useful operational boundary.

The production job acquisition schedule runs every six hours in global-acquisition-only mode. It refreshes the shared job corpus without silently launching user-specific planning, tailoring, or application decisions.

Personalization happens separately.

That separation became one of the most important pieces of the system.


2. I deliberately did not let one LLM make the whole decision

It is very easy to design an AI application like this:

Resume + Job Description
          โ†“
         LLM
          โ†“
       Fit Score
Enter fullscreen mode Exit fullscreen mode

It is simple.

It is also not the architecture I wanted.

Those are deliberately different stages with different authority.

Where AI gets authority โ€” and where it doesn't

ApplyLens AI deterministic, LLM, and human decision boundaries

Cheap, reproducible decisions remain deterministic. LLM reasoning is introduced where interpretation adds value, while consequential actions remain human-controlled.

Deterministic prefiltering

Before an expensive model needs to interpret anything, deterministic logic can evaluate things such as:

  • role/title relevance
  • seniority
  • location policy
  • freshness
  • exclusions
  • duplicate identity
  • previously seen jobs
  • early evidence rules

There is no reason to ask an LLM whether a posting is three weeks old or whether a title violates an explicit exclusion rule.

Those are deterministic questions.

LLM evaluation

The LLM becomes useful later, where interpretation matters.

For example:

  • interpreting job-description requirements
  • reasoning over less explicit fit signals
  • enriching structured intelligence
  • producing grounded tailoring assistance

But the model operates on a bounded candidate set rather than becoming the first and only filter.

Final application scoring

Final application-priority scoring remains a separate responsibility.

This matters because an LLM evaluation is not automatically equivalent to the final decision the application should make.

A relevance score, model interpretation, resume-match score, optimization score, and final application-priority score represent different things.

I wanted the architecture to preserve those distinctions instead of collapsing them into one mysterious number.


3. Different AI workloads should not automatically use the same model

Another thing I wanted to avoid was this:

MODEL = "whatever-model-is-currently-best"
Enter fullscreen mode Exit fullscreen mode

followed by sending every AI task through it.

ApplyLens treats AI work as workload-specific.

Different workloads include areas such as:

  • skill extraction
  • job-fit evaluation
  • JD intelligence
  • grounded RAG answers
  • resume fallback ranking
  • ambiguous resume adjudication
  • critic evaluation
  • tailoring generation
  • tailoring refinement
  • tailoring judging
  • manual Scan phrase generation

The application maintains explicit provider/model qualification and routing rather than assuming that one preferred provider should own every task.

For authenticated workflows, the effective route is resolved for the specific workload.

A user's generic provider preference does not silently override a workload-specific qualified route.

And if required routing, qualification, or credentials are unavailable, sensitive AI paths can fail closed rather than silently switching behavior.

That was an important lesson for me:

Multi-model support is not the same thing as model routing.

Supporting several APIs is easy.

Defining which workload is allowed to use which model, how that decision is resolved, and what happens when that route is unavailable is the actual systems problem.


4. Resume tailoring has an evidence boundary

Resume optimization creates another tempting failure mode for generative AI.

If the system's only goal is "make this resume look more relevant," a model can produce wording that sounds excellent while quietly introducing experience that never existed.

I didn't want that behavior.

The tailoring path is built around resume and job evidence.

Conceptually:

Resume evidence + JD evidence
            โ†“
     Rewrite direction
            โ†“
     Candidate generation
            โ†“
        Validation
            โ†“
    Replacement selection
            โ†“
      Human review/export
Enter fullscreen mode Exit fullscreen mode

Generated changes are expected to remain grounded in supplied evidence.

Unsupported tools, skills, metrics, methods, domains, or responsibilities are rejected or kept as directional guidance instead of being silently inserted into the resume.

And the source resume itself is not overwritten by generation.

That distinction is important to me because a resume assistant should help express real evidence better โ€” not manufacture new evidence.


5. The human remains the authority for consequential actions

This eventually became one of the core design principles behind ApplyLens.

AI can:

  • analyze
  • interpret
  • recommend
  • prioritize
  • explain
  • draft
  • critique
  • assist with tailoring

But ApplyLens does not give the AI unrestricted authority over consequential actions.

It does not silently:

  • submit applications to an ATS
  • mark jobs as applied
  • message recruiters
  • overwrite source resumes
  • turn an AI recommendation into application approval

Human review and application execution remain separate.

Even optional adjudication is intentionally constrained: commentary can be added without silently overriding the authoritative resume winner, final score, ranking, queue, or action.

I think this boundary becomes more important as AI systems become more agentic.

The interesting question is no longer only:

"Can the model do this?"

It is also:

"Should the model have authority to do this?"

Those are very different questions.


6. Agentic does not have to mean autonomous

I also experimented with agentic orchestration and LangGraph.

But I did not want to rebuild the entire application around an "agent owns everything" architecture just because agent frameworks became available.

ApplyLens keeps its existing deterministic and AI owners, while selected stages can be routed through bounded, explicitly gated LangGraph paths.

Those guarded paths cover areas such as:

  • deterministic prefilter/deduplication
  • JD intelligence
  • semantic evaluation
  • final scoring
  • prioritization
  • tailoring decisions
  • tailoring generation
  • conditional operator review

The important word there is bounded.

These paths are not permission for an agent to invent a new application state or bypass an existing authority.

Agent recommendations, trace persistence, human checkpoints, and authoritative application actions remain separate concerns.

Agentic workflows are inspectable, not opaque

ApplyLens AI Agentic Review showing workflow inspection and diagnostics

Run-scoped Agentic Review keeps workflow summaries, artifacts, traces, diagnostics, and advisory evidence inspectable without giving the agent unrestricted production authority.

Many of these capabilities are also deliberately default-off unless explicitly enabled.

That may sound less exciting than saying "fully autonomous AI agent."

I think it is better engineering.


7. Observability became part of the product

Once a pipeline has enough moving pieces, "it failed" is no longer useful information.

I wanted to be able to answer:

  • Which stage is currently running?
  • Which stages completed?
  • How many jobs survived each stage?
  • Which source degraded?
  • Was a failure transport-related, parsing-related, or pagination-related?
  • Did an AI workload hit cache?
  • Which provider/model route was resolved?
  • Did parsing fail?
  • Was there a retry?
  • What artifacts were produced?
  • What happened in an agentic workflow?
  • What is the scheduler doing?
  • Did a persisted operation actually succeed?

That led to several operational surfaces inside the application:

  • Pipeline Dashboard
  • Advanced Diagnostics
  • Agentic Operations
  • Run-scoped Agentic Review
  • Scheduler Health
  • run/artifact inspection
  • notification and operational state
  • admin and Super User controls

The operational layer is visible in the product

ApplyLens AI Executive Queue showing personalized job intelligence and acquisition state

ApplyLens AI Executive Queue โ€” shared acquisition, personalized recommendations, source intelligence, and pipeline freshness exposed directly through the product.

The runtime itself tracks a 16-stage status model:

startup
โ†’ scraping
โ†’ filtering
โ†’ dedupe
โ†’ ranking
โ†’ cache_filter
โ†’ details
โ†’ intelligence
โ†’ ai_evaluation_filter
โ†’ embedding_prefilter
โ†’ ai_evaluation
โ†’ resume_matching
โ†’ application_priority
โ†’ rag_export
โ†’ planning
โ†’ finalization
Enter fullscreen mode Exit fullscreen mode

Different runtime modes can substitute or bypass work where appropriate โ€” for example, an authenticated shared-corpus projection uses shared input rather than scraping โ€” but the stage model provides a common observability contract.

The lesson here was simple:

For production AI, observability is not an admin afterthought. It is part of the architecture.


8. Acquisition needed its own reliability model

Scraping multiple ATS platforms is not simply "send HTTP requests in a loop."

Different sources have different:

  • APIs
  • pagination rules
  • rate limits
  • HTML/data structures
  • detail endpoints
  • completeness characteristics
  • failure modes

So acquisition has bounded retry, timeout, pagination, and concurrency behavior.

Source health distinguishes states such as:

  • SUCCESS
  • EMPTY
  • PARTIAL
  • FAILED

Failures are classified rather than flattened into one generic exception.

The system also tracks data-quality signals such as URL, timestamp, and description completeness.

Why does that matter?

Because "this source returned zero relevant jobs" and "this source silently failed halfway through pagination" are not the same operational event.

If downstream AI consumes bad or incomplete acquisition data, the problem may look like a model-quality problem even though the failure happened much earlier.


9. PostgreSQL is authoritative; Redis is not

Another production decision was making state ownership explicit.

In the deployed architecture:

PostgreSQL 18 is authoritative for persistent application state.

It stores areas such as:

  • authentication/session state
  • profile resumes
  • onboarding preferences
  • AI settings
  • pipeline runs
  • seen-job state
  • saved scans
  • application actions
  • operator decisions
  • scheduler history
  • job/RAG documents
  • discovery state
  • metrics
  • agent traces and related operational state

Redis 7 is used for acceleration and coordination where appropriate โ€” caching, invalidation, and locking โ€” but it is not treated as authoritative application storage.

I wanted the system to remain understandable when caches disappear.

That sounds basic, but clear state ownership prevents a surprising number of production problems.


10. RAG is useful, but it still needs boundaries

ApplyLens also includes retrieval over the job corpus for search and grounded question answering.

The retrieval layer supports things such as:

  • job corpus search
  • query filtering
  • lexical retrieval
  • semantic retrieval where configured
  • result deduplication/ranking
  • grounded answer generation

Owner-facing retrieval is constrained to allowed job identities before evidence reaches the answerer.

An empty allowed set fails closed instead of silently widening into the entire shared corpus.

That is another example of a pattern that appears throughout the project:

Retrieval convenience should not override identity boundaries.


11. Production engineering was most of the work nobody sees

The visible AI features are only one part of the application.

A large amount of development time went into things that are much less impressive in a screenshot:

  • retry behavior
  • rate limiting
  • timeout bounds
  • pagination limits
  • source-health monitoring
  • cache behavior
  • deduplication
  • owner isolation
  • authentication
  • role-based access
  • registration approval
  • persisted run identity
  • concurrency controls
  • scheduler safety
  • process liveness
  • artifact tracking
  • backups
  • rollback procedures
  • deployment health checks

The production topology currently looks roughly like this:

Production topology

ApplyLens AI production deployment architecture

Production topology: Caddy fronts the FastAPI application, PostgreSQL remains authoritative, Redis provides cache and coordination, and systemd manages scheduled operational workloads.

The application is containerized with Docker Compose.

Production scheduling is handled through systemd, including the six-hour shared acquisition schedule.

Deployment has an explicit backup and rollback path rather than treating git pull && restart as a deployment strategy.

This project started as a scraper.

It ended up teaching me much more about operating an AI system than about scraping itself.


12. The frontend also had to grow with the architecture

ApplyLens is not only a pipeline with a thin HTML wrapper.

The application has a hybrid frontend architecture:

  • FastAPI server-rendered surfaces
  • classic JavaScript for several workflow-heavy pages
  • React 18 + TypeScript + Vite for dashboard/operational surfaces

The product now includes areas for:

  • executive overview
  • planning
  • decisions
  • application tracking
  • pipeline execution
  • optimization review
  • tailoring
  • scheduler health
  • agentic operations
  • advanced diagnostics
  • profile/resume management
  • AI provider/model settings
  • saved scans
  • admin workflows

Building the UI changed the backend design too.

Once users can inspect a score, save a decision, rerun a pipeline, or review an AI suggestion, concepts like identity, persistence, dirty state, idempotency, and write verification stop being backend implementation details.

They become product behavior.


What makes ApplyLens different for me

I don't think the interesting claim is:

"I built a job app with AI."

There are already many good products in that category.

The part I wanted to explore was the architecture underneath the experience:

Can shared acquisition, deterministic processing, bounded model reasoning, workload-specific routing, evidence-grounded generation, retrieval, agentic orchestration, observability, and human authority coexist in one system without collapsing everything into an LLM call?

ApplyLens became my attempt to answer that question.

And the answer I arrived at is:

Yes โ€” but only if the boundaries are designed deliberately.


What I learned

A few lessons became much clearer while building this.

1. The biggest AI decision is often where not to use AI

Freshness, explicit exclusions, identity, access control, deterministic scoring boundaries, and state transitions do not become better merely because a language model can participate in them.

2. Expensive intelligence should come after cheap certainty

Filter and deduplicate first.

Use model reasoning when the remaining ambiguity actually benefits from it.

3. Model routing is an application architecture problem

Provider/model selection, qualification, credentials, failure behavior, caching, and fallback rules need explicit ownership.

4. Agentic systems need authority boundaries

A traceable recommendation is not the same thing as permission to mutate production state.

5. Observability is part of AI quality

If you cannot tell whether the source failed, the cache was stale, parsing broke, the provider route changed, or the model produced poor output, every failure starts looking like "the AI is bad."

6. Human-in-the-loop should be an architectural property

It should not be a disclaimer added after autonomous behavior is already designed.


The stack

The main technology stack currently includes:

Python ยท FastAPI ยท PostgreSQL ยท Redis ยท React ยท TypeScript ยท Vite ยท Docker ยท LangGraph ยท LLM/GenAI workflows ยท RAG ยท systemd ยท Caddy

But the most valuable part of this project for me has not been learning another framework.

It has been learning how to decide which component should own which decision.


Explore the project

If you'd like to look deeper:

๐ŸŒ Access-controlled live deployment: https://applylensjobs.com

GitHub logo sriram-hariharan / job-scraper

Production AI job intelligence & application planning platform with multi-source acquisition, LLM evaluation, resume intelligence, RAG, and observability.

ApplyLens AI

AI-powered job discovery, application planning, resume scanning, and tailoring workspace.

Scrape jobs from modern ATS platforms, rank opportunities against saved resumes, review AI optimization guidance, generate tailored drafts, and track the full application workflow from one local operator app


Table of Contents


What This App Does

ApplyLens AI is a job-search operating system for serious application workflows. It combines job scraping, resume intelligence, AI-assisted review, and decision tracking into a single FastAPI web application.

The app is designed around a practical application loop:

  1. Discover and scrape jobs from multiple ATS platforms.
  2. Filter, deduplicate, and rank jobs against your savedโ€ฆ




I'm especially interested in conversations around AI engineering, ML engineering, production LLM systems, model routing, evaluation, RAG, agentic workflows, and human-in-the-loop system design.

If you've built something similar โ€” or disagree with any of these architectural choices โ€” I'd genuinely like to hear how you approached it.


AI disclosure: I used ChatGPT to help research, structure, and edit this article. The ApplyLens project, implementation, architecture decisions, and technical claims are my own, and I reviewed the final content against the project repository before publishing.

Top comments (1)

Collapse
 
launchgatecheck profile image
Launch Gate •

The SUCCESS / EMPTY / PARTIAL / FAILED distinction raises a useful corpus-refresh test: start with two known jobs, return only the first page on the next acquisition, then have page two time out. Does the missing job stay "last seen on the previous successful run", or become closed/stale as though the source had completed?

I'd follow that with a complete run where the job is genuinely absent. The pair checks whether incomplete acquisition can accidentally turn "not observed" into "no longer available", including in the personalized queue and RAG answers. I haven't run ApplyLens; this is a fixture suggested by the shared-corpus and source-health design in the article.