I Asked AI to Improve My Resume. It Started Asking Me for Numbers Instead.
Most people treat AI resume tools like a magic wand. They paste their draft, hit generate, and hope for better phrasing. But when you actually push a language model into doing meaningful optimization work, it quickly becomes clear that prose alone is insufficient. The model starts asking for quantitative signals because that is what it needs to make decisions that matter.
This realization changed how I approach automated resume optimization entirely. What follows is not a beginner's tutorial. It is a production-grade framework for building a system that forces your resume through real metrics before any AI touches it.
The Problem with Text-Only Resume Piping
A standard prompt looks something like this:
# DON'T do this - pure text input leads to generic output
response = client.chat.completions.create(
model="gpt-4o",
messages=[
{"role": "user", "content": "Improve my resume bullet points."}
]
)
The output will always be vague advice wrapped in corporate language. Phrases like "synergized cross-functional teams" or "spearheaded initiative" are exactly the kind of noise that makes resumes unreadable to both humans and applicant tracking systems. Without numbers anchoring each claim, the model has no signal to ground its improvements.
The Quantified Input Framework
The shift happens when you force every resume bullet into a structured numeric format before it ever reaches the language model:
from dataclasses import dataclass
from typing import Optional
@dataclass
class ResumeBullet:
action_verb: str # "Led", "Built", "Reduced"
metric_type: str # "percentage", "absolute", "ratio"
value_before: Optional[float] # baseline measurement
value_after: Optional[float] # outcome measurement
timeframe: Optional[str] # "Q3 2024", "6 months"
context: str # the what and why
@property
def has_numbers(self) -> bool:
# A bullet without quantification should be flagged
# before any LLM processing begins
return self.value_before is not None and self.value_after is not None
@property
def impact_ratio(self) -> Optional[float]:
if not self.has_numbers:
return None
return (self.value_after - self.value_before) / abs(self.value_before)
This structure forces a painful but necessary step: you must extract or estimate the real numbers behind every claim. That exercise alone improves your resume more than any AI paraphrase ever could.
Building the Scoring Pipeline
Once you have quantified bullets, you can build a deterministic scoring function that runs before the LLM stage:
import json
def calculate_bullet_score(bullet: ResumeBullet) -> dict:
"""
Produces a composite score from four independent dimensions.
Higher scores indicate stronger, more credible accomplishments.
"""
scores = {}
# Dimension 1: Quantification strength
if bullet.has_numbers:
scores["quantification"] = min(bullet.impact_ratio * 10, 100)
else:
scores["quantification"] = 0 # No numbers means this slot is empty
# Dimension 2: Action verb specificity
strong_verbs = {
"built": 90, "architected": 95, "designed": 85,
"reduced": 88, "optimized": 82, "launched": 78,
"led": 70, "managed": 55, "helped": 30
}
scores["verb_strength"] = strong_verbs.get(bullet.action_verb.lower(), 40)
# Dimension 3: Time-bound credibility
score_time = 50 if bullet.timeframe else 0
if bullet.timeframe and "quarter" in bullet.timeframe.lower():
score_time = 70
if bullet.timeframe and "month" in bullet.timeframe.lower():
score_time = 85
scores["temporal_precision"] = score_time
# Dimension 4: Scale of impact
scale_scores = {"team": 50, "department": 70, "company": 90, "individual": 30}
scores["scope_score"] = scale_scores.get(bullet.context.split()[0].lower(), 40)
composite = sum(scores.values()) / len(scores)
return {"scores": scores, "composite": round(composite, 1)}
This pipeline gives you something most resume tools never provide: a transparent audit trail showing exactly which bullet points are weak and why.
The Two-Stage Optimization System
Here is the architecture that actually works in practice:
def optimize_resume_stage_one(raw_bullets: list[dict]) -> list[ResumeBullet]:
"""
STAGE 1: Structural enforcement.
Converts freeform bullets into quantified objects.
Rejects or flags any bullet missing hard numbers.
"""
structured = []
for raw in raw_bullets:
bullet = ResumeBullet(
action_verb=raw.get("verb", ""),
metric_type=raw.get("type", ""),
value_before=raw.get("before"),
value_after=raw.get("after"),
timeframe=raw.get("timeframe"),
context=raw.get("context", "")
)
if not bullet.has_numbers:
# Flag for manual review instead of silently proceeding
print(f"[FLAG] Needs quantification: {bullet.context}")
structured.append(bullet)
return structured
def optimize_resume_stage_two(
bullets: list[ResumeBullet],
client,
job_description: str
) -> list[str]:
"""
STAGE 2: LLM rewriting.
The model now has concrete numbers to preserve and emphasize.
It rephrases around the data instead of inventing fluff.
"""
prompt = f"""
Optimize these quantified resume bullets for ATS readability.
Job description: {job_description}
RULES:
1. NEVER remove or soften existing numbers
2. Replace weak verbs with specific action terms
3. Keep each bullet under 2 lines
4. Front-load the metric whenever possible
Bullets:
{[f"{b.action_verb} | {b.value_before} → {b.value_after} | {b.context}"
for b in bullets]}
Return ONLY the optimized bullets as a JSON array.
"""
response = client.chat.completions.create(
model="claude-sonnet-4-20250514",
messages=[{"role": "user", "content": prompt}],
temperature=0.3 # Low temperature preserves factual accuracy
)
return json.loads(response.choices[0].message.content)
The Unexpected Discovery
When I ran this two-stage system against my own resume, the results were startling. Stage One flagged seven out of eleven bullets as missing hard numbers. Some of those gaps were honest oversights. Others were deliberate choices I had made because quantifying felt harder than writing.
The model did not need to invent metrics. It needed the original data point to exist in the first place. Once those numbers were present, Stage Two produced dramatically different quality output. The AI stopped padding language and started sharpening it.
# Example transformation with actual numbers
before = "Improved API response times significantly"
after_optimized = "Reduced p99 API latency from 840ms to 120ms by implementing Redis caching layer and query batch optimization"
# The second bullet is objectively better because it preserves the signal.
# AI amplifies signal. It cannot create it from silence.
Practical Takeaways
The lesson extends far beyond resume writing. Any time you ask an AI to improve something, the quality of your output is bounded by the quality of your structured input. Garbage in produces polished garbage. Numbers in produces sharper output.
Start by auditing every claim on your resume against this simple question: can I attach a measurable number to this statement? If the answer is no, you have found the exact bullet point that is weakening your entire document. Fix it there first. Then let the AI handle the language.
The system below captures the full pipeline end to end:
def full_resume_pipeline(
bullets: list[dict],
job_desc: str,
client
) -> dict:
"""
Complete optimization pipeline from raw input to scored output.
"""
stage_one = optimize_resume_stage_one(bullets)
scores = [calculate_bullet_score(b) for b in stage_one]
# Filter out bullets scoring below 40 before sending to LLM
qualified = [
b for b, s in zip(stage_one, scores)
if s["composite"] >= 40
]
low_scoring = [
b for b, s in zip(stage_one, scores)
if s["composite"] < 40
]
stage_two = optimize_resume_stage_two(qualified, client, job_desc)
return {
"qualified_bullets": stage_two,
"flagged_for_review": [b.context for b in low_scoring],
"score_summary": {s["composite"] for s in scores}
}
Quantify first. Optimize second. The order matters more than most people realize.
What is one bullet point on your resume right now that feels impactful but cannot survive being checked against a number?
Top comments (1)
The way you force quantification before LLM processing is smart - it stops the AI from inventing metrics. My own resume has three bullets that need numbers: 'Improved API response times' (no baseline), 'Reduced latency by 85%' (no target), and 'Launched feature for 10k users' (no growth metric). The scoring pipeline's 40-point threshold feels right - it catches weak claims without over-filtering.