DEV Community

Vin Lookup
Vin Lookup

Posted on

Metrics That Matter for a Free VIN Decode Product

Vanity counters look great on a landing page and teach you almost nothing about whether a free VIN decode works. "Decodes today" and "total lookups" rise when bots scrape you, when users retry after a timeout, and when the same VIN is pasted three times. For a product built on NHTSA vPIC, the metrics that pay rent are latency, cache behavior, and upstream health -- not applause numbers.

This post sketches a small TypeScript metrics surface for a decode path: what to count, what to histogram, and what to ignore so you do not optimize for theater.

What "working" means for a free decode

Users care about three outcomes:

  1. Fast enough -- the form feels responsive on a phone on mediocre LTE.
  2. Honest -- errors look like errors, not empty cars.
  3. Stable under load -- a burst of marketplace traffic does not melt your courtesy toward NHTSA.

Those map cleanly to instrumentation: end-to-end latency, cache hit ratio (and why you missed), and upstream error / timeout rates. Page views and "VINs decoded" are secondary; they are useful for capacity planning only after you separate humans from scrapers.

A typed metric bag, not a kitchen sink

Keep the event names boring and stable. Prefer a handful of counters and two histograms over dozens of bespoke labels.

export type DecodeOutcome =
  | "ok"
  | "vin_invalid"
  | "cache_hit"
  | "upstream_timeout"
  | "upstream_http"
  | "upstream_parse"
  | "circuit_open"
  | "queue_full";

export type DecodeMetrics = {
  recordLatencyMs: (ms: number, labels: { outcome: DecodeOutcome; cache: "hit" | "miss" | "bypass" }) => void;
  incr: (name: string, labels?: Record<string, string>) => void;
};

export function observeDecode(
  m: DecodeMetrics,
  startedAt: number,
  outcome: DecodeOutcome,
  cache: "hit" | "miss" | "bypass",
): void {
  m.recordLatencyMs(Date.now() - startedAt, { outcome, cache });
  m.incr("vin_decode_total", { outcome, cache });
  if (outcome.startsWith("upstream_")) {
    m.incr("vin_decode_upstream_errors_total", { outcome });
  }
}
Enter fullscreen mode Exit fullscreen mode

Call observeDecode once per user-visible attempt, after you know the final outcome. Do not increment "success" when you returned a cached 500 from last week.

Latency: measure the user path

Split latency if you can, but always publish one number users would feel:

  • handler_ms -- validate + cache lookup + optional upstream + serialize
  • upstream_ms -- only the NHTSA round trip when you actually call out

P50 is comforting; P95 and P99 are where mobile users live. Alert on P95 climbing while cache hit ratio falls -- that usually means TTLs expired together or negative cache is too short. Do not alert on "decodes per minute" alone; scrapers inflate it while real users wait.

export async function decodeWithMetrics(
  vin: string,
  deps: {
    cacheGet: (vin: string) => Promise<object | null>;
    upstream: (vin: string) => Promise<object>;
    metrics: DecodeMetrics;
  },
): Promise<{ body: object; outcome: DecodeOutcome }> {
  const t0 = Date.now();
  if (!/^[A-HJ-NPR-Z0-9]{17}$/.test(vin)) {
    observeDecode(deps.metrics, t0, "vin_invalid", "bypass");
    return { body: { error: "VIN_INVALID" }, outcome: "vin_invalid" };
  }

  const cached = await deps.cacheGet(vin);
  if (cached) {
    observeDecode(deps.metrics, t0, "cache_hit", "hit");
    return { body: cached, outcome: "cache_hit" };
  }

  try {
    const body = await deps.upstream(vin);
    observeDecode(deps.metrics, t0, "ok", "miss");
    return { body, outcome: "ok" };
  } catch (e) {
    const outcome =
      e instanceof Error && e.name === "TimeoutError"
        ? "upstream_timeout"
        : "upstream_http";
    observeDecode(deps.metrics, t0, outcome, "miss");
    throw e;
  }
}
Enter fullscreen mode Exit fullscreen mode

Cache hit ratio without lying

A hit ratio of 99% looks elite until you realize you are serving stale "vehicle not found" for newly assigned WMIs, or that bots hammer the same popular VIN. Report:

  • hit / miss / bypass separately (bypass = invalid VIN, circuit open, force-refresh)
  • negative-cache hits as their own series if you cache failures
  • unique VIN cardinality (approx) per window so you know whether hits are real reuse

If hit ratio rises while unique VINs fall, you are caching bots, not delighting buyers.

Upstream errors: classify, do not average away

NHTSA can time out, return 5xx, or hand you JSON that fails your parser. Collapsing everything into error_rate hides the fix. Prefer:

Series Meaning
upstream_timeout Your deadline fired
upstream_http Non-2xx or transport failure
upstream_parse Body shape unexpected
circuit_open You refused to call (good citizen mode)

Page those when the rate crosses a budget, not when a single VIN fails. One bad paste is not an incident.

Vanity metrics to demote

Useful for a weekly note, dangerous as SLOs:

  • Total lifetime decodes
  • "Cars identified"
  • Unfiltered request count (includes health checks and scrapers)
  • Frontend button clicks without pairing to outcomes

If marketing needs a big number, derive a human decode estimate: successful outcomes minus known bot ASNs minus identical VIN repeats inside a short window. Keep that estimate out of your pager.

Dashboards that fit on one screen

A practical free-product board:

  1. P50 / P95 handler latency (5m)
  2. Cache hit vs miss vs bypass
  3. Upstream error rates by class
  4. In-flight upstream concurrency (are you a good neighbor?)
  5. Queue depth / reject count if you fair-queue

Add "decodes OK" as a sixth row if you must -- below the fold.

Takeaway

Instrument a free VIN decode for latency, cache truthfulness, and upstream failure modes. Count outcomes with stable labels, histogram what users feel, and treat vanity totals as optional narrative, not success criteria. When those three stay healthy, the product feels fast and honest even if the marketing counter is boring.

I maintain VIN Lookup, a free VIN decode based on NHTSA data.

Top comments (0)