DEV Community

Cover image for Records first, chart second. The Sankey disagreed.
Michael Truong
Michael Truong

Posted on

Records first, chart second. The Sankey disagreed.

I thought my job search history could be represented as a conventional funnel: application, recruiter, hiring manager, technical, outcome.

The real histories did not work like that. Some skipped stages. Some repeated them. Inbound opportunities did not start with an application at all.

I decided to model the histories before trying to visualise them. The Sankey, a flow diagram showing how each application moved from one stage to another, was supposed to come later: records first, chart second.

The operational board was not funnel history

I had already built a job search system that kept each opportunity in Notion alongside its research and generated resume in a repo. As the search ramped up, overlapping hiring processes replaced a handful of quiet opportunities. With enough of them accumulating, I wanted to inspect patterns across the search: how opportunities entered, how far they progressed, and where they ended.

I use the Notion board to keep track of where each opportunity is, whether I am waiting on a company, and what I need to do next. What it does not preserve is the path each opportunity took to get there.

When a role ends in rejection, Phase often moves to Completed. That is useful on a kanban, but it loses the history I needed for analysis. The board no longer answers "how far did this application get before it ended?" I could see that a process closed, not the path it took.

Funnel analytics needs observed transitions between stages, grounded in substantiated events. Company A might run recruiter, then hiring manager, then technical. Company B might skip recruiter entirely. No global ordering of stage names can represent both without lying.

I needed to normalise those histories without forcing them back into a global funnel: how an opportunity entered, the ordered process events that followed, and any terminal outcome.

That is why lifecycle.yml exists as a separate canonical store: ordered events, one file per role posting, validated in CI. Notion stays operational. The repo holds history.

What the first Sankey passes exposed

Once I wired the first passes, the chart did not tell me much about how opportunities moved. It kept finding problems in what those events meant.

The first mismatch was mixing how an opportunity entered the funnel with what happened after. Inbound outreach and cold applications became difficult to read on the same chart. A LinkedIn InMail that never went through a portal looked like it had already "reached" an application stage before any recruiter conversation. Entry provenance is how the opportunity entered the funnel. Recruiter, hiring manager, and technical rounds are what happened after. Those are different layers.

Order was the next fight. Repeating a stage is normal: two technical rounds are two technical rounds, not one node with a count badge. The chart keys nodes by sequence position so it preserves order without inventing a global stage ladder. Direct exits are valid too. One application went directly from cold application to rejected with no process events in between. The YAML allows events: [] while a search is still open. The chart had to render Cold application → Rejected without padding imaginary recruiter steps.

Not every wrong branch was a modelling bug. The backfill had created a lifecycle file for another posting at the same company: researched, never submitted. Once the Sankey existed, that record looked wrong. There was no application process to chart, only notes where a hiring path should be. I deleted the lifecycle file rather than inventing another state for something that had never entered the funnel. The next day I applied the same boundary to the operational board and removed its Notion row too.

Separating data, projection, and presentation

Failures came from different parts of the system, so I kept three concerns separate.

Canonical YAML stores how an opportunity entered and the ordered process events that followed. A terminal outcome such as rejected, stalled, withdrawn, or accepted ends that history. If the research only substantiates entry so far, the file can stop there.

Projection code turns each history into Sankey edges. Histories without a terminal outcome append an Active branch on the chart only. That state never gets written back to lifecycle.yml. It answers "where does this open path end on the chart right now?" not "what is the canonical outcome?"

Presentation is layout: labels, column alignment, tooltips. I avoided a layout mode that forces every terminal into a shared rightmost column. Rejections after one recruiter screen and rejections after three rounds looked equivalent when all sinks lined up. Outcome nodes now terminate at natural depth, so a quick rejection after application does not share a column with a rejection after several rounds.

That separation made triage faster. When a branch looked wrong, I could ask whether the canonical data was wrong, the projection code misread it, or the layout was misleading. Sometimes the problem was that an opportunity should never have been included in the lifecycle history at all.

Records alone hid the category errors. Spreadsheets and YAML validators catch schema mistakes; they do not show you that two rejection depths collapsed into one visual column. The Sankey challenged my representation of these hiring histories in specific, fixable ways.

Takeaway: If you are modelling any multi-step human process with irregular ordering, separate entry from stages, project in-flight state instead of storing it, and let the chart argue with your schema early.

Top comments (0)