DEV Community

Cover image for From 60-Second Timeouts to 200ms: A Two-Day GraphQL Pilot on a Fintech Dashboard
Elias Melgaço
Elias Melgaço

Posted on Originally published at Medium

From 60-Second Timeouts to 200ms: A Two-Day GraphQL Pilot on a Fintech Dashboard

A performance war story — and an honest accounting of where the 300x actually came from.

The context

A few years ago I was the senior frontend engineer at a US fintech that processes billions of dollars in invoices and payments. Like many fast-moving startups at that stage, we had no error monitoring or APM in place yet. Our first signal that something was wrong came the old-fashioned way: a customer told us their dashboard was unusable.

The problem

Digging in, we found that one specific endpoint — the one that loaded vendors — was routinely timing out at 60+ seconds for this customer. When it timed out, the screen simply never rendered the data. Worse: there was no clear loading or error state in the UI. From the user's perspective, the product wasn't slow; it was broken.

The endpoint was served by Django, and the vendors query traversed multiple relationships. Without profiling in place (see "no APM" above), we couldn't produce a flame graph — but the symptoms matched the classic pattern of inefficient ORM relationship loading, the N+1 shape: response time scaling with the customer's data size, a large vendor base on exactly the account that was suffering, and a query structure the ORM was unlikely to be batching well. I'm deliberately calling this an informed diagnosis rather than a proven one; part of the lesson of this story is that we were operating without the instrumentation that would have made it provable.

The constraints (which shaped the solution more than the tech did)

Three organizational facts mattered more than any benchmark:

  1. The backend team was swamped. A proper fix inside the Django codebase would have to wait in line behind a full roadmap.
  2. A mobile project was being born, and the plan was for it to consume GraphQL. Whatever we built here could become its foundation.
  3. I was given carte blanche to solve it — with one requirement from the CTO: if I was going to build a Node service, it had to run on Fastify, not Express.

My initial sketch was Express + Apollo Server — the default pairing most of us reach for. The Fastify requirement pushed me to Mercurius, and I want to be fair to that constraint, because it turned out to have solid engineering behind it rather than being arbitrary:

  • Fastify itself: meaningfully lower per-request overhead and higher throughput than Express, which at the time was in near-maintenance mode; a plugin architecture with real encapsulation; schema-based validation and serialization out of the box.
  • Mercurius is not "Apollo for Fastify." It's a GraphQL adapter integrated into Fastify's lifecycle and plugin system, with JIT query compilation (via graphql-jit) for hot queries and — the part that mattered most here — a native loaders mechanism implementing the DataLoader pattern, which became my primary tool against N+1 at the GraphQL layer.

A constraint I didn't choose gave me a lighter, faster stack than the one I would have picked by default. I've tried to remember that lesson every time I'm on the other side of the table, imposing a constraint on someone else.

The solution

The architecture was deliberately boring:

  • A small Fastify + Mercurius GraphQL service, owned by the frontend team.
  • Connected to a read replica of our Postgres database — no writes, no coupling to the Django application code, and near-zero blast radius: the worst realistic failure mode was this one dashboard degrading, not the platform.
  • The service lived inside our existing infrastructure perimeter and reused the platform's authentication — it added a query path, not a new attack surface.
  • A schema exposing exactly what the dashboard needed: vendors and their relationships, resolved with SQL written for the access pattern instead of generated by an ORM.
import Fastify from 'fastify'
import mercurius from 'mercurius'

const app = Fastify()

app.register(mercurius, {
  schema,
  resolvers,
  loaders, // batching layer: N+1 prevention at the GraphQL layer
})
Enter fullscreen mode Exit fullscreen mode

GraphQL has its own famous N+1 trap — naive resolvers fire one query per field per parent. Since the entire point of this service was escaping an N+1-shaped problem, recreating it one layer up would have been embarrassing. Mercurius' loaders (per-request batching and caching, the DataLoader pattern) meant resolving hundreds of vendors' relationships produced a handful of batched queries instead of hundreds of individual ones.

One trade-off I accepted with eyes open: reading from a replica means eventual consistency. For a vendors dashboard, data that is seconds behind the primary is indistinguishable from perfect; for a payments flow it would have been disqualifying. Same technique, different context, different answer.

From decision to production took two days.

Architecture before and after the GraphQL pilot

The result — and where it actually came from

The vendors endpoint went from 60+ seconds (frequently timing out) to roughly 200 milliseconds — about a 300x improvement. The dashboard loaded. The customer noticed the same day.

Now let me be precise about something, because this is where posts like this usually lose experienced readers: the 300x did not come from GraphQL, and it did not come from Fastify being faster than Express. Framework overhead differences are measured in microseconds to low milliseconds; they explain nothing at the scale of a 60-second response. The performance win came from one thing: replacing ORM-generated query patterns with SQL written for the access pattern, executed against a replica. A hand-tuned REST endpoint — in Fastify, or in Python bypassing the ORM — would have achieved essentially the same latency.

So why GraphQL? Because the performance fix and the architectural choice were solving two different problems:

  • The SQL rewrite solved the latency.
  • The GraphQL layer was the strategic wrapper around it: it gave the frontend team autonomy to shape queries per screen without backend tickets, and it doubled as the proof of concept for the API layer the upcoming mobile app would consume.

Conflating those two is how "we adopted X and got 300x" myths get written. I'd rather you leave this post knowing exactly which lever did what.

Vendors endpoint response time, before vs after

The honest trade-off analysis

The road not taken: a new, lean Python endpoint bypassing the Django ORM (raw SQL or a carefully tuned queryset with select_related/prefetch_related), plus frontend work — paginated/infinite loading on the vendors dropdown and React Query for caching and request state. That path keeps the stack homogeneous and avoids operating a new service.

Why we didn't take it:

  • Time-to-fix under real capacity. The Python path required attention from a team fully committed elsewhere. The GraphQL service could be built and owned end-to-end by frontend, immediately.
  • The mobile roadmap tipped the scale. If GraphQL was coming anyway, the pilot did double duty; a one-off endpoint solved one screen.
  • A new service is a real cost — accepted knowingly. More infrastructure, another deploy pipeline, another thing to monitor (ironic, given how this story started). In a team with spare backend capacity and no GraphQL on the roadmap, the boring Python fix is probably the better call.

The uncomfortable truth of engineering decisions at startups: the "best" solution is a function of team capacity and roadmap, not just of the code. The same problem, in a different org chart, has a different right answer.

What I'd tell you if you're facing something similar

  1. Instrument before you're forced to. We learned about a 60-second endpoint from a customer, not a dashboard — and it left us diagnosing by inference instead of by flame graph. We fixed that afterwards (Sentry); the lesson stuck.
  2. Never let a request fail silently. A UI that hangs with no feedback turns a slow endpoint into a "broken product" in the user's mind, regardless of what the backend is doing.
  3. A read replica is a beautiful de-risking tool — as long as you've checked that eventual consistency is acceptable for the read path in question.
  4. Treat constraints as design input. The CTO's "use Fastify" requirement felt like friction and turned out to be a well-founded gift.
  5. Attribute your wins precisely. If you fix a data-access problem and wrap it in new architecture, know — and say — which one produced the numbers. Your credibility with senior engineers depends on it.

I'm a senior frontend engineer with 14+ years building high-traffic web applications in TypeScript, React, and Next.js, currently focused on data-heavy, accessible interfaces and AI-assisted development workflows. If you've faced a similar build-vs-wait decision, I'd genuinely like to hear how you called it.

Originally published on Medium.

Top comments (0)