DEV Community

Roberto Luna
Roberto Luna

Posted on

Exposing Email Cron Errors and Adding Real‑Time Observability to the VS API

Exposing Email Cron Errors and Adding Real‑Time Observability to the VS API

TL;DR:

I refactored the email cron to stop swallowing SMTP/Brevo failures and to surface them through our observability layer. The change adds a scoped error reporter, propagates failures to the cron‑trigger controller, and ships a new test suite that guarantees errors aren’t hidden again.


The Problem

Our nightly runEmailCron() job ( apps/api/src/email/email.cron.ts ) was designed to send a batch of reminder emails. When the underlying SMTP provider (Brevo) rejected a message, the sendMail() helper caught the exception and silently marked the notification as sent. The symptom was a growing backlog of undelivered emails with no trace in logs:

ERROR: Cron job "email" completed successfully
Enter fullscreen mode Exit fullscreen mode

Even worse, the cron-trigger.controller (apps/api/src/cron/cron-trigger.controller.ts) considered the job successful because the caught error never bubbled up. This made debugging impossible and left our stakeholders unaware of delivery issues.


What I Tried First

The initial implementation (pre‑fix) looked like this:

// email.service.ts (original)
export async function sendMail(to: string, html: string) {
  try {
    await transporter.sendMail({ to, html });
  } catch (e) {
    // Swallow the error – we don’t want the cron to crash
    logger.warn('Email delivery failed, marking as sent');
    return;
  }
}
Enter fullscreen mode Exit fullscreen mode

I kept the try/catch because I feared the cron would abort and block other unrelated jobs. The idea was “fail fast, but keep the pipeline moving”. Unfortunately, the catch block never logged the real error and the calling cron kept counting the email as delivered.


The Implementation

1. Introduce a Scoped Error Reporter

We already have an observability module (apps/api/src/common/observability.ts) that forwards errors to Sentry/Datadog. I added a mail‑specific scope so we can filter and aggregate email failures separately.

// apps/api/src/email/email.service.ts
import { AsyncLocalStorage } from "node:async_hooks";
import { reportError } from "../common/observability.js";

export const mailFailureScope = new AsyncLocalStorage<string>();

export async function sendMail(to: string, html: string) {
  try {
    await transporter.sendMail({ to, html });
  } catch (err) {
    // Attach the current scope (e.g., cron run id) for better tracing
    const scope = mailFailureScope.getStore() ?? "unknown";
    reportError(err, { scope, to });
    // Re‑throw so the caller knows the job failed
    throw err;
  }
}
Enter fullscreen mode Exit fullscreen mode

Why AsyncLocalStorage? It propagates a request‑level identifier (cronRunId) across async boundaries without manually threading it through every function.

2. Wire the Scope into the Cron

// apps/api/src/email/email.cron.ts
import {
  sendCondoFeeReminder,
  sendCondoFeeOverdue,
  // …
  mailFailureScope,
} from "./email.service.js";
import { reportError } from "../common/observability.js";

export async function runEmailCron() {
  const runId = `email-${Date.now()}`;
  // Store the runId for the whole execution
  return mailFailureScope.run(runId, async () => {
    try {
      await Promise.all([
        sendCondoFeeReminder(),
        sendCondoFeeOverdue(),
        // other email jobs…
      ]);
    } catch (e) {
      // The error is already reported by sendMail; just surface it
      reportError(e, { runId, job: "emailCron" });
      // Propagate to the controller
      throw e;
    }
  });
}
Enter fullscreen mode Exit fullscreen mode

Now any failure inside sendMail automatically carries the runId, and the cron itself re‑throws the error after reporting.

3. Make the Cron‑Trigger Controller Respect Failures

// apps/api/src/cron/cron-trigger.controller.ts
async function runJob(
  job: string,
  fn: () => Promise<unknown>
): Promise<JobResult> {
  try {
    const res = await fn();
    // If the job returns a falsy value we still consider it success
    return { job, status: "ok", result: res };
  } catch (err) {
    // Previously we ignored the error and returned ok
    // Now we mark the job as failed
    return { job, status: "error", error: err.message };
  }
}
Enter fullscreen mode Exit fullscreen mode

The controller now returns status: "error" when any job throws, which propagates to the GitHub Actions workflow.

4. Add a Test Suite to Guard Against Silent Failures

// apps/api/src/__tests__/email.cron.errores.test.ts
import { runEmailCron } from "../../email/email.cron.js";
import { sendMail } from "../../email/email.service.js";

jest.mock("../../email/email.service", () => ({
  sendMail: jest.fn(),
}));

test("runEmailCron propagates SMTP errors", async () => {
  // Force sendMail to reject
  (sendMail as jest.Mock).mockRejectedValue(new Error("SMTP 550"));
  await expect(runEmailCron()).rejects.toThrow("SMTP 550");
});
Enter fullscreen mode Exit fullscreen mode

The test guarantees that if sendMail throws, the cron does not swallow it.

5. Update the GitHub Actions Workflow

# .github/workflows/cron-diario.yml
name: ⏰ Cron diario — API PlayaMXCRM

on:
  schedule:
    - cron: "7 7 * * *"   # 07:07 UTC, compensates GitHub delay

jobs:
  run-daily:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v3
      - name: Run API cron jobs
        run: |
          npm ci
          npm run start:cron   # invokes cron-trigger.controller
Enter fullscreen mode Exit fullscreen mode

Running the jobs via GitHub Actions ensures the API container (Render free tier) wakes up, executes the cron, and exits with a non‑zero code if any job fails. This gives us immediate feedback in the Actions UI.


Key Takeaway

Never swallow errors in background jobs. Use a scoped observability helper (AsyncLocalStorage + reportError) to surface failures, and let the orchestrator (cron‑trigger controller) treat thrown errors as job failures. This pattern gives you real‑time alerts, reliable test coverage, and a clean separation between “job ran” and “job succeeded”.


What's Next

  1. Alerting: Hook reportError into PagerDuty so a failed email batch triggers an on‑call incident.
  2. Retry Logic: Implement exponential back‑off retries inside sendMail for transient SMTP errors.
  3. Metrics: Emit a Prometheus counter (email_cron_success_total, email_cron_failure_total) for dashboards.
  4. Feature Flag: Wrap the new error‑exposing behavior behind a flag to roll out gradually in production.

Roberto Luna Osorio – Full Stack Developer & Project Lead

Playa del Carmen, México

vibecoding #buildinpublic #nodejs #typescript #cron #observability #github-actions


Part of my Build in Public series — sharing the real process of building Building PlayaMXCRM from Playa del Carmen, México.

Repo: zaerohell/VS · 2026-10-09

#playadev #buildinpublic

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •
You need to verify your account.
Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to