DEV Community

GodfreySterling1574
GodfreySterling1574

Posted on

Choose Heartbeats Over a FastAPI Django Error Tracking API for 15-Minute SaaS Imports

Choose an explicit result heartbeat over a simple error tracking API when a FastAPI or Django SaaS must answer, "Did the scheduled import produce a result?" Use searchable exception events and grouped issues as supporting diagnosis, not as the primary detector. For an import expected every 15 minutes, alert on an overdue, uniquely identified schedule occurrence only after its completion deadline has passed, then attach any correlated exception group to that alert; this choice produces a higher-quality signal because silence becomes observable while retries, duplicate deliveries, and repeated stack traces do not manufacture extra incidents.

Short answer: choose heartbeat-first detection when missing work matters more than counting thrown errors. Choose exception-first detection only when every meaningful failure is guaranteed to throw, reach the tracker, and remain distinguishable from retry noise. That is a demanding invariant for scheduled imports.

Should FastAPI and Django SaaS teams choose an error tracking API?

An exception proves that one execution path reported an error. It does not prove that a schedule fired, that a worker received the job, that the process remained alive long enough to report telemetry, or that a syntactically successful import produced a usable result. A quiet queue, a disabled scheduler, expired credentials handled as an empty response, and a worker terminated before export can all leave the exception channel silent. No news is ambiguous.

A result heartbeat answers a narrower question. Record one durable observation per expected schedule occurrence, with fields such as tenant_id, integration_id, scheduled_for, completed_at, outcome, records_seen, and trace_id. The scheduler or monitor can then compare expected occurrences with recorded outcomes. The heartbeat is not a log line whose presence depends on search retention; it is evidence tied to a business-level unit of work.

The ADR therefore has three invariants. First, (tenant_id, integration_id, scheduled_for) is unique, so a retry cannot create a second logical completion. Second, an occurrence is not late until its explicit deadline, which prevents a slow but valid run from paging the operator. Third, every state change is appendable to an audit trail or otherwise reconstructable: scheduled, started, completed, failed, and alert acknowledged must retain actor or process identity and timestamps according to the organization's retention policy.

Exactly once is the mindset, not a claim about transport magic. Delivery may happen more than once; the database must make the business effect idempotent.

Silence needs a record.

Decision record and failure boundaries

The decision is to make the schedule ledger authoritative for alert state and keep searchable exception events as diagnostic context. The detector reads durable occurrences, not application log volume. Logs remain event streams, as described by the Twelve-Factor App, and a W3C Trace Context trace-id links the occurrence to traces and exception events without turning either telemetry stream into the system of record.

Criterion Result heartbeat and schedule ledger Grouped exception events
Detects scheduler silence Yes, because an expected occurrence can become overdue No event exists to group
Handles successful zero-result runs Explicit outcome and count distinguish valid zero from suspicious zero Usually silent unless application code invents an exception
Retry behavior Unique occurrence key collapses repeated delivery into one business outcome Repeated throws may increase event volume inside one or several groups
Diagnostic depth Requires a trace or error link for stack-level detail Stack, request context, and grouping are the main strengths
Auditability State transitions map directly to scheduled business work Event history describes failures, not the complete schedule
Primary noise risk Deadlines set tighter than real completion latency Retry storms, unstable fingerprints, and non-actionable exceptions

There are boundaries. If the schedule ledger is unavailable, the detector must fail closed operationally: emit a distinct monitor-health signal rather than asserting that every tenant import is late. If trace export is unavailable, result recording must still succeed because observability cannot sit in the transaction's critical commit path. If a worker completes after an alert opens, the same occurrence should resolve or annotate the existing alert instead of opening a fresh one.

Compliance also constrains event content. GDPR Article 5 requires data minimization and limits storage to what is necessary, while Article 32 requires security measures appropriate to risk. Those principles favor opaque tenant and integration identifiers, bounded retention, access controls, and redaction before exception payloads leave the application boundary. Region is consequently a deployment and data-flow decision, not a checkbox inferred from a vendor's marketing page: document where event bodies, backups, indexes, and support access reside for the US and EU paths.

The trade-off is real. A schedule ledger adds a schema, a reconciliation query, retention work, and a monitor whose own health must be tested; it is not suitable for jobs with no stable expected cadence, and an event-only alternative may be proportionate when missed work has no operational consequence. Heartbeats also provide less diagnostic detail than exception events, so this design deliberately operates two signals instead of pretending that one replaces the other.

The critical path in Go

The following handler is deliberately small. A FastAPI, Django, Rails, or Laravel worker can call an equivalent internal endpoint after committing its import result; the receiver uses a uniqueness constraint to make retries harmless. Authentication, authorization, rate limits, and schema migration details belong around this core rather than inside the example.

package heartbeat

import (
    "context"
    "database/sql"
    "errors"
    "time"
)

type Completion struct {
    TenantID     string
    IntegrationID string
    ScheduledFor time.Time
    CompletedAt  time.Time
    Outcome      string
    RecordsSeen  int64
    TraceID      string
}

type Store struct {
    DB *sql.DB
}

func (s Store) RecordCompletion(ctx context.Context, c Completion) error {
    if c.TenantID == "" || c.IntegrationID == "" || c.ScheduledFor.IsZero() {
        return errors.New("missing occurrence identity")
    }
    if c.CompletedAt.Before(c.ScheduledFor) {
        return errors.New("completion precedes schedule")
    }

    _, err := s.DB.ExecContext(ctx, `
        INSERT INTO import_occurrences
          (tenant_id, integration_id, scheduled_for, completed_at,
           outcome, records_seen, trace_id)
        VALUES ($1, $2, $3, $4, $5, $6, $7)
        ON CONFLICT (tenant_id, integration_id, scheduled_for)
        DO UPDATE SET
          completed_at = EXCLUDED.completed_at,
          outcome = EXCLUDED.outcome,
          records_seen = EXCLUDED.records_seen,
          trace_id = EXCLUDED.trace_id
        WHERE import_occurrences.completed_at <= EXCLUDED.completed_at`,
        c.TenantID, c.IntegrationID, c.ScheduledFor, c.CompletedAt,
        c.Outcome, c.RecordsSeen, c.TraceID,
    )
    return err
}
Enter fullscreen mode Exit fullscreen mode

The unique database key is the decisive control. A random event identifier would deduplicate network delivery but would not establish that two retries represent the same scheduled business occurrence. Keeping the natural occurrence identity also supports reconciliation: an operator can enumerate the expected schedule, left-join recorded results, and explain why each alert existed.

That distinction matters.

Do not let the upsert conceal conflicting outcomes. In a production schema, append each attempt to an immutable attempt table and maintain a separately derived current occurrence state; reject an invalid state transition, or record it for review, rather than overwriting evidence. The compact example shows the idempotency boundary, while the audit model preserves what happened.

Noise control is a policy, not a grouping algorithm

Start with deadlines derived from the import contract: cadence, allowed start delay, and maximum completion duration. For the 15-minute example, the number defines the expected schedule, not automatically the paging threshold. A run that starts every 15 minutes but routinely needs 18 minutes demands overlap control and a deadline grounded in its service objective. Inventing a five-minute grace period would look precise while lacking evidence.

No universal grace period exists.

Then separate states that require different responses. overdue means no terminal result exists after the deadline. failed means a terminal result explicitly reports failure. suspicious_zero means the import completed with zero records under a rule backed by domain history or an upstream contract. Only the first two necessarily indicate execution failure; a zero can be perfectly valid for a quiet tenant.

One alert should represent one affected integration and schedule occurrence, even if three worker retries throw the same exception. Aggregate repeated late occurrences only when the escalation policy says the operator action is identical. Preserve counts and timestamps in the audit record, because aggressive grouping can make a prolonged outage look like one old event while ungrouped retries can make one fault look like a fleet-wide emergency.

Test the policy with deterministic absence. Pause schedule publication, prevent worker consumption, return a valid empty payload, force a database rollback, duplicate a completion request, and delay trace export. Each test should predict one ledger state and one alert transition before deployment. This is also where framework choice recedes: middleware can capture exceptions in FastAPI, Django, Rails, and Laravel, but only the application's scheduling contract can define a missing result.

Why reject exception-first detection?

Exception-first detection is rejected for this job because its negative case is unknowable: an empty search can mean healthy execution, total scheduler silence, exporter failure, sampling, retention expiry, or a query mismatch. Searchable events and grouped issues remain valuable after the heartbeat identifies the affected occurrence. They shorten diagnosis when stack traces, release metadata, or request context explain the failed attempt.

The rejected option has a valid boundary. It is sufficient for a best-effort background task when missed execution carries no contractual or financial consequence, all actionable failures are known to report an exception, and nobody needs proof that every schedule occurrence completed. It can also be the first stage during an early service rollout, provided the decision record explicitly accepts blind spots and names the event that will trigger migration to a schedule ledger.

For low-operations teams, the selection exercise should therefore test capabilities rather than collect feature grids. Verify stable fingerprint controls, searchable structured fields, retention and deletion behavior, regional data flows, export paths, role-based access, and trace-context preservation with representative payloads. Run the same duplicate, redaction, late-arrival, and outage cases against every candidate. Product labels change; these failure boundaries do not.

The final choice is firm: for scheduled B2B SaaS imports whose missing results require action, own a small idempotent schedule ledger and use exception tracking beside it. That architecture makes silence queryable, keeps the alert tied to a reconcilable business occurrence, and leaves exception grouping to the diagnostic task it can actually prove.

References

Top comments (0)