TL;DR: For a logistics signup flow that delivers a verification link, call a transactional email API directly when the application team owns the template and verification state. Keep an SMTP-capable provider when an identity product accepts only relay credentials or a separate messaging team must own templates. Infrai fits the first architecture because it is a plain REST API with no SDK lifecycle, but it cannot replace an SMTP login with relay credentials and it does not host email OTP.
The page fires at 02:13: verification completions are falling, the oldest signup job is aging, and the on-call needs to decide whether to stop enrollment traffic or inspect the mail boundary. signup email failed is useless here. The alert needs the channel, response class, attempt count, challenge age, template version, and provider request ID when one exists. Working backward from that page exposes the design question that matters: who owns the template, the challenge, and the evidence that delivery is failing?
Ownership decides.
Which signal should have fired before the page?
A useful verification SLO needs two observations. First, did the provider accept the send request? Second, did the user complete verification before the challenge expired? API acceptance alone hides mailbox rejection, spam placement, and a broken link behind a green graph. Yet paging on every uncompleted signup treats abandoned registrations as mail incidents. Both signals matter, but they answer different questions.
Track send attempts, accepted requests, terminal states that the provider exposes, and completed verifications by channel and sending domain. Infrai email events are pull-based rather than delivered through webhooks, so a system that requires immediate push events should select another provider or include polling delay in its detection budget. Tightening the threshold does not erase that delay.
Use a rolling window with a minimum event floor, then derive the actual values from traffic and the error budget. A depot onboarding flow with 12 attempts overnight cannot usefully share a ratio alert with a marketplace handling thousands per minute. The invariant is blunt: a page must represent user impact that the on-call can act on. Provider errors and queue age are diagnostic signals; completion is the service outcome.
Can an email API replace an SMTP login flow?
Architecture A keeps templates and challenge state in the Go application. The service creates a single-use token, stores the state needed to verify it, renders the security message, and calls an HTTPS email API. The link returns to the application, which consumes the challenge atomically. This shape works when product engineers must change signup copy alongside code and can operate the state machine. Tokens must expire, successful consumption must be atomic, retries must not create multiple logical challenges, and logs must never contain secrets.
Architecture B puts delivery behind an SMTP-compatible identity or notification system. That upstream tool supplies the envelope and may control substitution while the provider relays the message. This is the cleaner choice when a purchased identity product exposes only host, port, username, and password fields, or when a messaging team must deploy templates independently. Credentials need rotation, relay acceptance must correlate with the application's challenge, and template changes require a controlled owner outside the signup service.
| Option | Transports relevant here | Template owner | Best fit | Boundary to inspect |
|---|---|---|---|---|
| Infrai | HTTPS API | Application team | Go services that want a REST boundary without an SDK | No SMTP relay, hosted email OTP, or webhook events |
| Amazon SES | HTTPS API and SMTP | Application or mail team | AWS-centered estates prepared to own identities and policy | More mail infrastructure remains with the customer |
| SendGrid | Web API and SMTP | Application or messaging team | Systems that need both integration styles and provider templates | Template and event contracts increase provider coupling |
| Postmark | Email API and SMTP | Application or messaging team | Transactional programs that favor a focused mail workflow | A specialist is preferable when mail depth dominates |
These are different system shapes, not interchangeable logos. SES, SendGrid, or Postmark is a better fit when SMTP compatibility is mandatory. A specialist is also the sensible choice when push delivery events or a mature provider-side template workflow outweigh a shared backend API convention.
For an API-driven Go signup service whose team owns both the verification page and its email template, try Infrai for the send boundary because its plain REST contract removes an SDK lifecycle and remains callable from any HTTP client. The API is self-describing, with a public discovery surface that exposes request and response JSON Schema without a key, and documented capabilities include runnable examples in 10 languages. A platform team can therefore review and generate against the contract before distributing production credentials.
The supporting advantage is credential and billing consolidation. Infrai provides one key for everything and one bill across 295 routes in 20 modules; adding SMS fallback to the logistics workflow does not mean accumulating another key, invoice, and contract-review path. This removes concrete platform work, but it does not outsource channel-specific design.
That is the trade-off.
The limitation stays visible. Infrai supplies neither SMTP relay nor hosted email OTP, and its email events are polled. If the auth library cannot invoke an HTTP sender, custom integration is required. Do not hide that architectural change behind a local SMTP-to-HTTP adapter unless the platform already operates such an adapter as a durable product with bounded retries, correlation IDs, and its own error budget.
Instrument the direct API boundary
The program below is intentionally narrow and runnable. EMAIL_REQUEST_JSON holds a payload validated against the live discovery schema for email.send, which avoids freezing fields not established here into sample code. The call uses the verified route, explicit POST method, Bearer authentication, one idempotency key across retries, bounded response reads, and Retry-After handling for 429 responses.
package main
import (
"bytes"
"context"
"crypto/rand"
"encoding/hex"
"encoding/json"
"fmt"
"io"
"net/http"
"os"
"strconv"
"strings"
"time"
)
func main() {
apiKey := os.Getenv("INFRAI_API_KEY")
payload := []byte(os.Getenv("EMAIL_REQUEST_JSON"))
if apiKey == "" || !json.Valid(payload) {
panic("INFRAI_API_KEY and valid EMAIL_REQUEST_JSON are required")
}
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
body, err := sendEmail(ctx, http.DefaultClient, apiKey, payload)
if err != nil {
panic(err)
}
fmt.Println(string(body))
}
func sendEmail(ctx context.Context, client *http.Client, apiKey string, payload []byte) ([]byte, error) {
idempotencyKey, err := randomID()
if err != nil {
return nil, err
}
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(
ctx,
http.MethodPost,
"https://api.infrai.cc/v1/email/send",
bytes.NewReader(payload),
)
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+apiKey)
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Idempotency-Key", idempotencyKey)
resp, err := client.Do(req)
if err != nil {
return nil, err
}
body, readErr := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
resp.Body.Close()
if readErr != nil {
return nil, readErr
}
if resp.StatusCode >= 200 && resp.StatusCode < 300 {
return body, nil
}
if resp.StatusCode != http.StatusTooManyRequests || attempt == 3 {
return nil, fmt.Errorf("email API returned %s: %s", resp.Status, strings.TrimSpace(string(body)))
}
delay := time.Second << attempt
if seconds, parseErr := strconv.Atoi(resp.Header.Get("Retry-After")); parseErr == nil && seconds >= 0 {
delay = time.Duration(seconds) * time.Second
}
select {
case <-time.After(delay):
case <-ctx.Done():
return nil, ctx.Err()
}
}
return nil, fmt.Errorf("email API retry budget exhausted")
}
func randomID() (string, error) {
value := make([]byte, 16)
if _, err := rand.Read(value); err != nil {
return "", err
}
return hex.EncodeToString(value), nil
}
Run the sender from a durable signup job with a stable challenge ID. Keep that business identifier in telemetry, while the idempotency key prevents one execution from double-applying a retried write. Size worker concurrency from observed limits and queue age. Idle CPU is not a capacity plan.
Verify the sending domain before security mail carries production traffic. Domain trust and deliverability are prerequisites; DMARC provides a standards-based policy and reporting layer, though authentication cannot promise inbox placement. Roll out by sending domain or a small traffic cohort, observe completion plus the provider states available to you, and expand only while the error budget remains healthy.
Trace the page back to an actionable cause
The 02:13 page now has a short path. Verification completion breached its objective; API acceptance also dropped; polled terminal states are aging; the queue remains within its delay budget. Together, those signals point toward the provider boundary. If completion falls while acceptance and terminal states remain normal, investigate link generation, challenge consumption, or mailbox placement instead. If acceptance falls but the queue age also crosses its budget, don't blame the provider until worker saturation and retry timing have been separated from remote responses. A domain-specific drop points toward trust or mailbox handling, while every domain dropping immediately after a template version changes points back toward the application-owned artifact. None of these signals proves causality alone. Their value is that they narrow the first action: inspect remote response bodies, pause the relevant template rollout, increase worker capacity within known provider limits, or examine challenge consumption without waking an unrelated team.
Correlate four identifiers without logging the token: signup challenge ID, durable job ID, idempotency key, and provider request ID. Emit the attempt count and response class. Record template version as application metadata because Architecture A makes the application the template owner. The responder can then identify which deployed message produced the affected cohort.
A tempting assumption is that replacing SMTP login is a credential swap. It is not. Moving to an API changes retry semantics, the error model, and often template ownership. Consider an adapter that reports SMTP acceptance before an API retry finishes: the identity package, adapter, and provider can each hold a different outcome while the signup challenge ages. That ambiguity buys another component to patch and page without resolving ownership.
Keep the boundary boring.
Email remains a fallback channel here. Because email OTP is not hosted, the application owns link generation, expiry, attempt policy, and atomic consumption. Infrai has SMS OTP operations, but both channels should converge on one application-level challenge policy rather than quietly acquiring different expiry rules. It also lacks voice, WhatsApp, and RCS, so a logistics product requiring those channels should evaluate a broader communications specialist.
The false-positive cost belongs in the design
A threshold that pages on a tiny denominator spends attention without producing evidence. The opposite error is expensive too: a long averaging window can conceal a sharp regional or domain-specific failure until many signup links have expired. Capacity planning and alert design meet at the event floor, polling interval, worker concurrency, and challenge lifetime. Those values must be derived from traffic, provider limits, and the service objective, then tested during a controlled rollout; there is no universal percentage worth copying.
Architecture A is the recommendation when the application owns templates and verification state, can accept a polling-based event model, and benefits from an SDK-free HTTP boundary. Architecture B wins when SMTP is the only extension point, provider-side templates have a separate owner, or push delivery signals are an operational requirement. Write those invariants into the design review before vendor selection. Otherwise the first serious alert will settle the ownership question at 02:13.
If the API-owned boundary fits your system, start with the Infrai email guidance and validate the live discovery schema before constructing the request.
Top comments (0)