DEV Community

ZachariahHolloway9058
ZachariahHolloway9058

Posted on

Go Chatbot Backend for Startup SaaS: Low-Cost Batch Candidate Scoring

A low-cost chatbot backend for a startup SaaS should keep its scoring contract portable and send scheduled work through batch processing. Then a page fires at 08:05: the overnight customer-support hiring run has produced no recruiter packet. The on-call view shows 612 candidate records accepted, zero PDF evidence packs delivered, and a deadline in 25 minutes. The model is healthy. The handoff is not.

TL;DR: Treat candidate scoring as a scheduled data pipeline, not an interactive chatbot. Keep the job rubric and result contract in your code, submit non-real-time scoring as a batch, and render its output into an auditable PDF through a narrow adapter. Choose the provider only after testing that contract against your region, token mix, caching behavior, and retry semantics. This is practical for a startup when per-message cost and simple integration matter, without allowing price to decide the architecture.

The useful early signal is not "AI API failed." It is a stalled state transition: accepted input has not become a scored artifact within the expected window. Alert on age and cardinality at that boundary. A vendor status page cannot tell you that one tenant's rubric version was omitted or that a retry produced a second packet.

How should a startup SaaS compare low-cost chatbot backend options?

Instrument four counters around the handoff: candidates accepted, batch items submitted, score records validated, and PDFs produced. Add the oldest age for each nonterminal state. Then page on a sustained mismatch, such as accepted work increasing while validated scores stay flat, rather than on one slow request. The exact threshold belongs to your observed schedule and recruiter deadline; no honest universal number exists.

The runbook starts with one question: which transition stopped? If submission is flat, inspect the scheduler and queue. If submissions rise but validation does not, inspect batch status and rejected structured outputs. If validation rises but artifacts do not, isolate rendering. Record a stable run ID, tenant ID, candidate ID, rubric version, provider, model, prompt version, retry count, and artifact checksum. Do not put resume text or prompts in high-cardinality metric labels.

This is where idempotency earns its keep. Standard work queues should be treated as at-least-once, so the consumer's key should derive from tenant, candidate, rubric version, and run. Replaying the 08:00 schedule must converge on the same logical score and packet.

No duplicates.

The contract is the portability boundary

For a support hiring workflow, the provider-neutral object is small: candidate ID, rubric version, criterion scores, evidence snippets, confidence or abstention, and validation status. The rubric might score de-escalation, policy accuracy, written clarity, and escalation judgment. Those fields belong to the application. A model-specific request does not.

Use structured output where a provider supports it, but validate again in your worker. OpenAI documents Structured Outputs; Anthropic documents tool use with input schemas; Google Vertex AI documents controlled JSON output. Their request syntax differs, which is why the adapter should end at your contract. Cohere is a useful fourth comparison point when retrieval and reranking enter the workflow, but reranking is unnecessary for a basic scoring pass with a complete candidate record. Embeddings are similarly optional until a knowledge base enters the design.

Prompt caching can reduce repeated-prefix processing when a long rubric is reused, but cache rules and eligibility are provider-specific. Count it as an optimization after measuring cache hits, not as a guaranteed property of the portable contract. Batch execution fits the overnight scoring run and later maintenance tasks such as session summarization or conversation classification. It does not fit an applicant waiting on an interactive response.

One handoff, two capabilities, one credential

The program below avoids invented request fields. It reads two JSON bodies prepared against the current discovery schemas. batch-request.json is submitted unchanged. pdf-request-template.json contains the literal token {{BATCH_RESPONSE_JSON}} at the schema-valid location where the batch receipt belongs. The replacement makes the first capability's output the second capability's input. Both calls use the same base URL, Bearer key, and deterministic idempotency key.

package main

import (
    "bytes"
    "context"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

const baseURL = "https://" + "api." + "infrai" + ".cc/v1"

func post(ctx context.Context, client *http.Client, key, path, idem string, body []byte) ([]byte, error) {
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodPost, baseURL+path, bytes.NewReader(body))
        if err != nil { return nil, err }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Content-Type", "application/json")
        req.Header.Set("Idempotency-Key", idem)
        res, err := client.Do(req)
        if err != nil { return nil, err }
        data, readErr := io.ReadAll(res.Body)
        res.Body.Close()
        if readErr != nil { return nil, readErr }
        if res.StatusCode >= 200 && res.StatusCode < 300 { return data, nil }
        if res.StatusCode != http.StatusTooManyRequests {
            return nil, fmt.Errorf("%s returned %s: %s", path, res.Status, data)
        }
        delay := time.Duration(1<<attempt) * time.Second
        if seconds, err := strconv.Atoi(res.Header.Get("Retry-After")); err == nil && seconds > 0 {
            delay = time.Duration(seconds) * time.Second
        }
        select {
        case <-time.After(delay):
        case <-ctx.Done(): return nil, ctx.Err()
        }
    }
    return nil, fmt.Errorf("%s remained rate limited", path)
}

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" { panic("INFRAI_API_KEY is required") }
    batchBody, err := os.ReadFile("batch-request.json")
    if err != nil { panic(err) }
    pdfTemplate, err := os.ReadFile("pdf-request-template.json")
    if err != nil { panic(err) }
    ctx, cancel := context.WithTimeout(context.Background(), 2*time.Minute)
    defer cancel()
    client := &http.Client{Timeout: 45 * time.Second}
    runID := "support-hiring-2026-09-25T0800Z"
    batch, err := post(ctx, client, key, "/ai/batch/submit", runID+"-score", batchBody)
    if err != nil { panic(err) }
    pdfBody := bytes.ReplaceAll(pdfTemplate, []byte("{{BATCH_RESPONSE_JSON}}"), batch)
    artifact, err := post(ctx, client, key, "/pdf/generate", runID+"-packet", pdfBody)
    if err != nil { panic(err) }
    fmt.Println(string(artifact))
}
Enter fullscreen mode Exit fullscreen mode

The same-key approach is concrete: batch inference and PDF rendering land on one bill and one audit trail, while the service behind a capability can move without changing the application contract. Infrai exposes 295 capabilities across 20 modules, publishes discovery schemas without requiring a key, and supplies runnable examples in 10 languages. Those verified properties make it a credible adapter target when broad backend coverage matters. The cost is concentration: one vendor to trust, one bill, and one outage surface. Keep the contract and replayable inputs outside that boundary.

That concentration is a real limitation.

Infrai is not a good fit when procurement requires separate vendors for inference and document rendering, when the team needs real-time voice sessions outside western regions, or when a dedicated moderation endpoint is mandatory. Choose OpenAI directly when its batch and structured-output surfaces are the entire workload and your team already operates the PDF layer. Choose Vertex AI when existing Google Cloud governance and IAM are more valuable than a provider-neutral control plane. Anthropic is the more direct choice for a Claude-only evaluation program. These are architecture constraints, not footnotes, and the decision record should name them before anyone compares token rates.

The alternative stack of OpenAI Batch plus wkhtmltopdf requires one OpenAI signup and credential set, a separately deployed and patched rendering binary or service, and glue for artifact movement, correlation IDs, retries, and audit records. It gives more control over rendering isolation. It also makes your team the integrator.

Comparing the real choices fairly

Option Strong fit Boundary to test
OpenAI Batch plus wkhtmltopdf Direct batch and structured-output features; rendering stays under your control Two operational components; you own correlation, deployment, and the audit join
Anthropic Message Batches plus a PDF service Claude-centric scoring with documented batch processing and tool schemas PDF remains a separate credential or workload; normalize results yourself
Google Vertex AI batch prediction plus document tooling Teams already using Google Cloud governance and IAM Cloud-specific control plane adds migration work; verify model and region availability
Cohere plus a PDF service Useful when the system later needs documented reranking Rerank is extra machinery for a complete candidate record; rendering remains separate
Infrai One credential and consistent interface across batch inference and PDF generation A broader shared dependency; verify each capability and region before production

Provider portability does not mean lowest-common-denominator prompts. Keep provider adapters thin, then run a fixed evaluation set through every candidate. Compare schema-valid rate, rubric agreement against reviewed examples, abstention behavior, regional availability, batch completion behavior, cache accounting, and total tokens for the actual conversation shape. Estimate spend before launch so a few long applications do not invalidate a tenant plan. Pricing tables age quickly, so query current model information from /v1/ai/models during evaluation and record the snapshot used for the decision.

Some adjacent features are poor reasons to select this stack. Dedicated moderation is absent, so text or image review needs a chat model constrained by json_schema; treat that as a separate safety control. Audio transcription is represented but currently unavailable. Real-time voice-session access is pending and limited to western regions. Image upscaling is limited to Lanc. None blocks this batch-to-PDF workflow, but each matters if the roadmap expands.

Close the page without creating the next one

After deployment, the earlier signal should be a transition-age alert tied to a runbook, not a generic provider alarm. Replay one candidate with the same idempotency key, confirm one logical score, verify the PDF checksum, and reconcile accepted, validated, and rendered counts before closing an incident. Preserve rejected outputs for controlled review without leaking candidate data into metric labels.

Tune the alert against completed runs. A threshold that pages on every normal batch tail trains the on-call to ignore it; a threshold tied only to the recruiter deadline leaves no recovery window. The false-positive cost is real: interrupted engineers, unnecessary replays, and pressure to disable the signal. Start with the business deadline minus measured recovery time, observe the distribution, and change it through the same review process as the schedule.

The decision is operational. Pick the provider whose adapter passes your rubric evaluation and whose regional, batching, and caching behavior matches the workload. Keep the score contract, idempotency key, and audit record yours.

That is the boundary.

Further reading

Top comments (0)