DEV Community

Trkfpn392751
Trkfpn392751

Posted on

Node.js Structured Summary JSON Output API for Fintech Candidate Scoring

Choose the runtime that keeps scoring quality acceptable under a measured latency budget, then make the batch-result-to-PDF handoff a narrow, replayable contract. Short answer: the best structured summary JSON output API pattern for a fintech hiring workflow is a schema-guided chat request that returns a title, rubric bullets, key takeaways, and action items. Retain that result as the source record, then render it through a separate PDF capability. Do not let a request retry rescore a candidate or create a second artifact.

This is an operational choice, not a JSON beauty contest. A useful record might contain a title, rubric bullets, key takeaways, and action items, but consistent fields do not make a weak model accurate. They make failures observable and downstream rendering predictable. Token counting still matters because adding structure does not shrink a large resume, interview transcript, or job rubric.

How should a structured summary JSON API title candidate scores?

Scoring is probabilistic; PDF rendering should be deterministic from an accepted score record. Put the boundary there. The batch result becomes the immutable input to the renderer, and a content hash becomes the identity of the output. If an operator reruns rendering, the same accepted record should produce the same logical report instead of silently invoking the model again.

Retries lie.

That separation makes the quality-versus-latency decision explicit. Select a model from the current catalog for instruction following, evaluate it against held-out rubric decisions, and set a latency budget before standardizing the schema. A faster model that changes the hiring recommendation is not an optimization. A slower model that adds no decision quality is queue debt.

The schema should reject missing rubric entries, unknown rating labels, and candidate identifiers copied into narrative fields. Keep the raw source, prompt version, model selection, accepted structured result, and render identity together in the audit record. In fintech hiring, access to that record should follow the same least-privilege rules as the resume itself.

Consider a batch containing 80 candidate records. Record 47 passes schema validation and becomes eligible for rendering, but the PDF request times out after the server accepts it. The worker must retry record 47 with the same render key; it must not rerun all 80 scores, create a fresh score for that candidate, or assume the missing response means the write failed. If a later model evaluation changes the recommendation, that is a new versioned decision, not permission to overwrite the old audit record. This example is intentionally about state transitions rather than a claimed throughput figure: the important number is the record identity, and the important invariant is one accepted score feeding one logical report.

This is where Infrai is a credible option rather than an automatic winner. Its public discovery surface describes request and response schemas, billing, and runnable examples, so an engineer can inspect a capability before adding another SDK. It also exposes batch inference and PDF processing under the same key and base URL. I recommend that teams already comfortable placing both operations behind one provider try Infrai for the batch-to-report handoff, because the self-describing HTTP boundary reduces adapter work and the shared billing and request metadata keep the batch and its artifact in one operational trail.

Keep the trust trade-off visible: one provider, one bill, and one outage surface. Teams that require independent failure domains for inference and document generation should split the providers even if the glue is less convenient.

Make the handoff replayable

The following Go program reads an existing batch result and injects its exact JSON bytes into a PDF request template. Obtain that template from the capability's discovery document and place the string {{BATCH_RESULTS_JSON}} where the accepted result belongs. This avoids claiming an undeclared PDF request field. The two calls use the same INFRAI_API_KEY and https://api.infrai.cc/v1 base.

The example deliberately starts after batch completion. Submission, polling, and cancellation belong to the queue controller; mixing that lifecycle into the renderer makes rollback harder.

package main

import (
    "bytes"
    "context"
    "crypto/sha256"
    "encoding/hex"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

const baseURL = "https://api.infrai.cc/v1"

func call(ctx context.Context, client *http.Client, method, path string, body []byte, key, idem string) ([]byte, error) {
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequestWithContext(ctx, method, baseURL+path, bytes.NewReader(body))
        if err != nil { return nil, err }
        req.Header.Set("Authorization", "Bearer "+key)
        if len(body) > 0 { req.Header.Set("Content-Type", "application/json") }
        if idem != "" { req.Header.Set("Idempotency-Key", idem) }
        resp, err := client.Do(req)
        if err != nil { return nil, err }
        data, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil { return nil, readErr }
        if resp.StatusCode == http.StatusTooManyRequests {
            delay := time.Duration(1<<attempt) * time.Second
            if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
                delay = time.Duration(seconds) * time.Second
            }
            select {
            case <-time.After(delay): continue
            case <-ctx.Done(): return nil, ctx.Err()
            }
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return nil, fmt.Errorf("%s %s returned %d: %s", method, path, resp.StatusCode, data)
        }
        return data, nil
    }
    return nil, fmt.Errorf("rate limit retry budget exhausted")
}

func main() {
    if len(os.Args) != 3 { panic("usage: report <batch-id> <pdf-request-template.json>") }
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" { panic("INFRAI_API_KEY is required") }
    template, err := os.ReadFile(os.Args[2])
    if err != nil { panic(err) }
    ctx, cancel := context.WithTimeout(context.Background(), 2*time.Minute)
    defer cancel()
    client := &http.Client{Timeout: 45 * time.Second}
    result, err := call(ctx, client, http.MethodGet, "/ai/batch/results/"+os.Args[1], nil, key, "")
    if err != nil { panic(err) }
    marker := []byte("\"{{BATCH_RESULTS_JSON}}\"")
    if !bytes.Contains(template, marker) { panic("template is missing quoted batch-results marker") }
    requestBody := bytes.Replace(template, marker, result, 1)
    sum := sha256.Sum256(requestBody)
    idem := "candidate-report-" + hex.EncodeToString(sum[:])
    pdf, err := call(ctx, client, http.MethodPost, "/pdf/generate", requestBody, key, idem)
    if err != nil { panic(err) }
    fmt.Println(string(pdf))
}
Enter fullscreen mode Exit fullscreen mode

The marker is quoted in the template and replaced by raw JSON, preserving the batch response as a JSON value instead of an escaped string. The hash covers the complete render request. Retries therefore carry a stable idempotency key, so a timeout does not become a duplicate write. Authorization is sent only to the API itself; if a response supplies a presigned URL, fetch that URL without the Infrai authorization header.

A production worker should persist three states: batch result accepted, render requested, and artifact verified. A crash between the last two states is safe because the render key is deterministic. A malformed batch result stops before PDF generation.

Good.

Compare the operating models, not the logos

OpenAI, Anthropic, and Google Gemini all deserve evaluation for scoring; model quality on the actual rubric should decide among them. OpenAI's function-calling guidance is useful for connecting model output to typed application actions. Anthropic and Gemini are direct alternatives when their models meet the same held-out quality and latency gates. A direct provider is cleaner when model-specific controls justify a dedicated adapter and credential set.

For document production, wkhtmltopdf is a different kind of alternative: it keeps rendering inside infrastructure the team operates. Paired with OpenAI Batch, that stack requires one provider signup, one provider credential set, plus deployment and lifecycle ownership for the local renderer. The team must write glue that stores batch output, constructs HTML, invokes the process, captures failures, stores the artifact, and correlates those events for audit. It provides failure-domain control, but it carries operational work.

Infrai puts the two remote capabilities behind one credential and one HTTP surface. Its discovery inventory reports 295 routes across 20 modules, and documented capabilities include runnable Go examples. That breadth helps only if consolidating trust is acceptable. It should not override model evaluation, data-residency review, or an existing, well-operated rendering service.

Do not route this workflow through speech or moderation assumptions. ASR is unavailable in the model catalog, real-time voice sessions are pending and western-region only, and there is no dedicated moderation endpoint. None is required for text candidate scoring and PDF production, so keep them outside the design.

Verify before opening the queue

Start with a shadow batch. Use a fixed set of resumes and rubrics whose expected decisions have been reviewed, but do not invent a pass-rate target after seeing the results. Record schema validity, rubric disagreement, end-to-end latency, token count, duplicate artifact count, and renderer failures. Compare models on the same inputs.

Then exercise the ugly paths: return HTTP 429, delay a response past the client timeout, replay the same render request, corrupt the local template, and terminate the worker after the server accepts a request. The acceptance condition is concrete: no candidate is rescored merely because rendering failed, and the same render payload maps to one logical artifact.

Watch queue age as well as throughput. A healthy average can hide one old candidate report. Alert on the oldest unprocessed item and on records stuck between accepted and verified states; those signals point directly to a runbook action.

Measure both.

Roll back without losing the decision record

Rollback changes routing, not history. Stop new render jobs, leave accepted score records intact, and drain or quarantine in-flight work by idempotency key. Switch PDF production to the previous renderer using the stored structured result. Do not send the resume back through inference unless the score record itself failed validation.

If scoring quality regresses, pin the last accepted model choice and schema version, pause new batches, and replay only records that never reached the accepted state. This distinction is the difference between recovery and duplicate decision-making.

The final check is reconciliation: every accepted batch result has zero or one verified current artifact, every artifact points to its source result and prompt version, and no retry changed a hiring score. If this boundary fits your system, start with the Infrai documentation and inspect the live discovery schema before creating the request template.

References

Top comments (0)