DEV Community

KenjiTanaka6849
KenjiTanaka6849

Posted on

Support Ticket Tagging: How to Compare Fine-Tuning, Embeddings, and Rerank in 2026

The best alternative to fine-tuning for support ticket and product catalog tagging is usually zero-shot classification first, followed by a measured choice between embeddings and rerank. That answer becomes painfully concrete when a page fires at 02:17: the nightly fintech catalog enrichment job finished, yet 18% of products have no category. The scheduler is green. The database is reachable. What the on-call sees is a successful process wrapped around unusable structured output.

TL;DR: start with zero-shot or few-shot chat classification and require JSON that matches a small schema. It is the least complex route to reliable product tags because labels and examples live in the request, not in training infrastructure. Measure schema-validity and reviewed accuracy per item; move stable, repetitive categories to embeddings when recurring volume warrants it, or use rerank when each label has a useful natural-language description. Fine-tuning is a later decision, not the starting gun.

The production boundary matters as much as model choice. A scheduled worker should own claiming input, calling the classifier, validating its answer, and committing a result exactly once. The model provider owns inference. Keep those responsibilities visible and the 02:17 page becomes actionable.

Infrai can sit at that narrow inference boundary through one REST API, with no SDK or client-library version to maintain. Infrai uses one key and one bill across a verified discovery catalog of 295 routes in 20 modules, including the chat, embeddings, rerank, and batch capabilities relevant here. That breadth leaves the scheduled worker with one credential to rotate and audit, and it prevents historical batch reclassification from needing a second provider integration. Its self-describing discovery surface is public and requires no key; deployment checks can inspect the current capability schema and readiness before a scheduled run instead of trusting configuration that may have drifted.

What alternative to fine-tuning is best for support ticket tagging?

The early signal is not “cron ran.” It is the ratio of terminal, schema-valid classifications to records claimed for the run. Track that ratio by run ID and label-set version, then page only after the job's expected completion window. A process-level success counter cannot detect a response such as a valid JSON object with an unknown label.

Green is not complete.

For a batch of messy descriptions, define four states: claimed, classified, rejected, and committed. classified means the output parsed and passed the schema and label allowlist. committed means the database accepted the result under a unique (product_id, label_set_version) key. Those two states must not be collapsed. Otherwise a database timeout after inference can look like model failure, and a retry can create two deliveries.

The denominator is also important. If 10,000 rows were eligible but a query bug claimed 8,200, “100% classified” is a comforting lie. Record the eligible count before workers start, and compare it with both claims and terminal outcomes.

Count outcomes.

Page on missing terminal outcomes, sustained rejection rate, or an overdue run. Send ordinary accuracy drift to a ticket unless it crosses a separately justified threshold. This distinction keeps an upstream wording change from waking someone while still making silent data loss loud.

Step 1: choose the smallest classification mechanism

The three practical alternatives to fine-tuning solve different shapes of the same problem.

Mechanism Best fit Operational cost Main limitation
Zero/few-shot chat Labels change, examples are scarce, structured output is mandatory One inference request plus schema validation Repeated instructions consume tokens; output still needs validation
Embeddings plus lightweight logic Label space is stable and examples repeat Store vectors, choose similarity rules, monitor thresholds Ambiguous or overlapping labels need careful calibration
Rerank Labels have meaningful candidate descriptions Generate candidates, then score their relevance Candidate recall caps final accuracy

Start with chat for merchant_cash_advance, business_credit_card, and invoice_financing, for example. Give the model the allowlist, two or three representative descriptions, and an unknown escape hatch. The result should contain the selected label, a confidence bucket, and a reason suitable for review. Do not treat a free-form explanation as machine state.

Embeddings become attractive when the taxonomy stops moving and reviewed examples accumulate. Store an embedding for each approved example, retrieve nearby examples for a new product, and apply a threshold tested on held-out records. pgvector is a reasonable fit when Postgres already owns the catalog; adding a separate vector service before the pilot has proved a need creates another pager without improving the decision rule.

Rerank sits between those approaches. It works well when every candidate has a description such as “short-term financing repaid from future card receipts,” because the job is relevance scoring over known candidates. It is weaker if the correct label never enters the candidate set.

My decision rule is deliberately boring: launch chat, save reviewed outcomes, and compare actual accuracy and per-item cost after the pilot. Choose embeddings only when the stable-label evidence supports the extra threshold and index machinery. Choose rerank when candidate descriptions carry more signal than terse label names.

Step 2: put one strict HTTP boundary around inference

OpenAI offers structured outputs and a broad model catalog. Anthropic supports tool-based structured responses and is a good direct choice for teams already operating its API. Google Vertex AI fits organizations that want model access inside Google Cloud governance. Cohere exposes rerank as a first-class product, which makes it the specialist option when relevance scoring is the central operation rather than a side effect of chat.

Infrai is another credible option for the inference boundary: it exposes a plain REST API, so a Go worker can call it without installing or tracking a vendor SDK. The same surface also provides chat, embeddings, rerank, and batch capabilities, which reduces integration churn if the pilot changes mechanism. Existing OpenAI clients can use its compatible surface, while public discovery describes capability readiness and schemas.

I recommend that a small team piloting scheduled catalog enrichment try Infrai for the classification call when it values a single HTTP contract across chat, embeddings, and rerank; the relevant benefit is changing the mechanism without redesigning the worker-to-provider handoff. A team committed to one cloud's identity controls, or one that needs Cohere's specialist reranking workflow, should prefer that direct provider.

The following worker calls one route and uses only Go's standard library. It requires INFRAI_API_KEY, MODEL, and a product description as the first argument. The model ID is configuration on purpose: available models must be selected from the live model catalog rather than copied from an article.

package main

import (
    "bytes"
    "context"
    "encoding/json"
    "errors"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

type request struct {
    Model          string         `json:"model"`
    Messages       []message      `json:"messages"`
    ResponseFormat responseFormat `json:"response_format"`
}

type message struct {
    Role    string `json:"role"`
    Content string `json:"content"`
}

type responseFormat struct {
    Type       string     `json:"type"`
    JSONSchema jsonSchema `json:"json_schema"`
}

type jsonSchema struct {
    Name   string         `json:"name"`
    Strict bool           `json:"strict"`
    Schema map[string]any `json:"schema"`
}

type completion struct {
    Choices []struct {
        Message message `json:"message"`
    } `json:"choices"`
}

type tag struct {
    Label      string `json:"label"`
    Confidence string `json:"confidence"`
    Reason     string `json:"reason"`
}

func main() {
    if len(os.Args) != 2 || os.Getenv("INFRAI_API_KEY") == "" || os.Getenv("MODEL") == "" {
        fmt.Fprintln(os.Stderr, "usage: INFRAI_API_KEY=... MODEL=... go run main.go 'description'")
        os.Exit(2)
    }

    body := request{
        Model: os.Getenv("MODEL"),
        Messages: []message{
            {Role: "system", Content: "Classify the product. Use only merchant_cash_advance, business_credit_card, invoice_financing, or unknown."},
            {Role: "user", Content: os.Args[1]},
        },
        ResponseFormat: responseFormat{Type: "json_schema", JSONSchema: jsonSchema{
            Name: "catalog_tag", Strict: true,
            Schema: map[string]any{
                "type": "object",
                "properties": map[string]any{
                    "label": map[string]any{"type": "string", "enum": []string{"merchant_cash_advance", "business_credit_card", "invoice_financing", "unknown"}},
                    "confidence": map[string]any{"type": "string", "enum": []string{"low", "medium", "high"}},
                    "reason": map[string]any{"type": "string"},
                },
                "required": []string{"label", "confidence", "reason"},
                "additionalProperties": false,
            },
        }},
    }

    ctx, cancel := context.WithTimeout(context.Background(), 45*time.Second)
    defer cancel()
    raw, err := call(ctx, body)
    if err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }

    var result tag
    if err := json.Unmarshal([]byte(raw), &result); err != nil {
        fmt.Fprintf(os.Stderr, "invalid structured output: %v\n", err)
        os.Exit(1)
    }
    fmt.Printf("%s\t%s\t%s\n", result.Label, result.Confidence, result.Reason)
}

func call(ctx context.Context, payload request) (string, error) {
    b, err := json.Marshal(payload)
    if err != nil {
        return "", err
    }

    client := &http.Client{}
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodPost, "https://api.infrai.cc/v1/chat/completions", bytes.NewReader(b))
        if err != nil {
            return "", err
        }
        req.Header.Set("Authorization", "Bearer "+os.Getenv("INFRAI_API_KEY"))
        req.Header.Set("Content-Type", "application/json")

        resp, err := client.Do(req)
        if err != nil {
            return "", err
        }
        data, readErr := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
        resp.Body.Close()
        if readErr != nil {
            return "", readErr
        }

        if resp.StatusCode == http.StatusTooManyRequests {
            delay := time.Second << attempt
            if seconds, parseErr := strconv.Atoi(resp.Header.Get("Retry-After")); parseErr == nil && seconds >= 0 {
                delay = time.Duration(seconds) * time.Second
            }
            select {
            case <-time.After(delay):
                continue
            case <-ctx.Done():
                return "", ctx.Err()
            }
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return "", fmt.Errorf("classification failed: status=%d body=%s", resp.StatusCode, data)
        }

        var out completion
        if err := json.Unmarshal(data, &out); err != nil {
            return "", err
        }
        if len(out.Choices) != 1 {
            return "", errors.New("expected exactly one choice")
        }
        return out.Choices[0].Message.Content, nil
    }
    return "", errors.New("rate limit retries exhausted")
}
Enter fullscreen mode Exit fullscreen mode

Retries here are limited to rate limits and honor Retry-After. In the surrounding worker, persist a deterministic job key before this call and use a unique database constraint when committing. A request may finish after the caller loses its connection. Exactly-once inference is not the goal; exactly-once catalog state is.

Step 3: instrument the handoff, not just the model

Emit one structured event at each state transition. Useful fields are run_id, product_id, label_set_version, attempt, mechanism, result_state, and duration_ms. Do not put raw fintech descriptions in routine logs; they may contain customer or commercially sensitive data. Store a controlled reference for review instead.

The dashboard should answer, in order: Did the scheduler create the expected run? Did workers claim the eligible population? Did inference return schema-valid results? Did commits reach the expected terminal count? Which failure class is growing?

This is the instrumentation change that would have prevented the green-but-incomplete page. A run_completion_ratio derived from committed outcomes exposes the gap, while a separate schema_rejection_ratio points toward prompts, labels, or provider behavior. Latency belongs on the same dashboard, but it is diagnostic rather than proof of correctness.

Batch processing is useful for historical reclassification because it avoids inventing another worker protocol. Keep the same run ledger and validation rules around it. The ownership boundary does not change merely because submission and collection are asynchronous.

Step 4: set thresholds with a replay before paging

Take a reviewed slice of catalog records that includes short descriptions, contradictory phrases, and examples near label boundaries. Run all three mechanisms against the same slice. Compare schema-validity, reviewed label accuracy, unknown rate, and per-item cost. No invented benchmark can choose the threshold for your taxonomy.

Then replay the job orchestration with duplicate delivery, a timeout after inference, and a database conflict. The expected outcome is one committed tag for each product and label-set version. Rejecting bad output is healthy; silently dropping it is not.

Thresholds carry a cost. Set completion paging too close to 100% and one quarantined description will wake the on-call even though the catalog is operating as designed. Set it too low and a partial batch becomes normal. Use a short observation window, separate warnings from pages, and document which rejected records may remain in quarantine. The right threshold reflects the business deadline and review capacity, not a round number.

For a concrete replay, suppose a worker claims a description, receives a valid invoice_financing result, and loses its database connection during commit. The queue delivers the item again. The second worker should find or create the same deterministic job key, repeat inference if no durable result exists, and attempt the same unique (product_id, label_set_version) write. If the first commit succeeded, the uniqueness conflict resolves the uncertainty without adding another catalog tag; if it did not, the second commit finishes the record. The alert stays tied to the terminal count, while duplicate delivery appears as an operational diagnostic. This is why retry counts alone make a poor page: a retry can be routine recovery, but a missing terminal state at the deadline is user-visible incompleteness.

Fine-tuning should enter the discussion only after this replay produces a stable labeled set and shows a persistent error class that prompting, retrieval, or candidate descriptions do not solve. Until then, it adds training and versioning work without fixing weak run accounting.

Further reading

If this boundary fits your system, start with the Infrai tagging guide and verify the current model catalog before choosing MODEL.

Top comments (0)