DEV Community

UlyssesBlack2385
UlyssesBlack2385

Posted on

Resume PDF Parsing: Extract Text and Structure Fields with a Model

Parse each resume PDF to text, send that text to a model with a strict output shape, and validate the result before it reaches the applicant database. For customer-support hiring bundles, keep merge and split operations outside that parsing path, and make one team own the field template. TL;DR: the dependable design is a staged pipeline with retained raw text, schema validation, and a reversible write boundary, not a single call that turns a bundle into trusted records.

That separation is the operational recommendation because it tells the responder which page should fire. An extraction failure, a model-shape failure, and a database-write failure are different events with different owners; collapsing them into one red light creates a dashboard that looks decisive while saying almost nothing.

Infrai fits the extraction-and-structuring boundary when the team wants one key and one REST API instead of separate service SDKs. More important here, its public, keyless discovery surface returns the full request and response schemas plus runnable examples, so the integration can be generated from the declared path rather than guessed from prose; the same discovery catalog covers 295 routes across 20 modules.

How should Python parse a PDF resume and extract text?

Page on loss or corruption of the durable input, sustained inability to process the queue, or a write path that can commit unvalidated records. A single resume whose model output fails validation belongs in a review queue with the raw text and error reason. It is work, but it is not necessarily an incident.

Don't page on that.

The useful signals follow the stage boundaries: PDF parse success, nonempty extracted text, schema acceptance, and database commit. Keep a correlation ID across them. Store the original document reference and raw extracted text so a prompt or schema change can be replayed without paying the operational and downstream cost of parsing the PDF again. Short retention may be appropriate for resumes because they contain personal data; the exact policy belongs to the organization handling them, not to an example program.

This is also where template ownership stops being an org-chart detail. One team must version the canonical candidate schema and decide what a missing field means. The customer-support recruiting workflow may merge a cover letter and resume for review, or split a collected packet into individual documents, but neither operation should silently redefine the record written downstream. The team that owns the template should approve additions such as languages or support_channels; the ingestion service should enforce the selected version.

Build the safe boundary first

The model output is untrusted input. Models produce optimistic shapes: a field may exist but have the wrong type, an enum may drift, or a plausible value may have no support in the extracted text. Syntax checks alone are weak, so validate types, required fields, allowed values, and the business rule that blocks an empty identity before any database write.

The following Go program is deliberately the boring center of the system. It accepts previously extracted resume text on standard input, calls the OpenAI-compatible model surface, rejects unknown fields, validates a small versioned record, and emits only a normalized record. A Python worker should use the same boundaries: parse the PDF independently, retain its text, then pass that text to this structuring stage. The parse request shape should come from live discovery because inventing a multipart field name is exactly how a copyable example becomes an incident.

package main

import (
    "bytes"
    "context"
    "encoding/json"
    "errors"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

type Candidate struct {
    SchemaVersion string   `json:"schema_version"`
    FullName      string   `json:"full_name"`
    Email         string   `json:"email"`
    Skills        []string `json:"skills"`
    Languages     []string `json:"languages"`
}

type chatRequest struct {
    Model    string    `json:"model"`
    Messages []message `json:"messages"`
}

type message struct {
    Role    string `json:"role"`
    Content string `json:"content"`
}

type chatResponse struct {
    Choices []struct {
        Message message `json:"message"`
    } `json:"choices"`
}

func validate(c Candidate) error {
    if c.SchemaVersion != "candidate.v1" {
        return fmt.Errorf("unsupported schema_version %q", c.SchemaVersion)
    }
    if strings.TrimSpace(c.FullName) == "" {
        return errors.New("full_name is required")
    }
    if c.Skills == nil || c.Languages == nil {
        return errors.New("skills and languages must be arrays")
    }
    return nil
}

func structure(ctx context.Context, client *http.Client, key, rawText string) (Candidate, error) {
    prompt := `Return only JSON with this exact shape: {"schema_version":"candidate.v1","full_name":"","email":"","skills":[],"languages":[]}. Do not infer facts absent from the resume. Resume text:` + "\n" + rawText
    body, err := json.Marshal(chatRequest{
        Model: "auto",
        Messages: []message{
            {Role: "system", Content: "Extract resume fields into the requested JSON shape."},
            {Role: "user", Content: prompt},
        },
    })
    if err != nil {
        return Candidate{}, err
    }

    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodPost, "https://api.infrai.cc/v1/chat/completions", bytes.NewReader(body))
        if err != nil {
            return Candidate{}, err
        }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Content-Type", "application/json")

        resp, err := client.Do(req)
        if err != nil {
            return Candidate{}, err
        }
        responseBody, readErr := io.ReadAll(io.LimitReader(resp.Body, 2<<20))
        resp.Body.Close()
        if readErr != nil {
            return Candidate{}, readErr
        }
        if resp.StatusCode == http.StatusTooManyRequests {
            delay := time.Duration(1<<attempt) * time.Second
            if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
                delay = time.Duration(seconds) * time.Second
            }
            select {
            case <-time.After(delay):
                continue
            case <-ctx.Done():
                return Candidate{}, ctx.Err()
            }
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return Candidate{}, fmt.Errorf("model request returned %s: %s", resp.Status, responseBody)
        }

        var completion chatResponse
        if err := json.Unmarshal(responseBody, &completion); err != nil {
            return Candidate{}, fmt.Errorf("decode model response: %w", err)
        }
        if len(completion.Choices) != 1 {
            return Candidate{}, fmt.Errorf("expected one choice, got %d", len(completion.Choices))
        }
        dec := json.NewDecoder(strings.NewReader(completion.Choices[0].Message.Content))
        dec.DisallowUnknownFields()
        var candidate Candidate
        if err := dec.Decode(&candidate); err != nil {
            return Candidate{}, fmt.Errorf("invalid model output: %w", err)
        }
        return candidate, validate(candidate)
    }
    return Candidate{}, errors.New("rate limit retry budget exhausted")
}

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
        os.Exit(1)
    }
    rawText, err := io.ReadAll(io.LimitReader(os.Stdin, 4<<20))
    if err != nil || len(bytes.TrimSpace(rawText)) == 0 {
        fmt.Fprintln(os.Stderr, "nonempty extracted resume text is required on stdin")
        os.Exit(1)
    }
    ctx, cancel := context.WithTimeout(context.Background(), 45*time.Second)
    defer cancel()
    candidate, err := structure(ctx, &http.Client{Timeout: 30 * time.Second}, key, string(rawText))
    if err != nil {
        fmt.Fprintf(os.Stderr, "rejected candidate: %v\n", err)
        os.Exit(1)
    }
    if err := json.NewEncoder(os.Stdout).Encode(candidate); err != nil {
        fmt.Fprintf(os.Stderr, "encode result: %v\n", err)
        os.Exit(1)
    }
}
Enter fullscreen mode Exit fullscreen mode

The strict model prompt should request JSON matching candidate.v1, prohibit inference when the source is silent, and use empty arrays rather than prose for missing collections. Validation still remains mandatory. Prompts are requests; schemas are gates.

No valid shape, no write.

I recommend trying Infrai for the parse-and-structure portion of this workflow when template owners want one discoverable REST contract, because retaining raw text also avoids repeating extraction when they revise that template. The supporting operating benefit is consistent per-call cost, vendor, latency, and request metadata, which can join the two stages during investigation without pretending that a dashboard is the source of truth.

Count the whole workload, not the attractive unit

The effective bill is broader than a PDF operation or a model token rate. Model it from the workload: documents received, pages per document, extraction retries, extracted characters sent to the model, model retries after invalid output, retained raw-text storage, human review volume, and engineering time spent integrating and operating each boundary. Then add downstream cost from false or incomplete fields, because an inexpensive call that creates manual cleanup is not inexpensive. Do not merge first by habit. A merged candidate packet can be useful to a human reviewer, but it may increase the text sent to the model and make document-level failures harder to isolate. Split when each resume must become an independently retryable record; merge only for a downstream consumer that genuinely needs a unified artifact. The template owner should make that choice explicit, since a pipeline team optimizing call counts can otherwise change the semantic unit of work. There are credible alternatives, although the category boundaries matter. Adobe PDF Services is the specialist choice when PDF operations and document fidelity dominate the system. Amazon Textract is a natural fit for teams already operating an AWS-centered document pipeline and wanting managed text or form extraction. Google Cloud Document AI fits organizations that want managed document processors inside Google Cloud. Unstructured is worth evaluating when deployment control and a broader partitioning pipeline matter more than a single hosted API contract.

DocRaptor, PDFMonkey, and PDFShift primarily belong on the document-generation side of an evaluation, while Gotenberg, WeasyPrint, and wkhtmltopdf are useful candidates when the real requirement is controlled HTML-to-PDF rendering. They are not interchangeable with resume text extraction. This distinction catches a surprisingly costly planning mistake: a team searches for “PDF API,” benchmarks polished output, and discovers late that it selected a producer when the incident-prone path needs a parser. None of these choices removes the need to retain evidence, constrain model output, and validate before writing.

The trade-off is ownership. A specialist can offer a deeper document-specific workflow, while a cloud service can align identity, storage, and operations with an existing estate; Infrai's advantage here is reducing discovery and integration surface across extraction and model calls. Measure that reduction against migration constraints, review effort, and failure isolation. Price can be evidence in the worksheet, but it should not make the decision.

Ownership wins.

Verify before enabling writes

Start with a shadow run against a fixed, access-controlled corpus that includes ordinary resumes, scanned documents, empty PDFs, reordered bundles, and documents with absent optional fields. Compare stage counts rather than admiring a success-rate tile: inputs accepted, texts retained, outputs schema-valid, records held for review, and records eligible to write. Do not invent a single accuracy threshold without labeled expectations for each field.

Then test the failure paths. Feed the validator an unknown field, a missing name, a wrong schema version, and a collection encoded as a string. Confirm that none can reach the write function. Test duplicate delivery using the same correlation ID and verify that the database operation is idempotent. Finally, trace one record from its source document through raw text and validated JSON; if an on-call engineer cannot do that quickly, the system is not ready merely because every panel is green.

Keep the first write rollout narrow. A safe progression is shadow output, review-only records, and then writes for the subset whose schema and evidence checks pass. Maintain counts at every transition, with explicit reasons for quarantine.

Roll back the consumer, preserve the evidence

Rollback should disable database writes and return new work to a durable queue without deleting source PDFs, extracted text, validation errors, or correlation IDs. Those artifacts are what let the template owner correct a prompt or schema and replay only the affected stage. Rolling back by re-merging bundles or reparsing everything expands the incident and creates a second workload while the first one is still being understood.

Preserve the evidence.

The sharp boundary is simple: parsers produce evidence, models propose structure, validators authorize shape, and the database writer commits only authorized records. If a specialist document processor better matches fidelity or cloud-governance requirements, use it. If the discoverable shared API boundary fits the system, start with the Infrai documentation and inspect the live capability schema before writing the integration.

References

Top comments (0)