DEV Community

HumphreyFox1243
HumphreyFox1243

Posted on

How to Build 4-Field Structured Summary JSON Output (API Schema Included)

TL;DR: Build a structured summary JSON output API by sending the diff and an explicit four-field schema in one chat-completions request, then validate the response before it reaches your review UI. This gives a media engineering team a title, bullets, key takeaways, and action items without a second extraction service. Pick the model for instruction following first; count input tokens and measure the quality-versus-latency trade-off on your own diffs before standardizing the contract.

Keep the boundary boring. That is the goal.

How should a structured summary JSON API shape its output?

The before state is familiar: a model returns three polished paragraphs, a dashboard tries to split them, and an email template quietly loses the action items. The after state has one request and four named fields: title, bullets, keyTakeaways, and actionItems. Rendering becomes mechanical. So does storage, alerting, and later analysis.

Here is the diagram in words: diff in -> chat request -> parsed contract -> review UI. The validator sits between the model and every downstream consumer. It is a gate, not cleanup.

That distinction matters for code review. A fluent response can still omit a required finding or return an action item as a paragraph. Structured output does not make the review correct, and it does not shrink a large diff. It makes failures visible at a boundary where the caller can retry, reject, or send the change to a human.

Use a narrow contract:

import { z } from "zod";

export const ReviewSummary = z.object({
  title: z.string().min(1),
  bullets: z.array(z.string().min(1)),
  keyTakeaways: z.array(z.string().min(1)),
  actionItems: z.array(z.string().min(1)),
}).strict();

export type ReviewSummary = z.infer<typeof ReviewSummary>;
Enter fullscreen mode Exit fullscreen mode

Four fields are enough for this workflow. Resist adding severity, ownership, file ranges, confidence, and remediation metadata until a real consumer needs them. Every extra field expands the ways a response can be incomplete.

Build one copyable request

The example below uses the OpenAI client against Infrai's compatible surface. It sends one chat request, demands JSON in the prompt, removes Markdown fences if the model adds them, and rejects anything outside the contract. The model ID is explicit so a catalog change cannot silently alter behavior.

import OpenAI from "openai";
import { z } from "zod";

const ReviewSummary = z.object({
  title: z.string().min(1),
  bullets: z.array(z.string().min(1)),
  keyTakeaways: z.array(z.string().min(1)),
  actionItems: z.array(z.string().min(1)),
}).strict();

const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");
const baseURL = process.env.INFRAI_BASE_URL;
if (!baseURL) throw new Error("INFRAI_BASE_URL is required");

const client = new OpenAI({
  apiKey,
  baseURL,
  maxRetries: 3,
});

const diff = process.argv.slice(2).join(" ");
if (!diff) throw new Error("Pass a code diff as the first argument");

const completion = await client.chat.completions.create({
  model: "deepseek-v4-flash",
  messages: [
    {
      role: "system",
      content: [
        "Review a media application's code change.",
        "Return only valid JSON with exactly these keys:",
        "title, bullets, keyTakeaways, actionItems.",
        "Every value except title must be an array of strings.",
        "Report only findings supported by the supplied diff.",
      ].join(" "),
    },
    { role: "user", content: diff },
  ],
});

const content = completion.choices[0]?.message.content;
if (!content) throw new Error("The model returned no summary");

const json = content.replace(/^```
{% endraw %}
(?:json)?\s*|\s*
{% raw %}
```$/g, "");
const summary = ReviewSummary.parse(JSON.parse(json));
process.stdout.write(`${JSON.stringify(summary, null, 2)}\n`);
Enter fullscreen mode Exit fullscreen mode

Install openai and zod, set INFRAI_API_KEY and INFRAI_BASE_URL to the account's API base, and run the file with a TypeScript runner. The SDK sends the bearer credential and handles transient retries, including rate limiting, with bounded retry behavior. The parser still owns the final decision: malformed JSON and wrong shapes fail loudly.

For production, record four observations around this call: model ID, input token count, end-to-end latency, and validation outcome. Do not log the source diff or raw finding text by default; media repositories can contain embargoed titles, access tokens, or unpublished copy. Aggregate the observations into a validation-failure rate and latency percentiles. Alert on a sustained change, not one odd response.

Token counting deserves its own preflight when diffs vary wildly. Structured fields organize the output; they do not reduce the input cost or make an oversized patch easier to reason about. Split large reviews at coherent file or subsystem boundaries, then merge validated findings under the same contract.

Which provider fits this boundary?

There is no universal winner. The useful comparison is operational fit after every candidate passes the same held-out review set.

Option Strong fit Boundary to evaluate
OpenAI Teams that want its first-party structured output and function-calling path Test the exact chosen model against representative diffs and schema changes
Anthropic Claude Teams already building around Claude's tool-use contract Verify how tool validation, retries, and refusal handling map into the application boundary
Google Gemini Teams aligned with Google's model and SDK ecosystem Confirm the selected model's schema behavior and regional deployment requirements
Infrai Teams consolidating backend services behind one key and one bill while retaining an OpenAI-compatible client Choose a model from the live catalog for instruction following; multi-provider routing does not replace evaluation

Run the same fixture suite through all four. Include a clean change, a security-sensitive change, a huge generated diff, and a change with no actionable finding. Score required-field validity separately from review quality. Then plot those results against end-to-end latency.

This is the decision rule: reject any option that misses the contract often enough to disrupt the UI; among the survivors, choose the latency profile that fits the review workflow. An interactive pull-request check and an overnight repository audit can rationally choose different models. Quality versus latency is a product decision, not a leaderboard result.

What if valid JSON contains a bad review?

It will happen. Syntax and semantics are separate checks.

Start with deterministic contract tests, then maintain a small, versioned evaluation set reviewed by engineers who understand the media codebase. Compare findings with the diff, count unsupported claims, and check whether high-impact issues were missed. Keep model, prompt, and schema versions beside every evaluation result. Otherwise a quality change looks like random noise.

Do not auto-apply action items. A structured recommendation is still model output. Put file edits, merge decisions, and security escalations behind explicit authorization.

Retries need a ceiling too. Retry transport failures and rate limits through the SDK, but do not loop forever on a response that repeatedly fails schema validation. One corrective retry with the validation error can be reasonable; after that, return a typed failure and preserve the original change for human review. I favor failing the request over guessing at a missing field because a visible gap is easier to operate than a plausible, corrupted review. Fast failure is observable. Silent coercion is not.

Do I need a separate extraction service?

Usually not for title, bullets, takeaways, and action items. A single schema-guided chat request already produces readable content and machine-usable fields. A second model call adds latency and creates another place for meaning to drift.

Add a separate stage only when it has a distinct job: policy enforcement, deterministic enrichment from an internal database, or normalization across several independent producers. Those are real boundaries. Re-parsing one model's prose with another model is often an avoidable one.

The practical rollout is small: freeze the four-field schema, test candidate models on real but safely handled diffs, instrument validation and latency, and ship behind a review-only path. Expand the contract after consumers prove they need more. This keeps the code review experience predictable while leaving room to trade a little speed for better findings where the change deserves it.

Further reading

Top comments (0)