DEV Community

Cover image for Reflections on the Relationship Between Agent Application Development and Rendering
Tidiane Stano
Tidiane Stano

Posted on

Reflections on the Relationship Between Agent Application Development and Rendering

Introduction

When developers transition from traditional frontend and backend development to building AI Agent applications, they frequently encounter confusion around how generated data pairs with custom UI rendering. This article dissects the core mental model for Agent development: leveraging Large Language Models (LLMs) to drive engineering workflows, and clarifies how LLM capabilities, client-server communication protocols, data generation and UI rendering connect with one another.

The core idea of Agent application development lies in using LLMs to drive engineering processes. The critical challenge becomes building a working rendering pipeline controlled by LLM outputs. This article breaks the problem down into three major sections: understanding LLM capabilities, defining client-server protocol boundaries, and implementing data generation alongside rendering logic. It includes practical code examples and standard patterns that developers can directly adopt in production projects.

1. Understanding LLM Capabilities and Core Functions

Comprehension, Reasoning and Generation

At its most basic level, an LLM is capable of three core operations: understanding input content, performing reasoning, and generating natural language text. A simple API call demonstrates this behavior.

llm.invoke("Hello")
Enter fullscreen mode Exit fullscreen mode

The model returns a natural language reply. Under the hood, this execution includes three sequential steps:

  1. Comprehend the semantics of the user input message.
  2. Execute reasoning to determine an appropriate response.
  3. Generate continuous text output as the final reply.

This fundamental capability enables LLMs to handle conversational interaction, summarization and content creation. For Agent systems, the more important extension is tool invocation via Function Call.

Function Call Mechanism

Function Call is not direct function execution inside the LLM. The LLM is essentially a text completion engine. Through training, the model learns to produce structured JSON text to express intent for invoking external tools.

Before Function Call was widely supported, developers relied heavily on prompt engineering. Engineers would craft detailed instructions to force the model to output content in fixed formats, then parse results with regular expressions. This method was fragile. Minor deviations in model output could break parameter parsing entirely. Some early constraint frameworks such as Guidance were used to mitigate this instability.

Modern LLMs, including DeepSeek and many other mainstream models, natively support Function Call. This capability expands LLM boundaries from plain text processing to structured tool orchestration. The model can describe tools, define parameters and trigger external logic. This is the foundation that makes practical AI Agent workflows possible.

2. Client and Server Protocol Design

SSE (Server-Sent Events)

Server-Sent Events, standardized by W3C, is a one-way real-time push protocol built on HTTP. It uses the text/event-stream content type for long-lived connections. Servers continuously push data chunks to clients over this persistent channel. Common use cases include real-time notifications, progress streaming and streamed AI model output.

SSE defines a fixed message format for event streaming.

event: progress
id: 12
retry: 3000
data: {"percent":60}
Enter fullscreen mode Exit fullscreen mode
  • event: Identifies the type of event message.
  • id: Event identifier for resuming interrupted streams.
  • retry: Time interval for reconnection attempts in milliseconds.
  • data: The payload body of the event. Blank lines separate individual event entries.

Clients must include the Accept: text/event-stream HTTP header to negotiate this stream mode. Developers can also design custom binary or JSON streaming formats instead of using the standard SSE specification.

SSE fits Agent applications naturally. Agent workflows run on the server, and LLM inference consumes significant time. Streaming incremental output to the client delivers a responsive user experience.

The native browser EventSource API has notable limitations.

  1. Native EventSource only supports GET requests. It cannot attach request bodies, so developers cannot send message payloads to the server.
  2. It strictly enforces one-way communication. The standard usage pattern where users submit messages and receive continuous downstream output cannot be implemented directly with native EventSource.

The standard workaround uses POST requests with the Accept: text/event-stream header. POST requests can carry JSON payloads, files and image content. The server detects the header and responds with streamed SSE content.

const response = await fetch("/api/chat", {
  method: "POST",
  headers: {
    "Content-Type": "application/json",
    "Accept": "text/event-stream",
  },
  body: JSON.stringify({ message: "Hello" }),
});
// Read bytes incrementally from response.body
// Decode UTF-8, parse SSE events and feed to UI rendering
Enter fullscreen mode Exit fullscreen mode

Microsoft’s @microsoft/fetch-event-source open-source library solves these limitations. It uses fetch under the hood and supports POST-based SSE streaming. The following example shows a reusable message sending function with custom rendering callbacks.

import { fetchEventSource } from "@microsoft/fetch-event-source";

async function sendMessage(message, render) {
  const controller = new AbortController();
  let answer = "";
  let completed = false;

  await fetchEventSource("/api/chat", {
    method: "POST",
    headers: {
      "Content-Type": "application/json",
    },
    body: JSON.stringify({ message }),
    signal: controller.signal,
    openWhenHidden: true,
    async onopen(response) {
      const contentType = response.headers.get("content-type") ?? "";
      if (!response.ok || !contentType.includes("text/event-stream")) {
        throw new Error(`HTTP ${response.status}: stream initialization failed`);
      }
    },
    onmessage(event) {
      const data = JSON.parse(event.data);
      if (event.event === "delta") {
        answer += data.text;
        render(answer);
      }
      if (event.event === "done") {
        completed = true;
        controller.abort();
      }
    },
    onclose() {
      if (!completed) {
        throw new Error("Stream terminated prematurely");
      }
    },
    onerror(error) {
      throw error;
    },
  });
}

await sendMessage("Analyze this report", (text) => {
  document.querySelector("#answer").textContent = text;
});
Enter fullscreen mode Exit fullscreen mode

The core idea of SSE can be extended to custom streaming protocols. Instead of standard SSE, developers can implement NDJSON streams. Each line contains one serialized JSON object.

{"type":"progress","percent":60}
{"type":"completed","result":"Analysis finished"}
Enter fullscreen mode Exit fullscreen mode

This custom line-delimited JSON approach works well for Agent applications that need richer structured event types beyond simple text chunks.

3. Data Generation and UI Rendering

For LLMs and backend services to drive frontend components, the system must describe available UI components. Metadata defines component use cases, accepted parameters, supported events and user interaction callbacks.

A simplified React approval component demonstrates this pattern.

interface ApprovalProps {
  requestId: string;
  title: string;
  impact: string;
  onResolve: (decision: "approved" | "rejected") => void;
}

export function Approval({ title, impact, onResolve }: ApprovalProps) {
  return (
    <section>
      <h3>{title}</h3>
      <p>{impact}</p>
      <button onClick={() => onResolve("approved")}>Approve</button>
      <button onClick={() => onResolve("rejected")}>Reject</button>
    </section>
  );
}
Enter fullscreen mode Exit fullscreen mode

The TypeScript interface defines data structure and callback signatures. However, type definitions alone are insufficient. The LLM and backend cannot understand the scenario this component applies to or what events to trigger. Developers must attach component metadata to describe purpose, parameters and event mapping.

The approvalMeta object below provides machine-readable metadata for this UI component.

export const approvalMeta = {
  name: "Approval",
  description: "Display an operation and its impact, request user approval or rejection.",
  usewhen: "The workflow requires user confirmation before continuing.",
  events: {
    onResolve: {
      type: "approval.resolve",
      targetArgs: ["requestId"],
      payloadFromArgs: ["decision"],
    },
  },
} as const;
Enter fullscreen mode Exit fullscreen mode

This metadata explains that calling onResolve(decision) emits an approval.resolve event. It carries the requestId identifier and decision payload.

Developers maintain component types, descriptions and event mappings. Build tools combine TypeScript type information and metadata definitions to generate standardized machine-readable component specifications. The LLM reads these specifications and decides when to render which UI component inside the Agent workflow.

This separation creates a clean boundary:

  1. Backend / LLM: Generate structured data and event instructions.
  2. Client: Receive structured payloads, render matching UI components, capture user interactions and send events back to server.

The LLM does not directly manipulate DOM elements. It only outputs declarative instructions for UI rendering. The client interprets these instructions and manages native component rendering. This design decouples Agent logic from frontend implementation details.

4. System Architecture and Agent Workflow Summary

The full Agent rendering workflow can be summarized in a sequential pipeline:

  1. User sends a request from frontend client via POST streaming request.
  2. Server receives input, invokes LLM reasoning.
  3. LLM decides whether to generate plain text or emit component rendering instructions through Function Call.
  4. Server pushes structured events to client over SSE or custom stream protocol.
  5. Client parses incoming events. If the payload describes a UI component, the client instantiates and renders the matching component.
  6. User interacts with rendered UI, triggering component callbacks.
  7. Client sends interaction events back to server.
  8. Server resumes Agent reasoning using the user interaction result.

This architecture keeps rendering logic fully contained on the frontend. The LLM remains responsible only for high-level workflow decisions and data generation. This separation avoids tight coupling between model outputs and UI frameworks.

In distributed Agent systems, developers often manage multi-model traffic routing through an API gateway. 4sapi provides unified endpoint management that simplifies orchestration when multiple LLMs serve Agent workloads.

5. Common Pitfalls and Engineering Best Practices

Protocol Selection Tradeoffs

SSE is ideal for server-to-client streaming, but it cannot handle bidirectional real-time communication natively. If an Agent needs frequent bidirectional message exchange, WebSocket may be more suitable. For most text and UI streaming scenarios, POST-based SSE is simpler to implement and easier to debug.

Component Metadata Maintenance

Metadata must stay lightweight and machine parseable. Avoid embedding long natural language descriptions. Every component should explicitly define trigger conditions, input schemas and output event formats. Outdated metadata causes the LLM to select incorrect components, breaking rendering behavior.

Separation of Concerns

Do not let LLMs directly output raw HTML or JSX. Direct code generation creates severe security risks and tight coupling. The preferred pattern is declarative rendering instructions. The LLM outputs component identifiers and parameters. The client renders pre-built registered components.

Stream Error Handling

Streaming connections can drop unexpectedly. Client code must implement reconnection logic, abort controllers and partial state recovery. Each event should carry an identifier so clients can resume interrupted streams without losing state.

Conclusion

Building Agent applications requires rethinking the traditional frontend-backend data flow. The core principle is a clear separation: LLMs handle reasoning and structured data generation, server protocols deliver streamed payloads, and the client consumes instructions to render predefined UI components.

The Function Call capability enables LLMs to declare intent for UI components. SSE and custom streaming protocols reliably transmit incremental data from server to browser. Component metadata bridges the gap between LLM reasoning and frontend rendering. This mental model helps developers avoid common mistakes such as trying to make LLMs directly generate UI markup.

This pattern supports scalable Agent systems. New UI components can be added independently, and the LLM can use them once their metadata is registered. When running multi-model Agent deployments, unified gateway services streamline endpoint management and request routing.

International access: https://4sapi.com
Domestic access: https://4sapi.cn

Top comments (0)