DEV Community

Cover image for LLM Application Development
Rhuturaj Takle
Rhuturaj Takle

Posted on

LLM Application Development

LLM Application Development

A deep-dive walkthrough of building applications on top of large language models — covering how LLMs actually work and what that implies for application design, prompt engineering, the system/user/assistant message roles, few-shot and structured prompting, context windows and tokens, why hallucinations happen and how to reduce them, how to choose a model, conversation memory strategies, and AI agents: what they are, how the loop works, and how to keep them safe and bounded.


Table of Contents

  1. Introduction
  2. LLM Fundamentals
  3. Prompt Engineering
  4. System, User, and Assistant Messages
  5. Few-Shot Prompting
  6. Structured Prompting
  7. Context Windows
  8. Tokens
  9. Hallucinations
  10. Model Selection
  11. Conversation Memory
  12. AI Agents
  13. Evaluating LLM Applications
  14. Putting It Together: Anatomy of an LLM Application
  15. Common Pitfalls
  16. Quick Reference Table
  17. Conclusion

Introduction

Building with LLMs is a different kind of engineering from traditional software. A conventional function is deterministic: the same input gives the same output, and a bug is a defect you can locate and fix. An LLM is probabilistic: it produces plausible text, usually correct, sometimes confidently wrong, and the "program" you write is largely natural language. Most of the craft of LLM application development is learning to build reliable systems out of an unreliable-by-nature component — by giving the model the right context, constraining its output, validating what comes back, measuring quality, and bounding what it can do.

A typical LLM application is NOT "call the model, show the answer." It is:

  User input
    -> validate and limit it
    -> retrieve relevant context (documents, memory, data)
    -> assemble a prompt (instructions + context + conversation + question)
    -> call the model (possibly in a loop with tools)
    -> validate / parse the output
    -> log, evaluate, and return the result
Enter fullscreen mode Exit fullscreen mode

This guide is the conceptual companion to this series' AI APIs with .NET guide. That guide covers the mechanics of calling the API (clients, streaming, tool-calling code, retries, rate limits, cost); this one covers how to design the application around the model. Where the mechanics matter, it points back there rather than repeating them.


1. LLM Fundamentals

An LLM predicts the next token — everything else follows from that

A large language model is a neural network trained on enormous amounts of text to do one thing: given a sequence of tokens, predict a probability distribution over the next token. It generates a response by repeatedly predicting a token, appending it, and predicting again.

Input:   "The capital of France is"
Model:   P(" Paris") = high,  P(" Lyon") = low, ...
Output:  " Paris"  ->  appended  ->  predict next token  ->  "." ...
Enter fullscreen mode Exit fullscreen mode

That simple mechanism produces surprisingly capable behavior — summarizing, translating, writing code, reasoning through problems — because predicting text well across the breadth of human writing requires modeling a great deal about the world. But it also explains the characteristic limitations, which you must design around:

- It generates PLAUSIBLE text, not VERIFIED text. Fluency is not accuracy.
  (Section 8: hallucinations)
- It is STATELESS. It has no memory between calls beyond what you send it.
  (Section 10: conversation memory)
- It has a KNOWLEDGE CUTOFF. It knows nothing about events after training,
  and nothing about your private data unless you supply it.
- It is NON-DETERMINISTIC by default. Sampling introduces variation.
- It has a bounded CONTEXT WINDOW. It can only consider so much text at once.
  (Section 6)
Enter fullscreen mode Exit fullscreen mode

How models get their behavior

1. PRE-TRAINING   Learn language and broad knowledge by predicting next tokens over a
                  huge text corpus. Produces a "base" model — capable but not
                  conversational or reliably helpful.

2. FINE-TUNING    Further train on curated examples (instruction/response pairs) so the
   (instruction   model follows instructions and answers in a useful format.
   tuning)

3. ALIGNMENT      Shape behavior toward helpful, honest, and safe responses, often using
   (e.g. RLHF)    human feedback or preference data as a training signal (this series'
                  AI/ML Fundamentals guide covers reinforcement learning).

(Newer "reasoning" models add training to produce extended internal reasoning
 before answering, trading latency and cost for accuracy on hard problems.)
Enter fullscreen mode Exit fullscreen mode

What this means for application builders

LLMs are strong at:  language tasks — summarizing, rewriting, extracting, classifying,
                     translating, drafting, explaining, and generating code — and at
                     handling messy, unstructured input.

LLMs are weak at:    exact arithmetic and counting, guaranteed factual recall, knowing
                     what they don't know, acting on current or private data without
                     being given it, and being perfectly repeatable.

Design rule:         Use the model for what it's good at (language, judgment, flexibility)
                     and use ORDINARY CODE for what code is good at (math, lookups,
                     validation, business rules, security). The best LLM applications
                     are hybrids.
Enter fullscreen mode Exit fullscreen mode

Embeddings: the other half of the toolkit

Alongside generation, models can produce embeddings — vectors where similar meanings sit close together (covered in this series' AI/ML Fundamentals guide). They power semantic search and retrieval-augmented generation (Sections 8, 10, 13), letting you find relevant text by meaning and hand it to the model as context.


2. Prompt Engineering

The prompt is your program — treat it with the same care as code

A prompt is the input you give the model, and small changes in wording can produce large changes in output. Prompt engineering is the practice of designing prompts that produce reliably good results.

A well-formed prompt has distinct parts

1. ROLE / CONTEXT   Who the model is acting as and the situation it's in.
2. TASK             Exactly what to do — specific and unambiguous.
3. INPUT DATA       The material to work on (clearly delimited from the instructions).
4. CONSTRAINTS      Rules: scope, tone, length, what NOT to do, what to do when unsure.
5. OUTPUT FORMAT    The exact shape of the response (a label, bullets, JSON schema...).
6. EXAMPLES         Optional demonstrations of good input/output pairs (Section 4).
Enter fullscreen mode Exit fullscreen mode

Principles that consistently help

BE SPECIFIC.        Vague in, vague out. "Summarize this" gets a generic summary;
                    "Summarize this support ticket in two sentences for an engineer,
                    focusing on the reproduction steps" gets something usable.

GIVE CONTEXT.       The model can't read your mind or your database. Say who the
                    audience is, what the goal is, and supply the facts it needs.

STATE THE FORMAT.   Say exactly what shape you want back, and why, if it helps.

SAY WHAT TO DO,     Positive instructions ("answer in plain language") generally work
NOT ONLY WHAT       better than a list of prohibitions, though explicit "do not"
NOT TO DO.          rules are still useful for hard constraints.

GIVE AN ESCAPE      Tell the model what to do when it can't comply: "If the answer
HATCH.              is not in the provided text, reply exactly: NOT_FOUND."
                    This is one of the most effective anti-hallucination measures.

BREAK DOWN HARD     A complex task often works better as several smaller prompts
TASKS.              (extract, then analyze, then format) than one giant instruction.

ASK FOR REASONING   For problems needing multi-step logic, asking a standard model
WHERE IT HELPS.     to work through the steps before giving a final answer often
                    improves accuracy. (Dedicated reasoning models do this internally,
                    and usually need less coaxing — follow your model's guidance.)
Enter fullscreen mode Exit fullscreen mode

A weak prompt versus a strong one

WEAK:
  Summarize this ticket.

STRONG:
  You are a support triage assistant.
  Summarize the ticket below in at most two sentences for an on-call engineer.
  Include: the affected feature, the observed error, and any reproduction steps.
  If a required detail is missing, write "unknown" for it. Do not speculate.

  <ticket>
  {ticket text}
  </ticket>
Enter fullscreen mode Exit fullscreen mode

Prompts are code: version, test, and review them

- Keep prompts in source control (files or resources), not scattered as inline strings.
- Use TEMPLATES with named placeholders rather than ad hoc string concatenation.
- Change ONE thing at a time, and measure the effect on a test set (Section 12) —
  not on one example you happen to like.
- A model upgrade can change how a prompt behaves. Re-run your evaluations when you
  switch models or versions.
- Treat prompt changes like code changes: review them, and be able to roll back.
Enter fullscreen mode Exit fullscreen mode

3. System, User, and Assistant Messages

Chat models take a structured conversation, not a single string

system     -> the developer's standing instructions: role, rules, tone, format, safety
              boundaries. Sets the frame for the whole conversation.
user       -> what the end user (or your application, on their behalf) is asking
assistant  -> the model's own earlier replies, which you replay as history
tool       -> results returned from functions the model asked you to run
Enter fullscreen mode Exit fullscreen mode

(The mechanics of building these message lists in .NET are in this series' AI APIs with .NET guide, Section 4.)

What belongs in the system message

System message  -> STABLE rules that apply to every turn:
                     identity ("You are the support assistant for ACME"),
                     scope ("Only answer questions about ACME products"),
                     tone, output format, what to do when unsure or out of scope,
                     and the handling of sensitive requests.

User message    -> the VARIABLE part: the actual question, plus per-request data
                   (retrieved documents, the text to summarize).
Enter fullscreen mode Exit fullscreen mode

Keeping stable instructions in the system message and variable data in the user message also helps with cost, since providers can often cache a repeated, unchanging prefix (see AI APIs with .NET, Section 12).

The system prompt is guidance, not a security boundary

Models generally give system instructions priority — but that is a TENDENCY, not an
enforced rule. Do NOT put anything in a system prompt that you can't afford to have
revealed, and do NOT rely on it alone to enforce security.

PROMPT INJECTION: text in the user input — or in a document, web page, or email you
feed the model — can contain instructions that try to override yours ("ignore previous
instructions and..."). This is a fundamental risk whenever untrusted text enters a prompt.

Defenses are layered, never a single trick:
  - Clearly DELIMIT untrusted content and tell the model to treat it as data, not instructions.
  - Give the model the LEAST privilege: only the tools and data the task needs.
  - Enforce permissions and validation in CODE, outside the model.
  - Require human confirmation for consequential actions.
  - Don't put secrets (API keys, credentials, private data) in prompts.
Enter fullscreen mode Exit fullscreen mode

Role discipline in code

- Never let user-supplied text be placed in the SYSTEM role. That hands the user
  your highest-priority channel.
- Replay the assistant's previous turns faithfully; rewriting them can confuse the model.
- Keep tool results in the tool role, and remember they may contain untrusted text too.
Enter fullscreen mode Exit fullscreen mode

4. Few-Shot Prompting

Teaching by example, inside the prompt

Zero-shot : instructions only, no examples.          "Classify this ticket."
One-shot  : instructions plus ONE example.
Few-shot  : instructions plus SEVERAL examples (commonly 2-10) showing the
            input -> output pattern you want.
Enter fullscreen mode Exit fullscreen mode

Examples are often more effective than description at communicating format, tone, and edge-case behavior — showing the model what a good answer looks like is easier than explaining it.

Example: a few-shot classifier built as message pairs

(string Text, string Label)[] examples =
[
    ("I was charged twice this month.",           "Billing"),
    ("The export button crashes the app.",        "Bug"),
    ("Please add dark mode to the dashboard.",    "Feature"),
    ("My invoice shows the wrong company name.",  "Billing"),
];

var messages = new List<ChatMessage>
{
    new SystemChatMessage("Classify the support ticket as exactly one of: Billing, Bug, Feature. Reply with the label only.")
};

// Each example becomes a user message followed by the assistant's ideal reply
foreach (var (text, label) in examples)
{
    messages.Add(new UserChatMessage(text));
    messages.Add(new AssistantChatMessage(label));
}

messages.Add(new UserChatMessage(newTicketText));     // the real input, answered in the same pattern

ChatCompletion completion = await client.CompleteChatAsync(messages);
Enter fullscreen mode Exit fullscreen mode

Presenting examples as alternating user/assistant messages is a natural fit for chat models: the model sees what "a correct reply" looks like in the same format it will be asked to produce.

Choosing good examples

- DIVERSE: cover the different categories and the tricky edge cases, not five near-duplicates.
- REPRESENTATIVE: drawn from real inputs, matching what the model will actually see.
- CORRECT: a wrong example teaches the wrong thing — the model imitates its examples faithfully.
- CONSISTENT in format: if examples vary in structure, the output will too.
- BALANCED: if 4 of 5 examples share one label, the model may be biased toward it.
- ORDER can matter: the most recent examples sometimes carry extra weight, so don't
  accidentally put the same label last every time.
Enter fullscreen mode Exit fullscreen mode

Trade-offs and variations

Costs:   every example consumes INPUT TOKENS on EVERY request. Ten long examples is
         a real, recurring cost (Section 7). Use as few as achieve the quality you need.

Dynamic few-shot: instead of fixed examples, store a bank of labeled examples with
         embeddings and, per request, retrieve the few most SIMILAR to the new input.
         Often beats a fixed set for varied inputs, at the price of more machinery.

When NOT to bother:  simple, well-specified tasks on a capable model may work fine
         zero-shot — test before adding examples. Some reasoning-focused models also
         respond better to clear instructions than to many examples; check your
         model's guidance.

Beyond a point:      if you need dozens of examples to get consistent behavior,
         that's a signal to consider fine-tuning (Section 9) instead.
Enter fullscreen mode Exit fullscreen mode

5. Structured Prompting

Giving the prompt itself clear structure

A long, unstructured prompt is hard for both humans and models to parse. Structured prompting means organizing the prompt into clearly labeled sections, using consistent delimiters, so the model can tell instructions, context, and input apart.

You are a contract-review assistant.

<instructions>
Review the clause below and identify risks for the CUSTOMER.
Respond only with the JSON described in <output_format>.
If the clause is not a contract clause, return {"risks": [], "note": "not a contract clause"}.
</instructions>

<context>
Customer is a small business. Jurisdiction: India. Focus on payment, termination, liability.
</context>

<clause>
{clause text}
</clause>

<output_format>
{"risks": [{"issue": string, "severity": "low" | "medium" | "high", "explanation": string}]}
</output_format>
Enter fullscreen mode Exit fullscreen mode

Why structure helps

- SEPARATION: delimiters (XML-style tags, Markdown headers, triple quotes) clearly mark
  where instructions end and untrusted data begins — which reduces confusion and gives
  you a (partial) defense against prompt injection.
- REFERENCE: you can refer to sections by name ("using only the text in <context>").
- MAINTAINABILITY: each section can be edited, tested, and templated independently.
- CONSISTENCY: the same skeleton across many tasks makes prompts predictable to
  write and review.
Enter fullscreen mode Exit fullscreen mode

Structure the output, too

Prompted JSON ("reply in JSON") works most of the time — and fails in production when
a stray sentence or missing field breaks your parser. For anything your code consumes,
prefer the API's STRUCTURED OUTPUTS feature (schema-constrained JSON) over prompt
wording alone, then still validate the values (see AI APIs with .NET, Section 7).

Remember: a schema guarantees SHAPE, not TRUTH.
Enter fullscreen mode Exit fullscreen mode

Prompt chaining: structure across multiple calls

Instead of one giant prompt, split a job into stages, each with a focused prompt, passing
output forward — and checking it between stages with ordinary code:

  1. EXTRACT   pull the key facts from the document        (structured output)
  2. VALIDATE  check them in code against rules/databases
  3. ANALYZE   reason over the validated facts
  4. DRAFT     write the final response in the required tone

Benefits: each step is simpler to prompt, test, and debug; failures are easier to
locate; you can use a cheaper model for easy stages; and code can enforce rules
between steps. Cost: more calls, more latency, more plumbing.
Enter fullscreen mode Exit fullscreen mode

Templates and safe assembly

// Build prompts from templates, not scattered string concatenation
const string Template = """
    <instructions>{0}</instructions>
    <document>
    {1}
    </document>
    """;

string prompt = string.Format(Template, instructions, documentText);
Enter fullscreen mode Exit fullscreen mode

One caution: if untrusted text can contain your own delimiter tags (for example a literal </document>), it can try to "close" the section early and smuggle in instructions. Strip or escape delimiter sequences from untrusted input, and don't treat delimiters as a security guarantee on their own.


6. Context Windows

The context window is the model's working memory — and it is finite

The context window is the maximum number of tokens the model can consider in a single request, and it includes everything: system prompt, conversation history, retrieved documents, tool definitions and results, your new question, and the reply it generates.

Context window  >=  system prompt
                  + conversation history
                  + retrieved documents / tool results
                  + tool definitions
                  + the user's new message
                  + the model's OUTPUT
Enter fullscreen mode Exit fullscreen mode

Window sizes vary by model and have grown dramatically, but they are never unlimited — and bigger isn't free or automatically better.

Bigger windows are not a license to stuff everything in

COST:     you pay for every input token, on every request. Resending a 100-page
          document with each question multiplies cost quickly.

LATENCY:  more input tokens generally means slower responses.

QUALITY:  models do not use all parts of a long context equally well. Research has shown
          that information buried in the MIDDLE of a very long prompt can be used less
          reliably than information near the beginning or end, and that adding lots of
          irrelevant text can degrade answers. The effect varies by model and has
          improved over time, but "more context" can still mean "worse answers."

GUIDELINE: include what's RELEVANT, not everything that's AVAILABLE.
Enter fullscreen mode Exit fullscreen mode

Strategies for working within the window

RETRIEVAL (RAG)     Store documents in chunks with embeddings; per question, retrieve only
                    the top few relevant chunks and include those. The standard approach
                    for "chat with your documents."

TRIMMING / SLIDING  Keep only the most recent conversation turns (Section 10).
WINDOW

SUMMARIZATION       Compress older context into a short summary and keep it plus recent turns.

SELECTIVE TOOL      Return only the fields the model needs from a tool, not an entire record.
OUTPUT

PLACEMENT           Put the most important instructions and the question where the model
                    attends well (typically at the start and/or restated at the end), and
                    keep reference material clearly delimited.

CHUNKING LONG       For documents that exceed the window, process in pieces (map) and
INPUT               combine the results (reduce).
Enter fullscreen mode Exit fullscreen mode

Always leave room for the answer

If input fills the window, there's no space left for output and the reply is cut off
(FinishReason = Length). Budget explicitly:

   available for input  =  context window  -  reserved output tokens  -  safety margin

Count tokens BEFORE sending (Section 7) and trim to that budget deliberately,
rather than discovering the limit via an error in production.
Enter fullscreen mode Exit fullscreen mode

7. Tokens

The unit the model reads, writes, and bills in

Models don't process characters or words — they process tokens, chunks of text produced by a tokenizer. A token might be a whole common word, part of a rarer word, a punctuation mark, or a few characters.

Rule of thumb (English):  1 token ≈ 4 characters ≈ 3/4 of a word
  "hamburger"  -> may split into several tokens
  " the"       -> usually a single token
  Numbers, code, and non-English text often use MORE tokens per unit of meaning.
Enter fullscreen mode Exit fullscreen mode

Why tokens matter to application builders

1. COST      Billing is per token, with INPUT and OUTPUT usually priced differently
             (output typically costs more).
2. LIMITS    The context window and rate limits (tokens per minute) are measured in tokens.
3. SPEED     Generation time grows with the number of OUTPUT tokens.
4. QUIRKS    Because models see tokens, not letters, tasks like "count the letters in
             this word," exact character manipulation, and some arithmetic are
             unexpectedly error-prone. Do those in code.
Enter fullscreen mode Exit fullscreen mode

Practical habits

- Measure real usage from every response and log it.
- Count tokens locally BEFORE sending (e.g. with a tokenizer library) to enforce budgets.
- Remember that few-shot examples, long system prompts, tool definitions, and conversation
  history are all re-sent — and re-billed — on every request.
- Set a maximum output token cap appropriate to the task.
- Be aware that the same text can tokenize differently across model families, so token
  counts aren't directly comparable between providers.
Enter fullscreen mode Exit fullscreen mode

The mechanics — counting tokens in .NET, logging usage, trimming history, and the cost levers — are covered in this series' AI APIs with .NET guide, Sections 8 and 12.


8. Hallucinations

Fluent, confident, and wrong

A hallucination is output that sounds plausible but is false or unsupported — an invented citation, a nonexistent API method, a made-up statistic, a fabricated quote. It's not a rare glitch; it's a direct consequence of how LLMs work: a model optimized to produce likely-sounding text will produce likely-sounding text whether or not it's true, and it has no built-in alarm for "I don't actually know this."

Common forms

- FABRICATED FACTS      invented names, dates, statistics, quotes
- FAKE SOURCES          citations, URLs, case law, or papers that don't exist
- INVENTED CODE/APIS    plausible-looking methods, parameters, or packages that aren't real
- UNFAITHFUL SUMMARIES  claims not actually present in the source document
- WRONG REASONING       confident, well-formatted, but flawed logic or arithmetic
- STALE ANSWERS         presenting out-of-date information as current
Enter fullscreen mode Exit fullscreen mode

Why it's dangerous in applications

Hallucinations are hardest to catch precisely because they LOOK right. Users trust fluent,
confident text. In medicine, law, finance, security, and customer support, a confidently
wrong answer can cause real harm — so the design goal is not "hope it's accurate" but
"build a system where errors are made unlikely, visible, and survivable."
Enter fullscreen mode Exit fullscreen mode

Mitigation: layers, not a single fix

1. GROUND THE MODEL IN SOURCE TEXT (RAG)
   Retrieve relevant documents and instruct: "Answer ONLY from the provided context."
   Far more reliable than relying on what the model memorized. This is the single
   most effective measure for factual applications.

2. GIVE IT PERMISSION TO SAY "I DON'T KNOW"
   "If the answer isn't in the context, say you don't know." Models often guess unless
   explicitly allowed to decline.

3. REQUIRE CITATIONS AND CHECK THEM
   Ask the model to quote or reference the supporting passage, then VERIFY in code that the
   quote actually appears in the retrieved source. A citation you don't verify is decoration.

4. LOWER THE TEMPERATURE
   For factual tasks, low randomness reduces (but doesn't eliminate) invention.

5. CONSTRAIN THE OUTPUT
   Structured outputs, enumerated labels, and validation against known values
   (does this product ID exist? is this a valid ISO currency code?) catch many errors in code.

6. USE TOOLS FOR FACTS AND MATH
   Don't let the model recall prices, balances, or dates, or do arithmetic — have it call a
   function that looks them up or computes them.

7. VERIFY HIGH-STAKES OUTPUT
   Second-pass checks (a separate "verifier" prompt or model), rule-based checks, and — where
   consequences are serious — a HUMAN in the loop before the answer is acted on.

8. COMMUNICATE UNCERTAINTY TO USERS
   Show sources, label AI-generated content, and don't present output as authoritative.

9. MEASURE IT
   Build an evaluation set with known answers and track a faithfulness / groundedness rate
   over time (Section 12).
Enter fullscreen mode Exit fullscreen mode

A grounded prompt pattern

Answer the question using ONLY the information in <context>.
Quote the sentence(s) that support your answer in a "sources" list.
If <context> does not contain the answer, reply exactly: "I don't have enough information to answer that."
Do not use outside knowledge.

<context>
{retrieved chunks}
</context>

<question>
{user question}
</question>
Enter fullscreen mode Exit fullscreen mode

An honest bottom line

Hallucinations can be REDUCED substantially but not ELIMINATED. Even with retrieval, a model
can misread or misrepresent its sources. Design for it: validate, cite, verify, and keep
humans in the loop where being wrong is costly.
Enter fullscreen mode Exit fullscreen mode

9. Model Selection

There is no "best model" — only the best model for your task, budget, and constraints

Providers offer many models spanning small/fast/cheap to large/capable/expensive, plus specialized variants (reasoning, coding, vision, embedding), and open-weight alternatives. Choosing well is an engineering decision, and it should be made with data, not hype.

The dimensions that matter

QUALITY          How well does it perform on YOUR task? Public benchmarks are a rough guide;
                 your own evaluation set is the real test (Section 12).

LATENCY          Time to first token and total generation time. Matters enormously for
                 interactive chat; barely at all for overnight batch jobs.

COST             Price per input/output token. Multiply by your real traffic and prompt sizes.
                 Large price gaps exist between model tiers.

CONTEXT WINDOW   Can it hold the documents/history your use case needs? (Section 6)

CAPABILITIES     Tool/function calling, structured outputs, vision or audio input,
                 streaming, fine-tuning support, reasoning mode — only what you need.

RELIABILITY      Instruction-following consistency, rate limits and quotas, and provider
                 uptime/support.

DATA & COMPLIANCE  Where data is processed and stored, retention and training policies,
                 regional availability, certifications, private networking (a major
                 reason to use a managed cloud offering such as Azure OpenAI).

OPEN vs. CLOSED  Open-weight models can be self-hosted (control, privacy, no per-token fee,
                 but you own the infrastructure and operations). Hosted proprietary models
                 are simpler to operate and often state-of-the-art, at a per-use price.
Enter fullscreen mode Exit fullscreen mode

A practical selection process

1. DEFINE the task and what "good" means (accuracy target, latency budget, cost ceiling).
2. BUILD an evaluation set of realistic inputs with expected outputs or grading criteria.
3. START with a capable model to find out what's achievable.
4. TRY cheaper/smaller models on the same set — you may find a small model is good enough.
5. COMPARE on quality, latency, and cost together, not quality alone.
6. DECIDE, then re-evaluate periodically: models and prices change fast.
Enter fullscreen mode Exit fullscreen mode

Common architecture patterns

MODEL ROUTING     Send easy requests to a cheap model, hard ones to a stronger one, based on
                  task type, input complexity, or a quick classifier. Often the largest
                  cost saver.

CASCADE           Try the cheap model first; escalate only if its output fails validation
                  or confidence checks.

RIGHT MODEL PER   Use a small model for classification and routing, a stronger one for final
STAGE             synthesis, a dedicated embedding model for retrieval.

SPECIALIZE LAST   Prefer prompting, then retrieval (RAG), and only then fine-tuning:
                  prompting is cheapest to iterate; RAG gives fresh, citable knowledge;
                  fine-tuning is best for consistent style, format, or narrow behavior —
                  not for teaching the model new facts reliably.
Enter fullscreen mode Exit fullscreen mode

Avoid painting yourself into a corner

- Keep model names, parameters, and prices in CONFIGURATION, not code.
- Put calls behind an interface (for example Microsoft.Extensions.AI's IChatClient)
  so providers can be swapped.
- Plan for DEPRECATION: providers retire model versions on a schedule, so pin versions where
  you need stability, track retirement notices, and re-run your evaluation set before migrating.
- Beware "prompt lock-in": prompts tuned to one model may need re-tuning on another.
Enter fullscreen mode Exit fullscreen mode

10. Conversation Memory

The model remembers nothing — memory is something your application builds

Because every request is stateless, "memory" is your responsibility. There are two broad kinds, usually combined.

SHORT-TERM MEMORY   The current conversation: the recent turns resent with each request.
LONG-TERM MEMORY    Information that persists ACROSS conversations: user preferences, facts
                    about the user, past decisions, accumulated knowledge — stored outside
                    the model and retrieved when relevant.
Enter fullscreen mode Exit fullscreen mode

Short-term strategies

FULL HISTORY        Resend every turn. Simplest and most faithful — until cost and the
                    context window make it impractical.

SLIDING WINDOW      Keep only the last N turns (or last N tokens). Simple and cheap, but the
                    model forgets anything older.

SUMMARIZATION       Condense older turns into a running summary, and keep the summary plus
                    recent turns verbatim. Preserves the gist at bounded size, at the cost
                    of an extra model call and some detail loss.

HYBRID              Summary of old turns + recent turns verbatim + retrieved long-term facts.
                    The usual production answer.
Enter fullscreen mode Exit fullscreen mode

A summary-plus-recent-window memory in C

public record Turn(string Role, string Text);

public class ConversationMemory
{
    private readonly List<Turn> _recent = new();
    private readonly int _maxRecentTurns;
    private string _summary = "";

    public ConversationMemory(int maxRecentTurns = 8) => _maxRecentTurns = maxRecentTurns;

    public void Add(string role, string text) => _recent.Add(new Turn(role, text));

    // Fold the oldest turns into the running summary when the window is exceeded
    public async Task CompactAsync(ChatClient client, CancellationToken ct = default)
    {
        if (_recent.Count <= _maxRecentTurns) return;

        int toSummarize = _recent.Count - _maxRecentTurns;
        string transcript = string.Join("\n", _recent.Take(toSummarize).Select(t => $"{t.Role}: {t.Text}"));

        var result = await client.CompleteChatAsync(
            [
                new SystemChatMessage("Update the running summary of this conversation. Keep names, decisions, " +
                                      "user preferences, and open questions. Maximum 150 words."),
                new UserChatMessage($"Current summary:\n{_summary}\n\nNew turns:\n{transcript}")
            ],
            cancellationToken: ct);

        _summary = result.Value.Content[0].Text;
        _recent.RemoveRange(0, toSummarize);
    }

    public List<ChatMessage> BuildMessages(string systemPrompt)
    {
        var messages = new List<ChatMessage> { new SystemChatMessage(systemPrompt) };

        if (_summary.Length > 0)
            messages.Add(new SystemChatMessage($"Summary of the earlier conversation:\n{_summary}"));

        foreach (var turn in _recent)
            messages.Add(turn.Role == "user" ? new UserChatMessage(turn.Text) : new AssistantChatMessage(turn.Text));

        return messages;
    }
}
Enter fullscreen mode Exit fullscreen mode

Call CompactAsync after each exchange (or in the background), and BuildMessages when assembling each request. Persist the memory object (a database or cache keyed by conversation ID) so conversations survive restarts and scale across servers.

Long-term memory

Typical design:
  1. EXTRACT   After conversations, pull durable facts worth keeping
               ("prefers metric units", "works in the finance team", "decided to use PostgreSQL").
  2. STORE     Save them in a database, with an EMBEDDING for semantic lookup.
  3. RETRIEVE  On each new request, fetch the few memories relevant to the current
               question and add them to the prompt.
  4. MAINTAIN  Update facts that change, merge duplicates, and expire stale ones.
Enter fullscreen mode Exit fullscreen mode

This is retrieval-augmented generation applied to the user's own history. The vectors can live in a vector database or in a vector-capable index in a database you already run.

Memory raises real design and privacy questions

- WHAT to remember: store durable, useful facts — not every utterance.
- CONSENT AND TRANSPARENCY: tell users what is remembered, and let them view, correct,
  and delete it. Many privacy regulations require this for personal data.
- SENSITIVE DATA: don't persist secrets, credentials, or unnecessary personal data.
  Minimizing what you store minimizes what can leak.
- ISOLATION: memory must be strictly scoped per user/tenant — a retrieval bug that
  surfaces one user's memories to another is a serious data breach.
- ACCURACY: summaries and extracted facts can themselves be wrong (hallucinated),
  and a wrong memory persists and compounds. Prefer storing the user's own words,
  and let users correct entries.
- INJECTION: anything written into memory from untrusted content can later re-enter
  prompts — treat stored memories as untrusted data, not instructions.
Enter fullscreen mode Exit fullscreen mode

11. AI Agents

From "answer a question" to "pursue a goal"

An AI agent is an LLM-driven system that decides what to do next: given a goal, it can plan, call tools, observe the results, and repeat until it judges the goal accomplished. Where a plain chatbot answers once, an agent runs a loop.

        ┌─────────────────────────────────────────────────┐
        │                                                 │
   GOAL ─> THINK (what's the next step?) ─> ACT (call a tool)
        ^                                         │
        │                                         v
        └──────── OBSERVE (read the tool result) ─┘
                          │
                  goal met? ──> FINAL ANSWER
Enter fullscreen mode Exit fullscreen mode

An agent is built from a small set of parts:

MODEL      the "brain" that reasons and chooses actions
TOOLS      functions it can call: search, query a database, call an API, run code, send email
MEMORY     the running transcript (short-term) plus any retrieved long-term knowledge
LOOP       the orchestration code that feeds tool results back until done
GUARDRAILS limits, permissions, and approvals that bound what it can do
Enter fullscreen mode Exit fullscreen mode

Workflows versus agents: choose the simplest thing that works

WORKFLOW   The steps are fixed IN YOUR CODE; the LLM fills in individual steps
           (classify -> retrieve -> draft -> validate). Predictable, testable, cheaper.

AGENT      The LLM chooses the steps dynamically at runtime. Flexible, handles open-ended
           tasks — but less predictable, harder to test, slower, and more expensive.

Guidance: if you can describe the steps in advance, build a WORKFLOW. Reach for an agent
only when the path genuinely can't be known up front. Many "agent" use cases are better
served by a well-designed workflow, and the most reliable systems often combine both:
a workflow that invokes a bounded agent for one open-ended step.
Enter fullscreen mode Exit fullscreen mode

A minimal agent loop in C

The tool-calling mechanics are covered in AI APIs with .NET, Section 6; an agent is that loop with deliberate limits around it.

public class SimpleAgent
{
    private const int MaxSteps = 8;
    private const int MaxTotalTokens = 50_000;

    private readonly ChatClient _client;
    private readonly ChatCompletionOptions _options;       // includes the Tools the agent may use
    private readonly Func<ChatToolCall, CancellationToken, Task<string>> _executeTool;

    public SimpleAgent(ChatClient client, ChatCompletionOptions options,
                       Func<ChatToolCall, CancellationToken, Task<string>> executeTool)
        => (_client, _options, _executeTool) = (client, options, executeTool);

    public async Task<string> RunAsync(string goal, string systemPrompt, CancellationToken ct)
    {
        var messages = new List<ChatMessage> { new SystemChatMessage(systemPrompt), new UserChatMessage(goal) };
        int tokensUsed = 0;

        for (int step = 0; step < MaxSteps; step++)               // HARD cap on iterations
        {
            ChatCompletion completion = await _client.CompleteChatAsync(messages, _options, ct);

            tokensUsed += completion.Usage.TotalTokenCount;
            if (tokensUsed > MaxTotalTokens)                      // HARD cap on spend
                return "Stopped: token budget exceeded.";

            if (completion.FinishReason == ChatFinishReason.Stop)
                return completion.Content[0].Text;                // the agent decided it is done

            if (completion.FinishReason != ChatFinishReason.ToolCalls)
                return $"Stopped: {completion.FinishReason}";

            messages.Add(new AssistantChatMessage(completion));   // keep the tool request in history

            foreach (ChatToolCall call in completion.ToolCalls)
            {
                string result;
                try   { result = await _executeTool(call, ct); }  // validation, authorization, and approval live HERE
                catch (Exception ex) { result = $"Tool error: {ex.Message}"; }   // errors go back as text so it can adapt

                messages.Add(new ToolChatMessage(call.Id, result));
            }
        }

        return "Stopped: step limit reached.";
    }
}
Enter fullscreen mode Exit fullscreen mode

Designing tools well

- FEW, FOCUSED tools with clear names and descriptions beat many overlapping ones. The
  description is how the model decides when to use it — write it like documentation.
- Return SMALL, relevant results. Everything returned re-enters the context window.
- Return helpful ERROR messages as text ("Customer 123 not found. Try searching by email.")
  so the agent can recover instead of looping blindly.
- Make tools IDEMPOTENT where possible, so a retry or repeated call can't double-charge
  or double-send.
- Separate READ tools (safe) from WRITE/ACT tools (consequential), and gate the latter.
Enter fullscreen mode Exit fullscreen mode

Guardrails: an agent amplifies both capability and risk

BOUND THE LOOP        Maximum steps, maximum tokens, maximum wall-clock time.
LEAST PRIVILEGE       Give the agent only the tools and data the task requires, scoped to the
                      current user's permissions — enforced in CODE, not in the prompt.
VALIDATE EVERY CALL   Tool arguments come from a model steered by input that may be hostile.
                      Validate them like any untrusted input.
HUMAN APPROVAL        Require confirmation before irreversible or high-impact actions
                      (sending money, deleting data, emailing customers, deploying code).
SANDBOX               Run code-execution tools in an isolated environment.
PROMPT INJECTION      An agent that reads web pages, emails, or documents can be hijacked by
                      instructions hidden in them. Treat all tool output as untrusted data,
                      and never let it expand the agent's privileges.
OBSERVABILITY         Log every step — prompts, tool calls, results, token usage — so you
                      can debug, audit, and replay what the agent did and why.
Enter fullscreen mode Exit fullscreen mode

Common failure modes

- LOOPS: repeating the same failing action. (Caps, and error messages that suggest alternatives.)
- COMPOUNDING ERRORS: a small early mistake cascades through later steps.
- WRONG TOOL / WRONG ARGUMENTS: especially with vague tool descriptions.
- CONTEXT BLOAT: the transcript grows with every step, raising cost and degrading quality.
- COST AND LATENCY: many sequential model calls add up fast.
- OVER-AUTONOMY: acting confidently on a misunderstanding.
Enter fullscreen mode Exit fullscreen mode

Multi-agent systems

Complex problems are sometimes split among several specialized agents (a researcher, a coder,
a reviewer) coordinated by an orchestrator. This can help with separation of concerns and
parallel work, but adds cost, latency, and new failure modes (miscommunication, loops between
agents). Start with ONE agent or a plain workflow, and add agents only when a single one
demonstrably struggles.
Enter fullscreen mode Exit fullscreen mode

Frameworks

In .NET, Microsoft.Extensions.AI provides IChatClient abstractions and automatic function-invocation middleware, and Microsoft maintains higher-level orchestration frameworks (such as Semantic Kernel and the Microsoft Agent Framework) that package agent loops, memory, and multi-agent patterns. These evolve quickly, so check their current documentation. Whatever you adopt, understand the underlying loop above — it's what the frameworks implement, and it is what you'll be debugging.


12. Evaluating LLM Applications

"It seemed to work when I tried it" is not evidence

Because outputs vary and failures are subtle, LLM applications need systematic evaluation — the equivalent of a test suite, adapted to probabilistic output.

BUILD AN EVALUATION SET   A collection of realistic inputs with expected outputs or grading
                          criteria — including edge cases, adversarial inputs, and examples
                          of past failures. Start with a few dozen; grow it with every bug.

CHOOSE METRICS            Exact match (classification, extraction), schema-validity rate,
                          groundedness/faithfulness to sources, task success rate, latency,
                          cost per request, refusal/escalation rate.

GRADING METHODS           - Programmatic checks (exact match, regex, JSON validation, unit tests
                            for generated code) — cheapest and most reliable where possible.
                          - Model-as-judge: another LLM grades outputs against criteria — scalable
                            but imperfect and biased; calibrate it against human judgments.
                          - Human review — the gold standard; use it for calibration and
                            for high-stakes samples.

RUN IT CONTINUOUSLY       Re-run on every prompt change, model change, or retrieval change, so
                          regressions show up before users find them. Compare versions side by side.

MONITOR IN PRODUCTION     Log inputs, outputs, tool calls, token usage, and user feedback
                          (thumbs up/down). Sample real traffic back into your evaluation set.
                          (Mind privacy when logging user content.)
Enter fullscreen mode Exit fullscreen mode

Evaluation is also what makes every decision in this guide testable instead of a matter of opinion: whether few-shot examples help, whether a cheaper model is good enough, whether a prompt change reduced hallucinations, whether the agent actually completes tasks.


13. Putting It Together: Anatomy of an LLM Application

A retrieval-augmented support assistant

 USER QUESTION
      │
      v
 [1] INPUT GUARDS        length cap, rate limit, basic abuse checks            (code)
      │
      v
 [2] RETRIEVE            embed the question -> fetch top-k relevant chunks
      │                  + long-term memory facts for this user                (embeddings + store)
      v
 [3] ASSEMBLE PROMPT     system rules + <context> chunks + conversation
      │                  memory (summary + recent turns) + the question        (template)
      v
 [4] CALL MODEL          right-sized model, low temperature, token cap,
      │                  structured output, retries/timeouts                   (LLM API)
      v
 [5] VALIDATE            schema check, verify citations appear in sources,
      │                  business rules, refusal/"don't know" handling         (code)
      v
 [6] RESPOND + LOG       stream/return answer with sources; log usage,
                         latency, feedback; update conversation memory         (code)
Enter fullscreen mode Exit fullscreen mode

Every section of this guide appears in that pipeline: prompt engineering and structured prompting shape step 3; tokens and context windows govern what fits; the message roles organize it; model selection and parameters configure step 4; hallucination mitigation spans steps 2, 3, and 5; conversation memory feeds steps 2 and 3; and an agent is what you get when step 4 can request tools and loop. Evaluation wraps the whole thing.


14. Common Pitfalls

Pitfall Why it hurts Better approach
Treating the model as a database of facts It generates plausible text, not verified facts, and has a knowledge cutoff Ground answers in retrieved sources; use tools for facts and math (Section 8)
Vague prompts Vague input produces generic, inconsistent output Specify role, task, constraints, and output format (Section 2)
No "I don't know" path Models guess unless told they may decline Give an explicit escape hatch and required fallback wording (Sections 2, 8)
Putting user text in the system message Hands users your highest-priority instruction channel Keep system messages developer-controlled; put user data in user messages (Section 3)
Trusting the system prompt for security Instruction priority is a tendency, not an enforced rule; prompts can leak Enforce permissions and validation in code; never put secrets in prompts (Section 3)
Ignoring prompt injection from documents/web/tool output Hidden instructions in untrusted content can hijack the model or agent Delimit untrusted data, least privilege, human approval for consequential actions (Sections 3, 11)
Wrong or unbalanced few-shot examples The model imitates its examples faithfully, including their errors and biases Use diverse, correct, balanced, representative examples; test with and without (Section 4)
Relying on "reply in JSON" wording alone Free-text JSON eventually breaks the parser Use structured outputs, then still validate values (Section 5)
Stuffing the whole document into every request High cost, slow responses, and often worse answers Retrieve only the relevant chunks (Section 6)
Forgetting to reserve output space A full context window truncates the reply Budget input as window minus reserved output minus margin (Section 6)
Using the model for exact arithmetic, counting, or lookups Token-based models are error-prone at these Do them in code; expose them as tools (Sections 7, 11)
Unverified citations Models invent plausible-looking sources Check in code that quoted passages exist in the retrieved source (Section 8)
Choosing a model by hype or a single benchmark Benchmarks may not reflect your task, cost, or latency needs Build an evaluation set; compare quality, latency, and cost together (Section 9)
Hard-coding one model and one provider Deprecations, price changes, and better options force rewrites Configuration and an abstraction layer; plan for version retirement (Section 9)
Unbounded conversation history Cost and latency grow every turn until the window overflows Sliding window, summarization, or retrieval-backed memory (Section 10)
Storing everything users say Privacy risk, accumulating errors, and noise Store durable facts, give users control, scope strictly per user (Section 10)
Building an agent when a workflow would do Unpredictable, slower, costlier, and harder to test Use fixed workflows for known steps; agents only for genuinely open-ended work (Section 11)
Agents with no limits Runaway loops, runaway cost, and unbounded blast radius Step, token, and time caps; least privilege; approval gates; full logging (Section 11)
Shipping with no evaluation Regressions and hallucinations are discovered by users Maintain an evaluation set; re-run on every change; monitor production (Section 12)

Quick Reference Table

Concept Key idea Practical takeaway
LLM Predicts the next token; fluent, probabilistic, stateless Use it for language; use code for math, lookups, rules
Prompt engineering Specific role, task, constraints, format Treat prompts as versioned, tested code
System message Developer's stable instructions Guidance, not a security boundary
User message The variable request and per-request data Never let user text into the system role
Assistant message The model's earlier replies Replayed as history; the model is stateless
Few-shot prompting Examples inside the prompt Diverse, correct, balanced; costs tokens every request
Structured prompting Delimited, labeled prompt sections Separates instructions from data; supports chaining
Structured output Schema-constrained response Guarantees shape, not truth — still validate
Context window Max tokens per request (input + output) Include what's relevant, not everything available
Tokens The unit of cost, limits, and speed Count before sending; cap output; log usage
Hallucination Plausible but false output Ground in sources, allow "I don't know", verify citations
RAG Retrieve relevant chunks, then generate The primary defense against hallucination and stale knowledge
Model selection Quality vs. latency vs. cost vs. constraints Evaluate on your own task; route by difficulty
Short-term memory Recent turns resent each call Sliding window, summarization, or hybrid
Long-term memory Facts persisted across conversations Extract, store, retrieve — with privacy and per-user isolation
Workflow Fixed steps, LLM fills each Prefer when the path is known
AI agent LLM loop: think, act with tools, observe Bound it, least-privilege it, require approval for risky actions
Evaluation A test suite for probabilistic output The only way to know a change helped

Conclusion

Building reliable applications on LLMs comes down to a handful of recurring ideas. The model is a powerful but probabilistic text predictor — stateless, bounded by a context window, billed in tokens, and prone to producing fluent falsehoods — so your job is to surround it with good engineering: precise, structured, versioned prompts; the right context supplied through retrieval rather than hope; constrained, validated outputs; a memory layer you design deliberately; and strict, code-enforced limits on what it is allowed to do. The roles in a chat (system, user, assistant, tool) give you a way to organize that, few-shot examples and structured prompts give you control over behavior and format, and hallucination mitigation is a layered discipline of grounding, citing, verifying, and keeping humans in the loop where mistakes are costly.

Two principles tie it together. First, use the simplest architecture that works: a single well-crafted prompt before a chain, a fixed workflow before an agent, a small model before a large one, retrieval before fine-tuning. Each step up the ladder buys flexibility at the price of cost, latency, and unpredictability. Second, measure instead of guessing: an evaluation set turns every question in this guide — does this example help, is this cheaper model good enough, did that prompt change reduce hallucinations, does the agent really finish the task — into something you can answer with data. Teams that do both ship LLM features that people can trust; teams that skip them ship impressive demos that fail in production.


Found this useful? Feel free to star the repo, open an issue with corrections, or share the "it sounded completely right and was completely wrong" story that made the case for grounding and evaluation click better than any theory.

Top comments (1)

Some comments may only be visible to logged-in visitors. Sign in to view all comments. Some comments have been hidden by the post's author - find out more