LLM Application Development
A deep-dive walkthrough of building applications on top of large language models — covering how LLMs actually work and what that implies for application design, prompt engineering, the system/user/assistant message roles, few-shot and structured prompting, context windows and tokens, why hallucinations happen and how to reduce them, how to choose a model, conversation memory strategies, and AI agents: what they are, how the loop works, and how to keep them safe and bounded.
Table of Contents
- Introduction
- LLM Fundamentals
- Prompt Engineering
- System, User, and Assistant Messages
- Few-Shot Prompting
- Structured Prompting
- Context Windows
- Tokens
- Hallucinations
- Model Selection
- Conversation Memory
- AI Agents
- Evaluating LLM Applications
- Putting It Together: Anatomy of an LLM Application
- Common Pitfalls
- Quick Reference Table
- Conclusion
Introduction
Building with LLMs is a different kind of engineering from traditional software. A conventional function is deterministic: the same input gives the same output, and a bug is a defect you can locate and fix. An LLM is probabilistic: it produces plausible text, usually correct, sometimes confidently wrong, and the "program" you write is largely natural language. Most of the craft of LLM application development is learning to build reliable systems out of an unreliable-by-nature component — by giving the model the right context, constraining its output, validating what comes back, measuring quality, and bounding what it can do.
A typical LLM application is NOT "call the model, show the answer." It is:
User input
-> validate and limit it
-> retrieve relevant context (documents, memory, data)
-> assemble a prompt (instructions + context + conversation + question)
-> call the model (possibly in a loop with tools)
-> validate / parse the output
-> log, evaluate, and return the result
This guide is the conceptual companion to this series' AI APIs with .NET guide. That guide covers the mechanics of calling the API (clients, streaming, tool-calling code, retries, rate limits, cost); this one covers how to design the application around the model. Where the mechanics matter, it points back there rather than repeating them.
1. LLM Fundamentals
An LLM predicts the next token — everything else follows from that
A large language model is a neural network trained on enormous amounts of text to do one thing: given a sequence of tokens, predict a probability distribution over the next token. It generates a response by repeatedly predicting a token, appending it, and predicting again.
Input: "The capital of France is"
Model: P(" Paris") = high, P(" Lyon") = low, ...
Output: " Paris" -> appended -> predict next token -> "." ...
That simple mechanism produces surprisingly capable behavior — summarizing, translating, writing code, reasoning through problems — because predicting text well across the breadth of human writing requires modeling a great deal about the world. But it also explains the characteristic limitations, which you must design around:
- It generates PLAUSIBLE text, not VERIFIED text. Fluency is not accuracy.
(Section 8: hallucinations)
- It is STATELESS. It has no memory between calls beyond what you send it.
(Section 10: conversation memory)
- It has a KNOWLEDGE CUTOFF. It knows nothing about events after training,
and nothing about your private data unless you supply it.
- It is NON-DETERMINISTIC by default. Sampling introduces variation.
- It has a bounded CONTEXT WINDOW. It can only consider so much text at once.
(Section 6)
How models get their behavior
1. PRE-TRAINING Learn language and broad knowledge by predicting next tokens over a
huge text corpus. Produces a "base" model — capable but not
conversational or reliably helpful.
2. FINE-TUNING Further train on curated examples (instruction/response pairs) so the
(instruction model follows instructions and answers in a useful format.
tuning)
3. ALIGNMENT Shape behavior toward helpful, honest, and safe responses, often using
(e.g. RLHF) human feedback or preference data as a training signal (this series'
AI/ML Fundamentals guide covers reinforcement learning).
(Newer "reasoning" models add training to produce extended internal reasoning
before answering, trading latency and cost for accuracy on hard problems.)
What this means for application builders
LLMs are strong at: language tasks — summarizing, rewriting, extracting, classifying,
translating, drafting, explaining, and generating code — and at
handling messy, unstructured input.
LLMs are weak at: exact arithmetic and counting, guaranteed factual recall, knowing
what they don't know, acting on current or private data without
being given it, and being perfectly repeatable.
Design rule: Use the model for what it's good at (language, judgment, flexibility)
and use ORDINARY CODE for what code is good at (math, lookups,
validation, business rules, security). The best LLM applications
are hybrids.
Embeddings: the other half of the toolkit
Alongside generation, models can produce embeddings — vectors where similar meanings sit close together (covered in this series' AI/ML Fundamentals guide). They power semantic search and retrieval-augmented generation (Sections 8, 10, 13), letting you find relevant text by meaning and hand it to the model as context.
2. Prompt Engineering
The prompt is your program — treat it with the same care as code
A prompt is the input you give the model, and small changes in wording can produce large changes in output. Prompt engineering is the practice of designing prompts that produce reliably good results.
A well-formed prompt has distinct parts
1. ROLE / CONTEXT Who the model is acting as and the situation it's in.
2. TASK Exactly what to do — specific and unambiguous.
3. INPUT DATA The material to work on (clearly delimited from the instructions).
4. CONSTRAINTS Rules: scope, tone, length, what NOT to do, what to do when unsure.
5. OUTPUT FORMAT The exact shape of the response (a label, bullets, JSON schema...).
6. EXAMPLES Optional demonstrations of good input/output pairs (Section 4).
Principles that consistently help
BE SPECIFIC. Vague in, vague out. "Summarize this" gets a generic summary;
"Summarize this support ticket in two sentences for an engineer,
focusing on the reproduction steps" gets something usable.
GIVE CONTEXT. The model can't read your mind or your database. Say who the
audience is, what the goal is, and supply the facts it needs.
STATE THE FORMAT. Say exactly what shape you want back, and why, if it helps.
SAY WHAT TO DO, Positive instructions ("answer in plain language") generally work
NOT ONLY WHAT better than a list of prohibitions, though explicit "do not"
NOT TO DO. rules are still useful for hard constraints.
GIVE AN ESCAPE Tell the model what to do when it can't comply: "If the answer
HATCH. is not in the provided text, reply exactly: NOT_FOUND."
This is one of the most effective anti-hallucination measures.
BREAK DOWN HARD A complex task often works better as several smaller prompts
TASKS. (extract, then analyze, then format) than one giant instruction.
ASK FOR REASONING For problems needing multi-step logic, asking a standard model
WHERE IT HELPS. to work through the steps before giving a final answer often
improves accuracy. (Dedicated reasoning models do this internally,
and usually need less coaxing — follow your model's guidance.)
A weak prompt versus a strong one
WEAK:
Summarize this ticket.
STRONG:
You are a support triage assistant.
Summarize the ticket below in at most two sentences for an on-call engineer.
Include: the affected feature, the observed error, and any reproduction steps.
If a required detail is missing, write "unknown" for it. Do not speculate.
<ticket>
{ticket text}
</ticket>
Prompts are code: version, test, and review them
- Keep prompts in source control (files or resources), not scattered as inline strings.
- Use TEMPLATES with named placeholders rather than ad hoc string concatenation.
- Change ONE thing at a time, and measure the effect on a test set (Section 12) —
not on one example you happen to like.
- A model upgrade can change how a prompt behaves. Re-run your evaluations when you
switch models or versions.
- Treat prompt changes like code changes: review them, and be able to roll back.
3. System, User, and Assistant Messages
Chat models take a structured conversation, not a single string
system -> the developer's standing instructions: role, rules, tone, format, safety
boundaries. Sets the frame for the whole conversation.
user -> what the end user (or your application, on their behalf) is asking
assistant -> the model's own earlier replies, which you replay as history
tool -> results returned from functions the model asked you to run
(The mechanics of building these message lists in .NET are in this series' AI APIs with .NET guide, Section 4.)
What belongs in the system message
System message -> STABLE rules that apply to every turn:
identity ("You are the support assistant for ACME"),
scope ("Only answer questions about ACME products"),
tone, output format, what to do when unsure or out of scope,
and the handling of sensitive requests.
User message -> the VARIABLE part: the actual question, plus per-request data
(retrieved documents, the text to summarize).
Keeping stable instructions in the system message and variable data in the user message also helps with cost, since providers can often cache a repeated, unchanging prefix (see AI APIs with .NET, Section 12).
The system prompt is guidance, not a security boundary
Models generally give system instructions priority — but that is a TENDENCY, not an
enforced rule. Do NOT put anything in a system prompt that you can't afford to have
revealed, and do NOT rely on it alone to enforce security.
PROMPT INJECTION: text in the user input — or in a document, web page, or email you
feed the model — can contain instructions that try to override yours ("ignore previous
instructions and..."). This is a fundamental risk whenever untrusted text enters a prompt.
Defenses are layered, never a single trick:
- Clearly DELIMIT untrusted content and tell the model to treat it as data, not instructions.
- Give the model the LEAST privilege: only the tools and data the task needs.
- Enforce permissions and validation in CODE, outside the model.
- Require human confirmation for consequential actions.
- Don't put secrets (API keys, credentials, private data) in prompts.
Role discipline in code
- Never let user-supplied text be placed in the SYSTEM role. That hands the user
your highest-priority channel.
- Replay the assistant's previous turns faithfully; rewriting them can confuse the model.
- Keep tool results in the tool role, and remember they may contain untrusted text too.
4. Few-Shot Prompting
Teaching by example, inside the prompt
Zero-shot : instructions only, no examples. "Classify this ticket."
One-shot : instructions plus ONE example.
Few-shot : instructions plus SEVERAL examples (commonly 2-10) showing the
input -> output pattern you want.
Examples are often more effective than description at communicating format, tone, and edge-case behavior — showing the model what a good answer looks like is easier than explaining it.
Example: a few-shot classifier built as message pairs
(string Text, string Label)[] examples =
[
("I was charged twice this month.", "Billing"),
("The export button crashes the app.", "Bug"),
("Please add dark mode to the dashboard.", "Feature"),
("My invoice shows the wrong company name.", "Billing"),
];
var messages = new List<ChatMessage>
{
new SystemChatMessage("Classify the support ticket as exactly one of: Billing, Bug, Feature. Reply with the label only.")
};
// Each example becomes a user message followed by the assistant's ideal reply
foreach (var (text, label) in examples)
{
messages.Add(new UserChatMessage(text));
messages.Add(new AssistantChatMessage(label));
}
messages.Add(new UserChatMessage(newTicketText)); // the real input, answered in the same pattern
ChatCompletion completion = await client.CompleteChatAsync(messages);
Presenting examples as alternating user/assistant messages is a natural fit for chat models: the model sees what "a correct reply" looks like in the same format it will be asked to produce.
Choosing good examples
- DIVERSE: cover the different categories and the tricky edge cases, not five near-duplicates.
- REPRESENTATIVE: drawn from real inputs, matching what the model will actually see.
- CORRECT: a wrong example teaches the wrong thing — the model imitates its examples faithfully.
- CONSISTENT in format: if examples vary in structure, the output will too.
- BALANCED: if 4 of 5 examples share one label, the model may be biased toward it.
- ORDER can matter: the most recent examples sometimes carry extra weight, so don't
accidentally put the same label last every time.
Trade-offs and variations
Costs: every example consumes INPUT TOKENS on EVERY request. Ten long examples is
a real, recurring cost (Section 7). Use as few as achieve the quality you need.
Dynamic few-shot: instead of fixed examples, store a bank of labeled examples with
embeddings and, per request, retrieve the few most SIMILAR to the new input.
Often beats a fixed set for varied inputs, at the price of more machinery.
When NOT to bother: simple, well-specified tasks on a capable model may work fine
zero-shot — test before adding examples. Some reasoning-focused models also
respond better to clear instructions than to many examples; check your
model's guidance.
Beyond a point: if you need dozens of examples to get consistent behavior,
that's a signal to consider fine-tuning (Section 9) instead.
5. Structured Prompting
Giving the prompt itself clear structure
A long, unstructured prompt is hard for both humans and models to parse. Structured prompting means organizing the prompt into clearly labeled sections, using consistent delimiters, so the model can tell instructions, context, and input apart.
You are a contract-review assistant.
<instructions>
Review the clause below and identify risks for the CUSTOMER.
Respond only with the JSON described in <output_format>.
If the clause is not a contract clause, return {"risks": [], "note": "not a contract clause"}.
</instructions>
<context>
Customer is a small business. Jurisdiction: India. Focus on payment, termination, liability.
</context>
<clause>
{clause text}
</clause>
<output_format>
{"risks": [{"issue": string, "severity": "low" | "medium" | "high", "explanation": string}]}
</output_format>
Why structure helps
- SEPARATION: delimiters (XML-style tags, Markdown headers, triple quotes) clearly mark
where instructions end and untrusted data begins — which reduces confusion and gives
you a (partial) defense against prompt injection.
- REFERENCE: you can refer to sections by name ("using only the text in <context>").
- MAINTAINABILITY: each section can be edited, tested, and templated independently.
- CONSISTENCY: the same skeleton across many tasks makes prompts predictable to
write and review.
Structure the output, too
Prompted JSON ("reply in JSON") works most of the time — and fails in production when
a stray sentence or missing field breaks your parser. For anything your code consumes,
prefer the API's STRUCTURED OUTPUTS feature (schema-constrained JSON) over prompt
wording alone, then still validate the values (see AI APIs with .NET, Section 7).
Remember: a schema guarantees SHAPE, not TRUTH.
Prompt chaining: structure across multiple calls
Instead of one giant prompt, split a job into stages, each with a focused prompt, passing
output forward — and checking it between stages with ordinary code:
1. EXTRACT pull the key facts from the document (structured output)
2. VALIDATE check them in code against rules/databases
3. ANALYZE reason over the validated facts
4. DRAFT write the final response in the required tone
Benefits: each step is simpler to prompt, test, and debug; failures are easier to
locate; you can use a cheaper model for easy stages; and code can enforce rules
between steps. Cost: more calls, more latency, more plumbing.
Templates and safe assembly
// Build prompts from templates, not scattered string concatenation
const string Template = """
<instructions>{0}</instructions>
<document>
{1}
</document>
""";
string prompt = string.Format(Template, instructions, documentText);
One caution: if untrusted text can contain your own delimiter tags (for example a literal </document>), it can try to "close" the section early and smuggle in instructions. Strip or escape delimiter sequences from untrusted input, and don't treat delimiters as a security guarantee on their own.
6. Context Windows
The context window is the model's working memory — and it is finite
The context window is the maximum number of tokens the model can consider in a single request, and it includes everything: system prompt, conversation history, retrieved documents, tool definitions and results, your new question, and the reply it generates.
Context window >= system prompt
+ conversation history
+ retrieved documents / tool results
+ tool definitions
+ the user's new message
+ the model's OUTPUT
Window sizes vary by model and have grown dramatically, but they are never unlimited — and bigger isn't free or automatically better.
Bigger windows are not a license to stuff everything in
COST: you pay for every input token, on every request. Resending a 100-page
document with each question multiplies cost quickly.
LATENCY: more input tokens generally means slower responses.
QUALITY: models do not use all parts of a long context equally well. Research has shown
that information buried in the MIDDLE of a very long prompt can be used less
reliably than information near the beginning or end, and that adding lots of
irrelevant text can degrade answers. The effect varies by model and has
improved over time, but "more context" can still mean "worse answers."
GUIDELINE: include what's RELEVANT, not everything that's AVAILABLE.
Strategies for working within the window
RETRIEVAL (RAG) Store documents in chunks with embeddings; per question, retrieve only
the top few relevant chunks and include those. The standard approach
for "chat with your documents."
TRIMMING / SLIDING Keep only the most recent conversation turns (Section 10).
WINDOW
SUMMARIZATION Compress older context into a short summary and keep it plus recent turns.
SELECTIVE TOOL Return only the fields the model needs from a tool, not an entire record.
OUTPUT
PLACEMENT Put the most important instructions and the question where the model
attends well (typically at the start and/or restated at the end), and
keep reference material clearly delimited.
CHUNKING LONG For documents that exceed the window, process in pieces (map) and
INPUT combine the results (reduce).
Always leave room for the answer
If input fills the window, there's no space left for output and the reply is cut off
(FinishReason = Length). Budget explicitly:
available for input = context window - reserved output tokens - safety margin
Count tokens BEFORE sending (Section 7) and trim to that budget deliberately,
rather than discovering the limit via an error in production.
7. Tokens
The unit the model reads, writes, and bills in
Models don't process characters or words — they process tokens, chunks of text produced by a tokenizer. A token might be a whole common word, part of a rarer word, a punctuation mark, or a few characters.
Rule of thumb (English): 1 token ≈ 4 characters ≈ 3/4 of a word
"hamburger" -> may split into several tokens
" the" -> usually a single token
Numbers, code, and non-English text often use MORE tokens per unit of meaning.
Why tokens matter to application builders
1. COST Billing is per token, with INPUT and OUTPUT usually priced differently
(output typically costs more).
2. LIMITS The context window and rate limits (tokens per minute) are measured in tokens.
3. SPEED Generation time grows with the number of OUTPUT tokens.
4. QUIRKS Because models see tokens, not letters, tasks like "count the letters in
this word," exact character manipulation, and some arithmetic are
unexpectedly error-prone. Do those in code.
Practical habits
- Measure real usage from every response and log it.
- Count tokens locally BEFORE sending (e.g. with a tokenizer library) to enforce budgets.
- Remember that few-shot examples, long system prompts, tool definitions, and conversation
history are all re-sent — and re-billed — on every request.
- Set a maximum output token cap appropriate to the task.
- Be aware that the same text can tokenize differently across model families, so token
counts aren't directly comparable between providers.
The mechanics — counting tokens in .NET, logging usage, trimming history, and the cost levers — are covered in this series' AI APIs with .NET guide, Sections 8 and 12.
8. Hallucinations
Fluent, confident, and wrong
A hallucination is output that sounds plausible but is false or unsupported — an invented citation, a nonexistent API method, a made-up statistic, a fabricated quote. It's not a rare glitch; it's a direct consequence of how LLMs work: a model optimized to produce likely-sounding text will produce likely-sounding text whether or not it's true, and it has no built-in alarm for "I don't actually know this."
Common forms
- FABRICATED FACTS invented names, dates, statistics, quotes
- FAKE SOURCES citations, URLs, case law, or papers that don't exist
- INVENTED CODE/APIS plausible-looking methods, parameters, or packages that aren't real
- UNFAITHFUL SUMMARIES claims not actually present in the source document
- WRONG REASONING confident, well-formatted, but flawed logic or arithmetic
- STALE ANSWERS presenting out-of-date information as current
Why it's dangerous in applications
Hallucinations are hardest to catch precisely because they LOOK right. Users trust fluent,
confident text. In medicine, law, finance, security, and customer support, a confidently
wrong answer can cause real harm — so the design goal is not "hope it's accurate" but
"build a system where errors are made unlikely, visible, and survivable."
Mitigation: layers, not a single fix
1. GROUND THE MODEL IN SOURCE TEXT (RAG)
Retrieve relevant documents and instruct: "Answer ONLY from the provided context."
Far more reliable than relying on what the model memorized. This is the single
most effective measure for factual applications.
2. GIVE IT PERMISSION TO SAY "I DON'T KNOW"
"If the answer isn't in the context, say you don't know." Models often guess unless
explicitly allowed to decline.
3. REQUIRE CITATIONS AND CHECK THEM
Ask the model to quote or reference the supporting passage, then VERIFY in code that the
quote actually appears in the retrieved source. A citation you don't verify is decoration.
4. LOWER THE TEMPERATURE
For factual tasks, low randomness reduces (but doesn't eliminate) invention.
5. CONSTRAIN THE OUTPUT
Structured outputs, enumerated labels, and validation against known values
(does this product ID exist? is this a valid ISO currency code?) catch many errors in code.
6. USE TOOLS FOR FACTS AND MATH
Don't let the model recall prices, balances, or dates, or do arithmetic — have it call a
function that looks them up or computes them.
7. VERIFY HIGH-STAKES OUTPUT
Second-pass checks (a separate "verifier" prompt or model), rule-based checks, and — where
consequences are serious — a HUMAN in the loop before the answer is acted on.
8. COMMUNICATE UNCERTAINTY TO USERS
Show sources, label AI-generated content, and don't present output as authoritative.
9. MEASURE IT
Build an evaluation set with known answers and track a faithfulness / groundedness rate
over time (Section 12).
A grounded prompt pattern
Answer the question using ONLY the information in <context>.
Quote the sentence(s) that support your answer in a "sources" list.
If <context> does not contain the answer, reply exactly: "I don't have enough information to answer that."
Do not use outside knowledge.
<context>
{retrieved chunks}
</context>
<question>
{user question}
</question>
An honest bottom line
Hallucinations can be REDUCED substantially but not ELIMINATED. Even with retrieval, a model
can misread or misrepresent its sources. Design for it: validate, cite, verify, and keep
humans in the loop where being wrong is costly.
9. Model Selection
There is no "best model" — only the best model for your task, budget, and constraints
Providers offer many models spanning small/fast/cheap to large/capable/expensive, plus specialized variants (reasoning, coding, vision, embedding), and open-weight alternatives. Choosing well is an engineering decision, and it should be made with data, not hype.
The dimensions that matter
QUALITY How well does it perform on YOUR task? Public benchmarks are a rough guide;
your own evaluation set is the real test (Section 12).
LATENCY Time to first token and total generation time. Matters enormously for
interactive chat; barely at all for overnight batch jobs.
COST Price per input/output token. Multiply by your real traffic and prompt sizes.
Large price gaps exist between model tiers.
CONTEXT WINDOW Can it hold the documents/history your use case needs? (Section 6)
CAPABILITIES Tool/function calling, structured outputs, vision or audio input,
streaming, fine-tuning support, reasoning mode — only what you need.
RELIABILITY Instruction-following consistency, rate limits and quotas, and provider
uptime/support.
DATA & COMPLIANCE Where data is processed and stored, retention and training policies,
regional availability, certifications, private networking (a major
reason to use a managed cloud offering such as Azure OpenAI).
OPEN vs. CLOSED Open-weight models can be self-hosted (control, privacy, no per-token fee,
but you own the infrastructure and operations). Hosted proprietary models
are simpler to operate and often state-of-the-art, at a per-use price.
A practical selection process
1. DEFINE the task and what "good" means (accuracy target, latency budget, cost ceiling).
2. BUILD an evaluation set of realistic inputs with expected outputs or grading criteria.
3. START with a capable model to find out what's achievable.
4. TRY cheaper/smaller models on the same set — you may find a small model is good enough.
5. COMPARE on quality, latency, and cost together, not quality alone.
6. DECIDE, then re-evaluate periodically: models and prices change fast.
Common architecture patterns
MODEL ROUTING Send easy requests to a cheap model, hard ones to a stronger one, based on
task type, input complexity, or a quick classifier. Often the largest
cost saver.
CASCADE Try the cheap model first; escalate only if its output fails validation
or confidence checks.
RIGHT MODEL PER Use a small model for classification and routing, a stronger one for final
STAGE synthesis, a dedicated embedding model for retrieval.
SPECIALIZE LAST Prefer prompting, then retrieval (RAG), and only then fine-tuning:
prompting is cheapest to iterate; RAG gives fresh, citable knowledge;
fine-tuning is best for consistent style, format, or narrow behavior —
not for teaching the model new facts reliably.
Avoid painting yourself into a corner
- Keep model names, parameters, and prices in CONFIGURATION, not code.
- Put calls behind an interface (for example Microsoft.Extensions.AI's IChatClient)
so providers can be swapped.
- Plan for DEPRECATION: providers retire model versions on a schedule, so pin versions where
you need stability, track retirement notices, and re-run your evaluation set before migrating.
- Beware "prompt lock-in": prompts tuned to one model may need re-tuning on another.
10. Conversation Memory
The model remembers nothing — memory is something your application builds
Because every request is stateless, "memory" is your responsibility. There are two broad kinds, usually combined.
SHORT-TERM MEMORY The current conversation: the recent turns resent with each request.
LONG-TERM MEMORY Information that persists ACROSS conversations: user preferences, facts
about the user, past decisions, accumulated knowledge — stored outside
the model and retrieved when relevant.
Short-term strategies
FULL HISTORY Resend every turn. Simplest and most faithful — until cost and the
context window make it impractical.
SLIDING WINDOW Keep only the last N turns (or last N tokens). Simple and cheap, but the
model forgets anything older.
SUMMARIZATION Condense older turns into a running summary, and keep the summary plus
recent turns verbatim. Preserves the gist at bounded size, at the cost
of an extra model call and some detail loss.
HYBRID Summary of old turns + recent turns verbatim + retrieved long-term facts.
The usual production answer.
A summary-plus-recent-window memory in C
public record Turn(string Role, string Text);
public class ConversationMemory
{
private readonly List<Turn> _recent = new();
private readonly int _maxRecentTurns;
private string _summary = "";
public ConversationMemory(int maxRecentTurns = 8) => _maxRecentTurns = maxRecentTurns;
public void Add(string role, string text) => _recent.Add(new Turn(role, text));
// Fold the oldest turns into the running summary when the window is exceeded
public async Task CompactAsync(ChatClient client, CancellationToken ct = default)
{
if (_recent.Count <= _maxRecentTurns) return;
int toSummarize = _recent.Count - _maxRecentTurns;
string transcript = string.Join("\n", _recent.Take(toSummarize).Select(t => $"{t.Role}: {t.Text}"));
var result = await client.CompleteChatAsync(
[
new SystemChatMessage("Update the running summary of this conversation. Keep names, decisions, " +
"user preferences, and open questions. Maximum 150 words."),
new UserChatMessage($"Current summary:\n{_summary}\n\nNew turns:\n{transcript}")
],
cancellationToken: ct);
_summary = result.Value.Content[0].Text;
_recent.RemoveRange(0, toSummarize);
}
public List<ChatMessage> BuildMessages(string systemPrompt)
{
var messages = new List<ChatMessage> { new SystemChatMessage(systemPrompt) };
if (_summary.Length > 0)
messages.Add(new SystemChatMessage($"Summary of the earlier conversation:\n{_summary}"));
foreach (var turn in _recent)
messages.Add(turn.Role == "user" ? new UserChatMessage(turn.Text) : new AssistantChatMessage(turn.Text));
return messages;
}
}
Call CompactAsync after each exchange (or in the background), and BuildMessages when assembling each request. Persist the memory object (a database or cache keyed by conversation ID) so conversations survive restarts and scale across servers.
Long-term memory
Typical design:
1. EXTRACT After conversations, pull durable facts worth keeping
("prefers metric units", "works in the finance team", "decided to use PostgreSQL").
2. STORE Save them in a database, with an EMBEDDING for semantic lookup.
3. RETRIEVE On each new request, fetch the few memories relevant to the current
question and add them to the prompt.
4. MAINTAIN Update facts that change, merge duplicates, and expire stale ones.
This is retrieval-augmented generation applied to the user's own history. The vectors can live in a vector database or in a vector-capable index in a database you already run.
Memory raises real design and privacy questions
- WHAT to remember: store durable, useful facts — not every utterance.
- CONSENT AND TRANSPARENCY: tell users what is remembered, and let them view, correct,
and delete it. Many privacy regulations require this for personal data.
- SENSITIVE DATA: don't persist secrets, credentials, or unnecessary personal data.
Minimizing what you store minimizes what can leak.
- ISOLATION: memory must be strictly scoped per user/tenant — a retrieval bug that
surfaces one user's memories to another is a serious data breach.
- ACCURACY: summaries and extracted facts can themselves be wrong (hallucinated),
and a wrong memory persists and compounds. Prefer storing the user's own words,
and let users correct entries.
- INJECTION: anything written into memory from untrusted content can later re-enter
prompts — treat stored memories as untrusted data, not instructions.
11. AI Agents
From "answer a question" to "pursue a goal"
An AI agent is an LLM-driven system that decides what to do next: given a goal, it can plan, call tools, observe the results, and repeat until it judges the goal accomplished. Where a plain chatbot answers once, an agent runs a loop.
┌─────────────────────────────────────────────────┐
│ │
GOAL ─> THINK (what's the next step?) ─> ACT (call a tool)
^ │
│ v
└──────── OBSERVE (read the tool result) ─┘
│
goal met? ──> FINAL ANSWER
An agent is built from a small set of parts:
MODEL the "brain" that reasons and chooses actions
TOOLS functions it can call: search, query a database, call an API, run code, send email
MEMORY the running transcript (short-term) plus any retrieved long-term knowledge
LOOP the orchestration code that feeds tool results back until done
GUARDRAILS limits, permissions, and approvals that bound what it can do
Workflows versus agents: choose the simplest thing that works
WORKFLOW The steps are fixed IN YOUR CODE; the LLM fills in individual steps
(classify -> retrieve -> draft -> validate). Predictable, testable, cheaper.
AGENT The LLM chooses the steps dynamically at runtime. Flexible, handles open-ended
tasks — but less predictable, harder to test, slower, and more expensive.
Guidance: if you can describe the steps in advance, build a WORKFLOW. Reach for an agent
only when the path genuinely can't be known up front. Many "agent" use cases are better
served by a well-designed workflow, and the most reliable systems often combine both:
a workflow that invokes a bounded agent for one open-ended step.
A minimal agent loop in C
The tool-calling mechanics are covered in AI APIs with .NET, Section 6; an agent is that loop with deliberate limits around it.
public class SimpleAgent
{
private const int MaxSteps = 8;
private const int MaxTotalTokens = 50_000;
private readonly ChatClient _client;
private readonly ChatCompletionOptions _options; // includes the Tools the agent may use
private readonly Func<ChatToolCall, CancellationToken, Task<string>> _executeTool;
public SimpleAgent(ChatClient client, ChatCompletionOptions options,
Func<ChatToolCall, CancellationToken, Task<string>> executeTool)
=> (_client, _options, _executeTool) = (client, options, executeTool);
public async Task<string> RunAsync(string goal, string systemPrompt, CancellationToken ct)
{
var messages = new List<ChatMessage> { new SystemChatMessage(systemPrompt), new UserChatMessage(goal) };
int tokensUsed = 0;
for (int step = 0; step < MaxSteps; step++) // HARD cap on iterations
{
ChatCompletion completion = await _client.CompleteChatAsync(messages, _options, ct);
tokensUsed += completion.Usage.TotalTokenCount;
if (tokensUsed > MaxTotalTokens) // HARD cap on spend
return "Stopped: token budget exceeded.";
if (completion.FinishReason == ChatFinishReason.Stop)
return completion.Content[0].Text; // the agent decided it is done
if (completion.FinishReason != ChatFinishReason.ToolCalls)
return $"Stopped: {completion.FinishReason}";
messages.Add(new AssistantChatMessage(completion)); // keep the tool request in history
foreach (ChatToolCall call in completion.ToolCalls)
{
string result;
try { result = await _executeTool(call, ct); } // validation, authorization, and approval live HERE
catch (Exception ex) { result = $"Tool error: {ex.Message}"; } // errors go back as text so it can adapt
messages.Add(new ToolChatMessage(call.Id, result));
}
}
return "Stopped: step limit reached.";
}
}
Designing tools well
- FEW, FOCUSED tools with clear names and descriptions beat many overlapping ones. The
description is how the model decides when to use it — write it like documentation.
- Return SMALL, relevant results. Everything returned re-enters the context window.
- Return helpful ERROR messages as text ("Customer 123 not found. Try searching by email.")
so the agent can recover instead of looping blindly.
- Make tools IDEMPOTENT where possible, so a retry or repeated call can't double-charge
or double-send.
- Separate READ tools (safe) from WRITE/ACT tools (consequential), and gate the latter.
Guardrails: an agent amplifies both capability and risk
BOUND THE LOOP Maximum steps, maximum tokens, maximum wall-clock time.
LEAST PRIVILEGE Give the agent only the tools and data the task requires, scoped to the
current user's permissions — enforced in CODE, not in the prompt.
VALIDATE EVERY CALL Tool arguments come from a model steered by input that may be hostile.
Validate them like any untrusted input.
HUMAN APPROVAL Require confirmation before irreversible or high-impact actions
(sending money, deleting data, emailing customers, deploying code).
SANDBOX Run code-execution tools in an isolated environment.
PROMPT INJECTION An agent that reads web pages, emails, or documents can be hijacked by
instructions hidden in them. Treat all tool output as untrusted data,
and never let it expand the agent's privileges.
OBSERVABILITY Log every step — prompts, tool calls, results, token usage — so you
can debug, audit, and replay what the agent did and why.
Common failure modes
- LOOPS: repeating the same failing action. (Caps, and error messages that suggest alternatives.)
- COMPOUNDING ERRORS: a small early mistake cascades through later steps.
- WRONG TOOL / WRONG ARGUMENTS: especially with vague tool descriptions.
- CONTEXT BLOAT: the transcript grows with every step, raising cost and degrading quality.
- COST AND LATENCY: many sequential model calls add up fast.
- OVER-AUTONOMY: acting confidently on a misunderstanding.
Multi-agent systems
Complex problems are sometimes split among several specialized agents (a researcher, a coder,
a reviewer) coordinated by an orchestrator. This can help with separation of concerns and
parallel work, but adds cost, latency, and new failure modes (miscommunication, loops between
agents). Start with ONE agent or a plain workflow, and add agents only when a single one
demonstrably struggles.
Frameworks
In .NET, Microsoft.Extensions.AI provides IChatClient abstractions and automatic function-invocation middleware, and Microsoft maintains higher-level orchestration frameworks (such as Semantic Kernel and the Microsoft Agent Framework) that package agent loops, memory, and multi-agent patterns. These evolve quickly, so check their current documentation. Whatever you adopt, understand the underlying loop above — it's what the frameworks implement, and it is what you'll be debugging.
12. Evaluating LLM Applications
"It seemed to work when I tried it" is not evidence
Because outputs vary and failures are subtle, LLM applications need systematic evaluation — the equivalent of a test suite, adapted to probabilistic output.
BUILD AN EVALUATION SET A collection of realistic inputs with expected outputs or grading
criteria — including edge cases, adversarial inputs, and examples
of past failures. Start with a few dozen; grow it with every bug.
CHOOSE METRICS Exact match (classification, extraction), schema-validity rate,
groundedness/faithfulness to sources, task success rate, latency,
cost per request, refusal/escalation rate.
GRADING METHODS - Programmatic checks (exact match, regex, JSON validation, unit tests
for generated code) — cheapest and most reliable where possible.
- Model-as-judge: another LLM grades outputs against criteria — scalable
but imperfect and biased; calibrate it against human judgments.
- Human review — the gold standard; use it for calibration and
for high-stakes samples.
RUN IT CONTINUOUSLY Re-run on every prompt change, model change, or retrieval change, so
regressions show up before users find them. Compare versions side by side.
MONITOR IN PRODUCTION Log inputs, outputs, tool calls, token usage, and user feedback
(thumbs up/down). Sample real traffic back into your evaluation set.
(Mind privacy when logging user content.)
Evaluation is also what makes every decision in this guide testable instead of a matter of opinion: whether few-shot examples help, whether a cheaper model is good enough, whether a prompt change reduced hallucinations, whether the agent actually completes tasks.
13. Putting It Together: Anatomy of an LLM Application
A retrieval-augmented support assistant
USER QUESTION
│
v
[1] INPUT GUARDS length cap, rate limit, basic abuse checks (code)
│
v
[2] RETRIEVE embed the question -> fetch top-k relevant chunks
│ + long-term memory facts for this user (embeddings + store)
v
[3] ASSEMBLE PROMPT system rules + <context> chunks + conversation
│ memory (summary + recent turns) + the question (template)
v
[4] CALL MODEL right-sized model, low temperature, token cap,
│ structured output, retries/timeouts (LLM API)
v
[5] VALIDATE schema check, verify citations appear in sources,
│ business rules, refusal/"don't know" handling (code)
v
[6] RESPOND + LOG stream/return answer with sources; log usage,
latency, feedback; update conversation memory (code)
Every section of this guide appears in that pipeline: prompt engineering and structured prompting shape step 3; tokens and context windows govern what fits; the message roles organize it; model selection and parameters configure step 4; hallucination mitigation spans steps 2, 3, and 5; conversation memory feeds steps 2 and 3; and an agent is what you get when step 4 can request tools and loop. Evaluation wraps the whole thing.
14. Common Pitfalls
| Pitfall | Why it hurts | Better approach |
|---|---|---|
| Treating the model as a database of facts | It generates plausible text, not verified facts, and has a knowledge cutoff | Ground answers in retrieved sources; use tools for facts and math (Section 8) |
| Vague prompts | Vague input produces generic, inconsistent output | Specify role, task, constraints, and output format (Section 2) |
| No "I don't know" path | Models guess unless told they may decline | Give an explicit escape hatch and required fallback wording (Sections 2, 8) |
| Putting user text in the system message | Hands users your highest-priority instruction channel | Keep system messages developer-controlled; put user data in user messages (Section 3) |
| Trusting the system prompt for security | Instruction priority is a tendency, not an enforced rule; prompts can leak | Enforce permissions and validation in code; never put secrets in prompts (Section 3) |
| Ignoring prompt injection from documents/web/tool output | Hidden instructions in untrusted content can hijack the model or agent | Delimit untrusted data, least privilege, human approval for consequential actions (Sections 3, 11) |
| Wrong or unbalanced few-shot examples | The model imitates its examples faithfully, including their errors and biases | Use diverse, correct, balanced, representative examples; test with and without (Section 4) |
| Relying on "reply in JSON" wording alone | Free-text JSON eventually breaks the parser | Use structured outputs, then still validate values (Section 5) |
| Stuffing the whole document into every request | High cost, slow responses, and often worse answers | Retrieve only the relevant chunks (Section 6) |
| Forgetting to reserve output space | A full context window truncates the reply | Budget input as window minus reserved output minus margin (Section 6) |
| Using the model for exact arithmetic, counting, or lookups | Token-based models are error-prone at these | Do them in code; expose them as tools (Sections 7, 11) |
| Unverified citations | Models invent plausible-looking sources | Check in code that quoted passages exist in the retrieved source (Section 8) |
| Choosing a model by hype or a single benchmark | Benchmarks may not reflect your task, cost, or latency needs | Build an evaluation set; compare quality, latency, and cost together (Section 9) |
| Hard-coding one model and one provider | Deprecations, price changes, and better options force rewrites | Configuration and an abstraction layer; plan for version retirement (Section 9) |
| Unbounded conversation history | Cost and latency grow every turn until the window overflows | Sliding window, summarization, or retrieval-backed memory (Section 10) |
| Storing everything users say | Privacy risk, accumulating errors, and noise | Store durable facts, give users control, scope strictly per user (Section 10) |
| Building an agent when a workflow would do | Unpredictable, slower, costlier, and harder to test | Use fixed workflows for known steps; agents only for genuinely open-ended work (Section 11) |
| Agents with no limits | Runaway loops, runaway cost, and unbounded blast radius | Step, token, and time caps; least privilege; approval gates; full logging (Section 11) |
| Shipping with no evaluation | Regressions and hallucinations are discovered by users | Maintain an evaluation set; re-run on every change; monitor production (Section 12) |
Quick Reference Table
| Concept | Key idea | Practical takeaway |
|---|---|---|
| LLM | Predicts the next token; fluent, probabilistic, stateless | Use it for language; use code for math, lookups, rules |
| Prompt engineering | Specific role, task, constraints, format | Treat prompts as versioned, tested code |
| System message | Developer's stable instructions | Guidance, not a security boundary |
| User message | The variable request and per-request data | Never let user text into the system role |
| Assistant message | The model's earlier replies | Replayed as history; the model is stateless |
| Few-shot prompting | Examples inside the prompt | Diverse, correct, balanced; costs tokens every request |
| Structured prompting | Delimited, labeled prompt sections | Separates instructions from data; supports chaining |
| Structured output | Schema-constrained response | Guarantees shape, not truth — still validate |
| Context window | Max tokens per request (input + output) | Include what's relevant, not everything available |
| Tokens | The unit of cost, limits, and speed | Count before sending; cap output; log usage |
| Hallucination | Plausible but false output | Ground in sources, allow "I don't know", verify citations |
| RAG | Retrieve relevant chunks, then generate | The primary defense against hallucination and stale knowledge |
| Model selection | Quality vs. latency vs. cost vs. constraints | Evaluate on your own task; route by difficulty |
| Short-term memory | Recent turns resent each call | Sliding window, summarization, or hybrid |
| Long-term memory | Facts persisted across conversations | Extract, store, retrieve — with privacy and per-user isolation |
| Workflow | Fixed steps, LLM fills each | Prefer when the path is known |
| AI agent | LLM loop: think, act with tools, observe | Bound it, least-privilege it, require approval for risky actions |
| Evaluation | A test suite for probabilistic output | The only way to know a change helped |
Conclusion
Building reliable applications on LLMs comes down to a handful of recurring ideas. The model is a powerful but probabilistic text predictor — stateless, bounded by a context window, billed in tokens, and prone to producing fluent falsehoods — so your job is to surround it with good engineering: precise, structured, versioned prompts; the right context supplied through retrieval rather than hope; constrained, validated outputs; a memory layer you design deliberately; and strict, code-enforced limits on what it is allowed to do. The roles in a chat (system, user, assistant, tool) give you a way to organize that, few-shot examples and structured prompts give you control over behavior and format, and hallucination mitigation is a layered discipline of grounding, citing, verifying, and keeping humans in the loop where mistakes are costly.
Two principles tie it together. First, use the simplest architecture that works: a single well-crafted prompt before a chain, a fixed workflow before an agent, a small model before a large one, retrieval before fine-tuning. Each step up the ladder buys flexibility at the price of cost, latency, and unpredictability. Second, measure instead of guessing: an evaluation set turns every question in this guide — does this example help, is this cheaper model good enough, did that prompt change reduce hallucinations, does the agent really finish the task — into something you can answer with data. Teams that do both ship LLM features that people can trust; teams that skip them ship impressive demos that fail in production.
Found this useful? Feel free to star the repo, open an issue with corrections, or share the "it sounded completely right and was completely wrong" story that made the case for grounding and evaluation click better than any theory.
Top comments (1)
Some comments may only be visible to logged-in visitors. Sign in to view all comments. Some comments have been hidden by the post's author - find out more