<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Blessed Josiah</title>
    <description>The latest articles on DEV Community by Blessed Josiah (@joxiahdev).</description>
    <link>https://dev.to/joxiahdev</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F264588%2F0c86a935-e50c-4218-91db-02f193676739.jpeg</url>
      <title>DEV Community: Blessed Josiah</title>
      <link>https://dev.to/joxiahdev</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/joxiahdev"/>
    <language>en</language>
    <item>
      <title>Fine-Tune, Deploy and Use LLM As AI Agent</title>
      <dc:creator>Blessed Josiah</dc:creator>
      <pubDate>Tue, 22 Sep 2026 08:10:00 +0000</pubDate>
      <link>https://dev.to/joxiahdev/fine-tune-deploy-and-use-llm-as-ai-agent-42cn</link>
      <guid>https://dev.to/joxiahdev/fine-tune-deploy-and-use-llm-as-ai-agent-42cn</guid>
      <description>&lt;p&gt;In this video I continue the fine-tuning series on my &lt;a href="https://www.youtube.com/@joxiahdev" rel="noopener noreferrer"&gt;channel&lt;/a&gt;. This time I go through the whole pipeline, not just the training part: renting GPUs on &lt;a href="https://runpod.io?ref=3cfpyagl" rel="noopener noreferrer"&gt;Runpod&lt;/a&gt;, fine-tuning a model with &lt;a href="https://unsloth.ai/docs/new/studio/install" rel="noopener noreferrer"&gt;Unsloth Studio&lt;/a&gt;, deploying it as an inference endpoint using Runpod's Serverless, and then actually using that endpoint inside a &lt;a href="https://pydantic.dev/docs/ai/overview/" rel="noopener noreferrer"&gt;Pydantic AI&lt;/a&gt; Agent.&lt;/p&gt;

&lt;p&gt;This is meant to cover the full path: rent the GPU, train the model, deploy it, and get it hooked up to something that can actually call it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Setting Up the Pod
&lt;/h2&gt;

&lt;p&gt;I go in-depth on creating a Pod (a dedicated GPU instance container), picking a GPU, setting up storage, and getting SSH access working. The whole setup also works through the accompanying Jupyter Notebook, but I show the terminal option too.&lt;/p&gt;

&lt;p&gt;One thing I didn't call out clearly enough in the video: run &lt;code&gt;apt update &amp;amp;&amp;amp; apt upgrade -y&lt;/code&gt; right after you SSH in, before installing anything else. A fresh Pod's package index is often out of date, so skipping this can mean installs failing or pulling older versions of tools than you'd expect.&lt;/p&gt;

&lt;h3&gt;
  
  
  Storage Options You Should Know About
&lt;/h3&gt;

&lt;p&gt;Before you get into training, it's worth understanding how storage works on a Pod, because it's easy to get caught out. You've got 3 options:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;container disk&lt;/strong&gt; — files disposed when pod stops&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;volume disk&lt;/strong&gt; — files persist until pod terminated&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;network volume&lt;/strong&gt; — files persist beyond Pod's termination&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Volume disk is usually what's mounted on your &lt;code&gt;/workspace&lt;/code&gt; directory.&lt;/p&gt;

&lt;h2&gt;
  
  
  A Gotcha With Unsloth Studio's Install Path
&lt;/h2&gt;

&lt;p&gt;If you install Unsloth Studio from Jupyter Notebook in the &lt;code&gt;/workspace&lt;/code&gt; directory without setting &lt;code&gt;UNSLOTH_STUDIO_HOME&lt;/code&gt; first, like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;UNSLOTH_STUDIO_HOME&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;/workspace
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Unsloth Studio will quietly ignore &lt;code&gt;/workspace&lt;/code&gt; and fall back to &lt;code&gt;/root/.unsloth/studio&lt;/code&gt; instead, which is outside the volume disk and sitting on the container disk. You won't notice anything's wrong until you stop the pod to save some money, come back the next day, and your training setup is gone. Ask me how I know!&lt;/p&gt;

&lt;h2&gt;
  
  
  Keeping Runpod Costs Down
&lt;/h2&gt;

&lt;p&gt;I didn't go in-depth on optimizing Runpod usage in the video itself, but a couple of things are worth knowing:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Runpod charges you for renting the pod and GPUs the whole time it's running, so it's worth preparing your dataset before you spin the pod up, rather than doing that work while the GPU meter's running.&lt;/li&gt;
&lt;li&gt;For Runpod's serverless endpoint, you can reduce cold start by keeping at least one GPU worker active, but that comes at a cost, so just be aware of the trade-off.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  From Model to Agent
&lt;/h2&gt;

&lt;p&gt;Once the model was trained and deployed, the last piece was actually using it. I set up a Pydantic AI &lt;a href="https://colab.research.google.com/drive/1ZXlspobgVa_lRUcDKlizOnuPg_fp7hg_?usp=sharing" rel="noopener noreferrer"&gt;agent&lt;/a&gt; that calls the fine-tuned model through the endpoint. I went with Pydantic AI's OpenAI provider for this, since vLLM (the inference framework running behind the endpoint) supports the OpenAI response format.&lt;/p&gt;

&lt;h2&gt;
  
  
  A Quick Note on Fine-Tuning vs. RAG
&lt;/h2&gt;

&lt;p&gt;Fine-tuning is best when you're trying to modify the behavior of a model, like response style, tone etc. You can also teach it new knowledge, but even after fine-tuning, models can still hallucinate or fall back on old data. For anything that needs to stay accurate and up to date, pairing your fine-tuned model with Retrieval Augmented Generation (RAG) is still the safer route, since RAG pulls in fresh, sourced info at request time instead of relying on whatever got baked in during training.&lt;/p&gt;

&lt;p&gt;If you enjoyed this, feel free to subscribe to the &lt;a href="https://www.youtube.com/@joxiahdev" rel="noopener noreferrer"&gt;channel&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Thanks and happy coding!&lt;/p&gt;

</description>
      <category>llm</category>
      <category>ai</category>
      <category>finetuning</category>
      <category>agents</category>
    </item>
    <item>
      <title>How to Fine-Tune an LLM Locally with Unsloth Studio</title>
      <dc:creator>Blessed Josiah</dc:creator>
      <pubDate>Mon, 14 Sep 2026 08:00:00 +0000</pubDate>
      <link>https://dev.to/joxiahdev/how-to-fine-tune-an-llm-with-unsloth-studio-1kno</link>
      <guid>https://dev.to/joxiahdev/how-to-fine-tune-an-llm-with-unsloth-studio-1kno</guid>
      <description>&lt;p&gt;Fine-tuning is the process of further training a language model on new data so it learns new information, behaviors, or style — updating the model's own weights, rather than just showing it information at the moment you ask a question.&lt;/p&gt;

&lt;p&gt;There are two common ways to get an LLM to work with new information: RAG and fine-tuning. RAG (Retrieval-Augmented Generation) retrieves relevant information and feeds it to the model as context at the time of the query. Fine-tuning instead trains the model on new examples so that knowledge becomes part of the model itself.&lt;/p&gt;

&lt;p&gt;In this video, I start a new series on fine-tuning, where I go over everything you need to know to fine-tune a model using &lt;a href="https://unsloth.ai/docs/new/studio/install" rel="noopener noreferrer"&gt;Unsloth Studio&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prepare Dataset
&lt;/h2&gt;

&lt;p&gt;First, we need to source our data. For this video, I manually created a PDF with facts about the just-concluded 2026 FIFA World Cup. You could also use the &lt;a href="https://en.wikipedia.org/wiki/2026_FIFA_World_Cup" rel="noopener noreferrer"&gt;Wikipedia page&lt;/a&gt;, but be mindful of noise in that data — you'll need to manually clean it so you don't end up training your model on noise instead of facts.&lt;/p&gt;

&lt;p&gt;After cleaning your PDF, you can head to Claude or ChatGPT and ask for a multi-turn conversational dataset in the ChatML format:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"messages"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"role"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"system"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"content"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"role"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"user"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"content"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"role"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"assistant"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"content"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"messages"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"role"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"system"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"content"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"role"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"user"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"content"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"role"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"assistant"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"content"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Or take advantage of Unsloth Studio's Data Recipes and generate a ChatML dataset right from the application. I explain how to do this in depth in the video — the prompt and response schema I used are below.&lt;/p&gt;

&lt;h3&gt;
  
  
  Response Schema (JSON Array)
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"array"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"minItems"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"items"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"object"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"properties"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"role"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"string"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"enum"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"system"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"user"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"assistant"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"content"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"string"&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"required"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"role"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"content"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Prompt
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Generate a conversation from the chunk below: a system message, then exactly two user/assistant exchanges.

Start with a system message setting the assistant's role as a helpful assistant knowledgeable about the 2026 FIFA World Cup. Then the user asks a factual question about something in the chunk, and the assistant answers accurately. The user then asks a follow-up question that builds on the assistant's answer — asking for more detail, a related fact, or clarification — and the assistant answers that too.

Rules:
- Never invent facts not in the chunk. If the chunk is too narrow for a genuine follow-up, find a second distinct fact from the same chunk to ask about instead — do not fabricate details to fill the follow-up.
- Answers must be self-contained (no "the text says...").
- Use full names, not pronouns.
- After the system message, roles must alternate: user, assistant, user, assistant.

Output only a JSON array matching this schema — no wrapper object, no extra text:
[{"role": "system", "content": "..."}, {"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}, {"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]

Chunk:
'{{chunk_text}}'
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I used the &lt;code&gt;unsloth/Qwen2.5-Coder-3B-Instruct-GGUF&lt;/code&gt; model for this step, since it reliably returned a valid JSON array. I tried the &lt;code&gt;unsloth/gemma-4-E2B-it-GGUF&lt;/code&gt; model first, but it couldn't handle the JSON array output correctly.&lt;/p&gt;

&lt;h2&gt;
  
  
  Base Model
&lt;/h2&gt;

&lt;p&gt;I used the &lt;code&gt;unsloth/Llama-3.1-8B-Instruct&lt;/code&gt; model for two reasons:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;It needs about &lt;code&gt;8.6GB&lt;/code&gt; of GPU memory, which fits comfortably within my machine's 32GB of unified memory (which serves as VRAM on Apple Silicon).&lt;/li&gt;
&lt;li&gt;It's built for conversational chat, since it ships with a chat template out of the box.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;I initially tried the base version, &lt;code&gt;unsloth/Llama-3.1-8B&lt;/code&gt;, and ran into this error:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2r2rifjwkph2o25z6iz8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2r2rifjwkph2o25z6iz8.png" alt="Chat template error page from my training" width="800" height="483"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://unsloth.ai/docs/get-started/fine-tuning-llms-guide/what-model-should-i-use#instruct-or-base-model" rel="noopener noreferrer"&gt;This page&lt;/a&gt; from Unsloth's documentation explains why: base models don't come with a chat template at all, while Instruct models do.&lt;/p&gt;

&lt;p&gt;You can still fine-tune a base model for conversational chat, but you'll need to manually create a &lt;a href="https://unsloth.ai/docs/basics/chat-templates#applying-chat-templates-with-unsloth" rel="noopener noreferrer"&gt;chat template&lt;/a&gt; for it first. Unsloth Studio doesn't yet support this out of the box — there's an open &lt;a href="https://github.com/unslothai/unsloth/issues/5755" rel="noopener noreferrer"&gt;GitHub issue&lt;/a&gt; tracking it as a feature request, so hopefully it lands soon.&lt;/p&gt;

&lt;h2&gt;
  
  
  Hyperparameters
&lt;/h2&gt;

&lt;p&gt;For hyperparameter configuration, check Unsloth's &lt;a href="https://unsloth.ai/docs/get-started/fine-tuning-llms-guide/lora-hyperparameters-guide" rel="noopener noreferrer"&gt;LoRA Hyperparameters Guide&lt;/a&gt; for guidance on getting optimal values for your training run.&lt;/p&gt;

&lt;p&gt;One of the most important settings is &lt;code&gt;eval_steps&lt;/code&gt;. Without it, evaluation never runs alongside training. I set mine to &lt;code&gt;0.1&lt;/code&gt; in the video, meaning evaluation runs every 10% of the way through training. This gives you an &lt;code&gt;eval loss&lt;/code&gt; graph, which shows how your model is actually performing — not just how well it fits its own training data.&lt;/p&gt;

&lt;p&gt;Here's how to read it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;If the graph stays flat and never comes down, that's &lt;strong&gt;underfitting&lt;/strong&gt; — the model isn't learning the patterns in your dataset.&lt;/li&gt;
&lt;li&gt;If it comes down and then goes back up, that's &lt;strong&gt;overfitting&lt;/strong&gt; — the model has stopped generalizing and started memorizing the training data instead.&lt;/li&gt;
&lt;li&gt;What you want is the &lt;code&gt;eval loss&lt;/code&gt; graph trending smoothly downward, alongside the &lt;code&gt;training loss&lt;/code&gt; graph.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you found this useful, drop the video a like and subscribe to the &lt;a href="https://www.youtube.com/@joxiahdev" rel="noopener noreferrer"&gt;channel&lt;/a&gt; for more content.&lt;/p&gt;

&lt;p&gt;Thanks, happy coding!&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
    </item>
    <item>
      <title>How To Build A Production-Ready Design System With Claude Code, Mobbin, And Penpot</title>
      <dc:creator>Blessed Josiah</dc:creator>
      <pubDate>Mon, 31 Aug 2026 08:00:00 +0000</pubDate>
      <link>https://dev.to/joxiahdev/how-i-built-a-production-ready-design-system-with-claude-code-mobbin-and-penpot-40il</link>
      <guid>https://dev.to/joxiahdev/how-i-built-a-production-ready-design-system-with-claude-code-mobbin-and-penpot-40il</guid>
      <description>&lt;p&gt;I started a new series on my &lt;a href="https://www.youtube.com/@joxiahdev" rel="noopener noreferrer"&gt;channel&lt;/a&gt; on &lt;a href="https://code.claude.com/docs/en/overview" rel="noopener noreferrer"&gt;Claude Code&lt;/a&gt; to help you understand and use it confidently, and in this video, I show you how to build a production-ready design system.&lt;/p&gt;

&lt;p&gt;The workflow uses &lt;a href="https://mobbin.com" rel="noopener noreferrer"&gt;Mobbin&lt;/a&gt;, a platform with a huge library of real application screenshots, to research how real, shipped products solve the same design problems. Claude Code takes those references and connects to &lt;a href="https://penpot.app/" rel="noopener noreferrer"&gt;Penpot&lt;/a&gt; — a free, open-source design platform you can self-host — to turn that research into an actual, working design system.&lt;/p&gt;

&lt;p&gt;Using a collection of &lt;a href="https://github.com/bjoxiah/buildloop" rel="noopener noreferrer"&gt;skills&lt;/a&gt; I built for my workflow, the &lt;code&gt;/design&lt;/code&gt; skill returns two files when invoked: &lt;code&gt;DESIGN.md&lt;/code&gt; and &lt;code&gt;tokens.json&lt;/code&gt;. Claude Code derives these tokens from its research, then uses them to build consistent, traceable designs directly in Penpot.&lt;/p&gt;

&lt;p&gt;The Mobbin MCP server does require a paid plan. If you'd rather skip that, you can point Claude Code at your own collection of reference screenshots instead, and still have it connect to Penpot to build the designs from there.&lt;/p&gt;

&lt;p&gt;If you need a quick brush-up on Claude Code itself, check out my &lt;a href="https://www.youtube.com/watch?v=NbIGQbNF-Ks" rel="noopener noreferrer"&gt;previous video&lt;/a&gt; in the Mastering Claude Code Series.&lt;/p&gt;

&lt;p&gt;While you're there, consider subscribing to the &lt;a href="https://www.youtube.com/@joxiahdev" rel="noopener noreferrer"&gt;channel&lt;/a&gt;&lt;/p&gt;

</description>
      <category>claudecode</category>
      <category>uiux</category>
      <category>ai</category>
      <category>designsystem</category>
    </item>
    <item>
      <title>How To Evaluate An AI Agent</title>
      <dc:creator>Blessed Josiah</dc:creator>
      <pubDate>Tue, 04 Aug 2026 08:00:00 +0000</pubDate>
      <link>https://dev.to/joxiahdev/how-to-evaluate-an-ai-agent-1oel</link>
      <guid>https://dev.to/joxiahdev/how-to-evaluate-an-ai-agent-1oel</guid>
      <description>&lt;p&gt;In classical test-driven development, we deal with deterministic outcomes. We write assertions against values and types we already know in advance — &lt;code&gt;assert result == expected&lt;/code&gt;, and we move on.&lt;/p&gt;

&lt;p&gt;AI agents don't play by those rules. Because the underlying LLM is non-deterministic, giving it the same input twice can produce two differently-worded outputs. That alone breaks the classical assertion model.&lt;/p&gt;

&lt;p&gt;But there's a second problem, beyond just consistency. Say we have a customer support agent, and we want it to be helpful, empathetic, professional, and friendly when it talks to customers. Even if the output &lt;em&gt;were&lt;/em&gt; consistent, how do we test for qualities like that? There's no fixed value or type to assert against — "empathetic" isn't a type or an exact string, it's a subjective judgment call.&lt;/p&gt;

&lt;p&gt;So we're dealing with two distinct problems: unpredictable output, and qualities that are inherently subjective. Testing an AI agent means solving for both.&lt;/p&gt;

&lt;h2&gt;
  
  
  Enter Pydantic Evals
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://pydantic.dev/docs/ai/evals/evals/" rel="noopener noreferrer"&gt;Pydantic Evals&lt;/a&gt; is Pydantic AI's recommended way to test the output of your AI agents. It's an evaluation library that gives you a set of evaluators purpose-built for exactly this kind of testing — some for the deterministic parts of your agent's behavior, and some for the parts that require judgment.&lt;/p&gt;

&lt;h2&gt;
  
  
  Testing Deterministic Behavior: IsInstance
&lt;/h2&gt;

&lt;p&gt;Not everything about an AI agent is unpredictable. If your agent is set up with a structured output type, you can — and should — still test that the shape of the response is correct, even if the &lt;em&gt;content&lt;/em&gt; varies.&lt;/p&gt;

&lt;p&gt;For example, say we have an agent that returns a structured output type:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pydantic_ai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Agent&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;dataclasses&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;dataclass&lt;/span&gt;

&lt;span class="c1"&gt;# Response Model
&lt;/span&gt;&lt;span class="nd"&gt;@dataclass&lt;/span&gt;
&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Response&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;
    &lt;span class="n"&gt;sentiment&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;

&lt;span class="c1"&gt;# Agent Setup
&lt;/span&gt;&lt;span class="n"&gt;support_agent&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Agent&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;openai/gpt-5.5&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;customer_support_agent&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;instructions&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;
        You are a customer support agent for an online store.
    &lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;output_type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;Response&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;We can define a test case that checks the output is an instance of that &lt;code&gt;Response&lt;/code&gt; type, using the &lt;code&gt;IsInstance&lt;/code&gt; evaluator:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pydantic_evals&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Case&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pydantic_evals.evaluators&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;IsInstance&lt;/span&gt;

&lt;span class="n"&gt;test_case&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Case&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;check_agent_response&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;inputs&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;where is my order?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;evaluators&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="nc"&gt;IsInstance&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;type_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Response&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is a classic assertion in spirit — we know exactly what type we expect back, so we test for it directly. No judgment call required.&lt;/p&gt;

&lt;h2&gt;
  
  
  Testing Non-Deterministic Behavior: LLMJudge
&lt;/h2&gt;

&lt;p&gt;Structural checks like &lt;code&gt;IsInstance&lt;/code&gt; don't help us with the harder problem: was the response actually &lt;em&gt;good&lt;/em&gt;? Was it empathetic? Professional? Did it address the customer's question?&lt;/p&gt;

&lt;p&gt;This is where &lt;code&gt;LLMJudge&lt;/code&gt; comes in. It uses an LLM as a judge to evaluate your agent's output against a set of criteria you define — scoring qualitative aspects of the response that a type check could never catch.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pydantic_evals.evaluators&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;LLMJudge&lt;/span&gt;

&lt;span class="n"&gt;quality_case&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Case&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;check_response_quality&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;inputs&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;where is my order?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;evaluators&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="nc"&gt;LLMJudge&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;rubric&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Response should be empathetic, professional, and directly address the customer&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s question about their order status.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Under the hood, &lt;code&gt;LLMJudge&lt;/code&gt; sends your agent's output and your rubric to a judge model, which returns a score (and often a reason) reflecting how well the response meets your criteria. Rather than asserting on an exact value, you're asserting on a &lt;em&gt;standard&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;Together, &lt;code&gt;IsInstance&lt;/code&gt; and &lt;code&gt;LLMJudge&lt;/code&gt; cover both problems from earlier: structural correctness for the deterministic parts, and rubric-based scoring for everything else.&lt;/p&gt;

&lt;h2&gt;
  
  
  Watch the Full Walkthrough
&lt;/h2&gt;

&lt;p&gt;I go deeper into &lt;code&gt;LLMJudge&lt;/code&gt; — including setting up rubrics, reading scores, and combining multiple evaluators in a single test suite — in this video, part of the Master Pydantic AI series on my &lt;a href="https://youtube.com/@joxiahdev" rel="noopener noreferrer"&gt;YouTube channel&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;If you find it useful, I'd really appreciate a like and a subscribe on the channel.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>pydanticai</category>
      <category>agents</category>
    </item>
    <item>
      <title>Build Durable Agents With Pydantic AI And Temporal</title>
      <dc:creator>Blessed Josiah</dc:creator>
      <pubDate>Sat, 18 Jul 2026 11:00:00 +0000</pubDate>
      <link>https://dev.to/joxiahdev/build-durable-agents-with-pydantic-ai-and-temporal-3e67</link>
      <guid>https://dev.to/joxiahdev/build-durable-agents-with-pydantic-ai-and-temporal-3e67</guid>
      <description>&lt;p&gt;Agentic applications can fail mid-run: a network call drops, a server process crashes, an LLM request times out. Without durability built in, these failures mean lost progress and repeated (wasted) LLM calls when the process has to start over.&lt;/p&gt;

&lt;p&gt;Pydantic AI provides a native integration with &lt;strong&gt;Temporal&lt;/strong&gt;, a durable execution engine. Temporal records each step of a workflow's execution as an event history, so if a process crashes, it can resume from where it left off instead of starting over.&lt;/p&gt;

&lt;p&gt;Pydantic AI exposes this through the &lt;code&gt;TemporalAgent&lt;/code&gt; class, which wraps an existing agent and converts its tool calls, MCP calls, and model requests into Temporal Activities.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pydantic_ai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Agent&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pydantic_ai.durable_exec.temporal&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;TemporalAgent&lt;/span&gt;

&lt;span class="n"&gt;planning_agent&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Agent&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;anthropic:claude-fable-5&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;temporal_agent&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;TemporalAgent&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;wrapped&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;planning_agent&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You call &lt;code&gt;.run(...)&lt;/code&gt; on &lt;code&gt;temporal_agent&lt;/code&gt; the same way you would on the underlying agent. The difference is under the hood: each tool call, MCP call, or model request now runs as a Temporal Activity.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three Core Concepts
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Activity&lt;/strong&gt; — a function that performs a unit of work: an LLM call, an API request, a file write. If it fails, Temporal can retry it based on a configured retry policy.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;temporalio&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;activity&lt;/span&gt;

&lt;span class="nd"&gt;@activity.defn&lt;/span&gt;
&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;create_sandbox&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;sandbox&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;AsyncSandbox&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;template&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;node-expo-builder&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;sandbox&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;sandbox_id&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Workflow&lt;/strong&gt; — orchestrates Activities, defining what runs and in what order. Workflow code must be deterministic (no direct network calls, no random values, no reading the system clock), since Temporal replays it to reconstruct state.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;temporalio&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;workflow&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;timedelta&lt;/span&gt;

&lt;span class="nd"&gt;@workflow.defn&lt;/span&gt;
&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;ExampleWorkflow&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nd"&gt;@workflow.run&lt;/span&gt;
    &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;sandbox_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;workflow&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;execute_activity&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;create_sandbox&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
            &lt;span class="n"&gt;start_to_close_timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nf"&gt;timedelta&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;minutes&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;sandbox_id&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Worker&lt;/strong&gt; — polls a task queue for scheduled work and executes it. A Workflow or Activity must be registered with a Worker before it can run.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;asyncio&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;temporalio.client&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Client&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;temporalio.worker&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Worker&lt;/span&gt;

&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;Client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;connect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;localhost:7233&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="n"&gt;worker&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Worker&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;task_queue&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;coding-agent&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;workflows&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;ExampleWorkflow&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="n"&gt;activities&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;create_sandbox&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;worker&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;__name__&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;__main__&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;asyncio&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The task queue name is the only link between the code that starts a Workflow and the Worker that executes it — they never call each other directly.&lt;/p&gt;

&lt;p&gt;Workers run Activities and Workflows differently: Activities run their function body when scheduled, and may be retried on failure per their retry policy. Workflows are replayed against their event history each time they need to advance, which is what lets a Workflow resume correctly on a new Worker after a crash.&lt;/p&gt;

&lt;h2&gt;
  
  
  Signal and Query
&lt;/h2&gt;

&lt;p&gt;Once a Workflow is running, Signal and Query let you interact with it from outside.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Signal&lt;/strong&gt; sends data into a running Workflow, asynchronously, with no return value:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nd"&gt;@workflow.signal&lt;/span&gt;
&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;approve_plan&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;_plan_approved&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;

&lt;span class="nd"&gt;@workflow.signal&lt;/span&gt;
&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;reject_plan&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;feedback&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;""&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;_plan_approved&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;
    &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;_plan_feedback&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;feedback&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Signals only work while a Workflow is still running — once it's completed, failed, or been terminated, it can no longer receive one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Query&lt;/strong&gt; reads a Workflow's current state, synchronously and without side effects:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nd"&gt;@workflow.query&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;state&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;WorkflowState&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nc"&gt;WorkflowState&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;status&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;_status&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;project_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;_project_name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;preview_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;_preview_url&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Queries can also run against a &lt;strong&gt;completed&lt;/strong&gt; Workflow, within your namespace's retention period. This doesn't extend to terminated Workflows, which aren't safely queryable.&lt;/p&gt;

&lt;h2&gt;
  
  
  Human In The Loop
&lt;/h2&gt;

&lt;p&gt;Because Temporal keeps a full event history, a Workflow can pause and resume without losing state — a pattern suited to human approval steps.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;temporalio&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;workflow&lt;/span&gt;

&lt;span class="nd"&gt;@workflow.defn&lt;/span&gt;
&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;ExampleWorkflow&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nd"&gt;@workflow.run&lt;/span&gt;
    &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;_plan_approved&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
            &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;workflow&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;wait_condition&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;lambda&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;_plan_approved&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;_plan_approved&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="k"&gt;break&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;workflow.wait_condition&lt;/code&gt; pauses execution until the given condition is true. Here, the Workflow waits until a Signal (like &lt;code&gt;approve_plan&lt;/code&gt; above) sets &lt;code&gt;self._plan_approved&lt;/code&gt;. If the process running this Workflow crashes while it's paused, a new Worker can pick it up and it resumes waiting at the same point.&lt;/p&gt;

&lt;p&gt;This example has no timeout, so it waits indefinitely. In production, you'd typically pass a &lt;code&gt;timeout&lt;/code&gt; to &lt;code&gt;wait_condition&lt;/code&gt; and handle the case where no decision ever comes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Resources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Watch Part 4 of this series: &lt;a href="https://youtu.be/J0_GeI8Srzc" rel="noopener noreferrer"&gt;https://youtu.be/J0_GeI8Srzc&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Full code for this example: &lt;a href="https://github.com/bjoxiah/pydantic-ai-series/tree/agent-workflow" rel="noopener noreferrer"&gt;https://github.com/bjoxiah/pydantic-ai-series/tree/agent-workflow&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is part of an ongoing series on my &lt;a href="https://youtube.com/@joxiahdev" rel="noopener noreferrer"&gt;channel&lt;/a&gt;, &lt;strong&gt;Mastering Pydantic AI&lt;/strong&gt;.&lt;/p&gt;

</description>
      <category>pydanticai</category>
      <category>temporal</category>
      <category>ai</category>
      <category>agents</category>
    </item>
    <item>
      <title>Building Modular AI Agent Features with Pydantic AI Capabilities</title>
      <dc:creator>Blessed Josiah</dc:creator>
      <pubDate>Sun, 14 Jun 2026 08:30:00 +0000</pubDate>
      <link>https://dev.to/joxiahdev/building-modular-ai-agent-features-with-pydantic-ai-capabilities-39d5</link>
      <guid>https://dev.to/joxiahdev/building-modular-ai-agent-features-with-pydantic-ai-capabilities-39d5</guid>
      <description>&lt;p&gt;If you're building AI Agents with Pydantic AI, understanding &lt;strong&gt;Capabilities&lt;/strong&gt; is invaluable - it's the recommended way to add modular, reusable features to your agents.&lt;/p&gt;

&lt;p&gt;This tutorial is part of my ongoing Pydantic AI series on &lt;a href="https://youtube.com/@joxiahdev" rel="noopener noreferrer"&gt;YouTube&lt;/a&gt;, where I build a full no-code AI agent platform from scratch.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is a Capability?
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;A capability in Pydantic AI is a modular unit of behavior that can be passed to an AI agent.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A capability can give your agent:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Custom toolsets&lt;/li&gt;
&lt;li&gt;Instructions&lt;/li&gt;
&lt;li&gt;Model settings&lt;/li&gt;
&lt;li&gt;Lifecycle hooks&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Think of it as a plug-and-play feature module - build it once, attach it to any agent.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to Create a Capability
&lt;/h2&gt;

&lt;p&gt;Capabilities are created using the &lt;code&gt;Capability&lt;/code&gt; or &lt;code&gt;AbstractCapability&lt;/code&gt; class:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pydantic_ai.capabilities&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;AbstractCapability&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Using &lt;code&gt;AbstractCapability&lt;/code&gt; gives you full control over instructions, tools, and behavior. It's Pydantic AI's recommended pattern if you're building a library or platform on top of the framework.&lt;/p&gt;

&lt;h2&gt;
  
  
  Example: A Research Capability
&lt;/h2&gt;

&lt;p&gt;In this tutorial, I build a &lt;strong&gt;Research Capability&lt;/strong&gt; powered by the Tavily Search API, and an &lt;strong&gt;Email Capability&lt;/strong&gt; powered by Resend.&lt;/p&gt;

&lt;p&gt;Here's the research capability:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pydantic_ai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;FunctionToolset&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pydantic_ai.capabilities&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;AbstractCapability&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;dataclasses&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;dataclass&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pydantic_ai.common_tools.tavily&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;tavily_search_tool&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;settings&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;settings&lt;/span&gt;


&lt;span class="nd"&gt;@dataclass&lt;/span&gt;
&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;ResearchCapability&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;AbstractCapability&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;get_instructions&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;You can use the Tavily search tool for research&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;get_toolset&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;toolset&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;FunctionToolset&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="n"&gt;toolset&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add_tool&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;tavily_search_tool&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;settings&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;tavily_api_key&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;toolset&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's it - &lt;code&gt;get_instructions()&lt;/code&gt; tells the agent what it can do, and &lt;code&gt;get_toolset()&lt;/code&gt; gives it the tools to do it. Attach this to any agent, and it instantly gains web research abilities.&lt;/p&gt;

&lt;h2&gt;
  
  
  Going Further: RAG + GraphRAG
&lt;/h2&gt;

&lt;p&gt;I also built a &lt;strong&gt;Company Knowledge Capability&lt;/strong&gt; that combines:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;pgvector&lt;/strong&gt; for semantic search over your own data&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Neo4j + Graphiti&lt;/strong&gt; for knowledge graph retrieval (GraphRAG)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This lets an agent answer questions from your company's documents, website content, or any knowledge base with both vector search and relationship-aware graph queries.&lt;/p&gt;

&lt;h2&gt;
  
  
  Watch the Full Tutorial
&lt;/h2&gt;

&lt;p&gt;This post covers the concept — the full video walks through building both capabilities live, plus the RAG/GraphRAG ingestion pipeline, observability with Logfire, and wiring it all into a working agent.&lt;/p&gt;

&lt;p&gt;🎥 &lt;strong&gt;&lt;a href="https://youtu.be/ILHtYme4O60" rel="noopener noreferrer"&gt;Watch on YouTube →&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;
💻 &lt;strong&gt;&lt;a href="https://github.com/bjoxiah/pydantic-ai-series/tree/no-code-agent" rel="noopener noreferrer"&gt;Full source code →&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;While you're there, subscribe for more software and AI related content!&lt;/p&gt;

</description>
      <category>ai</category>
      <category>pydanticai</category>
      <category>graphrag</category>
      <category>rag</category>
    </item>
  </channel>
</rss>
