<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Sebastian Dracopol</title>
    <description>The latest articles on DEV Community by Sebastian Dracopol (@dracopol).</description>
    <link>https://dev.to/dracopol</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4158504%2F5925b133-30c8-456c-af02-730f033220d1.jpg</url>
      <title>DEV Community: Sebastian Dracopol</title>
      <link>https://dev.to/dracopol</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/dracopol"/>
    <language>en</language>
    <item>
      <title>Trust nothing your AI assistant tells you it did published: false</title>
      <dc:creator>Sebastian Dracopol</dc:creator>
      <pubDate>Sat, 03 Oct 2026 10:58:01 +0000</pubDate>
      <link>https://dev.to/dracopol/trust-nothing-your-ai-assistant-tells-you-it-did-published-false-3hkp</link>
      <guid>https://dev.to/dracopol/trust-nothing-your-ai-assistant-tells-you-it-did-published-false-3hkp</guid>
      <description>&lt;p&gt;I use an AI every day for real work: reading mail, linking documents to files, drafting letters, tracking deadlines, researching on the web. It saves me hours a week. It also lies to me several times a week.&lt;/p&gt;

&lt;p&gt;It doesn't lie on purpose. It does something worse: it produces answers that look finished. They're well formatted, confident and plausible, and some fraction of them are wrong in ways you only notice if you go and check.&lt;/p&gt;

&lt;p&gt;After a year of this, my working assumption is simple. Every output is wrong until something outside the model says otherwise. This post explains why, and what "checking" actually means in practice.&lt;/p&gt;

&lt;h2&gt;
  
  
  The failure modes nobody puts in the demo
&lt;/h2&gt;

&lt;p&gt;Hallucination is the famous one, but in day-to-day work it isn't the most common problem. These are the failures I actually see.&lt;/p&gt;

&lt;h3&gt;
  
  
  Reporting actions that didn't happen
&lt;/h3&gt;

&lt;p&gt;The assistant calls a tool to add a comment, send a draft or update a record. The call times out, returns an error, or succeeds on the wrong object. The assistant then writes "Done, comment added" because that's what the plan said would happen next.&lt;/p&gt;

&lt;p&gt;The narration comes from the plan, not from the result. If you read the narration, you'll believe it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Treating "nothing found" as a fact
&lt;/h3&gt;

&lt;p&gt;A search returns zero results, and the assistant concludes the thing doesn't exist. The real cause can be any of these:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the page was a single-page app that hadn't loaded yet;&lt;/li&gt;
&lt;li&gt;the site served a captcha or a rate-limit page;&lt;/li&gt;
&lt;li&gt;the session had expired and the page silently showed a login screen;&lt;/li&gt;
&lt;li&gt;the query had a typo or used the wrong field.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;An empty result is information about the search, not about the world. Last week my assistant got an HTTP 429 from a search engine on the first try. If I hadn't been watching, the next step would have been "no relevant results".&lt;/p&gt;

&lt;h3&gt;
  
  
  Answering from memory when the source changed
&lt;/h3&gt;

&lt;p&gt;Ask about a rule, a setting or an API, and the model answers from training data. Training data has a date. Interfaces change, laws get amended, menus move.&lt;/p&gt;

&lt;p&gt;A small example: I asked how to change a page's username on a large social network. The assistant searched, found three blog posts and confidently gave me a menu path. The path no longer existed in the current interface. The blog posts were consistent with each other and all out of date. The answer only became correct when the assistant opened the actual settings screen and read what was there.&lt;/p&gt;

&lt;h3&gt;
  
  
  Recommending things it never looked at
&lt;/h3&gt;

&lt;p&gt;This one is subtle and expensive. I asked for a plan to improve how a set of my public profiles ranked in search. The assistant produced a reasonable list of which profiles to strengthen and how. It had never opened them.&lt;/p&gt;

&lt;p&gt;When it finally did, several of those profiles contained exactly the content I didn't want to amplify. The plan was internally logical and completely wrong for my situation. In another case it told me to file a request that had already been filed weeks earlier. The record was one search away in my own knowledge base, and it didn't search.&lt;/p&gt;

&lt;h3&gt;
  
  
  Grading its own homework
&lt;/h3&gt;

&lt;p&gt;Ask a model "are you sure?" and it re-reads its own answer, finds it coherent, and says yes. Coherence is not correctness. Re-reading your own text checks the prose, not the facts.&lt;/p&gt;

&lt;h3&gt;
  
  
  Losing detail over long sessions
&lt;/h3&gt;

&lt;p&gt;In long working sessions, the assistant's view of earlier steps gets summarized to fit its context window. The summary says "updated the deadline". It doesn't say which deadline, to what date, or whether the update succeeded. Later steps build on the summary as if it were the record.&lt;/p&gt;

&lt;h2&gt;
  
  
  What "verification" has to mean
&lt;/h2&gt;

&lt;p&gt;Most advice on this stops at "double-check AI output". That's useless unless you define what a check is. These are the rules I follow now.&lt;/p&gt;

&lt;h3&gt;
  
  
  A check is a tool call, not a thought
&lt;/h3&gt;

&lt;p&gt;Verification means touching something outside the model: reading the record back from the API, opening the page, running the query, fetching the current official text. "I reviewed my answer and it looks correct" isn't verification. It's the same model with the same blind spots, reading the same text again.&lt;/p&gt;

&lt;p&gt;If a claim can't be tied to a tool output, it gets marked as unverified. It doesn't get marked as true.&lt;/p&gt;

&lt;h3&gt;
  
  
  Actions are confirmed by the system that performed them
&lt;/h3&gt;

&lt;p&gt;"Email sent" counts when there's a message ID returned by the mail API, or when the message shows up in the Sent folder when you read it back. "Record updated" counts when the record, read back, shows the new value. The assistant's own summary of what it did doesn't count.&lt;/p&gt;

&lt;h3&gt;
  
  
  Negative results need a positive control
&lt;/h3&gt;

&lt;p&gt;When a search finds nothing, run the same search for something you know exists. If that also comes back empty, your method is broken, not the world. This one habit has saved me from a long list of false conclusions about missing data.&lt;/p&gt;

&lt;h3&gt;
  
  
  Read the source that's in force today
&lt;/h3&gt;

&lt;p&gt;For anything that cites a rule, a setting or an interface, the check is against the current version of the source. Not a blog post about it, and not the model's memory of it. Store verified texts once so you don't fetch the same thing a hundred times, but record when you verified them.&lt;/p&gt;

&lt;h3&gt;
  
  
  Look before recommending
&lt;/h3&gt;

&lt;p&gt;Before the assistant recommends doing something to an object (a profile, a document, a record), it opens that object and reads what's there now. A recommendation built on assumptions about content nobody looked at is a guess.&lt;/p&gt;

&lt;h3&gt;
  
  
  Search your own records first
&lt;/h3&gt;

&lt;p&gt;If you have a knowledge base, the assistant searches it before saying anything about your own past work. "You should file X" when X was filed last month costs credibility fast, and it's entirely avoidable.&lt;/p&gt;

&lt;h3&gt;
  
  
  Separate "I couldn't check" from "I checked and it's fine"
&lt;/h3&gt;

&lt;p&gt;These two states look the same in a friendly summary, and they mean completely different things. A tool that failed to load is not evidence that the action happened, and it's not evidence that it didn't. Report it as what it is.&lt;/p&gt;

&lt;h3&gt;
  
  
  External effects need a human
&lt;/h3&gt;

&lt;p&gt;Anything that leaves your control (sending, publishing, filing, paying, deleting) needs explicit approval from a person. The assistant can do 95% of the work and stop one click before the irreversible part.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why bother, if it doubles the work
&lt;/h2&gt;

&lt;p&gt;It does roughly double the effort per workflow at first. The alternative is worse: an assistant that's right 90% of the time and wrong in ways that look identical to being right. You end up checking everything manually anyway, or you stop checking and find out the hard way.&lt;/p&gt;

&lt;p&gt;With verification built into the workflow, I can read a short report that says what was done, what was confirmed and how, and what couldn't be checked. Then I decide. That's the point where the time savings become real, because I spend my attention only on the parts that need it.&lt;/p&gt;

&lt;h2&gt;
  
  
  A short checklist
&lt;/h2&gt;

&lt;p&gt;If you take one thing from this post, use these questions on any AI output that matters:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;What did it claim to do, and where is the system's confirmation?&lt;/li&gt;
&lt;li&gt;Which statements came from a tool output, and which from the model?&lt;/li&gt;
&lt;li&gt;Did any "nothing found" result get a positive control?&lt;/li&gt;
&lt;li&gt;Was the cited rule or interface checked against today's source?&lt;/li&gt;
&lt;li&gt;Did it look at the thing it's recommending changes to?&lt;/li&gt;
&lt;li&gt;What couldn't be verified, and is that clearly labeled?&lt;/li&gt;
&lt;li&gt;Does anything here leave my control without my approval?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;None of this is specific to one model or vendor. Every assistant I've used fails in these ways. The ones that are useful in real work are the ones wrapped in checks that don't take the model's word for anything.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;I help small businesses and independent professionals build AI workflows they can trust. More at &lt;a href="https://dracopol.com/en" rel="noopener noreferrer"&gt;dracopol.com&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>productivity</category>
      <category>programming</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>How I built my own AI tools for a small NGO instead of buying software</title>
      <dc:creator>Sebastian Dracopol</dc:creator>
      <pubDate>Sat, 03 Oct 2026 10:40:23 +0000</pubDate>
      <link>https://dev.to/dracopol/how-i-built-my-own-ai-tools-for-a-small-ngo-instead-of-buying-software-4eng</link>
      <guid>https://dev.to/dracopol/how-i-built-my-own-ai-tools-for-a-small-ngo-instead-of-buying-software-4eng</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnh5kuo8l7eajai7gyzuu.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnh5kuo8l7eajai7gyzuu.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;Small nonprofits have the same paperwork as companies and none of the budget. When I started running one, the problems were familiar: hundreds of emails a month, documents arriving as PDFs, Word files and phone scans, deadlines spread across inboxes, and a small team trying to remember which reply belonged to which file.&lt;/p&gt;

&lt;p&gt;The software on the market was either expensive, generic, or built for sales teams. I've spent about twenty years in project and delivery management, so I did what I'd do on any project: broke the work into pieces and built the pieces myself. This time an AI assistant does most of the work, and I supervise.&lt;/p&gt;

&lt;p&gt;This post covers the architecture, the parts that failed, and the rules I added after they failed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start from the work, not from the model
&lt;/h2&gt;

&lt;p&gt;"We have an AI subscription, what can we do with it?" is the wrong first question. I started by writing down every repetitive task in a normal week and roughly how long each one took. Four came out on top:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Reading incoming mail and working out which file it belongs to.&lt;/li&gt;
&lt;li&gt;Turning documents into something searchable.&lt;/li&gt;
&lt;li&gt;Tracking deadlines and replies so nothing gets lost.&lt;/li&gt;
&lt;li&gt;Finding "that document from three months ago".&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;None of these needs intelligence in the impressive sense. They need consistency, memory and patience, which an assistant can provide if you give it structure and tools.&lt;/p&gt;

&lt;h2&gt;
  
  
  The architecture
&lt;/h2&gt;

&lt;p&gt;The setup is an AI assistant (I use Claude) connected to a set of small tools through the &lt;a href="https://modelcontextprotocol.io/" rel="noopener noreferrer"&gt;Model Context Protocol&lt;/a&gt;. MCP is an open standard: each tool runs as a small server that exposes a few functions, and the assistant decides when to call them. A tool does one job. The assistant combines them.&lt;/p&gt;

&lt;p&gt;These are the servers I run.&lt;/p&gt;

&lt;h3&gt;
  
  
  Mail and ticket connector
&lt;/h3&gt;

&lt;p&gt;This one reads the shared inbox, searches threads, downloads attachments and works with the project tracker (Jira). Each incoming message gets linked to the right ticket, with its attachments and a comment that summarizes what changed.&lt;/p&gt;

&lt;p&gt;It can also create drafts and send email, and that's exactly why sending is locked behind human approval. The assistant drafts. A person reads and approves. More on this below, because it was the most important design decision.&lt;/p&gt;

&lt;p&gt;The same server holds a small deadline register: add, update, supersede, remove and check. Every outgoing request gets a due date calculated from the rules that apply to it. When a date passes, the file comes back for a decision.&lt;/p&gt;

&lt;h3&gt;
  
  
  Document converter
&lt;/h3&gt;

&lt;p&gt;Everything that arrives (PDF, DOCX, XLSX, PPTX, HTML, ZIP) goes through a converter to Markdown before the model touches it. Plain text is cheaper to process, easier to search and easier to quote exactly.&lt;/p&gt;

&lt;p&gt;The exception is images. Modern models read a photo of a document better than OCR does, so images go straight to the model.&lt;/p&gt;

&lt;h3&gt;
  
  
  Private knowledge base
&lt;/h3&gt;

&lt;p&gt;All files, tickets, documents and replies are indexed in a vector database running on my own machine, with metadata such as source, ticket key and date. The assistant has &lt;code&gt;find&lt;/code&gt;, &lt;code&gt;store&lt;/code&gt;, &lt;code&gt;update&lt;/code&gt; and &lt;code&gt;reindex&lt;/code&gt; operations, plus search by entity.&lt;/p&gt;

&lt;p&gt;The rule is simple: for any question about our own work, the assistant searches the knowledge base first, and only after that the tracker or the web. Answers have to cite the chunk they came from. This one rule removed most hallucinated answers about our own history, because the model stops guessing when the real document is one call away.&lt;/p&gt;

&lt;h3&gt;
  
  
  Browser automation
&lt;/h3&gt;

&lt;p&gt;The assistant is told never to try to log in on the anonymous profile. Single-page apps get an explicit wait after each search, because an empty result two seconds after submit is a false negative, not a sign the data doesn't exist. Before long runs, a health check navigates to the authenticated sources and confirms the session is still alive. Having cookies doesn't prove you're logged in.&lt;/p&gt;

&lt;h2&gt;
  
  
  Skills: writing the process down
&lt;/h2&gt;

&lt;p&gt;The tools were maybe 40% of the value. The rest came from writing down how we work as reusable instructions the assistant can load, which Claude calls "skills". Each one is a Markdown file with a trigger and a procedure:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;one processes the morning's mail end to end: link, compare the reply with the original request point by point, update the deadline, recommend the next step;&lt;/li&gt;
&lt;li&gt;one audits an outgoing letter before it's sent;&lt;/li&gt;
&lt;li&gt;one runs a decision through five opposing perspectives (contrarian, first principles, expansionist, outsider, executor) plus a peer review before we commit;&lt;/li&gt;
&lt;li&gt;one looks for what's wrong with the previous answer and fixes it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Writing them forced us to make our process explicit, and that was worth doing for its own sake. A good share of the "AI problems" I fixed turned out to be "we never agreed how to do this" problems.&lt;/p&gt;

&lt;h2&gt;
  
  
  Verification
&lt;/h2&gt;

&lt;p&gt;AI assistants are confident even when they're wrong. Early on, mine:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;reported "comment added" when the tool call had failed silently;&lt;/li&gt;
&lt;li&gt;cited rules from memory instead of the version in force;&lt;/li&gt;
&lt;li&gt;treated an empty search result as proof that something didn't exist;&lt;/li&gt;
&lt;li&gt;filled gaps with plausible details.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So verification became part of the design instead of an afterthought.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Every claimed action needs the tool's response.&lt;/strong&gt; "Email sent" has to come with the message ID the API returned. If there's no tool output, the action didn't happen.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;References are checked against the current source.&lt;/strong&gt; Anything that cites a rule is checked against the official, up-to-date text, not the model's memory. The verified texts are stored once and reused, so the same source isn't fetched again for every document.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Negative results need a control.&lt;/strong&gt; If a search returns nothing, the same search must find a case we know exists. Otherwise "not found" is a property of the method, not a fact.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;There's a verification loop with a cap.&lt;/strong&gt; After important work, a verification pass looks for concrete problems (unconfirmed actions, unsourced claims, contradictions, unstated assumptions), fixes them in place and runs again until nothing changes. The cap is 5 passes for tool-backed actions and 2 for pure text judgment, so it can't loop forever.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Humans approve external effects.&lt;/strong&gt; Sending, publishing and filing need explicit approval. The assistant can prepare everything up to the last click.&lt;/p&gt;

&lt;p&gt;Building this roughly doubled the effort for each workflow. It's also the only reason I trust the output.&lt;/p&gt;

&lt;h2&gt;
  
  
  What changed
&lt;/h2&gt;

&lt;p&gt;The measurable gain is time. Sorting a morning's mail, matching it to files and updating deadlines used to take a morning. Now it's a few minutes of review.&lt;/p&gt;

&lt;p&gt;The less obvious gain is memory. Nothing depends on one person remembering where a document is or when something is due.&lt;/p&gt;

&lt;p&gt;The costs are real too. The assistant still needs supervision. Tools break when a website or an API changes. Some days doing the task by hand is faster than explaining it.&lt;/p&gt;

&lt;h2&gt;
  
  
  If you want to build something similar
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;List your repetitive tasks and time them. Automate the top two, not everything.&lt;/li&gt;
&lt;li&gt;Keep your data where you control it. A local vector database is cheap and solves most privacy questions, which matters if you handle personal data under GDPR.&lt;/li&gt;
&lt;li&gt;Build small tools that do one thing, and let the assistant combine them.&lt;/li&gt;
&lt;li&gt;Write the process down as instructions. That's where most of the value is.&lt;/li&gt;
&lt;li&gt;Don't skip verification or human approval for anything that leaves the building.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;I now help small businesses and independent professionals build the same kind of setup for their own work. You can find me at &lt;a href="https://dracopol.com/en" rel="noopener noreferrer"&gt;dracopol.com&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>mcp</category>
      <category>automation</category>
      <category>ngo</category>
    </item>
  </channel>
</rss>
