<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Touma Asakura</title>
    <description>The latest articles on DEV Community by Touma Asakura (@coffee00125).</description>
    <link>https://dev.to/coffee00125</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4128808%2Fa3ba2ff3-53fd-4158-ae12-b90e2476bdaa.png</url>
      <title>DEV Community: Touma Asakura</title>
      <link>https://dev.to/coffee00125</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/coffee00125"/>
    <language>en</language>
    <item>
      <title>The Hard Part of AI Engineering Isn’t the Model</title>
      <dc:creator>Touma Asakura</dc:creator>
      <pubDate>Sun, 27 Sep 2026 19:58:26 +0000</pubDate>
      <link>https://dev.to/coffee00125/the-hard-part-of-ai-engineering-isnt-the-model-4okp</link>
      <guid>https://dev.to/coffee00125/the-hard-part-of-ai-engineering-isnt-the-model-4okp</guid>
      <description>&lt;p&gt;When I first started looking seriously at AI engineering, I thought the main problem would be choosing the right model.&lt;/p&gt;

&lt;p&gt;GPT, Claude, Gemini, open-source models, context windows, benchmarks, reasoning ability…&lt;/p&gt;

&lt;p&gt;But after building with them for a while, I started noticing something.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The model is often not the hardest part.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The difficult part is everything around it.&lt;/p&gt;

&lt;p&gt;You can have an extremely capable model and still build a terrible AI system.&lt;/p&gt;

&lt;p&gt;For example, imagine an AI support agent.&lt;/p&gt;

&lt;p&gt;At first it looks simple:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User → LLM → Answer
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then you actually try to put it into production.&lt;/p&gt;

&lt;p&gt;Now it needs to read documentation, search previous conversations, call APIs, understand permissions, remember context, handle tool failures, retry requests, validate outputs, and sometimes admit that it doesn’t know what to do.&lt;/p&gt;

&lt;p&gt;Suddenly it looks more like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User
 ↓
Context / Retrieval
 ↓
LLM
 ↓
Validation
 ↓
Tools
 ↓
Verification
 ↓
Response
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The LLM is just one component in the middle.&lt;/p&gt;

&lt;p&gt;And unlike most components in traditional software, it is probabilistic.&lt;/p&gt;

&lt;p&gt;If I write:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I expect &lt;code&gt;30&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;But if I ask a model:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;model&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;What action should I take?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I am not getting a guaranteed result.&lt;/p&gt;

&lt;p&gt;I’m getting a prediction.&lt;/p&gt;

&lt;p&gt;That difference seems small until the model is allowed to do something real.&lt;/p&gt;

&lt;p&gt;Imagine:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Model decides → refund customer
Model decides → delete record
Model decides → send email
Model decides → execute code
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now a hallucination is no longer just a weird sentence.&lt;/p&gt;

&lt;p&gt;It becomes a system failure.&lt;/p&gt;

&lt;p&gt;That’s why I think one of the most important patterns in AI engineering is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Model proposes
System validates
Tool executes
System verifies
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model should be able to reason.&lt;/p&gt;

&lt;p&gt;But reasoning and authority should not be the same thing.&lt;/p&gt;

&lt;p&gt;Another thing I’ve started thinking about differently is &lt;strong&gt;prompt engineering&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;People talk a lot about writing the perfect prompt.&lt;/p&gt;

&lt;p&gt;I’m not sure that’s the biggest problem anymore.&lt;/p&gt;

&lt;p&gt;A perfect prompt with the wrong context is still useless.&lt;/p&gt;

&lt;p&gt;If a model needs five documents to answer a question, the real engineering challenge is finding the correct five documents out of thousands.&lt;/p&gt;

&lt;p&gt;So the problem becomes less:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“How should I phrase the instruction?”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;and more:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“What information should the model see right now?”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is a much deeper problem.&lt;/p&gt;

&lt;p&gt;The same thing happens with memory.&lt;/p&gt;

&lt;p&gt;Giving an agent “memory” sounds impressive until you ask basic questions:&lt;/p&gt;

&lt;p&gt;What should it remember?&lt;/p&gt;

&lt;p&gt;For how long?&lt;/p&gt;

&lt;p&gt;What happens when information becomes outdated?&lt;/p&gt;

&lt;p&gt;What if two memories contradict each other?&lt;/p&gt;

&lt;p&gt;When should something be forgotten?&lt;/p&gt;

&lt;p&gt;At that point, memory stops being a cool AI feature and starts becoming a serious data engineering problem.&lt;/p&gt;

&lt;p&gt;And then there is debugging.&lt;/p&gt;

&lt;p&gt;With normal software, if something fails, I can usually inspect the code path.&lt;/p&gt;

&lt;p&gt;With an AI system, the failure could come from almost anywhere.&lt;/p&gt;

&lt;p&gt;Maybe retrieval returned the wrong document.&lt;/p&gt;

&lt;p&gt;Maybe the prompt was ambiguous.&lt;/p&gt;

&lt;p&gt;Maybe a tool failed.&lt;/p&gt;

&lt;p&gt;Maybe the model misunderstood the tool output.&lt;/p&gt;

&lt;p&gt;Maybe old memory polluted the context.&lt;/p&gt;

&lt;p&gt;Maybe the model just made a bad decision.&lt;/p&gt;

&lt;p&gt;So logs like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Request failed
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;are not enough anymore.&lt;/p&gt;

&lt;p&gt;We need to see the whole chain:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User input
Retrieved context
Model output
Tool calls
Tool responses
Validation
Retries
Final answer
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That’s where AI engineering starts feeling less like “using an API” and more like actual systems engineering.&lt;/p&gt;

&lt;p&gt;And weirdly, a lot of the solutions are not new.&lt;/p&gt;

&lt;p&gt;Distributed systems already taught us about retries and failures.&lt;/p&gt;

&lt;p&gt;Security taught us about permissions and trust boundaries.&lt;/p&gt;

&lt;p&gt;Databases taught us about state and consistency.&lt;/p&gt;

&lt;p&gt;Software architecture taught us not to give one component control over everything.&lt;/p&gt;

&lt;p&gt;AI just adds a new kind of component:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;something extremely useful that can also be confidently wrong.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Maybe that is the real challenge.&lt;/p&gt;

&lt;p&gt;Not building a system where the AI is always correct.&lt;/p&gt;

&lt;p&gt;But building a system that remains reliable when it isn’t.&lt;/p&gt;

&lt;p&gt;That, to me, is where AI engineering becomes genuinely interesting.&lt;/p&gt;

&lt;p&gt;Curious how others see this.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;When you build with LLMs, what causes more problems: the model itself, or the system around it?&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>softwareengineering</category>
      <category>webdev</category>
      <category>ai</category>
      <category>programming</category>
    </item>
  </channel>
</rss>
