<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: yocho — Hikaru Sato</title>
    <description>The latest articles on DEV Community by yocho — Hikaru Sato (@hikaru_sato_ce8ce2cff7c01).</description>
    <link>https://dev.to/hikaru_sato_ce8ce2cff7c01</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4072584%2F5b310885-b077-4a80-8c6f-73de113d6dec.jpg</url>
      <title>DEV Community: yocho — Hikaru Sato</title>
      <link>https://dev.to/hikaru_sato_ce8ce2cff7c01</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/hikaru_sato_ce8ce2cff7c01"/>
    <language>en</language>
    <item>
      <title>When AI-Generated Articles Fail, Keep the Body and Repair Only What Broke</title>
      <dc:creator>yocho — Hikaru Sato</dc:creator>
      <pubDate>Mon, 14 Sep 2026 12:18:22 +0000</pubDate>
      <link>https://dev.to/hikaru_sato_ce8ce2cff7c01/when-ai-generated-articles-fail-keep-the-body-and-repair-only-what-broke-1f51</link>
      <guid>https://dev.to/hikaru_sato_ce8ce2cff7c01/when-ai-generated-articles-fail-keep-the-body-and-repair-only-what-broke-1f51</guid>
      <description>&lt;p&gt;The dangerous failure in AI publishing is not that a model gets one request wrong. It is losing track of what already succeeded, what was charged, and what is safe to run again.&lt;/p&gt;

&lt;p&gt;While building YOCHO, an AI-industry intelligence service, we changed the editorial pipeline around that problem. The body, summary, and title are separate phases. A failed summary does not throw away a valid body. A bad title does not trigger a full rewrite.&lt;/p&gt;

&lt;p&gt;You can see the current product at &lt;a href="https://yocho.ai" rel="noopener noreferrer"&gt;https://yocho.ai&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I built
&lt;/h2&gt;

&lt;p&gt;The pipeline is intentionally boring:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;source snapshot
  -&amp;gt; body generation
  -&amp;gt; body checkpoint
  -&amp;gt; summary generation
  -&amp;gt; title validation or title-only repair
  -&amp;gt; final publication transaction
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Before a generated result is considered publishable, we store a private draft receipt tied to the call ID. It contains the source digest, response digest, phase, cost receipt, validation result, and any title correction. It does not contain API keys, HTTP headers, or an unbounded provider response.&lt;/p&gt;

&lt;p&gt;The checkpoint lets a later phase reuse the exact body instead of asking the model to recreate it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why failure states need different treatment
&lt;/h2&gt;

&lt;p&gt;A known failure and an unknown failure are not the same thing.&lt;/p&gt;

&lt;p&gt;A known failure means that the provider response and its billing outcome are settled. If a summary fails validation in that state, we can retry only the summary once under the shared rate and cost controls.&lt;/p&gt;

&lt;p&gt;An unknown failure is a timeout, connection loss, or an unsettled billing result. We keep it on hold and do not automatically send the request again. The missing response might still exist on the provider side. Retrying blindly can create duplicate work and duplicate charges.&lt;/p&gt;

&lt;p&gt;This is less about being conservative for its own sake than about preserving a truthful state machine.&lt;/p&gt;

&lt;h2&gt;
  
  
  A concrete example: the title is wrong, the body is fine
&lt;/h2&gt;

&lt;p&gt;Corporate disclosure pages are a useful stress case. A generated title may contain a number that is not in the source, an ambiguous company name, metadata such as a URL or JSON fragment, or an English source headline where a Japanese title is required.&lt;/p&gt;

&lt;p&gt;The safe response is not to regenerate the article. First, construct a source-grounded title from the retrieved issuer name and listing title, then run the same title validator again. If a Japanese title still needs model assistance, run a title-only phase against the stored body and source identity.&lt;/p&gt;

&lt;p&gt;In the September 14, 2026 bounded rescue, the fixed target was 95 tasks. Stored bodies were reused, and 43 earlier title substitutions were corrected while preserving the body and summary. That number is deliberately not presented as completion of the entire 3,356-task manifest.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I learned
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Checkpoints are more valuable than optimistic retries
&lt;/h3&gt;

&lt;p&gt;Without a checkpoint, every failure becomes a full regeneration. With one, the system can ask a smaller question: “Is the summary the only thing that needs work?”&lt;/p&gt;

&lt;h3&gt;
  
  
  The final database transaction is part of the quality gate
&lt;/h3&gt;

&lt;p&gt;Provider-side validation is not enough. The publication transaction can still expose a stale source, a malformed title, or a mismatched identifier. The same contract must be applied again immediately before the status changes to &lt;code&gt;published&lt;/code&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Completion is not the same as an empty pending queue
&lt;/h3&gt;

&lt;p&gt;A queue with no pending items may still contain holds or in-flight work. We distinguish &lt;code&gt;queue_empty&lt;/code&gt; from &lt;code&gt;queue_drained&lt;/code&gt;, and we keep the oldest wait time and remaining budget visible.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is next
&lt;/h2&gt;

&lt;p&gt;The next question is not whether we can generate more articles. It is whether we can measure the complete path honestly: source freshness, time to publication, checkpoint reuse, known and unknown holds, confirmed cost, and the reasons final validation stopped a draft.&lt;/p&gt;

&lt;p&gt;That is the direction behind YOCHO: let AI extract, classify, and propose, while keeping the boundary where evidence becomes interpretation visible. The live project is at &lt;a href="https://yocho.ai" rel="noopener noreferrer"&gt;https://yocho.ai&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;This is an implementation and operations note as of September 14, 2026. It does not claim perfect factuality, search ranking, readership, or AI-answer inclusion.&lt;/p&gt;

</description>
      <category>ai</category>
    </item>
    <item>
      <title>What I Learned Turning an AI News Crawler into a Working Product</title>
      <dc:creator>yocho — Hikaru Sato</dc:creator>
      <pubDate>Wed, 09 Sep 2026 04:38:21 +0000</pubDate>
      <link>https://dev.to/hikaru_sato_ce8ce2cff7c01/what-i-learned-turning-an-ai-news-crawler-into-a-working-product-i6g</link>
      <guid>https://dev.to/hikaru_sato_ce8ce2cff7c01/what-i-learned-turning-an-ai-news-crawler-into-a-working-product-i6g</guid>
      <description>&lt;p&gt;Fetching an article is only the beginning of a news pipeline. The harder work is deciding whether the response is an article, whether it belongs in the product, and whether a reader can understand what was published.&lt;/p&gt;

&lt;p&gt;I am building yocho, an AI industry news and analysis service. This is an architecture note about the work between acquisition and publication, based on the implementation and operational checks available on September 9, 2026. It is not a benchmark or a claim that every item in the collection has been processed successfully.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I built
&lt;/h2&gt;

&lt;p&gt;The service brings category-based news browsing, Japanese editorial articles, and public analysis posts into one application. Its subject matter extends from AI models to chips, infrastructure, power, and business developments.&lt;/p&gt;

&lt;p&gt;The backend uses FastAPI and Python for database access and domain logic. PostgreSQL holds application data and the job queue; GCS holds raw archives. The Next.js frontend accesses the backend rather than connecting directly to the database.&lt;/p&gt;

&lt;p&gt;That division matters because publication rules need to stay consistent across a news page, a mixed news-and-analysis feed, and background processing. A frontend rendering successfully cannot establish that a document passed those rules.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why fetching is a separate milestone
&lt;/h2&gt;

&lt;p&gt;A successful response answers a narrow question: did the server return something? It does not establish that the response contains a usable article.&lt;/p&gt;

&lt;p&gt;A collection pipeline has several distinct decisions:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Discover a candidate URL.&lt;/li&gt;
&lt;li&gt;Fetch and retain the source material.&lt;/li&gt;
&lt;li&gt;Extract the article content.&lt;/li&gt;
&lt;li&gt;Determine relevance and category.&lt;/li&gt;
&lt;li&gt;Prepare the editorial representation.&lt;/li&gt;
&lt;li&gt;Apply publication conditions.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This list is a conceptual model, not a literal function call sequence. Its value is that each step can fail independently and needs its own observable outcome.&lt;/p&gt;

&lt;p&gt;For example, a URL may return a listing page containing many headlines. Treating that response as one article creates a plausible-looking record with the wrong unit of content. Another page can contain a real article while still being unsuitable for the service's editorial scope.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I learned about classification
&lt;/h2&gt;

&lt;p&gt;One concrete correction in the recent commit history separates energy storage from computer memory. The shared word “storage” is not enough to decide where an article belongs.&lt;/p&gt;

&lt;p&gt;A useful regression case therefore includes context: batteries and power infrastructure in one example, memory devices and computing in another. The expected result is a meaningful category distinction, rather than merely proving that a keyword was detected.&lt;/p&gt;

&lt;p&gt;The broader lesson is to test confusing neighbors. A classifier can look reasonable on obvious positive examples and still damage the reader's experience at category boundaries. For a developer applying this elsewhere, collect a small set of near misses before expanding the source list.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I learned about retries
&lt;/h2&gt;

&lt;p&gt;The current article automation generates the body and a dedicated feed summary, then commits publication and task completion in the same database transaction. That boundary prevents those two database outcomes from being recorded independently.&lt;/p&gt;

&lt;p&gt;It does not make the whole workflow one atomic operation. The external generation request happens outside that transaction. A provider can rate-limit a request, and accounting and retry state still need to survive that interruption.&lt;/p&gt;

&lt;p&gt;Operational checks recorded rate limiting and subsequent successful processing. That supports a bounded statement about recovery in those checks. It does not establish that the entire backlog is finished or that every future provider failure will recover automatically.&lt;/p&gt;

&lt;p&gt;For another queue-backed application, the practical question is: after an interruption, can the next worker distinguish work that needs retrying from work already published? Write down that state transition before increasing concurrency.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is next
&lt;/h2&gt;

&lt;p&gt;The public product now has more of the reading journey in place, but completeness and quality remain separate goals. More source coverage does not guarantee a useful stream in every category. A generated summary also needs its own review; assembling fragments from the article can produce repetition rather than a good introduction.&lt;/p&gt;

&lt;p&gt;I also want to evaluate the transition from news to a reader's own analysis. That is a product hypothesis, not an observed adoption result: preserving sources and interpretation separately may make reasoning easier to revisit. The test is whether someone can recover the original evidence and explain what would change their conclusion.&lt;/p&gt;

&lt;p&gt;Open &lt;a href="https://yocho.ai/" rel="noopener noreferrer"&gt;yocho&lt;/a&gt; and inspect one news item's title, category, editorial text, and source link as separate outputs. That same small audit is a useful starting point for testing your own crawler.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>architecture</category>
      <category>product</category>
      <category>startup</category>
    </item>
    <item>
      <title>Why I’m Building yocho: From Events to Hypotheses</title>
      <dc:creator>yocho — Hikaru Sato</dc:creator>
      <pubDate>Wed, 12 Aug 2026 03:42:03 +0000</pubDate>
      <link>https://dev.to/hikaru_sato_ce8ce2cff7c01/why-im-building-yocho-from-events-to-hypotheses-2h9e</link>
      <guid>https://dev.to/hikaru_sato_ce8ce2cff7c01/why-im-building-yocho-from-events-to-hypotheses-2h9e</guid>
      <description>&lt;p&gt;AI industry news is abundant, but information alone does not make the next change easier to understand.&lt;/p&gt;

&lt;p&gt;I’m building yocho.ai, an AI industry intelligence system that keeps four layers distinct:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Event: what happened&lt;/li&gt;
&lt;li&gt;Entity: who or what is involved&lt;/li&gt;
&lt;li&gt;Relationship: how changes connect&lt;/li&gt;
&lt;li&gt;Hypothesis: what may follow and why&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This boundary is practical. Facts can be corrected, relationships can be re-evaluated, and hypotheses can change as new evidence arrives. An AI system should help organize the evidence without silently turning its own interpretation into an official conclusion.&lt;/p&gt;

&lt;p&gt;I’m starting with a founder build log covering design decisions, implementation lessons, discarded approaches, and open questions.&lt;/p&gt;

&lt;p&gt;If you are interested in evidence-led AI industry intelligence, I’ll share more as the system develops.&lt;/p&gt;

&lt;p&gt;Learn more: &lt;a href="https://yocho.ai" rel="noopener noreferrer"&gt;https://yocho.ai&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  A concrete example
&lt;/h2&gt;

&lt;p&gt;For illustration, imagine Company A releases Model B. yocho keeps the layers separate:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Event:&lt;/strong&gt; Company A releases Model B.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Entity:&lt;/strong&gt; Company A, Model B, and competitor Company C.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Relationship:&lt;/strong&gt; Model B may change Company C’s pricing strategy.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hypothesis:&lt;/strong&gt; Lower inference costs may change how teams evaluate adoption.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A simple flow is: &lt;strong&gt;Source → Event → Entity / Relationship → Hypothesis → Review&lt;/strong&gt;. The hypothesis is provisional, not a fact. New evidence can support it, weaken it, or leave it unresolved.&lt;/p&gt;

&lt;h2&gt;
  
  
  What’s next
&lt;/h2&gt;

&lt;p&gt;I’m using &lt;a href="https://yocho.ai" rel="noopener noreferrer"&gt;yocho.ai&lt;/a&gt; as the canonical hub for the product’s design decisions, evidence, and implementation updates. The next useful step is to compare a hypothesis with the sources and relationships that support it.&lt;/p&gt;

</description>
      <category>opensource</category>
    </item>
  </channel>
</rss>
