<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: sunny yuen</title>
    <description>The latest articles on DEV Community by sunny yuen (@yuens1002).</description>
    <link>https://dev.to/yuens1002</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F72273%2F1f7cd29e-4d12-444f-97de-5cb562ad9894.png</url>
      <title>DEV Community: sunny yuen</title>
      <link>https://dev.to/yuens1002</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/yuens1002"/>
    <language>en</language>
    <item>
      <title>Hiring Is a Black Box on Both Ends</title>
      <dc:creator>sunny yuen</dc:creator>
      <pubDate>Mon, 14 Sep 2026 18:11:45 +0000</pubDate>
      <link>https://dev.to/yuens1002/hiring-is-a-black-box-on-both-ends-53dc</link>
      <guid>https://dev.to/yuens1002/hiring-is-a-black-box-on-both-ends-53dc</guid>
      <description>&lt;p&gt;I've been on both ends of the same black box this year.&lt;/p&gt;

&lt;p&gt;On one end: a veteran software engineer (6 years, full-stack, AI-native workflow) running an actual job hunt — 100+ applications submitted through a pipeline I built, with a fit score on every one. On the other end: employers, at least on paper, drowning — applications per open role up 111% since 2022, applications per recruiter up 412%, time-to-fill up 37% (Greenhouse's own benchmark data across 6,000+ companies).&lt;/p&gt;

&lt;p&gt;Both sides are drowning. Both sides are trying to figure out the same two things, in opposite directions:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Employers:&lt;/strong&gt; is this candidate a good fit — and can I even trust what they say?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Candidates:&lt;/strong&gt; is this job the right fit for me — and will anyone actually look at me?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Here's what our first-party data — plus the 2026 research now landing on both sides — says about what's actually happening inside that box.&lt;/p&gt;

&lt;h2&gt;
  
  
  The black box on the employer end
&lt;/h2&gt;

&lt;p&gt;The story the industry tells itself is that companies are being flooded with AI-generated applications and are fighting back with AI detection. That story is mostly wrong, and 2026's data keeps confirming it's wrong:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;No major ATS ships AI-authorship detection.&lt;/strong&gt; A 10-vendor survey (Workday, iCIMS, Greenhouse, Oracle, SAP, Lever, Workable, SmartRecruiters, Ashby, BambooHR) found zero platforms that analyze resume or cover-letter text for AI authorship. Greenhouse — the same vendor whose public policy calls AI-assisted writing "acceptable" — ships a paid fraud product ("Real Talent," with CLEAR identity verification and a fraud-risk score) that scores metadata and identity signals. It &lt;em&gt;explicitly does not read your resume text for AI authorship&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Talent leaders rank fraud as the #1 challenge of 2026&lt;/strong&gt; — ahead of "lack of qualified talent" for the first time (GoodTime's annual survey, 500+ talent acquisition (TA) leaders). But notice what that means: the thing employers are bracing for is &lt;strong&gt;impersonation and fabrication&lt;/strong&gt;, not "a candidate polished their resume with Claude."&lt;/li&gt;
&lt;li&gt;99.8% of TA teams are using, piloting, or planning AI agents themselves. The receiving side isn't building walls against AI. It's adopting AI and installing border control for liars.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So the line employers are actually policing is not "AI-assisted" vs. "human." It's &lt;strong&gt;"real" vs. "fabricated."&lt;/strong&gt; And here's the twist that matters for anyone on the candidate side:&lt;/p&gt;

&lt;h2&gt;
  
  
  A real candidate, AI-assisted, is on the &lt;em&gt;safe&lt;/em&gt; side of that line
&lt;/h2&gt;

&lt;p&gt;The research keeps landing in the same direction:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Tailored AI-assisted resumes got &lt;strong&gt;18% employer response&lt;/strong&gt; in a 204-application field experiment vs. &lt;strong&gt;10% for generic AI&lt;/strong&gt; and &lt;strong&gt;31% for fully human-written&lt;/strong&gt;. Tailoring nearly doubled generic AI's response rate. Generic AI is the worst-performing format of anything tested.&lt;/li&gt;
&lt;li&gt;When AI noise floods a channel, employers don't stop discriminating — they re-price toward &lt;strong&gt;harder-to-fake signals&lt;/strong&gt;. In one large field study, after an AI cover-letter tool launched, the correlation between letter quality and callbacks fell 51% — and employers shifted weight to work history, the one thing you can't prompt-engineer.&lt;/li&gt;
&lt;li&gt;Fabrication is the real disease: a controlled study of automated LLM hiring pipelines found unsupported claims (fabricated credentials, inflated qualifiers, invented experience) in &lt;strong&gt;96.7% of outputs&lt;/strong&gt;. That's what the receiving side is actually afraid of. And they're right to be.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The conclusion I've drawn from running my own pipeline through all this: &lt;strong&gt;the problem with AI-assisted applications isn't that they're AI-assisted. It's that most of them are ungrounded.&lt;/strong&gt; The receiving side isn't scanning for Claude. It's scanning for liars. A candidate whose claims are &lt;em&gt;verifiable&lt;/em&gt; is solving the actual problem employers are paying CLEAR and IPQualityScore to solve. A candidate with a generic AI resume is just adding to the noise both sides are drowning in.&lt;/p&gt;

&lt;h2&gt;
  
  
  What my own data says — and doesn't
&lt;/h2&gt;

&lt;p&gt;Here's the honest part. My pipeline (100+ applications, each scored for fit, each tailored): &lt;strong&gt;3 phone screens, 0 technical interviews, 15 rejections&lt;/strong&gt; so far. If AI-assisted tailoring were a magic bullet, those numbers would look different. They don't.&lt;/p&gt;

&lt;p&gt;Why? Because hiring is a black box on both ends:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;From the employer side, they can't easily tell a grounded, evidence-backed candidate from a confident-sounding fabrication — so they fall back on human review friction, identity checks, and slower processes that treat &lt;em&gt;everyone&lt;/em&gt; like a suspect.&lt;/li&gt;
&lt;li&gt;From the candidate side, you can't see why you were rejected. Was it the AI-assisted approach? Calibration? The market? Comp band? Just volume and time? &lt;strong&gt;Every one of those produces the exact same funnel numbers from the inside.&lt;/strong&gt; More applications doesn't resolve it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The industry's own data suggests the answer isn't "stop using AI," either. The receiving side pays a premium for the thing AI can't fake — verified history, real work, provable claims. And it pays a trust discount on everything else. Even on proof that's genuine: one study of 1,380 HR professionals found a real digital credential was trusted &lt;em&gt;less&lt;/em&gt; than the same certificate shown in person.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we're changing because of this
&lt;/h2&gt;

&lt;p&gt;This isn't commentary from the sidelines — this is my own pipeline, and here's what changed:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Shortlist calibration over volume.&lt;/strong&gt; The evidence says fit-calibrated, tailored submissions outperform volume; the pipeline now optimizes for the strongest-fit, most-evidence-backed applications per day, not raw throughput.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Evidence-forward over claim-forward.&lt;/strong&gt; Every claim the pipeline generates carries the evidence behind it (public repo, shipped work, queryable endpoint) rather than polished prose alone. If the receiving side is policing fabrication, the winning side of that line is the verifiable one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Disclose the mechanism, not the tool.&lt;/strong&gt; A recruiter's own AI can query a candidate endpoint that returns structured, evidence-grounded claims. The response rate data says that's the direction the channel is moving; the fraud-product data says the receiving side is explicitly building for it.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The takeaway
&lt;/h2&gt;

&lt;p&gt;Hiring in 2026 is a black box on both ends, and both ends are panicking about the same thing in opposite directions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Employers are building &lt;strong&gt;fraud infrastructure&lt;/strong&gt;, not AI detection. The line they're policing is trust, not tooling.&lt;/li&gt;
&lt;li&gt;Candidates are burning out on a channel where generic AI output is now the &lt;em&gt;worst&lt;/em&gt;-performing format, not the best.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The unglamorous conclusion: the AI-assisted candidate isn't fighting the system. The ungrounded one is. And that line is now enforced by paid fraud infrastructure. Which means the question isn't whether you used AI. It's whether what you submitted is true.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Job-hunt figures are first-party, from the pipeline I run; market figures are attributed to their sources in the text.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>hiring</category>
      <category>career</category>
      <category>jobsearch</category>
    </item>
    <item>
      <title>The Decision Runtime: Bounded Authority for the Chief of Staff Control Plane</title>
      <dc:creator>sunny yuen</dc:creator>
      <pubDate>Mon, 31 Aug 2026 13:00:00 +0000</pubDate>
      <link>https://dev.to/yuens1002/the-decision-runtime-bounded-authority-for-the-chief-of-staff-control-plane-5gib</link>
      <guid>https://dev.to/yuens1002/the-decision-runtime-bounded-authority-for-the-chief-of-staff-control-plane-5gib</guid>
      <description>&lt;p&gt;&lt;a href="https://dev.to/yuens1002/beyond-the-chatbot-why-the-future-of-workplace-ai-needs-a-chief-of-staff-control-plane-2jhm"&gt;Part one of this series&lt;/a&gt; argued that workplace AI needs a control plane — a Chief of Staff that triages, remembers, and governs, instead of a pile of siloed point-solution agents. That piece described the shape of the thing. It didn't answer a harder question: how does a Chief of Staff earn the right to act on your behalf, and prove afterward that it stayed inside that right?&lt;/p&gt;

&lt;p&gt;That question is what shaped the &lt;strong&gt;Decision Runtime&lt;/strong&gt;'s design. Not "give the agent more tools" — tools were never the constraint. The constraint was that acting on someone's behalf without a provable record of authority isn't delegation, it's just risk with extra steps. Everything below is a design decision made in service of one goal: an agent can do real work, and every piece of that work can be traced back to exactly what allowed it.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. The gap the design closes
&lt;/h3&gt;

&lt;p&gt;A Chief of Staff that can act but can't account for what it did, why, or under what authority isn't a Chief of Staff — it's a very fast intern with no paper trail. The design had to answer three questions any real operator would ask walking in cold, before it earned the right to touch anything consequential:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What requires owner attention right now?&lt;/li&gt;
&lt;li&gt;What did the agent already do without me?&lt;/li&gt;
&lt;li&gt;What outcome is still expected, and did it land?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A system that can't answer those on demand isn't trustworthy at any level of autonomy, bounded or not — so the runtime's job is to make those three questions always answerable from durable state, not from someone's memory of what the agent was supposed to be doing.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Bounded autonomy, not autonomy
&lt;/h3&gt;

&lt;p&gt;The obvious design would be a single permission flag: this agent may act, or it may not. We rejected that shape because real operating decisions aren't binary. Something can be worth doing without needing a human told about it. Something else can be worth flagging without touching priority at all. Collapsing those into one yes/no either over-grants (the agent silently reprioritizes things it should only ever report) or under-grants (every trivial action waits on a human because the one flag covering it also covers something consequential). So every unit of work the runtime tracks carries three &lt;em&gt;independent&lt;/em&gt; authority signals instead of one:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;execution&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;       &lt;span class="s"&gt;act&lt;/span&gt;
&lt;span class="na"&gt;attention&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;       &lt;span class="s"&gt;inform&lt;/span&gt;
&lt;span class="na"&gt;priority_effect&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;within_current_policy&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The second, harder design decision: the runtime itself decides none of this. It doesn't grant authority — it records and enforces whatever authority an external operating policy already grants. That line exists because a workflow-neutral runtime that also opined on what agents should be allowed to do would stop being reusable; the moment it bakes in one consumer's judgment about acceptable risk, every other consumer inherits that judgment too. A consumer defines the policy; the runtime's only job is making it impossible to quietly exceed it.&lt;/p&gt;

&lt;p&gt;The consequence of designing it that way: the operating mode is &lt;strong&gt;governed, not irreversible&lt;/strong&gt;. A deployment can start a policy conservative and expand it only through a reviewed change, and narrowing it back down is a policy edit, not a rollback or a rewrite. That reversibility is what makes "bounded" an actual property of the system rather than a one-time promise made at launch and never revisited.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Attributable work
&lt;/h3&gt;

&lt;p&gt;The cheap way to build this would be a log line per action: timestamp, actor, message. We didn't do that, because a log line answers "something happened" but not "was this actually authorized, and by what chain of reasoning." A prose log can be edited, is easy to leave incomplete under time pressure, and doesn't let you &lt;em&gt;traverse&lt;/em&gt; from an outcome back to the exact evidence and approval that produced it — you can only read it linearly and hope the relevant line is in there.&lt;/p&gt;

&lt;p&gt;So every event, work item, action attempt, approval, result, and artifact is instead a typed, append-only record connected by explicit edges: what caused it, who or what performed it, which tool invocation ran, what evidence it used, what decision or approval it required, and what it produced. And identity is never a field the caller fills in — the authenticated principal, the effective actor, and the authorization decision are derived server-side, on purpose, so a compromised or careless caller can't self-assign a different identity than the one it authenticated as.&lt;/p&gt;

&lt;p&gt;The design bet: a graph you can traverse answers "why did this happen, who or what performed it, and under which authority" on demand. A log you can only grep answers it only if you already know what you're looking for.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Portfolio-level reasoning
&lt;/h3&gt;

&lt;p&gt;Part one named the failure mode this is designed against: agent sprawl, where every system gets its own point-solution agent with no shared context. The design response is that the runtime deliberately doesn't care which repository or service an event came from — a subject is just a namespaced type and ID, not "the GitHub domain" or "the deploy domain." That's what makes one reasoning chain possible across systems that share nothing else:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;signal → affected bet → current constraint → recommendation → action
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A failed deployment, a customer reply, and an open PR can all bear on the same operating decision even though they originate in three unrelated systems. The alternative — a purpose-built integration between every pair of systems that need to inform each other — doesn't scale past a handful of systems; a shared, typed substrate is the design choice that avoids that combinatorics problem entirely.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. A real Decision Queue
&lt;/h3&gt;

&lt;p&gt;A manually maintained queue rots two ways: someone forgets to add something that mattered, or nobody prunes what stopped mattering, and the queue turns into noise the owner learns to skim past. Both failure modes come from the same root cause — the queue is a second copy of state that a human has to keep in sync with reality by hand.&lt;/p&gt;

&lt;p&gt;The design fix is to make the queue a &lt;em&gt;projection&lt;/em&gt; of runtime state instead of an independent list: rebuildable from durable records, not hand-curated, surfacing only what the records say genuinely needs the operating owner —&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Approval required&lt;/li&gt;
&lt;li&gt;A material assumption changed&lt;/li&gt;
&lt;li&gt;A priority change was proposed&lt;/li&gt;
&lt;li&gt;Protected attention is needed&lt;/li&gt;
&lt;li&gt;A review deadline was reached&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Everything else stays out of the queue by construction. No one has to remember to filter it, because there's no separate list to fall out of sync with the thing it's supposed to represent.&lt;/p&gt;

&lt;h3&gt;
  
  
  6. Closed feedback loops
&lt;/h3&gt;

&lt;p&gt;The easiest design mistake here would be treating a decision as finished the moment the action fires — which is exactly how automation quietly drifts from reality: an assumption goes stale, nobody re-checks it, and the system keeps acting on it anyway because nothing in the design forces a look back. So a decision in this model isn't a point in time, it's a loop that has to close:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;decision → action → observed result → metric or assumption update → continue, reconsider, or escalate
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Closing the loop is also what keeps the improvement inside the right boundary. The Chief of Staff can get better at operating &lt;em&gt;within&lt;/em&gt; a policy by learning from observed outcomes — but changing the policy itself stays outside that loop, under Git review and owner approval, not inside a model's own judgment about what it should be allowed to do next. Those are deliberately two different mechanisms, not one blurred together.&lt;/p&gt;

&lt;h3&gt;
  
  
  7. Event-driven operation
&lt;/h3&gt;

&lt;p&gt;The real test of the design in sections 1-6 is whether a new trigger source needs new runtime logic, or just a new registration. A GitHub dispatcher that turns a repository event directly into attributable work is one instance of the pattern:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;GitHub event → Chief of Staff review work item → exact-head verification → GitHub MCP actions → review result → runtime attribution and outcome
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Nothing about that chain is GitHub-specific at the runtime layer — the event, the work item, the attempt, and the outcome are the same typed primitives every other section describes. The same pattern extends to a deployment, an incident, a schedule, or a Slack message without touching the runtime itself, only adding the typed registration for that new source. If a new trigger required new runtime code, that would mean the earlier design decisions weren't actually general — this is the part of the design meant to prove that they are.&lt;/p&gt;

&lt;h3&gt;
  
  
  What it is not
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Not a store for private model chain-of-thought.&lt;/li&gt;
&lt;li&gt;Not a replacement for the agent's own memory layer.&lt;/li&gt;
&lt;li&gt;Not a general-purpose job queue.&lt;/li&gt;
&lt;li&gt;Not a source of new authority — it enforces authority, it doesn't grant it.&lt;/li&gt;
&lt;li&gt;Not a replacement for GitHub issues or PRs.&lt;/li&gt;
&lt;li&gt;Not the agent making decisions by itself.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It stores structured reasoning products — facts, evidence, assumptions, recommendations, classifications, decisions, and outcomes — not hidden internal reasoning.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why this matters
&lt;/h3&gt;

&lt;p&gt;The real shift isn't "the agent can now do more." It's that the Chief of Staff stops being purely reactive. It can hold a durable understanding of what matters, notice when reality changes, act within a policy someone else set and can audit, and bring back only the decisions that actually need a human — instead of every decision, or none of them.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Open source, same repo as part one: &lt;a href="https://github.com/yuens1002/openclaw-control-plane" rel="noopener noreferrer"&gt;github.com/yuens1002/openclaw-control-plane&lt;/a&gt;. The Decision Runtime API, its MCP bridge, and the architecture/authentication docs are all public. If you're building something similar, I'd genuinely like to hear where the authority model breaks down in your use case — that's exactly the kind of edge case a workflow-neutral runtime needs to survive.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>architecture</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Beyond the Chatbot: Why the Future of Workplace AI Needs a "Chief of Staff" Control Plane</title>
      <dc:creator>sunny yuen</dc:creator>
      <pubDate>Fri, 14 Aug 2026 19:09:02 +0000</pubDate>
      <link>https://dev.to/yuens1002/beyond-the-chatbot-why-the-future-of-workplace-ai-needs-a-chief-of-staff-control-plane-2jhm</link>
      <guid>https://dev.to/yuens1002/beyond-the-chatbot-why-the-future-of-workplace-ai-needs-a-chief-of-staff-control-plane-2jhm</guid>
      <description>&lt;p&gt;Today, teams deploying AI across business operations keep hitting the same wall: &lt;strong&gt;agent sprawl&lt;/strong&gt;. Organizations end up with siloed AI tools — one for customer support, another for lead scoring, another for code verification — operating in isolation with fragmented context, unpredictable failure modes, and zero centralized governance.&lt;/p&gt;

&lt;p&gt;To build reliable business automation, we need to rethink the architecture. The fix isn't a bigger, monolithic model; it's establishing a &lt;strong&gt;control plane that operates like an AI Chief of Staff&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. The problem: AI workers without central governance
&lt;/h2&gt;

&lt;p&gt;In a traditional organization, you wouldn't hire five specialized contractors and let them work without coordination, shared context, or executive oversight. Yet that's often exactly how multi-agent architectures get built today.&lt;/p&gt;

&lt;p&gt;When specialized agents run without a central control plane, you get:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Context fragmentation&lt;/strong&gt; — knowledge trapped inside point solutions instead of shared across the client lifecycle.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Uncontrolled side effects&lt;/strong&gt; — without centralized policy checks, automated tasks risk making unauthorized updates to CRMs, databases, or client-facing channels.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Operator fatigue&lt;/strong&gt; — humans forced to babysit multiple distinct interfaces instead of reviewing high-signal, decision-ready briefings.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  2. The architecture: control plane vs. execution plane
&lt;/h2&gt;

&lt;p&gt;The fix is a clear architectural separation between the &lt;strong&gt;control plane&lt;/strong&gt; (governance, triage, state) and the &lt;strong&gt;execution plane&lt;/strong&gt; (domain-specific workers).&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbhu5t98ipwbd64swowu6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbhu5t98ipwbd64swowu6.png" alt="Control plane architecture: a human operator sets goals for a central control plane (" width="800" height="573"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The control plane doesn't execute line-level tasks itself. Like a Chief of Staff, it handles four core responsibilities:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Intake triage &amp;amp; routing&lt;/strong&gt; — ingests incoming events (webhooks, scheduled triggers, user instructions) and dispatches generic event envelopes to the right domain worker.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;State &amp;amp; cross-functional memory&lt;/strong&gt; — maintains historical context, task logs, and outcome tracking across disjoint workstreams.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Human-in-the-loop (HITL) governance&lt;/strong&gt; — evaluates risk boundaries. Read-only tasks execute autonomously; state-mutating actions (contract generation, production deploys, financial transactions) route through approval gates.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Executive synthesis&lt;/strong&gt; — filters out execution noise, delivering concise, actionable briefings to the operator via Slack, email, or a centralized dashboard.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;In this architecture, the three example workers are a starting illustration, not a prescription — swap in whatever domains your business actually runs: an &lt;strong&gt;Acquisition Worker&lt;/strong&gt; (intake &amp;amp; scoring), a &lt;strong&gt;Support Worker&lt;/strong&gt; (ticket triage &amp;amp; response), and a &lt;strong&gt;Growth Analytics&lt;/strong&gt; worker (telemetry &amp;amp; logs) are common enough to stand in for the pattern, but the control plane doesn't care what's downstream as long as it speaks generic event envelopes and MCP tool calls.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Connecting the fleet: standardized protocols (MCP)
&lt;/h2&gt;

&lt;p&gt;A control plane is only as good as its interfaces. Leaning on the &lt;strong&gt;Model Context Protocol (MCP)&lt;/strong&gt; and a workflow-neutral runtime shell gets you:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Decoupled domain workers&lt;/strong&gt; — specialized workers (sales intake, customer support, bookkeeping) can be added, updated, or replaced without touching core database schemas or routing logic.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Standardized tool access&lt;/strong&gt; — third-party integrations (CRMs, Stripe, cloud hosting, Git workflows) are exposed as structured tool endpoints with built-in audit logs and rate limits.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Closed-loop feedback&lt;/strong&gt; — analytics workers track performance and attribution metrics, feeding operational data back to the control plane to continuously refine prompt strategies and scoring heuristics.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  4. From theory to infrastructure: open-sourcing the foundation
&lt;/h2&gt;

&lt;p&gt;Talking about control planes and agent swarms is easy; operationalizing them in production usually breaks down at the infrastructure layer — fragile Docker builds, environment drift, unverified runtime setups.&lt;/p&gt;

&lt;p&gt;To fix that, I've open-sourced the foundation layer: &lt;strong&gt;&lt;a href="https://github.com/yuens1002/openclaw-control-plane" rel="noopener noreferrer"&gt;OpenClaw Control Plane&lt;/a&gt;&lt;/strong&gt; — a workflow-neutral TypeScript monorepo that puts source-controlled operating discipline around an OpenClaw instance on Railway.&lt;/p&gt;

&lt;p&gt;To be precise about what it is: OpenClaw itself already ships &lt;code&gt;/setup&lt;/code&gt;, login, and the &lt;code&gt;/openclaw&lt;/code&gt; Control UI, and there's an existing community Railway template for the fastest generic install. This repo isn't a replacement for either of those — it's the governed install path &lt;em&gt;around&lt;/em&gt; them, for teams that want more than a quick demo:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A pinned, auditable dependency on the upstream OpenClaw Railway wrapper, with weekly update detection and no silent auto-upgrades to production.&lt;/li&gt;
&lt;li&gt;Live Railway proof checks — source repo, runtime settings, domain port, deployment source, &lt;code&gt;/setup&lt;/code&gt;, &lt;code&gt;/setup/healthz&lt;/code&gt;, and &lt;code&gt;/openclaw&lt;/code&gt; all get verified, not assumed.&lt;/li&gt;
&lt;li&gt;A workflow-neutral starter-kit boundary — no baked-in client pipeline, connector, database, or secret assumptions.&lt;/li&gt;
&lt;li&gt;A clean place to define setup-profile conventions that private agency/client repos can build on to automate provider, channel, plugin, and workflow attachment.&lt;/li&gt;
&lt;li&gt;Handoff and verification discipline for repeatable client onboarding.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It's early — M1 foundation work, with the default API and worker runner starting empty of registered workflows on purpose. Production connectors and client-specific assumptions are intentionally out of scope for this layer.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. What's next on the roadmap
&lt;/h2&gt;

&lt;p&gt;This open-source template is the foundation — the baseline runtime where the "Chief of Staff" lives. Coming next:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Modular domain worker specs&lt;/strong&gt; — reference implementations for inbound sales intake, acceptance-criteria (AC) verification, and analytics attribution.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Standardized MCP tool bridges&lt;/strong&gt; — pre-built connectors for CRMs, billing systems, and client infrastructure.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Operator dashboards&lt;/strong&gt; — lightweight HITL interfaces for single-tap task approvals and telemetry monitoring.&lt;/li&gt;
&lt;/ol&gt;




&lt;p&gt;&lt;em&gt;Build in public: check out &lt;a href="https://github.com/yuens1002/openclaw-control-plane" rel="noopener noreferrer"&gt;the repo&lt;/a&gt;, spin up a test instance on Railway, and let me know what you'd want a "Chief of Staff" control plane to handle first.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>architecture</category>
      <category>opensource</category>
    </item>
    <item>
      <title>I made my résumé something a machine can read fairly — here's how it's built, and how to stand up your own</title>
      <dc:creator>sunny yuen</dc:creator>
      <pubDate>Tue, 02 Jun 2026 13:10:02 +0000</pubDate>
      <link>https://dev.to/yuens1002/i-made-my-resume-something-a-machine-can-read-fairly-heres-how-its-built-and-how-to-stand-up-2ilg</link>
      <guid>https://dev.to/yuens1002/i-made-my-resume-something-a-machine-can-read-fairly-heres-how-its-built-and-how-to-stand-up-2ilg</guid>
      <description>&lt;p&gt;A model reads my work before any person does, and I had no say in what it concluded from a frozen PDF. I couldn't change that a machine reads first. I could change what I put in front of it. So I did — and here's the build.&lt;/p&gt;

&lt;h2&gt;
  
  
  The shape of it
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;A small backend exposes a profile as an API — &lt;code&gt;GET /info&lt;/code&gt;, &lt;code&gt;POST /query&lt;/code&gt; (grounded answers + sources), &lt;code&gt;POST /match&lt;/code&gt; (fit score), &lt;code&gt;POST /resume&lt;/code&gt; (tailored). It also speaks agent: an MCP endpoint and an A2A agent-card for machine callers.&lt;/li&gt;
&lt;li&gt;A web front (React + a tiny Hono server) renders it as a conversation for people, and as JSON-LD + &lt;code&gt;llms.txt&lt;/code&gt; + a crawlable &lt;code&gt;&amp;lt;noscript&amp;gt;&lt;/code&gt; for machines — so a non-JS fetch isn't an empty shell.&lt;/li&gt;
&lt;li&gt;Nothing hardcodes me. Identity comes from &lt;code&gt;/info&lt;/code&gt;. Fork the front, point two env vars at your backend, and it's yours. Both repos are MIT.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Why I bothered
&lt;/h2&gt;

&lt;p&gt;I'd rather be queryable and checkable than impressively static. The whole thing is grounded — ask it "what's the evidence?" and it answers with commit counts, tests, and live endpoints; the dated reasoning is browsable too. (The &lt;code&gt;resume-agent&lt;/code&gt; repo links to a live instance if you want to poke it.)&lt;/p&gt;

&lt;h2&gt;
  
  
  If you want to build one
&lt;/h2&gt;

&lt;p&gt;The smallest version is about ten minutes — fork the front, point it at any backend (even a stub &lt;code&gt;/info&lt;/code&gt;), and deploy. That's a real, queryable node. If you stand one up, I'd genuinely like to see it. You don't have to agree with where I think this goes; a working node is its own statement.&lt;/p&gt;

&lt;p&gt;This is the candidate side of a two-sided thing — an open protocol for hiring. The employer-side reference and the spec are open too, if the architecture pulls you further.&lt;/p&gt;

&lt;h2&gt;
  
  
  The repos
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;resume-agent&lt;/strong&gt; — the candidate-side backend, the profile-as-API. Links to a live instance for a demo: &lt;a href="https://github.com/yuens1002/resume-agent" rel="noopener noreferrer"&gt;https://github.com/yuens1002/resume-agent&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;resume-agent-web&lt;/strong&gt; — the forkable web front: &lt;a href="https://github.com/yuens1002/resume-agent-web" rel="noopener noreferrer"&gt;https://github.com/yuens1002/resume-agent-web&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;open-employment-protocol&lt;/strong&gt; — the reconciler that sits between the two sides: &lt;a href="https://github.com/yuens1002/open-employment-protocol" rel="noopener noreferrer"&gt;https://github.com/yuens1002/open-employment-protocol&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;employer-agent&lt;/strong&gt; — the employer side, scaffolded as a call to build: &lt;a href="https://github.com/yuens1002/employer-agent" rel="noopener noreferrer"&gt;https://github.com/yuens1002/employer-agent&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Built with a lot of AI help, deliberately unnamed — it wasn't one tool, it was the compounded work of everything that came before mine. Which is the kind of AI I'm building on: something you extend and pass forward.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>opensource</category>
      <category>ai</category>
      <category>webdev</category>
      <category>career</category>
    </item>
    <item>
      <title>You Don't Have to Learn Hermes From Scratch — I Brought My Existing Skills In</title>
      <dc:creator>sunny yuen</dc:creator>
      <pubDate>Thu, 28 May 2026 21:41:06 +0000</pubDate>
      <link>https://dev.to/yuens1002/you-dont-have-to-learn-hermes-from-scratch-i-brought-my-existing-skills-in-18f0</link>
      <guid>https://dev.to/yuens1002/you-dont-have-to-learn-hermes-from-scratch-i-brought-my-existing-skills-in-18f0</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/hermes-agent-2026-05-15"&gt;Hermes Agent Challenge&lt;/a&gt;: Write About Hermes Agent&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  I Didn't Start With Hermes
&lt;/h2&gt;

&lt;p&gt;Six months ago I started building a set of agent skills and personas for how I build software. Not generic prompts — opinionated role files. A &lt;code&gt;/backend-architect&lt;/code&gt; that owns schema and recommendation logic. A &lt;code&gt;/test-engineer&lt;/code&gt; that writes Vitest coverage and flags weak acceptance criteria. A &lt;code&gt;/project-manager&lt;/code&gt; that maintains planning docs and closes iterations cleanly.&lt;/p&gt;

&lt;p&gt;These roles have evolved across multiple projects. They have layering rules, discovery checklists, inheritance from a base engineering discipline file. They produce consistent, reviewable work because they're scoped — the backend architect doesn't touch test files, the test engineer doesn't redesign the schema, each persona has a defined mandate and exits cleanly.&lt;/p&gt;

&lt;p&gt;When I heard about Hermes Agent, my first instinct wasn't "let me learn a new system." It was: &lt;strong&gt;can I run my existing system inside this?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The answer is yes. That's what this article is about — what it looks like to bring a mature workflow into Hermes, what you gain, where it breaks down, and what I'd do differently.&lt;/p&gt;




&lt;h2&gt;
  
  
  What Hermes Is (and Isn't) to Someone Who Already Has a Workflow
&lt;/h2&gt;

&lt;p&gt;Hermes is an LLM-agnostic orchestration layer. It has its own skill system, its own &lt;code&gt;soul.md&lt;/code&gt; concept for persistent agent identity, built-in cron scheduling and MCP management. All of that is real and useful.&lt;/p&gt;

&lt;p&gt;But it's also a runtime. If you have skills that work, you can bring them in.&lt;/p&gt;

&lt;p&gt;I installed a local Hermes instance — few clicks, straightforward setup — and ran it inside VSCode's integrated terminal pointed at my existing persona files. No migration. No rewrite. My &lt;code&gt;/backend-architect&lt;/code&gt; runs in Hermes the same way it runs in Claude Code.&lt;/p&gt;

&lt;p&gt;Before settling on this, I'd tried a couple of other paths — a VPS instance with a Telegram interface for ideation, and attempting to build through a browser-based terminal. The VPS was fine for sketching ideas. The browser terminal made it clear that building production-grade tooling without proper local environment was the wrong path. Local Hermes in VSCode removed that friction.&lt;/p&gt;

&lt;p&gt;The thing that surprised me: you don't have to choose between "learn Hermes natively" and "keep using what works." You can do both at once. Hermes becomes the runtime. Your skills stay the structure.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Build: Production in 3 Days on a New System
&lt;/h2&gt;

&lt;p&gt;The project I built with this setup is &lt;a href="https://brew-guide-production.up.railway.app" rel="noopener noreferrer"&gt;Brew Guide&lt;/a&gt; — a community coffee knowledge base exposed as a public MCP server. It logs real brew experiments, builds consensus recommendations from that data, and returns technique guidance (bloom timing, pour stages, agitation style) via 5 MCP tools. Production endpoint, no auth, live now.&lt;/p&gt;

&lt;p&gt;I built it in 3 days — on Hermes, which I'd never used before.&lt;/p&gt;

&lt;p&gt;That number is the point. Not because the build was simple (Neon Postgres, Prisma migrations, 55 passing tests, Railway auto-deploy, strict TypeScript throughout), but because I didn't spend those 3 days learning Hermes. I spent them building. The workflow did the heavy lifting on an unfamiliar runtime because the workflow was already mature.&lt;/p&gt;

&lt;p&gt;The competition sprint — 7 deliverables including a scraper, technique JSONB on the schema, a landing page, and the competition article — reached "verified" on the first review pass. One iteration. That's the workflow functioning, not the AI being infallible.&lt;/p&gt;




&lt;h2&gt;
  
  
  What Breaks When You Switch Models
&lt;/h2&gt;

&lt;p&gt;Here's the more interesting part, and I want to be honest about it.&lt;/p&gt;

&lt;p&gt;During an earlier iteration, I ran the same skill set through a different model on the same codebase. Same persona files, same task scope, same Hermes orchestration. The goal: understand whether my workflow was portable across LLMs.&lt;/p&gt;

&lt;p&gt;The observable failure mode was &lt;strong&gt;tool call adherence&lt;/strong&gt;. The other model fumbled calls more often — retries, moments where it found its own path around the structured orchestration rather than following the skill specification. Tasks that took 30 minutes with Claude took most of a working day. The output required remediation: a Node version API call that crashed production, acceptance criteria tests that confirmed plumbing but not the scoring invariants the ACs required, docs that drifted from the code.&lt;/p&gt;

&lt;p&gt;I want to be careful about what I can and can't claim. I wasn't using Hermes' native model adapters at the time — the skills were running through the same interface I'd built for Claude. So I can't say definitively whether the gap was model capability or a Hermes-model integration issue. Both are plausible.&lt;/p&gt;

&lt;p&gt;What I can say: &lt;strong&gt;same instructions, same personas, dramatically different adherence to the spec&lt;/strong&gt;. My skills were written and refined on Claude's way of parsing structured instructions. When you hand those same instructions to a model with different parsing behaviour, adherence degrades — and tool call reliability is the first thing to break.&lt;/p&gt;

&lt;p&gt;This is the portability question Hermes is built to solve. It's a genuinely hard problem, and I didn't solve it. But surfacing where it breaks is the useful finding.&lt;/p&gt;




&lt;h2&gt;
  
  
  What I'd Build Differently
&lt;/h2&gt;

&lt;p&gt;The gap I felt throughout came from one step I skipped: I never migrated my skills to native Hermes format.&lt;/p&gt;

&lt;p&gt;Hermes has a &lt;code&gt;soul.md&lt;/code&gt; concept — a persistent document that shapes agent identity across sessions. Think of it as the context your agent carries into every conversation: its values, working style, constraints. My skills work without it, but they're missing an anchor. A &lt;code&gt;soul.md&lt;/code&gt; tuned to how I build — layering rules, persona boundaries, the engineering discipline that governs all my roles — would give Hermes native context that currently lives in my head, and make skills more robust across model handoffs.&lt;/p&gt;

&lt;p&gt;The second missing step: model-specific skill validation. My skills assume Claude's instruction-following behaviour. A proper migration would test each persona against multiple model families and adjust language and structure where adherence breaks down. That's what "native Hermes skills" gives you — not just ported files, but skills validated against the runtime you're actually using.&lt;/p&gt;

&lt;p&gt;The parts of Hermes where the DX is already smooth regardless: cron scheduling and &lt;code&gt;hermes mcp add&lt;/code&gt;. Setting up the weekly coffee literature automation as a cron job was trivial. Connecting the production MCP endpoint to any client is one command. These infrastructure pieces are where Hermes earns its keep without needing any skill migration at all.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Honest Verdict
&lt;/h2&gt;

&lt;p&gt;Six months of evolved Claude skills beats Hermes' out-of-the-box defaults — for the way I specifically build software.&lt;/p&gt;

&lt;p&gt;But that's the wrong comparison. The right question: &lt;strong&gt;does Hermes give you something you don't already have?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;For me, two things.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;LLM portability.&lt;/strong&gt; The ability to run your skills against a local model, a different provider, without rebuilding anything. For production work with tight quality requirements, I want Claude's reliability. For experiments, local automations, cost-sensitive builds: Hermes makes the option real without a rewrite tax.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The infrastructure layer.&lt;/strong&gt; Cron, MCP management, soul.md persistence. Not things you'd build for one project, but immediately useful once they exist.&lt;/p&gt;

&lt;p&gt;What I'd tell a developer starting out: &lt;strong&gt;bring your existing workflow in first.&lt;/strong&gt; Don't wait until you've learned Hermes natively before you build anything. Run your current skills, see what holds, see what breaks at the model boundary, and use that signal to decide what to migrate properly. The portability is the point — you don't earn it by starting over.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;The project this workflow produced:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Web: &lt;a href="https://brew-guide-production.up.railway.app" rel="noopener noreferrer"&gt;brew-guide-production.up.railway.app&lt;/a&gt; — no login, try it now&lt;/li&gt;
&lt;li&gt;MCP: &lt;code&gt;https://brew-guide-production.up.railway.app/mcp&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;GitHub: &lt;a href="https://github.com/yuens1002/brew-guide" rel="noopener noreferrer"&gt;yuens1002/brew-guide&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>hermesagentchallenge</category>
      <category>devchallenge</category>
      <category>agents</category>
      <category>devtools</category>
    </item>
    <item>
      <title>Every Great Cup Starts with the Right Question — I Built the Community Behind the Answer with Hermes Agent</title>
      <dc:creator>sunny yuen</dc:creator>
      <pubDate>Thu, 28 May 2026 21:29:01 +0000</pubDate>
      <link>https://dev.to/yuens1002/every-great-cup-starts-with-the-right-question-i-built-the-community-behind-the-answer-with-o04</link>
      <guid>https://dev.to/yuens1002/every-great-cup-starts-with-the-right-question-i-built-the-community-behind-the-answer-with-o04</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/hermes-agent-2026-05-15"&gt;Hermes Agent Challenge&lt;/a&gt;: Build With Hermes Agent&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;Real brewing knowledge lives in human experience — in roaster guides, in community notes, in what a barista learned from last Tuesday's pour. It doesn't accumulate anywhere. Every brew is forgotten. Ask any AI and you get statistical averages: 93°C, 1:16 ratio, four minutes. Technically defensible. Practically generic. Worse still for rare origins where training data is thin.&lt;/p&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;h3&gt;
  
  
  For coffee drinkers
&lt;/h3&gt;

&lt;p&gt;Visit &lt;a href="https://brew-guide-production.up.railway.app" rel="noopener noreferrer"&gt;brew-guide-production.up.railway.app&lt;/a&gt;. No account. No setup. No AI client required.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fsavrk9g4wy65zzbmahb1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fsavrk9g4wy65zzbmahb1.png" alt="Brew Guide — " width="799" height="440"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Pick your coffee origin, roast level, and brew method. What comes back isn't a generic recipe — it's community consensus: the grind, temperature, ratio, and brew time that real people have logged and rated for that origin, plus step-by-step technique guidance (bloom timing, pour stages, agitation style). If data is sparse for your origin, the confidence tier says so honestly and falls back to method defaults rather than making something up.&lt;/p&gt;

&lt;p&gt;This is for the person who just picked up a bag of Kenyan peaberry and wants to know how to do it justice. It works for anyone who cares about their cup — no technical knowledge required.&lt;/p&gt;

&lt;h3&gt;
  
  
  For developers and AI clients
&lt;/h3&gt;

&lt;p&gt;Connect to any MCP-capable client in one line:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;https://brew-guide-production.up.railway.app/mcp
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Ask your AI: &lt;em&gt;"recommend a pour over for Ethiopian light roast."&lt;/em&gt; What comes back is a traceable community consensus object: brew parameters, a confidence tier (high/medium/low), the source brews that contributed, and method-specific technique guidance. You can see where the knowledge came from and how certain the system is — a fundamentally different epistemic object from an AI-generated recipe.&lt;/p&gt;

&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;GitHub:&lt;/strong&gt; &lt;a href="https://github.com/yuens1002/brew-guide" rel="noopener noreferrer"&gt;yuens1002/brew-guide&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Five MCP tools — &lt;code&gt;get_brewing_methods&lt;/code&gt;, &lt;code&gt;recommend&lt;/code&gt;, &lt;code&gt;log_brew&lt;/code&gt;, &lt;code&gt;search_brews&lt;/code&gt;, &lt;code&gt;compare_brew&lt;/code&gt; — over Streamable HTTP transport. Public, no auth required.&lt;/p&gt;

&lt;h3&gt;
  
  
  My Tech Stack
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Technology&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;HTTP&lt;/td&gt;
&lt;td&gt;Hono 4 + &lt;code&gt;@hono/node-server&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MCP&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;@modelcontextprotocol/sdk&lt;/code&gt; + &lt;code&gt;@hono/mcp&lt;/code&gt; (Streamable HTTP)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Database&lt;/td&gt;
&lt;td&gt;Neon Postgres + Prisma ORM&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Runtime&lt;/td&gt;
&lt;td&gt;Node 24, TypeScript strict, ESM&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tests&lt;/td&gt;
&lt;td&gt;Vitest (55 tests, 0 errors)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Deploy&lt;/td&gt;
&lt;td&gt;Railway (auto-deploy from &lt;code&gt;main&lt;/code&gt;)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The recommendation engine is fully deterministic — no LLM on the hot path. &lt;code&gt;computeBestBrew()&lt;/code&gt; fetches up to 50 recent brews, scores each against your request params (origin, method, roast, variety, grind), applies recency decay (linear 1.0 → 0.1 over 365 days) and source trust weights, takes the top 5, and builds consensus via weighted average (numeric fields) or weighted mode (categorical). Sub-100ms. Reproducible. Auditable.&lt;/p&gt;

&lt;p&gt;The voting infrastructure is live (&lt;code&gt;thumbs_up&lt;/code&gt;/&lt;code&gt;thumbs_down&lt;/code&gt; on recommendations, &lt;code&gt;brew_recommendation_links&lt;/code&gt; tracking which brews followed which recommendations). Vote weighting inside &lt;code&gt;computeBestBrew&lt;/code&gt; is the one acknowledged gap — the checks-and-balances mechanism is designed, the math isn't wired yet. That's the next commit.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I Used Hermes Agent
&lt;/h2&gt;

&lt;p&gt;I built this in 3 days — on a system I'd never used before.&lt;/p&gt;

&lt;p&gt;That's the headline, and it's what I want to explain. I didn't start from scratch on Hermes. I installed a local instance, ran it inside VSCode's terminal, and pointed it at persona files I'd spent six months building. Three role files govern the entire build:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;/backend-architect&lt;/code&gt; — owns schema design, the recommendation engine, all DB logic&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;/test-engineer&lt;/code&gt; — owns Vitest coverage, catches weak ACs, flags regressions&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;/project-manager&lt;/code&gt; — owns planning docs, retrospectives, and this article&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Each persona has a focused mandate and a defined exit condition. The backend architect doesn't touch test files. The test engineer doesn't redesign the schema.&lt;/p&gt;

&lt;p&gt;What the workflow looks like in practice:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Write a plan with a deliverable table (D1–D7), each with owner, files, acceptance criteria, commit schedule&lt;/li&gt;
&lt;li&gt;Hand each deliverable to the relevant persona&lt;/li&gt;
&lt;li&gt;After each feature, the test engineer verifies coverage and flags gaps&lt;/li&gt;
&lt;li&gt;Review report before merge&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The competition sprint — scraper, technique JSONB, landing page, this article — reached "verified" on the first review pass. 55 tests, zero TypeScript errors, all deliverables complete. One iteration.&lt;/p&gt;

&lt;p&gt;The previous iteration of this codebase was run on a different model through the same Hermes setup. Similar scope, same skill set, same orchestration. That iteration required production hotfixes (a Node version API crash on Railway), remediation of weak acceptance criteria tests, and took the better part of a working day. The gap in tool call adherence — the other model fumbling calls and finding workarounds around the skill spec rather than through it — was the visible failure mode.&lt;/p&gt;

&lt;p&gt;Hermes as a runtime made that comparison possible without changing anything in the workflow. Same skills, same personas, different model. The portability is the point.&lt;/p&gt;

&lt;p&gt;Beyond orchestration, Hermes added two things directly: cron scheduling for the weekly coffee literature automation (&lt;code&gt;hermes-automation/&lt;/code&gt;), and &lt;code&gt;hermes mcp add&lt;/code&gt; for connecting the production endpoint to any client instantly. That MCP management DX is genuinely smooth.&lt;/p&gt;

&lt;p&gt;What I'd build differently next time: migrate the skills to native Hermes format with a &lt;code&gt;soul.md&lt;/code&gt; as a persistent identity anchor. The skills work as-is, but validating them against multiple model families — adjusting language and structure where tool call adherence degrades — is the proper portability work I didn't have time for. That's the experiment this project sets up.&lt;/p&gt;

</description>
      <category>hermesagentchallenge</category>
      <category>devchallenge</category>
      <category>agents</category>
      <category>mcp</category>
    </item>
  </channel>
</rss>
