<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: infracore</title>
    <description>The latest articles on DEV Community by infracore (@infracore).</description>
    <link>https://dev.to/infracore</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4111868%2F1a80ab53-5b26-4268-8a23-0540ffa5ed81.png</url>
      <title>DEV Community: infracore</title>
      <link>https://dev.to/infracore</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/infracore"/>
    <language>en</language>
    <item>
      <title>Audit an agent skill like a pull request before it gets file access</title>
      <dc:creator>infracore</dc:creator>
      <pubDate>Tue, 22 Sep 2026 12:11:26 +0000</pubDate>
      <link>https://dev.to/infracore/audit-an-agent-skill-like-a-pull-request-before-it-gets-file-access-3p8e</link>
      <guid>https://dev.to/infracore/audit-an-agent-skill-like-a-pull-request-before-it-gets-file-access-3p8e</guid>
      <description>&lt;p&gt;Installing a skill or MCP server hands someone else's instructions your files, credentials, and shell. Treat it like merging unreviewed code, not adding docs.&lt;/p&gt;

&lt;p&gt;A credible DIY baseline covers most one-off installs if you install rarely. Clone or unpack it outside your worktree, read the manifest and tool definitions first, then grep for shell execution, file writes, reads of env files and SSH keys, and outbound network calls. Run the first trial in a throwaway container or separate user with no live secrets.&lt;/p&gt;

&lt;p&gt;That leaves recurring gaps:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;static hits without context: a flag tells you &lt;code&gt;curl&lt;/code&gt; exists, not whether it exfiltrates data or fetches a schema&lt;/li&gt;
&lt;li&gt;install-time vs runtime drift: what you read today is not what runs after an auto-update&lt;/li&gt;
&lt;li&gt;review fatigue: nobody re-reads every dependency on every update&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A practical method that helps: a pre-run check bot for the boring second pass. Point it at the unpacked skill directory, open the code around each risky call, classify read-only versus write versus network behavior, and output a short verdict with file paths and line numbers. Keep that verdict next to the install so the next update gets a diff, not a fresh full review.&lt;/p&gt;

&lt;p&gt;What workaround actually sticks for you: blocking installs by default, sandboxing every new skill, or re-auditing on each version bump?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>mcp</category>
      <category>security</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Turn Screen Recordings Into Bug Reports Agents Can Act On</title>
      <dc:creator>infracore</dc:creator>
      <pubDate>Mon, 21 Sep 2026 12:56:12 +0000</pubDate>
      <link>https://dev.to/infracore/turn-screen-recordings-into-bug-reports-agents-can-act-on-1o5m</link>
      <guid>https://dev.to/infracore/turn-screen-recordings-into-bug-reports-agents-can-act-on-1o5m</guid>
      <description>&lt;p&gt;A 30-second clip of a flickering dropdown tells an agent almost nothing. Without steps, expected behavior, and environment, the model fills gaps with plausible but wrong fixes.&lt;/p&gt;

&lt;p&gt;A solid manual baseline works for most teams. Rewatch the recording once and fill a fixed template:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;repro steps with timestamps from the clip&lt;/li&gt;
&lt;li&gt;expected vs actual behavior in one sentence each&lt;/li&gt;
&lt;li&gt;browser, viewport, test account state, and feature flags&lt;/li&gt;
&lt;li&gt;console errors and failing request URLs&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Paste that Markdown with the video link into the agent prompt. That alone removes most back-and-forth about what broke and where.&lt;/p&gt;

&lt;p&gt;Keep clips under a minute and trim to one bug per file. If the repro needs login or seeded data, note the setup commands so the agent can reproduce locally.&lt;/p&gt;

&lt;p&gt;The remaining gap is assembly cost. Copying timestamps, selectors, and logs by hand is slow and easy to skip under pressure. Capture-to-Markdown tooling can export the same structure automatically after one recording instead of typing it.&lt;/p&gt;

&lt;p&gt;What do you include in UI bug reports for agents today, and what part do you still skip because it takes too long to capture?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>mcp</category>
      <category>opensource</category>
      <category>testing</category>
    </item>
    <item>
      <title>MCP Is an Adapter Layer, So Version the API First</title>
      <dc:creator>infracore</dc:creator>
      <pubDate>Sat, 19 Sep 2026 22:11:01 +0000</pubDate>
      <link>https://dev.to/infracore/mcp-is-an-adapter-layer-so-version-the-api-first-1m5j</link>
      <guid>https://dev.to/infracore/mcp-is-an-adapter-layer-so-version-the-api-first-1m5j</guid>
      <description>&lt;p&gt;If an MCP server is usually a thin layer over an API, the practical takeaway is simple: &lt;strong&gt;treat the API contract as the thing that can actually break you&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;A useful method is to review changes in this order:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;API surface:&lt;/strong&gt; endpoints, required params, response fields, enums, auth shape&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MCP mapping:&lt;/strong&gt; which API fields become tool inputs/outputs&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Agent behavior:&lt;/strong&gt; prompts, tool choice, retry logic, and error handling&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That order matters because MCP can make an integration look stable while the underlying API has already shifted. A renamed field, a newly required parameter, or an enum value removal may not show up as an obvious MCP problem at first. It often lands later as vague tool failure, bad completions, or agents taking the wrong branch.&lt;/p&gt;

&lt;p&gt;The solid DIY baseline is enough for many teams: keep an OpenAPI spec, diff it in CI, and manually classify changes as breaking or non-breaking before updating the MCP wrapper. If your API is small and the tool surface is narrow, that can be perfectly sufficient.&lt;/p&gt;

&lt;p&gt;The remaining gap is consistency when lots of small schema edits pile up. The MCP layer may stay thin, but the review burden does not. Safest habit: &lt;strong&gt;version and gate the API first, then regenerate or update the MCP adapter second&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;What's the most annoying break you've seen in practice: required fields changing, enum drift, auth changes, or something else at the MCP-to-API boundary?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>mcp</category>
      <category>opensource</category>
      <category>api</category>
    </item>
    <item>
      <title>Expose Crypto KAT Runners as MCP Tools Instead of Pasting Hex</title>
      <dc:creator>infracore</dc:creator>
      <pubDate>Fri, 18 Sep 2026 21:27:45 +0000</pubDate>
      <link>https://dev.to/infracore/expose-crypto-kat-runners-as-mcp-tools-instead-of-pasting-hex-122l</link>
      <guid>https://dev.to/infracore/expose-crypto-kat-runners-as-mcp-tools-instead-of-pasting-hex-122l</guid>
      <description>&lt;p&gt;When a Known Answer Test fails on a crypto primitive, the cause is often byte-ordering or padding. Pasting that raw hex plus the ACVP JSON into a chat window works once, but it does not scale to large vector files.&lt;/p&gt;

&lt;p&gt;A credible baseline stays manual: run the KAT binary locally, copy only the failing vector ID and byte offset into the prompt, fix, and re-run. For a single algorithm and a small file, that loop is sufficient and easier to audit than adding server code.&lt;/p&gt;

&lt;p&gt;The gap appears with massive ACVP payloads and repeated re-runs. One method is to put the runner behind Model Context Protocol over stdio:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;tool &lt;code&gt;run_crypto_kat(target, algorithm)&lt;/code&gt; returning structured status, failing case ID, and diff window&lt;/li&gt;
&lt;li&gt;tool &lt;code&gt;load_acvp_vectors(path)&lt;/code&gt; with filtering or paging instead of dumping the whole file&lt;/li&gt;
&lt;li&gt;resources for expected schema layout and tail of the latest log&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Keep tool outputs small and typed: verdict, file, case, expected vs actual slice, and next command. Let the agent request more context explicitly rather than pushing the whole file by default.&lt;/p&gt;

&lt;p&gt;What workaround do you use today to keep large KAT vectors out of the prompt while still giving the agent enough to fix the byte-level bug?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>mcp</category>
      <category>testing</category>
      <category>github</category>
    </item>
    <item>
      <title>A pre-flight gate for vague MCP store rejections</title>
      <dc:creator>infracore</dc:creator>
      <pubDate>Thu, 17 Sep 2026 22:36:17 +0000</pubDate>
      <link>https://dev.to/infracore/a-pre-flight-gate-for-vague-mcp-store-rejections-56gp</link>
      <guid>https://dev.to/infracore/a-pre-flight-gate-for-vague-mcp-store-rejections-56gp</guid>
      <description>&lt;p&gt;The failure class to design for is this: your MCP server works against local clients, then stalls in a plugin store review with feedback too vague to act on.&lt;/p&gt;

&lt;p&gt;Freeze what you submitted. Pin the manifest, tool list with JSON Schemas, auth scopes, and example transcripts to one commit. Then run the same pre-flight before every resubmit:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;validate all tool input/output schemas and examples with strict validators&lt;/li&gt;
&lt;li&gt;diff tools and required fields against your last submitted commit&lt;/li&gt;
&lt;li&gt;check scope minimization, stable names/descriptions, and consistent error shapes&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A credible DIY baseline is plain CI: Ajv or Zod plus a JSON diff and saved fixtures. For infrequent releases and a small tool surface, that catches schema drift and accidental breaking renames without new tooling.&lt;/p&gt;

&lt;p&gt;The remaining gap is reviewability. A green CI run does not show a reviewer exactly what changed and what you checked. Keep the commit pair and check logs together so a rejection can be answered with specifics instead of guesses. That bundle matters most when feedback is one vague line.&lt;/p&gt;

&lt;p&gt;What pre-submit check has actually caught a store rejection for you before you hit submit?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>mcp</category>
      <category>opensource</category>
      <category>devtools</category>
    </item>
    <item>
      <title>Check renewable sessions before forcing a browser re-login</title>
      <dc:creator>infracore</dc:creator>
      <pubDate>Wed, 16 Sep 2026 16:46:24 +0000</pubDate>
      <link>https://dev.to/infracore/check-renewable-sessions-before-forcing-a-browser-re-login-2ifi</link>
      <guid>https://dev.to/infracore/check-renewable-sessions-before-forcing-a-browser-re-login-2ifi</guid>
      <description>&lt;p&gt;Provider OAuth sessions can fail in a confusing way: a short-lived access token expires while a usable refresh token is still stored. A background quota probe then gets a 401, and the UI makes a renewable session look like it needs a new browser login.&lt;/p&gt;

&lt;p&gt;Fix it with a manual, explicitly quota-consuming session check. On operator click only - never from status polling - run one isolated request through the provider's supported path, using a fixed minimal prompt with no tools, no MCP servers, safe mode, and no conversation persistence. Hold the same exclusive process slot as tasks and login, then compare safe credential metadata before and after without logging or returning token material.&lt;/p&gt;

&lt;p&gt;Report a normalized outcome so operators stop guessing:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;session renewed&lt;/li&gt;
&lt;li&gt;session already current&lt;/li&gt;
&lt;li&gt;session valid but quota exhausted&lt;/li&gt;
&lt;li&gt;re-authentication required&lt;/li&gt;
&lt;li&gt;check failed&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The DIY baseline is enough for most teams: a single serialized button, clear disclosure that it may consume quota, and a separate Re-authenticate action. It breaks down when adapters guess across vendors, or when quota exhaustion is mislabeled as auth failure and clears the wrong backoff.&lt;/p&gt;

&lt;p&gt;Keep the control-plane operation vendor-neutral, advertise quota cost, and make unsupported adapters fail loudly. Only a rotated credential should clear a stale unauthenticated backoff and allow one immediate quota retry.&lt;/p&gt;

&lt;p&gt;How do you keep quota-exhausted separate from truly expired in your session UI today?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>mcp</category>
      <category>security</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Don't Trust Green Karate Runs: Diff Them Against OpenAPI</title>
      <dc:creator>infracore</dc:creator>
      <pubDate>Tue, 15 Sep 2026 16:39:31 +0000</pubDate>
      <link>https://dev.to/infracore/dont-trust-green-karate-runs-diff-them-against-openapi-50ff</link>
      <guid>https://dev.to/infracore/dont-trust-green-karate-runs-diff-them-against-openapi-50ff</guid>
      <description>&lt;p&gt;A passing Karate suite only proves the scenarios you wrote still pass. It says nothing about OpenAPI operations that were never called. That gap grows when specs add endpoints, deprecate fields, or rename routes faster than feature files are updated.&lt;/p&gt;

&lt;p&gt;A direct check is to diff exercised calls against the spec instead of trusting the green run. A small script is often enough for small services.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;List every path + method pair from your OpenAPI document as the expected set.&lt;/li&gt;
&lt;li&gt;Collect the actual set by parsing Karate feature files for url, path, and method, or by logging requests during a run.&lt;/li&gt;
&lt;li&gt;Join the two sets to find untested operations, then prioritize by auth scope and breaking-change risk.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Generate stubs only for the missing operations, with one happy path and one auth or validation failure each. Keep the coverage diff in version control so the next spec change shows which operations lost coverage.&lt;/p&gt;

&lt;p&gt;For larger specs, the harder part is keeping that mapping stable across refactors, parameterized paths, and versioned routes without hand-maintaining aliases.&lt;/p&gt;

&lt;p&gt;How do you currently detect OpenAPI endpoints your Karate suite never exercises?&lt;/p&gt;

</description>
      <category>testing</category>
      <category>api</category>
      <category>ai</category>
      <category>devops</category>
    </item>
    <item>
      <title>Keeping One MCP Server Working Across SDK Renames and Three CLIs</title>
      <dc:creator>infracore</dc:creator>
      <pubDate>Mon, 14 Sep 2026 23:47:59 +0000</pubDate>
      <link>https://dev.to/infracore/keeping-one-mcp-server-working-across-sdk-renames-and-three-clis-2g7k</link>
      <guid>https://dev.to/infracore/keeping-one-mcp-server-working-across-sdk-renames-and-three-clis-2g7k</guid>
      <description>&lt;p&gt;An MCP server breaks at three seams: the SDK class name, the upstream request shape, and each CLI's server config. Treating those as one upgrade is what makes a small rename expensive.&lt;/p&gt;

&lt;p&gt;Put a thin adapter between your tools and the vendor SDK. Keep all imports and request construction in one module, so a FastMCP to MCPServer or genai client bump touches one file. Add a startup check that lists tools, sends one tiny generation call, and prints the resolved model, endpoint, and SDK version on failure.&lt;/p&gt;

&lt;p&gt;DIY baseline that covers the rename case: pin exact versions in the lockfile, run that smoke script in CI, and keep one config snippet per CLI pointing at the same entrypoint. For a single-maintainer server, that is usually enough.&lt;/p&gt;

&lt;p&gt;The remaining gap is silent divergence: one CLI passes env or cwd differently, or the Interactions API returns 400 for a field shape the old SDK sent. Log the outbound payload bytes on non-2xx and diff them against the last known-good call before changing code.&lt;/p&gt;

&lt;p&gt;What workaround have you used when the same MCP server passes in one CLI but fails in another?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>mcp</category>
      <category>devtools</category>
      <category>api</category>
    </item>
    <item>
      <title>Wrapping foreign MCP and skill metadata without silent authority</title>
      <dc:creator>infracore</dc:creator>
      <pubDate>Sun, 13 Sep 2026 15:40:22 +0000</pubDate>
      <link>https://dev.to/infracore/wrapping-foreign-mcp-and-skill-metadata-without-silent-authority-363f</link>
      <guid>https://dev.to/infracore/wrapping-foreign-mcp-and-skill-metadata-without-silent-authority-363f</guid>
      <description>&lt;p&gt;When you pull an existing MCP server or an OpenAI/Claude-style skill into a local agent package model, the hard part is not the install. It is turning foreign tool metadata into something reviewable: provenance and version kept intact, network/filesystem/secret needs mapped to explicit permissions, and hooks left off until someone approves them.&lt;/p&gt;

&lt;p&gt;A credible baseline is still manual. Read the server or skill manifest, list every tool action, note undeclared filesystem or network reach, and reject or mark unsupported anything you cannot map. Auto-enable is how permanent owner-memory write and silent background behavior sneak in.&lt;/p&gt;

&lt;p&gt;Running imports through the same validator and permission model as native packages only helps if unmappable authority fails closed and compatibility status stays visible (native, wrapped, partial, rejected).&lt;/p&gt;

&lt;p&gt;How do you currently record provenance and permission gaps when wrapping a third-party MCP server so a later review can tell those four outcomes apart?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>mcp</category>
      <category>security</category>
      <category>github</category>
    </item>
    <item>
      <title>Stale entries in agent discovery manifests</title>
      <dc:creator>infracore</dc:creator>
      <pubDate>Sat, 12 Sep 2026 16:07:24 +0000</pubDate>
      <link>https://dev.to/infracore/stale-entries-in-agent-discovery-manifests-3mc6</link>
      <guid>https://dev.to/infracore/stale-entries-in-agent-discovery-manifests-3mc6</guid>
      <description>&lt;p&gt;Publishing &lt;code&gt;/.well-known/ai-catalog.json&lt;/code&gt; is one way to list MCP servers, A2A agents, and OpenAPI schemas for agent discovery. The shape is simple: specVersion, host, and an entries array.&lt;/p&gt;

&lt;p&gt;The failure mode is quieter than a missing file. A catalog can stay up while an entry still points at a moved MCP URL, a renamed skill, or an OpenAPI document that no longer matches the live API.&lt;/p&gt;

&lt;p&gt;DIY checks that catch this without new tooling:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;curl each entry URL in CI and require a non-error response&lt;/li&gt;
&lt;li&gt;diff listed names against what the live MCP or schema endpoint actually returns&lt;/li&gt;
&lt;li&gt;fail the job when the catalog and runtime disagree&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;When you change an MCP mount path or ship a breaking schema edit, what process updates the ARD manifest, and what check do you run so agents do not keep discovering the stale entry?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>mcp</category>
      <category>github</category>
      <category>api</category>
    </item>
    <item>
      <title>Hard caps on autonomous agent API spend</title>
      <dc:creator>infracore</dc:creator>
      <pubDate>Fri, 11 Sep 2026 23:05:10 +0000</pubDate>
      <link>https://dev.to/infracore/hard-caps-on-autonomous-agent-api-spend-1i02</link>
      <guid>https://dev.to/infracore/hard-caps-on-autonomous-agent-api-spend-1i02</guid>
      <description>&lt;p&gt;Autonomous agents in long research loops can exhaust an API budget before anyone notices. Provider dashboards and account-level alerts help after the fact, but they often fire too late for a single runaway session.&lt;/p&gt;

&lt;p&gt;A practical DIY baseline is a thin client wrapper that:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;tracks cumulative token or dollar cost per session&lt;/li&gt;
&lt;li&gt;refuses the next call once a hard session ceiling is hit&lt;/li&gt;
&lt;li&gt;logs the stop reason so the operator can inspect what burned the budget&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Provider key quotas and org spend limits still matter as a backstop when the wrapper fails open.&lt;/p&gt;

&lt;p&gt;When you run agents against paid model APIs, where do you enforce the hard stop: in the agent loop, at an API gateway, or only via the provider's billing controls?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>github</category>
      <category>api</category>
    </item>
    <item>
      <title>Hunting Down a 25-Second Network Ghost</title>
      <dc:creator>infracore</dc:creator>
      <pubDate>Thu, 10 Sep 2026 13:34:47 +0000</pubDate>
      <link>https://dev.to/infracore/hunting-down-a-25-second-network-ghost-4jpe</link>
      <guid>https://dev.to/infracore/hunting-down-a-25-second-network-ghost-4jpe</guid>
      <description>&lt;p&gt;Some HTTP MCP clients log a basic connectivity pre-check, finish TCP/TLS in well under a second, then sit idle for a tight ~25s band before logging success and sending the real request.&lt;/p&gt;

&lt;p&gt;When a capture shows the handshake fully ACKed and zero application bytes until the client closes, the stall is almost certainly in client logic: a fixed timer, a probe that never writes, or negotiation that only continues after a timeout.&lt;/p&gt;

&lt;p&gt;Useful isolation without owning the server:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Wall-clock from the "testing connectivity" log to the first client-sent application byte&lt;/li&gt;
&lt;li&gt;Whether an HTTP/2 preface, SETTINGS, or any request leaves during the gap&lt;/li&gt;
&lt;li&gt;Local config beside the server URL for protocol or feature-flag state that could gate the next step&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Which client pre-check designs have you seen that wait a fixed multi-second delay after the socket is already live?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>mcp</category>
      <category>testing</category>
      <category>github</category>
    </item>
  </channel>
</rss>
