<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: David Díaz</title>
    <description>The latest articles on DEV Community by David Díaz (@dd8888).</description>
    <link>https://dev.to/dd8888</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F552620%2F7e44501a-31d2-4e1a-a5e2-6e71bc5bc737.jpg</url>
      <title>DEV Community: David Díaz</title>
      <link>https://dev.to/dd8888</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/dd8888"/>
    <language>en</language>
    <item>
      <title>anything2explainer Packages Remotion Explainers as an Agent Skill</title>
      <dc:creator>David Díaz</dc:creator>
      <pubDate>Sat, 19 Sep 2026 19:40:10 +0000</pubDate>
      <link>https://dev.to/dd8888/anything2explainer-packages-remotion-explainers-as-an-agent-skill-19g9</link>
      <guid>https://dev.to/dd8888/anything2explainer-packages-remotion-explainers-as-an-agent-skill-19g9</guid>
      <description>&lt;p&gt;&lt;a href="https://github.com/Vincentwei1021/anything2explainer" rel="noopener noreferrer"&gt;anything2explainer&lt;/a&gt; packages a Claude Code and Codex skill that turns a topic, article or document into a narrated explainer video. Developers get a code-first production path: every frame is drawn with Remotion in React and TypeScript, rather than assembled from stock footage or produced by a generative video model, so individual shots remain editable as source code.&lt;/p&gt;

&lt;p&gt;The project targets a 1280×720, 30fps H.264 MP4 with synchronized voiceover, word-aligned subtitles, chapter cards, a top HUD and a chapter progress bar. It supports Chinese and English and delivers the research document, narration, storyboard, per-shot source code and QC reports alongside the finished video, according to the &lt;a href="https://github.com/Vincentwei1021/anything2explainer" rel="noopener noreferrer"&gt;repository specification&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  A production method, not a CLI
&lt;/h2&gt;

&lt;p&gt;The repository explicitly describes the project as something other than a CLI. It ships a compilable Remotion template, visual primitives, lighting components, voiceover and rendering tools, style and motion specifications, a multi-agent work protocol and a complete reference film intended to establish the quality target for a run, as detailed in the &lt;a href="https://github.com/Vincentwei1021/anything2explainer" rel="noopener noreferrer"&gt;project documentation&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Its prescribed workflow has nine stages: scaffold the project; research the subject; write the narration and frame-accurate timeline; storyboard each shot; build overlays and topic-specific primitives; produce a pilot group and 30-second preview; construct the remaining shots in parallel; render the full film and collect frame metrics; then run QC, apply fixes, re-verify the result and prepare delivery notes. Build agents receive groups of five to seven shots and write Remotion components for them, according to the &lt;a href="https://github.com/Vincentwei1021/anything2explainer" rel="noopener noreferrer"&gt;documented process&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Narration structure controls the edit. A blank line marks both a paragraph and a shot, while TTS word boundaries feed the subtitle table and frame-level timeline. The repository says shots retain a short hold after their final element appears, making the completed film longer than the raw speech by design &lt;a href="https://github.com/Vincentwei1021/anything2explainer" rel="noopener noreferrer"&gt;under its timing rules&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The process stops at four checkpoints: selection of length and language, narration approval, choice of voiceover engine and review of the first 30 seconds. The narration checkpoint occurs before voice generation because shot frame numbers are subsequently fixed to that timeline, while the preview is rendered before all shot groups are built &lt;a href="https://github.com/Vincentwei1021/anything2explainer" rel="noopener noreferrer"&gt;to allow an earlier style decision&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The repository estimates one to three hours of wall-clock time depending on video length. Its reference tier for a three-to-five-minute film specifies 40–50 shots, eight build agents and roughly two hours; these figures are the project’s planning estimates rather than an external benchmark &lt;a href="https://github.com/Vincentwei1021/anything2explainer" rel="noopener noreferrer"&gt;published by the maintainer&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Manual control remains available
&lt;/h2&gt;

&lt;p&gt;For agent use, the documented installation symlinks the repository into either the Claude Code or Codex skills directory. That setup is optional: developers can instead invoke the project template, TTS, storyboard, preview, render and frame-metrics scripts by hand &lt;a href="https://github.com/Vincentwei1021/anything2explainer" rel="noopener noreferrer"&gt;using the documented sequence&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The listed environment includes Node 18 or later, FFmpeg and Python packages selected partly by the TTS engine. The scripts were developed and verified on macOS; the repository says Linux should work, documents a verified Raspberry Pi 5 configuration and labels Windows untested &lt;a href="https://github.com/Vincentwei1021/anything2explainer" rel="noopener noreferrer"&gt;in its platform notes&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Analysis: inspectability versus pipeline ownership
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Analysis:&lt;/strong&gt; The project’s defensible distinction is not automatic video creation by itself, but the conversion of production decisions into inspectable artifacts: sourced research, narration, timing tables, storyboards and shot components. That structure may make a factual correction or visual revision easier to isolate than it would be in a single generated clip. The same design leaves the developer responsible for dependencies, TTS selection, agent execution and final review, all of which remain explicit parts of the &lt;a href="https://github.com/Vincentwei1021/anything2explainer" rel="noopener noreferrer"&gt;documented workflow&lt;/a&gt;. The unresolved trade-off is whether that added control justifies operating a multi-stage video toolchain rather than using a less editable service.&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>react</category>
      <category>typescript</category>
    </item>
    <item>
      <title>Awesome-Astra Maps Reported GPT-6 Astra Robotics Demos</title>
      <dc:creator>David Díaz</dc:creator>
      <pubDate>Thu, 17 Sep 2026 21:24:57 +0000</pubDate>
      <link>https://dev.to/dd8888/awesome-astra-maps-reported-gpt-6-astra-robotics-demos-30h7</link>
      <guid>https://dev.to/dd8888/awesome-astra-maps-reported-gpt-6-astra-robotics-demos-30h7</guid>
      <description>&lt;p&gt;&lt;a href="https://github.com/zjwzcx/Awesome-Astra-Embodied-AI" rel="noopener noreferrer"&gt;Awesome-Astra-Embodied-AI&lt;/a&gt;, a public GitHub repository, has assembled reported GPT-6 Astra robotics demonstrations across simulation, physical deployment, policy calls, real-to-sim replay and reinforcement-learning workflows. For developers, the practical consequence is a consolidated index whose case notes identify roles ranging from planning and trajectory generation to direct control and environment construction.&lt;/p&gt;

&lt;p&gt;The README lists 12 simulation cases, 10 real-world cases, one agentic policy call, six real-to-sim replay or data-rollout cases and six RL environment and training cases. &lt;a href="https://github.com/zjwzcx/Awesome-Astra-Embodied-AI" rel="noopener noreferrer"&gt;The repository presents those totals in its contents&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Different control boundaries under one label
&lt;/h2&gt;

&lt;p&gt;The simulated Unitree G1 cola-bottle case assigns high-level planning to Astra, while GEAR-SONIC converts the plan into a whole-body qpos trajectory for execution in Isaac Sim. A separate G1 navigation entry describes the same division between Astra’s navigation plan and GEAR-SONIC’s trajectory generation. &lt;a href="https://github.com/zjwzcx/Awesome-Astra-Embodied-AI" rel="noopener noreferrer"&gt;Both cases are described in the repository&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The FluxVLA case uses another architecture: Astra performs task inference and planning, but a pretrained FluxVLA policy executes the low-level embodied actions. The repository explicitly limits the entry’s use of “zero-shot” to Astra’s part of that pipeline. &lt;a href="https://github.com/zjwzcx/Awesome-Astra-Embodied-AI" rel="noopener noreferrer"&gt;That control split appears in the FluxVLA case description&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Other entries document constraints that narrow their claims. In the Dual-ALOHA demonstration, the claw task is a kinematic replay beginning from a pregrasped state, while the rope task uses native MuJoCo dynamics. Both use ideal-grasp assumptions and show specific planned motions rather than an online policy. &lt;a href="https://github.com/zjwzcx/Awesome-Astra-Embodied-AI" rel="noopener noreferrer"&gt;The README states those conditions alongside the case&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Physical control, reconstruction and training
&lt;/h2&gt;

&lt;p&gt;The 10 real-world reports include keyboard operation, marker grasping, plug insertion, mobile manipulation, cucumber slicing and Piper pick-and-place. Their interfaces also differ: one case has Astra output end-effector poses using third-person and wrist cameras, while another describes direct robot-arm control through Loop-ROS. &lt;a href="https://github.com/zjwzcx/Awesome-Astra-Embodied-AI" rel="noopener noreferrer"&gt;The real-world section lists these demonstrations and deployment details&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The collection extends beyond robot control. Its real-to-sim section includes a kitchen reconstructed from monocular RGB video, dexterous-hand motion reconstruction and a multi-view workflow combining robot actions, calibration, assets, system identification, MuJoCo and Blender. &lt;a href="https://github.com/zjwzcx/Awesome-Astra-Embodied-AI" rel="noopener noreferrer"&gt;Those reconstruction cases are listed in the repository&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;One training entry reports that Astra created a pen mesh, implemented a Sharpa-hand pen-spinning task in Isaac Lab, trained a PPO policy and produced a visualization video. &lt;a href="https://github.com/zjwzcx/Awesome-Astra-Embodied-AI" rel="noopener noreferrer"&gt;The case is presented as an RL environment and training workflow&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Analysis: Breadth complicates attribution
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Analysis:&lt;/strong&gt; The defensible interpretation is that “Astra for robotics” describes several integration patterns rather than one fixed controller architecture. The catalogue places high-level planning with GEAR-SONIC, planning with FluxVLA, direct pose output, simulated replay and environment construction under the same project heading. &lt;a href="https://github.com/zjwzcx/Awesome-Astra-Embodied-AI" rel="noopener noreferrer"&gt;Those differing roles are visible across the case descriptions&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The unresolved trade-off is breadth versus attribution. The range of cases gives developers multiple interfaces to investigate, but a successful planned replay cannot establish the same capability as direct physical control, and a task executed through a pretrained policy cannot be attributed to the planning layer alone. Evaluations based on these workflows therefore need to state which component handles perception, planning, trajectory generation and low-level execution, along with any simulator or grasp assumptions documented by the case. &lt;a href="https://github.com/zjwzcx/Awesome-Astra-Embodied-AI" rel="noopener noreferrer"&gt;The repository’s GEAR-SONIC, FluxVLA and Dual-ALOHA entries show why those boundaries differ&lt;/a&gt;.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>NiubiGEO Records Evidence Behind AI Brand Reports</title>
      <dc:creator>David Díaz</dc:creator>
      <pubDate>Tue, 15 Sep 2026 20:25:54 +0000</pubDate>
      <link>https://dev.to/dd8888/niubigeo-records-evidence-behind-ai-brand-reports-1djp</link>
      <guid>https://dev.to/dd8888/niubigeo-records-evidence-behind-ai-brand-reports-1djp</guid>
      <description>&lt;p&gt;NiubiGEO has published an Apache-2.0 workbench for testing how AI models describe a product, which competitors they mention and what sources they return. Developers can inspect the answers and execution conditions behind a report instead of relying only on a visibility score, according to the &lt;a href="https://github.com/Albert-Weasker/niubigeo" rel="noopener noreferrer"&gt;project repository&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What each test retains
&lt;/h2&gt;

&lt;p&gt;A NiubiGEO project starts with a domain. Users select one or more OpenRouter models and configure web search separately for each model before running a test. Models answer independently, and a failure from one model does not remove successful results from the others; failed models can also be retried separately. The saved results include the actual execution conditions for each run, as described in the &lt;a href="https://github.com/Albert-Weasker/niubigeo" rel="noopener noreferrer"&gt;repository documentation&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The initial domain test reports product descriptions, categories, competitors and associated keywords. Users can then confirm keywords and test them without naming the target brand. NiubiGEO’s methodology distinguishes recognition after a direct brand prompt from an unprompted appearance, and it treats a mention, a positive description and an explicit recommendation as different findings. Those distinctions are documented in the project’s &lt;a href="https://github.com/Albert-Weasker/niubigeo" rel="noopener noreferrer"&gt;examples and usage notes&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;NiubiGEO displays original answers, provider citations and ordinary URLs separately while retaining failures and unresolved findings. This evidence can help users locate a statement or source, but the project cautions that a citation alone does not explain why a model recommended a product. &lt;a href="https://github.com/Albert-Weasker/niubigeo" rel="noopener noreferrer"&gt;Its documentation&lt;/a&gt; also says that publishing an article does not guarantee an AI recommendation.&lt;/p&gt;

&lt;p&gt;Repeated measurements and scheduled monitoring create a history tied to the underlying answers. Earlier records remain available when the selected models change, while scheduled execution requires the monitoring worker to be running. The project explicitly warns that a few closely spaced tests demonstrate repeated testing rather than long-term growth. &lt;a href="https://github.com/Albert-Weasker/niubigeo" rel="noopener noreferrer"&gt;The monitoring notes&lt;/a&gt; do not present any single answer as a permanent ranking.&lt;/p&gt;

&lt;h2&gt;
  
  
  Deployment sets a clear cost boundary
&lt;/h2&gt;

&lt;p&gt;The local quick start requires Node.js 22 or later and the user’s own OpenRouter API key. Published cases can be explored without installation or an API key, but testing another product incurs model and search API charges; operators also cover their hosting costs. NiubiGEO separately advertises paid AI testing by real people and GEO optimization services through the &lt;a href="https://github.com/Albert-Weasker/niubigeo" rel="noopener noreferrer"&gt;repository&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The project positions self-hosting, model choice and access to source evidence as its distinguishing criteria. It suggests considering commercial platforms when hosted services, marketing workflows or an existing search dataset take priority, while stating that its vendor comparisons are selection suggestions rather than a controlled benchmark or ranking. That limitation appears in the &lt;a href="https://github.com/Albert-Weasker/niubigeo" rel="noopener noreferrer"&gt;comparison guidance&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Analysis: the audit trail is the useful part
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Analysis:&lt;/strong&gt; NiubiGEO’s most defensible value is auditability, not proof of a stable level of AI visibility. A stored response can establish what a selected model returned under recorded conditions, giving developers a concrete artifact to inspect when a product is omitted or described inaccurately. It cannot, by itself, establish why the result occurred or whether it will persist—limits consistent with the project’s warnings about recommendations, citations and short-term measurements in the &lt;a href="https://github.com/Albert-Weasker/niubigeo" rel="noopener noreferrer"&gt;repository&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The unresolved trade-off is breadth versus comparability. Adding models, search modes and scheduled runs expands the set of observations, but it also adds model and search charges and creates more execution conditions to compare. NiubiGEO records those conditions; it does not claim that doing so converts variable model answers into a permanent ranking. &lt;a href="https://github.com/Albert-Weasker/niubigeo" rel="noopener noreferrer"&gt;The project documentation&lt;/a&gt; instead frames the records as evidence for further investigation.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>github</category>
      <category>opensource</category>
      <category>testing</category>
    </item>
    <item>
      <title>UAJY Handbook RAG Chatbot Splits FAISS Search From Gemini</title>
      <dc:creator>David Díaz</dc:creator>
      <pubDate>Tue, 15 Sep 2026 19:55:40 +0000</pubDate>
      <link>https://dev.to/dd8888/uajy-handbook-rag-chatbot-splits-faiss-search-from-gemini-45o5</link>
      <guid>https://dev.to/dd8888/uajy-handbook-rag-chatbot-splits-faiss-search-from-gemini-45o5</guid>
      <description>&lt;p&gt;The public &lt;a href="https://github.com/BenyRonald77/uajy-academic-rag-chatbot" rel="noopener noreferrer"&gt;UAJY Academic Document RAG Chatbot repository&lt;/a&gt; packages a Streamlit assistant for the 2025/2026 academic handbook of Universitas Atma Jaya Yogyakarta’s Faculty of Industrial Technology. For developers, it provides an inspectable implementation of document ingestion, local vector search, cited answer generation, refusal controls and a 20-question evaluation harness in one codebase.&lt;/p&gt;

&lt;h2&gt;
  
  
  From handbook PDF to cited answer
&lt;/h2&gt;

&lt;p&gt;The offline ingestion pipeline extracts text and tables with pdfplumber, divides the material into semantic or paragraph-based chunks, creates 3,072-dimensional embeddings with &lt;code&gt;gemini-embedding-001&lt;/code&gt;, and persists the vectors in a FAISS index alongside page and heading metadata. The &lt;a href="https://github.com/BenyRonald77/uajy-academic-rag-chatbot" rel="noopener noreferrer"&gt;repository&lt;/a&gt; says the resulting index contains 350 chunks drawn from 112 handbook pages.&lt;/p&gt;

&lt;p&gt;At runtime, a Streamlit query triggers top-&lt;em&gt;K&lt;/em&gt; similarity retrieval from FAISS. The application builds a prompt from the retrieved context and conversation history, sends it to Gemini 2.5 Flash, and formats the answer with page numbers and chapter or section titles, according to the &lt;a href="https://github.com/BenyRonald77/uajy-academic-rag-chatbot" rel="noopener noreferrer"&gt;documented architecture&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The project describes two controls for unsupported questions: a similarity threshold filters retrieval results, while the system prompt instructs Gemini to refuse requests that the retrieved document context cannot answer. The &lt;a href="https://github.com/BenyRonald77/uajy-academic-rag-chatbot" rel="noopener noreferrer"&gt;README&lt;/a&gt; characterizes these controls as an anti-hallucination defense and says generated answers must rely exclusively on retrieved PDF chunks.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the published evaluation covers
&lt;/h2&gt;

&lt;p&gt;The repository includes an evaluation runner and a test file with 20 questions. Its &lt;a href="https://github.com/BenyRonald77/uajy-academic-rag-chatbot" rel="noopener noreferrer"&gt;benchmark summary&lt;/a&gt; reports Retrieval Recall@4 of 100% for 15 in-scope questions, refusal correctness of 100% for five out-of-scope questions, average retrieval latency of 0.42 seconds and total response latency of approximately 1.85 seconds. These figures describe the project’s published test suite.&lt;/p&gt;

&lt;h2&gt;
  
  
  Analysis: an inspectable design with bounded evidence
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Analysis:&lt;/strong&gt; The repository is best read as a concrete RAG implementation whose grounding mechanisms can be inspected, rather than conclusive evidence for its broader “production-grade” and “hallucination-free” labels. The published summary establishes reported behavior for 20 prompts; it does not, by itself, establish behavior beyond that suite. The project’s own disclaimer draws a similar boundary by describing the chatbot as an educational and information-search assistant while reserving official authority for the deanery and academic administration office. &lt;a href="https://github.com/BenyRonald77/uajy-academic-rag-chatbot" rel="noopener noreferrer"&gt;Project README&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The unresolved deployment trade-off is the boundary between local and hosted components. The FAISS index and CPU-based similarity search run locally, while embedding creation and answer generation use Google Gemini services through an API; setup therefore requires a Gemini API key. &lt;a href="https://github.com/BenyRonald77/uajy-academic-rag-chatbot" rel="noopener noreferrer"&gt;Repository architecture and setup&lt;/a&gt; &lt;strong&gt;Analysis:&lt;/strong&gt; This gives adopters a local retrieval layer, but not a fully local RAG runtime. The relevant engineering decision is whether that split fits the intended deployment environment—not whether local vector search alone is sufficient.&lt;/p&gt;

</description>
      <category>gemini</category>
      <category>opensource</category>
      <category>python</category>
      <category>rag</category>
    </item>
    <item>
      <title>Reef Connects Agent Feedback, Learning and Versioned Delivery</title>
      <dc:creator>David Díaz</dc:creator>
      <pubDate>Tue, 15 Sep 2026 19:55:19 +0000</pubDate>
      <link>https://dev.to/dd8888/reef-connects-agent-feedback-learning-and-versioned-delivery-b1o</link>
      <guid>https://dev.to/dd8888/reef-connects-agent-feedback-learning-and-versioned-delivery-b1o</guid>
      <description>&lt;p&gt;Human-Agent-Society has published Reef, open-source infrastructure that links agent inference, interaction feedback, learning and versioned delivery. For developers, the immediate consequence is that one system can manage updates to model weights or to an agent harness—including its prompts, rules and skills—rather than leaving learning and deployment as separate pipelines (&lt;a href="https://github.com/Human-Agent-Society/reef" rel="noopener noreferrer"&gt;Reef repository&lt;/a&gt;).&lt;/p&gt;

&lt;h2&gt;
  
  
  One loop, two update surfaces
&lt;/h2&gt;

&lt;p&gt;Reef divides a learning cycle into four stages. &lt;strong&gt;Serve&lt;/strong&gt; handles requests and records interactions; &lt;strong&gt;Observe&lt;/strong&gt; matches later feedback to those records; &lt;strong&gt;Grow&lt;/strong&gt; produces updates from eligible records; and &lt;strong&gt;Commit&lt;/strong&gt; applies a configured selection policy before publishing accepted versions. The repository maps those stages to its service, storage, training, evaluation, artifact-history and delivery modules (&lt;a href="https://github.com/Human-Agent-Society/reef" rel="noopener noreferrer"&gt;Reef architecture&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;A deployment recipe determines what changes. Weight-oriented recipes can train model parameters through Slime and SGLang, while harness-oriented recipes can modify prompts, rules and skills using a model endpoint rather than local training GPUs (&lt;a href="https://github.com/Human-Agent-Society/reef" rel="noopener noreferrer"&gt;Reef learning surfaces&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;For weight training, Reef documents OpenAI- and Anthropic-compatible inference endpoints. Each response carries an interaction record ID that a later report can reference; reports may include a numeric score plus textual or structured feedback. The project says that once a recipe has sufficient eligible feedback, it can run a training step and synchronize updated weights to the serving runtime without restarting Reef (&lt;a href="https://github.com/Human-Agent-Society/reef" rel="noopener noreferrer"&gt;Reef usage guide&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;The documented &lt;code&gt;harness-evolve&lt;/code&gt; path uses a model API instead of a local GPU training stack. In its coding tutorial, a failed report can trigger a candidate skill update; Reef compares that candidate with the current harness on three coding tasks and publishes it only if it wins (&lt;a href="https://github.com/Human-Agent-Society/reef" rel="noopener noreferrer"&gt;harness-evolve example&lt;/a&gt;).&lt;/p&gt;

&lt;h2&gt;
  
  
  Different infrastructure, shared release gate
&lt;/h2&gt;

&lt;p&gt;The two paths impose different prerequisites. Model-weight training needs a trainable model, feedback usable by the selected recipe and a supported GPU stack. Harness optimization needs a model endpoint, representative tasks and an evaluator, but no local training GPUs (&lt;a href="https://github.com/Human-Agent-Society/reef" rel="noopener noreferrer"&gt;Reef requirements&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;Both paths converge on the Commit stage. Reef applies the configured selection policy there and publishes accepted updates through its version-management and artifact-delivery components; the project also lists remaining live during updates as a built-in capability (&lt;a href="https://github.com/Human-Agent-Society/reef" rel="noopener noreferrer"&gt;Reef workflow&lt;/a&gt;).&lt;/p&gt;

&lt;h2&gt;
  
  
  Analysis: Evaluation becomes the release gate
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Analysis:&lt;/strong&gt; Reef’s consequential design choice is the separation of update generation from publication. A completed training step or harness edit does not automatically become the delivered version; candidate evaluation and selection intervene first. That makes the system more controlled than a loop that applies every generated change, although the control is only as useful as the policy and evaluator configured for it (&lt;a href="https://github.com/Human-Agent-Society/reef" rel="noopener noreferrer"&gt;Reef Commit stage&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;The unresolved trade-off is between automating more of the improvement cycle and retaining confidence in what gets accepted. Reef can standardize interaction records, update jobs, version history and delivery, but harness operators must still supply representative tasks and an evaluator. A narrow evaluation can therefore become a narrow release gate: the infrastructure enforces the decision, while the recipe operator determines what evidence counts (&lt;a href="https://github.com/Human-Agent-Society/reef" rel="noopener noreferrer"&gt;Reef harness requirements&lt;/a&gt;).&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>machinelearning</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Goldie Brings Mobile Store Assets Into Agent Workflows</title>
      <dc:creator>David Díaz</dc:creator>
      <pubDate>Sat, 05 Sep 2026 19:41:14 +0000</pubDate>
      <link>https://dev.to/dd8888/goldie-brings-mobile-store-assets-into-agent-workflows-2k8</link>
      <guid>https://dev.to/dd8888/goldie-brings-mobile-store-assets-into-agent-workflows-2k8</guid>
      <description>&lt;p&gt;Goldie packages App Store and Google Play screenshot production, plus App Store preview rendering, into a public CLI with a coding-agent skill. Argent replays the configured flows in an iOS simulator or Android emulator; Goldie adds bezels, backgrounds and headlines, joins preview clips, and checks the output against store upload rules. Developers can start that pipeline through a compatible agent or run the CLI themselves. &lt;a href="https://github.com/kacperkapusciak/goldie" rel="noopener noreferrer"&gt;Goldie's repository documents both routes&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Agent-assisted and manual routes
&lt;/h2&gt;

&lt;p&gt;The packaged skill works with agents that support the skills format. In the documented workflow, a developer asks for app-store screenshots from the application repository; the agent asks which stores to target, explores the app, writes the flows and configuration, and opens Goldie's browser-based studio. Follow-up instructions modify the same files. &lt;a href="https://github.com/kacperkapusciak/goldie" rel="noopener noreferrer"&gt;The README describes this agent workflow&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Manual operation starts with a &lt;code&gt;goldie.config.ts&lt;/code&gt; file whose scenes point to Argent flows under &lt;code&gt;.argent/flows&lt;/code&gt;. The &lt;code&gt;goldie doctor&lt;/code&gt; command checks tools, simulators and flows; &lt;code&gt;goldie all&lt;/code&gt; captures and frames screenshots, renders the preview, and verifies the output; &lt;code&gt;goldie studio&lt;/code&gt; opens the assets for adjustment. Generated files go under &lt;code&gt;out/screenshots/&amp;lt;device&amp;gt;/&amp;lt;locale&amp;gt;/&lt;/code&gt; and &lt;code&gt;out/previews/&amp;lt;device&amp;gt;/&amp;lt;locale&amp;gt;/&lt;/code&gt;. &lt;a href="https://github.com/kacperkapusciak/goldie" rel="noopener noreferrer"&gt;Those commands and paths are specified by the project&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The division of responsibility is explicit: Argent performs the flow replay, while Goldie handles the resulting store assets. Goldie describes itself as framework-agnostic because it drives the app through a simulator or emulator, and the project lists SwiftUI, UIKit, Jetpack Compose, Flutter, React Native and Kotlin Multiplatform among the applicable frameworks. &lt;a href="https://github.com/kacperkapusciak/goldie" rel="noopener noreferrer"&gt;The repository explains this simulator-level approach&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Host and store constraints
&lt;/h2&gt;

&lt;p&gt;Goldie requires Node 20 or newer and &lt;code&gt;ffmpeg&lt;/code&gt; on the system path. Creating App Store screenshots requires macOS with Xcode's iOS simulators, while the Android emulator used for Google Play screenshots can run on macOS, Linux or Windows. &lt;a href="https://github.com/kacperkapusciak/goldie" rel="noopener noreferrer"&gt;The installation requirements are documented in the README&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The studio can switch devices, backgrounds, templates, bezels and fonts, as well as edit copy for individual tiles. It saves those choices in &lt;code&gt;goldie.design.json&lt;/code&gt;, which the CLI then uses when rendering. &lt;a href="https://github.com/kacperkapusciak/goldie" rel="noopener noreferrer"&gt;Goldie's design documentation lists these controls and the saved file&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The documented output is 1320 × 2868 for iPhone screenshots, 886 × 1920 H.264 for an App Store preview, and 1080 × 1920 for Google Play screenshots. Apple previews must run for 15 to 30 seconds. &lt;a href="https://github.com/kacperkapusciak/goldie" rel="noopener noreferrer"&gt;The repository specifies these dimensions and the Apple duration window&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;For Android, the &lt;code&gt;pixel-10-pro&lt;/code&gt; device key renders Play Store phone screenshots from the same scenes used elsewhere, but a flow works across both platforms only when its selectors match. Goldie also renders a portrait promotional video from the emulator for separate posting to YouTube; Apple's 15-to-30-second rule does not apply to that Google Play workflow. &lt;a href="https://github.com/kacperkapusciak/goldie" rel="noopener noreferrer"&gt;The Android documentation distinguishes these outputs&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  UI changes remain a maintenance cost
&lt;/h2&gt;

&lt;p&gt;The project recommends release builds because debug builds can add LogBox banners to captures. It also warns that flows fail when the app changes, requiring developers to ask an agent to repair them or re-record them with Argent. &lt;a href="https://github.com/kacperkapusciak/goldie" rel="noopener noreferrer"&gt;Both limitations appear in the README's remarks&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Analysis:&lt;/strong&gt; Goldie's defensible value is process integration, not proof of design quality. It combines scripted app states with saved presentation settings and exposes their configuration to an agent-assisted workflow. The unresolved trade-off is selector maintenance: because the project explicitly warns that flows break as the application changes, teams still need to repair or re-record scenes as the UI evolves. &lt;a href="https://github.com/kacperkapusciak/goldie" rel="noopener noreferrer"&gt;That interpretation follows from Goldie's documented workflow and maintenance warning&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>goldie</category>
      <category>argent</category>
      <category>appstore</category>
      <category>googleplay</category>
    </item>
    <item>
      <title>Codex with ChatGPT Splits Planning From Execution</title>
      <dc:creator>David Díaz</dc:creator>
      <pubDate>Tue, 01 Sep 2026 19:39:25 +0000</pubDate>
      <link>https://dev.to/dd8888/codex-with-chatgpt-splits-planning-from-execution-3bp2</link>
      <guid>https://dev.to/dd8888/codex-with-chatgpt-splits-planning-from-execution-3bp2</guid>
      <description>&lt;p&gt;&lt;a href="https://github.com/XiaoDuoYa/codex-with-chatgpt" rel="noopener noreferrer"&gt;Codex with ChatGPT&lt;/a&gt;, a public GitHub project, routes planning and code review to the ChatGPT web app while retaining Codex as the agent that edits files, runs shell commands and executes tests. For developers, the practical aim is to use an existing ChatGPT web subscription for reasoning without replacing the Codex-based execution harness.&lt;/p&gt;

&lt;h2&gt;
  
  
  A read-only handoff between the two agents
&lt;/h2&gt;

&lt;p&gt;The project describes a local C2C Bridge between ChatGPT and a workspace. Codex and ChatGPT exchange small structured state messages for the plan, execution and review loop, while ChatGPT retrieves repository context through nine read-only MCP tools, including file reads, workspace search, Git status and diffs, and test-status and execution-output records. The README says full file bodies, diffs and logs are not placed in those control-plane messages. &lt;a href="https://github.com/XiaoDuoYa/codex-with-chatgpt" rel="noopener noreferrer"&gt;See the project README&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;That separation is central to its security design. The repository says the bridge has no write, delete, shell or commit tools, and that sensitive paths such as &lt;code&gt;.env&lt;/code&gt; files, keys, SSH material and credentials are denied by default. It also describes workspace-scoped access, OAuth 2.1 protection and one-time pairing codes for the publicly reachable MCP endpoint. These are project claims rather than an independent security assessment. &lt;a href="https://github.com/XiaoDuoYa/codex-with-chatgpt" rel="noopener noreferrer"&gt;The stated security model is documented in the README&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Setup still carries operational cost
&lt;/h2&gt;

&lt;p&gt;Installation is packaged as a Codex Skill, with a manual route that copies the skill into Codex’s skills directory and asks Codex to perform first-time setup. The project requires Git, Node.js 20 or later and &lt;code&gt;cloudflared&lt;/code&gt; for its public connection; it says users may need to sign in to ChatGPT and, for an optional stable hostname, authorize Cloudflare and use a domain already managed there. &lt;a href="https://github.com/XiaoDuoYa/codex-with-chatgpt" rel="noopener noreferrer"&gt;The installation and hostname flow are specified by the project&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The default connection uses a temporary Cloudflare URL, according to the README. When that address changes after a bridge restart, the project says Codex repairs the ChatGPT connector; a stable Cloudflare hostname is presented as an optional way to avoid that repair cycle. &lt;a href="https://github.com/XiaoDuoYa/codex-with-chatgpt" rel="noopener noreferrer"&gt;The project describes the temporary and stable tunnel options here&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The unresolved trade-off
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Analysis:&lt;/strong&gt; This is a narrow integration pattern, not evidence that splitting planning from execution produces better code. Its appeal is architectural: it gives a web-based planner access to selected live workspace evidence while keeping side effects in the coding harness. But the arrangement also adds a tunnel, OAuth pairing, connector configuration and another failure boundary to a development workflow. The read-only policy can limit what ChatGPT can do directly, yet developers still need to decide whether the allowed source context is appropriate to expose through the bridge. The project is explicitly an unofficial community effort, not an OpenAI-endorsed Codex integration. &lt;a href="https://github.com/XiaoDuoYa/codex-with-chatgpt" rel="noopener noreferrer"&gt;Its status and disclaimer are in the README&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>chatgpt</category>
      <category>codex</category>
      <category>modelcontextprotocol</category>
      <category>oauth21</category>
    </item>
    <item>
      <title>Tencent Releases WeMM-Embedding for Multimodal Retrieval</title>
      <dc:creator>David Díaz</dc:creator>
      <pubDate>Sun, 30 Aug 2026 19:39:57 +0000</pubDate>
      <link>https://dev.to/dd8888/tencent-releases-wemm-embedding-for-multimodal-retrieval-310d</link>
      <guid>https://dev.to/dd8888/tencent-releases-wemm-embedding-for-multimodal-retrieval-310d</guid>
      <description>&lt;p&gt;Tencent’s WeChat Vision team has published &lt;a href="https://github.com/Tencent/WeMM-Embedding" rel="noopener noreferrer"&gt;WeMM-Embedding&lt;/a&gt;, a family comprising 2B, 4B and 9B embedding models. Each variant supports text, images, videos, visual documents and interleaved multimodal inputs; audio is not supported. For developers, the immediate consequence is a single repository containing model options, inference examples, serving instructions and evaluation code for retrieval across the supported inputs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Inputs, vectors and serving paths
&lt;/h2&gt;

&lt;p&gt;The models obtain embeddings from the last-layer hidden state at a dedicated &lt;code&gt;&amp;lt;embedding&amp;gt;&lt;/code&gt; token, then apply L2 normalization. The project presents this as a unified representation method across its supported input types rather than separate embedding formats for each modality (&lt;a href="https://github.com/Tencent/WeMM-Embedding" rel="noopener noreferrer"&gt;repository&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;All three variants support reduced Matryoshka dimensions. The 2B model offers widths from 64 to 2,048, the 4B model from 64 to 2,560 and the 9B model from 64 to 4,096. For a supported smaller width, the repository instructs users to truncate the full vector and normalize it again (&lt;a href="https://github.com/Tencent/WeMM-Embedding" rel="noopener noreferrer"&gt;repository&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;Inference examples cover Transformers and Sentence Transformers. Tencent recommends &lt;code&gt;transformers==5.2.0&lt;/code&gt; for inference and reproducibility because newer releases may differ in preprocessing, while the serving instructions list vLLM 0.27.0 and SGLang 0.5.9 as tested versions (&lt;a href="https://github.com/Tencent/WeMM-Embedding" rel="noopener noreferrer"&gt;repository&lt;/a&gt;).&lt;/p&gt;

&lt;h2&gt;
  
  
  Project-reported benchmark results
&lt;/h2&gt;

&lt;p&gt;On MMEB-v2, which covers 78 datasets, Tencent reports average scores of 77.9, 79.2 and 80.6 for the 2B, 4B and 9B models respectively. The table uses Hit@1 for image and video tasks and NDCG@5 for visual-document tasks (&lt;a href="https://github.com/Tencent/WeMM-Embedding" rel="noopener noreferrer"&gt;repository&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;The broader MMEB-v3 evaluation contains 190 tasks, including the 78 MMEB-v2 tasks alongside text, agent, audio and MCMR tasks. Tencent reports V3-All scores of 56.0 for the 2B model, 58.2 for 4B and 59.5 for 9B. All three receive zero for audio because unsupported tasks are assigned zero, and the repository includes the MMEB-v3 evaluation code used for the reported results (&lt;a href="https://github.com/Tencent/WeMM-Embedding" rel="noopener noreferrer"&gt;repository&lt;/a&gt;).&lt;/p&gt;

&lt;h2&gt;
  
  
  An adjustable retrieval trade-off
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Analysis:&lt;/strong&gt; Adjustable vector width is a more directly testable design choice than the headline benchmark ranking. Teams can compare supported dimensions within one model family, but the repository substantiates performance retention with one specific result: on MMEB-v2, the 2B model at 256 dimensions retains 98.7% of its full-dimensional image and video performance. That does not resolve retention for text, visual documents, interleaved inputs or a production corpus. The leaderboard figures should likewise be treated as project-reported evidence until independently reproduced (&lt;a href="https://github.com/Tencent/WeMM-Embedding" rel="noopener noreferrer"&gt;repository&lt;/a&gt;).&lt;/p&gt;

</description>
      <category>tencent</category>
      <category>wechatvision</category>
      <category>multimodalembeddings</category>
      <category>multimodalretrieval</category>
    </item>
    <item>
      <title>Codex Skill Adds Guardrails to AI Presenter Video Workflows</title>
      <dc:creator>David Díaz</dc:creator>
      <pubDate>Fri, 28 Aug 2026 19:38:53 +0000</pubDate>
      <link>https://dev.to/dd8888/codex-skill-adds-guardrails-to-ai-presenter-video-workflows-4jk2</link>
      <guid>https://dev.to/dd8888/codex-skill-adds-guardrails-to-ai-presenter-video-workflows-4jk2</guid>
      <description>&lt;p&gt;The public GitHub project &lt;a href="https://github.com/cclank/lanshu-create-ai-presenter-video" rel="noopener noreferrer"&gt;lanshu-create-ai-presenter-video&lt;/a&gt; packages a Codex-oriented Skill for turning a topic or script plus an authorized, clearly adult presenter image into an AI presenter video workflow. Developers should care because the project deliberately does not bind its source code to a particular video, speech or lip-sync provider; instead, it selects from capabilities available in the current environment. &lt;a href="https://github.com/cclank/lanshu-create-ai-presenter-video" rel="noopener noreferrer"&gt;The repository&lt;/a&gt; describes the resulting process as covering script work, voice generation, presenter generation, lip-sync calibration, captions, keyword motion effects, editing, rendering and quality acceptance.&lt;/p&gt;

&lt;h2&gt;
  
  
  A workflow rather than a video API
&lt;/h2&gt;

&lt;p&gt;The Skill can be installed under &lt;code&gt;~/.codex/skills/lanshu-create-ai-presenter-video&lt;/code&gt;, and its repository includes job initialization, preflight and delivery-finalization scripts alongside reference documents for generation, editing, and QA recovery. &lt;a href="https://github.com/cclank/lanshu-create-ai-presenter-video" rel="noopener noreferrer"&gt;The project page&lt;/a&gt; lists Codex or another local Skill-compatible agent environment, Python 3.9+, FFmpeg/ffprobe, standard shell tools, and at least one callable capability each for video generation, speech generation and lip sync as runtime requirements.&lt;/p&gt;

&lt;p&gt;Its central production rule is that the completed narration becomes the timeline reference: presenter video, subtitles, shots, keyword treatments and transitions are positioned against the same audio track. &lt;a href="https://github.com/cclank/lanshu-create-ai-presenter-video" rel="noopener noreferrer"&gt;The repository documentation&lt;/a&gt; says this is intended to reduce lip-sync drift and discontinuities between segments. The documented defaults include 9:16 output at 1080×1920 and 30fps, with topic-led videos generally targeted at 45–75 seconds. &lt;a href="https://github.com/cclank/lanshu-create-ai-presenter-video" rel="noopener noreferrer"&gt;The same project page&lt;/a&gt; also specifies a roughly -16 LUFS publishing loudness target.&lt;/p&gt;

&lt;h2&gt;
  
  
  Consent and cost checks are part of the job
&lt;/h2&gt;

&lt;p&gt;The supplied job setup accepts a &lt;code&gt;--rights-confirmed&lt;/code&gt; flag and an &lt;code&gt;--adult-presenter-confirmed&lt;/code&gt; flag, after which the user is expected to review &lt;code&gt;job.json&lt;/code&gt; for manual checks and permission for remote uploads before running preflight validation. &lt;a href="https://github.com/cclank/lanshu-create-ai-presenter-video" rel="noopener noreferrer"&gt;The repository&lt;/a&gt; states that image-use rights and adult status should be confirmed before remote upload, and that voice-cloning authorization should be confirmed before cloning a voice.&lt;/p&gt;

&lt;p&gt;The project also requires an explanation of upload content, generation duration, pricing basis, trial plan and retry limit before the first paid generation. &lt;a href="https://github.com/cclank/lanshu-create-ai-presenter-video" rel="noopener noreferrer"&gt;Its documented cost boundary&lt;/a&gt; tells the workflow to query existing task IDs after an interruption to avoid duplicate charges, and to stop and summarize the problem after three consecutive paid candidate failures.&lt;/p&gt;

&lt;p&gt;The repository says it does not retain API keys, access tokens, signed download URLs or user media, and that task-level request records should have credentials and temporary URLs removed before they are committed. &lt;a href="https://github.com/cclank/lanshu-create-ai-presenter-video" rel="noopener noreferrer"&gt;The project page&lt;/a&gt; further says that preflight and delivery reports retain filenames rather than absolute paths on the developer's machine. It is published under the MIT License. &lt;a href="https://github.com/cclank/lanshu-create-ai-presenter-video" rel="noopener noreferrer"&gt;The repository license information&lt;/a&gt; permits use, modification and distribution under that license.&lt;/p&gt;

&lt;h2&gt;
  
  
  The unresolved boundary of “verified”
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Analysis:&lt;/strong&gt; This is a useful attempt to make consent, spend control and technical QA explicit in an agent-driven media pipeline. But its “verified” outcome is only as reliable as the operator's rights attestations, the selected providers and the checks actually available in the local environment. The project can require confirmations and prescribe lip-sync, presenter, voice and visual checks; &lt;a href="https://github.com/cclank/lanshu-create-ai-presenter-video" rel="noopener noreferrer"&gt;the repository&lt;/a&gt; does not claim to independently establish image ownership, voice authorization or the truthfulness of a finished video. Provider neutrality therefore improves portability, while leaving the hardest accountability question with the team running the workflow.&lt;/p&gt;

</description>
      <category>codex</category>
      <category>aivideo</category>
      <category>digitalhumans</category>
      <category>ffmpeg</category>
    </item>
    <item>
      <title>NorthCinder Adds Buyer Approval to MCP Shopping Workflows</title>
      <dc:creator>David Díaz</dc:creator>
      <pubDate>Thu, 27 Aug 2026 07:32:55 +0000</pubDate>
      <link>https://dev.to/dd8888/northcinder-adds-buyer-approval-to-mcp-shopping-workflows-5hmh</link>
      <guid>https://dev.to/dd8888/northcinder-adds-buyer-approval-to-mcp-shopping-workflows-5hmh</guid>
      <description>&lt;p&gt;NorthCinder is an open-source MCP server intended to let an AI app compare products from sources selected by the buyer, show supporting facts, and obtain approval before a purchase. For developers building agentic-commerce flows, its practical distinction is that recommendation and checkout are deliberately separate: a checkout requires fresh approval for one exact offer and one unit. &lt;a href="https://github.com/cinderline/northcinder" rel="noopener noreferrer"&gt;The project repository&lt;/a&gt; says the software is MIT-licensed and self-hosted rather than operated as a NorthCinder cloud service.&lt;/p&gt;

&lt;h2&gt;
  
  
  A local MCP layer, not a marketplace
&lt;/h2&gt;

&lt;p&gt;The project runs alongside an MCP-capable AI application. In its default local mode, NorthCinder runs the MCP server and search engine together in one process using a temporary loopback port; the repository specifies Node.js 20 or later and describes local mode as keyless. &lt;a href="https://github.com/cinderline/northcinder" rel="noopener noreferrer"&gt;NorthCinder's README&lt;/a&gt; says that the repository owner is not in the path between the AI app, the local engine, and the store connections.&lt;/p&gt;

&lt;p&gt;NorthCinder's built-in adapters cover Shopify, WooCommerce, eBay, Etsy, and read-only Amazon comparison. The project also says it reports stores that are unavailable or unconfigured rather than presenting a partial search as complete market coverage. &lt;a href="https://github.com/cinderline/northcinder" rel="noopener noreferrer"&gt;The repository documentation&lt;/a&gt; frames these connections as buyer-selected sources, not a universal product index.&lt;/p&gt;

&lt;p&gt;The recommendation output is designed to be narrow: normally no more than three useful choices, consisting of the strongest fit, a lower-risk choice, and a cheaper or meaningfully different option when available. It can also expose other finalists, rejected offers, and facts that could not be verified. &lt;a href="https://github.com/cinderline/northcinder" rel="noopener noreferrer"&gt;NorthCinder's README&lt;/a&gt; says rankings are rerun locally and that recommendations, approvals, and checkout attempts are written to a local audit log.&lt;/p&gt;

&lt;h2&gt;
  
  
  Checkout remains explicitly gated
&lt;/h2&gt;

&lt;p&gt;A recommendation is not authorization to buy. NorthCinder requires a signed, single-use approval containing the merchant, variant, price, known total, and spending cap for a specific offer and unit. &lt;a href="https://github.com/cinderline/northcinder" rel="noopener noreferrer"&gt;The project documentation&lt;/a&gt; says raw card details are rejected; a supported automated checkout can use an opaque payment token, while another path hands the buyer a cart link for completion in their own browser.&lt;/p&gt;

&lt;p&gt;The repository also states that seller payment does not improve ranking, sponsored offers remain labeled and below organic results, and unknown seller history is left unknown rather than inferred to be safe or unsafe. &lt;a href="https://github.com/cinderline/northcinder" rel="noopener noreferrer"&gt;NorthCinder's published project page&lt;/a&gt; presents these as inspectable rules, alongside ranking, trust, neutrality, and checkout documentation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Research is the unresolved constraint
&lt;/h2&gt;

&lt;p&gt;NorthCinder separates product and seller research from ranking. The MCP host is instructed to read research guides, create a research plan for the actual subject, and follow the returned checklist; when sources conflict or do not identify the exact item or seller, the result remains provisional. &lt;a href="https://github.com/cinderline/northcinder" rel="noopener noreferrer"&gt;The repository README&lt;/a&gt; further says that no host-and-model combination is currently qualified for routine research use.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Analysis:&lt;/strong&gt; This is a more credible boundary than treating a shopping agent as an autonomous buyer. Local ranking, disclosed coverage, and per-offer approval can make an agent's actions easier to inspect. But they do not solve the harder input problem: store access can be missing, and research remains provisional until the buyer verifies identity, sources, conflicts, and unknowns. That leaves NorthCinder most useful as a constrained comparison and checkout-control layer, not as evidence that product discovery itself is reliable. &lt;a href="https://github.com/cinderline/northcinder" rel="noopener noreferrer"&gt;The project repository&lt;/a&gt; makes both the coverage limitation and the provisional research status explicit.&lt;/p&gt;

</description>
      <category>modelcontextprotocol</category>
      <category>northcinder</category>
      <category>agenticcommerce</category>
      <category>localfirst</category>
    </item>
    <item>
      <title>dsh-web Bundles DSH Web Plugins Behind a Workshop Catalog</title>
      <dc:creator>David Díaz</dc:creator>
      <pubDate>Mon, 24 Aug 2026 19:39:32 +0000</pubDate>
      <link>https://dev.to/dd8888/dsh-web-bundles-dsh-web-plugins-behind-a-workshop-catalog-4gf4</link>
      <guid>https://dev.to/dd8888/dsh-web-bundles-dsh-web-plugins-behind-a-workshop-catalog-4gf4</guid>
      <description>&lt;p&gt;The dsh-web project has published an all-in-one plugin package for DeepSeek Harness Web, giving developers a single installation path for task automation, remote browser access, SSH administration, image analysis, skins and other interface extensions. The consequence is a much broader DSH Web workbench—and a larger set of packages and security boundaries for its operator to manage.&lt;/p&gt;

&lt;p&gt;The recommended package, &lt;code&gt;@linxin666/dsh-web-all&lt;/code&gt;, installs through DSH’s &lt;code&gt;web&lt;/code&gt; profile. Developers who do not want the complete bundle can install individual plugins instead. A companion catalog at dsh-market.com distributes plugins alongside skins and virtual pets.&lt;/p&gt;

&lt;p&gt;This is better understood as an attempt to make DSH Web an extension host than as a collection of interface decorations. The profile mechanism keeps plugins out of DSH’s source tree, but it does not remove the operational cost of combining code from npm, Git repositories and external maintainers. dsh-web reduces assembly work; it does not eliminate dependency or trust decisions.&lt;/p&gt;

&lt;h2&gt;
  
  
  One profile carries UI features and host operations
&lt;/h2&gt;

&lt;p&gt;The project says every bundled component mounts through DSH’s official profile mechanism. That is a sensible boundary: plugins can be installed, replaced or removed without patching the host, while skins remain asset directories loaded by a dedicated skin plugin. An upgrade to DSH therefore need not require rewriting a theme integration.&lt;/p&gt;

&lt;p&gt;The bundle reaches well beyond presentation. Its task board has five columns and can submit work to an actual DSH agent session, then update the card after execution. Cron schedules run through the DSH Web host rather than the browser, so closing the tab does not stop them. The documented limit is important: triggers missed while the host is stopped or the machine is asleep are skipped rather than queued for later execution.&lt;/p&gt;

&lt;p&gt;Other packages add paired mobile and PC access, an SSH operations panel, Git history visualization, conversation recovery and a &lt;code&gt;describe_image&lt;/code&gt; tool. The image tool sends referenced images to a configured OpenAI-compatible vision endpoint and places only the returned text in the conversation record. The bundle also integrates the external &lt;code&gt;dsh-better-sidebar&lt;/code&gt; plugin, while its older &lt;code&gt;aionui-panel&lt;/code&gt; has stopped receiving maintenance and is scheduled for removal.&lt;/p&gt;

&lt;p&gt;This breadth makes the aggregation package convenient, but ownership becomes less obvious. The documentation distinguishes bundle-prefixed configuration IDs from IDs used by separately installed copies of the same plugin. It says loading both sources avoids duplicate registration but provides no additional benefit. Operators still need to know which copy supplies a feature before changing its configuration or upgrade path.&lt;/p&gt;

&lt;p&gt;The installation notes expose the same tension. The repository documents pnpm layouts that can hide nested packages from DSH, build-script approval requirements for &lt;code&gt;cloudflared&lt;/code&gt;, &lt;code&gt;cpu-features&lt;/code&gt; and &lt;code&gt;ssh2&lt;/code&gt;, and a pnpm 11 release-age policy that can select an older package even when &lt;code&gt;@latest&lt;/code&gt; is requested. In the documented failure case, an older skin loader can leave DSH Web unable to start because a referenced package is missing. These are recoverable package-management problems, but an aggregation layer concentrates their effects.&lt;/p&gt;

&lt;h2&gt;
  
  
  dsh-market.com joins discovery directly to installation
&lt;/h2&gt;

&lt;p&gt;dsh-market.com is produced from the same repository and catalogs skins, pets and plugins. The Web GUI includes a workshop card that can browse the catalog, install skin and pet assets into the DSH home directory, and send plugins through the plugin manager. Skins can be previewed without committing them to disk.&lt;/p&gt;

&lt;p&gt;The catalog itself has a deliberately narrow architecture. A static build script generates it from &lt;code&gt;skin.json&lt;/code&gt;, &lt;code&gt;pet.json&lt;/code&gt; and &lt;code&gt;community.json&lt;/code&gt;; pushes to the main branch trigger deployment. Dynamic likes run through a Cloudflare Workers API backed by D1, with one vote allowed per device. Items are ranked by device-based popularity, and the top three in each category appear on the home page.&lt;/p&gt;

&lt;p&gt;That structure makes the catalog inputs identifiable and keeps social ranking separate from the static asset metadata. It does not make popularity a security signal. A device vote says that an item attracted approval, not that its package lifecycle, install scripts or maintenance status have been reviewed.&lt;/p&gt;

&lt;p&gt;The repository explains how authors contribute, how catalog data is built and how users install packages. It does not describe a formal security-review or package-approval process for workshop entries. That omission matters because the workshop is not merely a screenshot gallery: it shortens the route from discovering an extension to running it inside a developer workbench.&lt;/p&gt;

&lt;h2&gt;
  
  
  Remote access and SSH define the real trust boundary
&lt;/h2&gt;

&lt;p&gt;The remote plugin provides the clearest evidence that dsh-web must be assessed as operational software. It pairs a phone or another PC browser by QR code or link, using a one-time, time-limited token. Unpaired devices are denied workspace data, and stopping the service revokes paired devices. Public access can be added through a Cloudflare tunnel.&lt;/p&gt;

&lt;p&gt;The project also warns against marking a tunnel domain with &lt;code&gt;--trusted-host&lt;/code&gt; when using its pairing route. According to the documentation, that option lets the SDK’s &lt;code&gt;/api&lt;/code&gt; path bypass the plugin’s pairing gate. This is a precise and useful warning, but it means the advertised access control depends partly on how the surrounding DSH service is launched.&lt;/p&gt;

&lt;p&gt;Real-time updates use Server-Sent Events. The documentation says Cloudflare Quick Tunnels and Tailscale Serve do not carry those events, so the plugin falls back to polling. Messages still work, but new updates may arrive several seconds late. Tunnel selection therefore changes behavior even when ordinary HTTP requests appear healthy.&lt;/p&gt;

&lt;p&gt;The SSH plugin raises a more direct security concern. It offers an xterm.js terminal, SFTP transfers, localhost-only port forwarding, concurrent commands across filtered host groups and agent access to the same host configurations used by the panel. Those are substantive administration capabilities.&lt;/p&gt;

&lt;p&gt;Its disclosed storage model is correspondingly consequential: SSH passwords and private-key passphrases are kept in plaintext in &lt;code&gt;~/.dsh/dsh-ssh.json&lt;/code&gt;, protected by &lt;code&gt;0600&lt;/code&gt; file permissions. The project also warns that reconnecting can replay non-idempotent commands and that remote output is returned without redaction.&lt;/p&gt;

&lt;p&gt;dsh-web’s plugin architecture keeps extensions separate from DSH source and makes a broad workbench easier to assemble. Its unresolved trade-off sits at the workshop boundary: the closer discovery gets to one-click installation, the more the ecosystem needs trust and maintenance signals that are stronger than package availability and device likes.&lt;/p&gt;

</description>
      <category>deepseekharness</category>
      <category>pluginecosystems</category>
      <category>developertools</category>
      <category>sshsecurity</category>
    </item>
    <item>
      <title>Sprix SAGE makes agent routing a stateful scheduling problem</title>
      <dc:creator>David Díaz</dc:creator>
      <pubDate>Sat, 22 Aug 2026 19:37:56 +0000</pubDate>
      <link>https://dev.to/dd8888/sprix-sage-makes-agent-routing-a-stateful-scheduling-problem-1mdj</link>
      <guid>https://dev.to/dd8888/sprix-sage-makes-agent-routing-a-stateful-scheduling-problem-1mdj</guid>
      <description>&lt;p&gt;Sprix AI has released SAGE Router, a public research prototype that chooses whether an agent should keep working, recruit collaborators, or hand a task to another agent. The consequence is less glamorous but more useful than another agent-discovery layer: it treats multi-agent routing as a decision that must account for work already completed, the cost of moving context, and the dependencies still blocking a task.&lt;/p&gt;

&lt;p&gt;The repository calls this State-Aware Graph Exchange, or SAGE. It sits above the Agent2Agent protocol rather than replacing it. A2A can describe agents, tasks, artifacts and transport; SAGE’s narrower job is to decide a feasible execution arrangement and explain it. The prototype is explicitly not an execution client: it returns a routing decision but does not itself transmit tasks.&lt;/p&gt;

&lt;p&gt;That boundary is important. The project is not evidence that open agent networks have solved reliable delegation. It is a reasonably concrete attempt to identify the layer where such networks are likely to fail in practice: a directory can tell a system which agents exist, but not whether changing the active team midway through a constrained task is worth the disruption.&lt;/p&gt;

&lt;h2&gt;
  
  
  The route decision includes the cost of changing course
&lt;/h2&gt;

&lt;p&gt;SAGE compares three modes in one objective. In SELF, the incumbent retains the task alone. In COLLABORATE, the incumbent remains owner while a complementary group takes assigned work. In HANDOFF, a peer receives full ownership. The distinction is not just administrative. A handoff may gain specialist capability but lose accumulated task context; collaboration may cover missing skills but create coordination overhead.&lt;/p&gt;

&lt;p&gt;The router represents a task as weighted requirements, with dependencies expressed as a directed acyclic graph. It then assigns each remaining requirement to an executor, schedules dependent work, and estimates a critical path. Assigning separate independent requirements to different agents can permit parallel work; assignments concentrated on one agent are serialized. Budget and deadline constraints are checked at the team level after construction, rather than treated as a property of an individual agent profile.&lt;/p&gt;

&lt;p&gt;Its stated utility function combines a predicted probability of success with penalties for cost, latency, risk, context-transfer loss, coordination overhead and uncertainty. It also includes an exploration term. That is an ambitious scope for a small reference implementation, but the modeling choice is sound in one limited sense: a routing system that optimizes only apparent competence will systematically prefer impressive-looking agents even when switching them into a live task is expensive or infeasible.&lt;/p&gt;

&lt;p&gt;The project also filters permissions and compatibility before scoring candidates. This is more than an implementation detail. In agent marketplaces, eligibility needs to be a hard constraint; a ranking model should not be allowed to recommend a high-scoring agent that cannot meet a security requirement or support the needed input and output modes.&lt;/p&gt;

&lt;p&gt;SAGE’s more interesting claim is progress-aware replanning. Its execution state can include active executors, completed DAG nodes, failures and transferable context. A router that sees only the original prompt cannot distinguish a task that is still cheap to reassign from one where the incumbent has accumulated the decisive context. The latter is the case where a nominally better specialist may be the wrong choice.&lt;/p&gt;

&lt;h2&gt;
  
  
  Capability scores are not enough to build a team
&lt;/h2&gt;

&lt;p&gt;The repository models capability by combining global trust with trust conditioned on a particular requirement, then calibrating it against outcome evidence. This is intended to avoid a familiar error in agent selection: treating a general reputation score as transferable across work types. Success on coding tasks should not automatically establish reliability for research, planning or another requirement category.&lt;/p&gt;

&lt;p&gt;For a team, SAGE calculates requirement coverage from the combined capability of its members and assigns each requirement to the strongest calibrated member. It uses beam search over team prefixes instead of a purely greedy team-building procedure. The stated aim is complementarity: adding an agent should be rewarded for marginally covering requirements that the existing team does not cover, not simply for adding another highly ranked profile with overlapping strengths.&lt;/p&gt;

&lt;p&gt;This is the strongest design judgment in the project. Multi-agent frameworks often make collaboration sound like an unconditional upgrade from a single capable worker. In real scheduling terms, each additional participant adds messages, dependencies, conflicting estimates and an attribution problem when the result fails. A router needs a reason not to recruit.&lt;/p&gt;

&lt;p&gt;SAGE attempts to supply that reason through explicit coordination and transfer-loss penalties, plus a role assignment that exposes the communication topology. The output is designed to include assignments, predicted success, coverage, cost, latency, risk, utility and a human-readable rationale. Such observability is necessary if an operator is expected to approve or challenge an automated handoff.&lt;/p&gt;

&lt;p&gt;Still, the system’s learned component should be read carefully. The repository says a regularized online predictor has replaced an earlier fixed success equation, and that bid confidence, quoted cost and quoted latency can be calibrated against observed evidence. Online adaptation is useful when marketplace claims diverge from delivery. It also raises a difficult question the prototype cannot settle: whether feedback reflects agent quality, task difficulty, the router’s prior choices, or all three.&lt;/p&gt;

&lt;p&gt;The repository acknowledges that production use would require calibrated evaluators, authenticated identities, signed capability metadata, privacy and security review, persistent event-driven recovery, monitoring and task-specific validation. Those are not peripheral deployment chores. They determine whether the evidence used to update trust is meaningful.&lt;/p&gt;

&lt;h2&gt;
  
  
  A synthetic benchmark shows a trade-off, not readiness
&lt;/h2&gt;

&lt;p&gt;The included benchmark runs 2,500 tasks across five deterministic seeds in an external simulator. The simulator deliberately makes hidden capabilities, pair effects, nonlinear quality, realized cost and realized latency differ from SAGE’s prediction model. That separation is a better test design than evaluating the router with its own score as ground truth.&lt;/p&gt;

&lt;p&gt;In the reported comparison, Online SAGE reached a mean quality of 0.634 and common utility of 0.487, compared with 0.591 and 0.467 for Static SAGE. The online variant also consumed a larger share of budget on average, 0.434 versus 0.329. Its deadline-miss rate was 0.2%, while Static SAGE reported none. The repository does not conceal the cost-quality trade-off, which is preferable to presenting improved quality as a free gain.&lt;/p&gt;

&lt;p&gt;But these figures are synthetic results, not validation of a production routing policy. The project says as much: it calls itself an early-stage research preview, says the benchmark is not evidence of real-world superiority, and lists real executions, stronger learned-routing baselines, heterogeneous agent benchmarks, trace replay, calibration analysis and adversarial conditions as needed work.&lt;/p&gt;

&lt;p&gt;That restraint should frame the release. SAGE is credible as an algorithmic sketch of how an A2A network might make mid-execution delegation decisions. It is not yet a demonstration that a learned router can safely improve an agent marketplace under real incentives, unreliable bids and incomplete outcome labels.&lt;/p&gt;

&lt;p&gt;The unresolved trade-off is central: the more SAGE learns from observed outcomes and marketplace signals, the more it can route around weak or overpriced agents; the more it relies on those signals, the more its decisions depend on trustworthy identity, evaluation and privacy controls that open agent networks have yet to establish.&lt;/p&gt;

</description>
      <category>a2a</category>
      <category>agentorchestration</category>
      <category>multiagentsystems</category>
      <category>python</category>
    </item>
  </channel>
</rss>
