<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: DiFlowrin</title>
    <description>The latest articles on DEV Community by DiFlowrin (@diflowrin).</description>
    <link>https://dev.to/diflowrin</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F2806945%2F4f10992c-3ed8-48b2-bc6c-90857a6cadf3.jpeg</url>
      <title>DEV Community: DiFlowrin</title>
      <link>https://dev.to/diflowrin</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/diflowrin"/>
    <language>en</language>
    <item>
      <title>Best AI Research Tools 2026: What Works</title>
      <dc:creator>DiFlowrin</dc:creator>
      <pubDate>Sun, 27 Sep 2026 12:10:19 +0000</pubDate>
      <link>https://dev.to/diflowrin/best-ai-research-tools-2026-what-works-5e8p</link>
      <guid>https://dev.to/diflowrin/best-ai-research-tools-2026-what-works-5e8p</guid>
      <description>&lt;p&gt;Everyone has a "deep research" button now. ChatGPT has one. Gemini has one. Perplexity built its whole identity on it. And here is the thing nobody tells you at the door: one button is not enough. If you are looking for the &lt;strong&gt;best AI research tools 2026&lt;/strong&gt; has to offer, the honest answer is that there is no single winner. There is a toolbox. And there is a verification habit you cannot skip, because every one of these tools invents citations at a rate you can actually measure.&lt;/p&gt;

&lt;p&gt;So this article does three things. It shows you how to judge a research application instead of being charmed by one. It walks through the general deep-research agents and the academic tools, what each is good at and where each one falls down. And it ends with the uncomfortable part: the numbers on how often these systems fabricate links, and the three rules that keep you safe.&lt;/p&gt;

&lt;h2&gt;
  
  
  How do you judge an AI research app?
&lt;/h2&gt;

&lt;p&gt;Not by how nice the prose sounds. That is the trap. A deep-research agent writes beautifully, with confident paragraphs and tidy little citations, and the whole point of the exercise is that the citations might be fiction. What matters is how well the thing stands on its sources.&lt;/p&gt;

&lt;p&gt;There are eight criteria worth checking, and none of them is "does the report read well":&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Citation correctness.&lt;/strong&gt; Does the link resolve? Does it point at the source claimed? Does that source actually support the sentence it is attached to? Three separate tests, and a citation can fail any one of them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Recall.&lt;/strong&gt; How much of the relevant literature does the tool find at all?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Precision.&lt;/strong&gt; Of what it finds, how much is actually relevant?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reproducibility.&lt;/strong&gt; Ask the same question twice. Do you get the same answer? Often you do not.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Source transparency.&lt;/strong&gt; Does the tool tell you where it searched, and what kind of index it used?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Epistemic humility.&lt;/strong&gt; Does it admit when it does not know, or does it answer everything with the same smooth confidence?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Depth.&lt;/strong&gt; How many searches, how many sources, how many reasoning steps went into the report?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cost and privacy.&lt;/strong&gt; What do you pay, what are the caps, and what happens to the data you feed it?&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  The four-box model that sorts every tool
&lt;/h3&gt;

&lt;p&gt;A simple way to place any research tool: two axes. Fast search versus deep search on one side. A list of results versus generated prose on the other. A classic search engine gives you a fast list. An academic database gives you a fast, structured list. A deep-research agent spends minutes searching, reading, comparing, and then writes you a cited report. Neither quadrant is "better". They answer different questions, and confusing them is how people end up treating a two-minute orientation report as a finished literature review.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;ℹ️ Note:&lt;/strong&gt; The distinction that explains almost everything else in this article is open web versus academic index. Open-web tools find current news, company pages, government documents, anything fresh, but source quality varies wildly. Academic indexes are narrower and structured: metadata, abstracts, citation graphs, peer-reviewed literature. They miss unpublished work and often cannot read paywalled full text.&lt;/p&gt;

&lt;h2&gt;
  
  
  The general deep-research agents, honestly
&lt;/h2&gt;

&lt;p&gt;These are the tools with the famous buttons. They search the open web, read what they find, and write you a report with citations. Each one has a personality, and each personality has a failure mode.&lt;/p&gt;

&lt;h3&gt;
  
  
  ChatGPT vs Gemini vs Claude vs Perplexity for research
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;ChatGPT&lt;/strong&gt; is the strongest overall choice in independent testing and the most cautious about inventing information. The evidence for that second part is the hallucination testing covered below: in the largest URL-validity study, OpenAI Deep Research invented 3.5% of its links, against 13.3% for Gemini's equivalent. Older launch figures still circulate and should be read as history, not as today's product. When OpenAI launched Deep Research on February 2, 2025, the model behind it scored 26.6% on Humanity's Last Exam, against 9.1% for o1 and 3.3% for GPT-4o, according to &lt;a href="https://www.firecrawl.dev/blog/best-ai-for-research" rel="noopener noreferrer"&gt;Firecrawl's comparison of AI research tools&lt;/a&gt;. That benchmark measures correct answers on very hard questions, not caution. The early-2025 quotas (5 Deep Research queries a month on the free tier, 250 on Pro) are also out of date, because OpenAI has reshuffled its plans since; check the current limits on your plan before you count on them. It suits you if you want a structured, broad report and can tolerate the wait.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Gemini&lt;/strong&gt; covers a lot of ground, especially if you live inside Google's ecosystem. The weakness: in comparative testing it has been associated with a higher number of invented or unreliable links. Breadth is real. So is the cleanup bill.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Claude&lt;/strong&gt; is the best reasoner and the best writer of the group. It reportedly fabricates fewer links than several competitors. The catch is cost: long, multi-step research queries burn through usage fast.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Perplexity&lt;/strong&gt; is the fastest of the bunch, typically two to four minutes for a cited report, and its citations are the easiest to check because they sit inline and link straight to the retrieved page. The trade-off is shallow analysis. Use it for rapid orientation and source discovery, not as a finished review. Perplexity launched its Deep Research mode on February 14, 2025, twelve days after OpenAI, with a reported 21.1% on Humanity's Last Exam at launch, a February 2025 figure that says little about the current version.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Grok&lt;/strong&gt; looks polished and plugs into live social data from X. For general research it is unreliable. For breaking news it is a signal, never an authority.&lt;/p&gt;

&lt;h3&gt;
  
  
  How accurate are AI deep research citations?
&lt;/h3&gt;

&lt;p&gt;Here is where it stops being a matter of taste. A Tow Center for Digital Journalism study tested eight AI search engines on 1,600 source-identification tasks: each system got an excerpt from a real article and had to name the headline, date, publisher, and URL. The systems failed more than 60% of the time overall. Perplexity had the lowest failure rate at 37%. Grok-3 Search had the highest at 94%, and returned 154 links to 404 pages across 200 tests. More than half of the answers from Gemini and Grok 3 cited fabricated or broken URLs, which are two different failures lumped into one number: a broken link may once have worked, an invented one never did. And the systems rarely hedged: ChatGPT signaled uncertainty only 15 times across all 200 of its responses, even though 134 of them were wrong.&lt;/p&gt;

&lt;p&gt;A separate, newer study took a different angle: instead of asking whether a claim is supported, it checked whether the URLs themselves exist. The paper, &lt;a href="https://arxiv.org/html/2604.03173" rel="noopener noreferrer"&gt;Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents&lt;/a&gt;, analyzed over 220,000 URLs across commercial models. It found hallucinated URL rates of 3% to 13% in retrieval-augmented settings, with 5% to 18% of URLs failing to resolve at all. A hallucinated URL, by their definition, is a dead link with no record in the Wayback Machine, meaning it probably never existed.&lt;/p&gt;

&lt;p&gt;The detail that should change your habits: deep-research agents generated far more citations per query than search-augmented models, and were less reliable, not more. Gemini 2.5 Pro Deep Research produced an average of 113.1 URLs per query with a 13.3% hallucination rate. OpenAI Deep Research produced 41.2 URLs per query with a 3.5% hallucination rate. Pooled together, the deep-research agents hallucinated 10.7% of URLs, against 4.8% for eight search-augmented models. More citations does not mean better citations. Sometimes it just means more marginal sources dressed up as evidence.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;⚠️ Warning:&lt;/strong&gt; These figures describe specific test setups, a news-attribution task in one case, URL validity in the other. They are not universal accuracy scores for every question you might ask. Treat them as what they are: proof that fabrication is common enough to measure, in every tool, including the paid ones.&lt;/p&gt;

&lt;h2&gt;
  
  
  The academic research tools, by the job they do
&lt;/h2&gt;

&lt;p&gt;General agents search the open web. Academic tools search structured scholarly indexes, and that changes what they are good at. Over 5.14 million academic articles are now published annually, according to &lt;a href="https://cypris.ai/insights/11-best-ai-tools-for-scientific-literature-review-in-2026" rel="noopener noreferrer"&gt;Cypris's review of literature tools&lt;/a&gt;, so nobody reads everything anymore. The question is which machine reads it for you, and how you check its work. No single tool wins every job, so here they are by role.&lt;/p&gt;

&lt;h3&gt;
  
  
  AI-powered literature search
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Consensus&lt;/strong&gt; is the strongest general academic search tool. It searches over 250 million research papers, partners with more than 170 university libraries, and shows a visual evidence meter meant to show which way the science leans. Treat that meter with care: it is essentially a count of papers saying yes versus no, which is vote counting, a method evidence synthesis abandoned decades ago because it ignores study size, quality, and effect size. Around 10 million researchers, students, and clinicians use it as an entry point into the literature. The free tier caps you at 3 Deep Searches a month; Pro runs $15 a month or $120 a year, and the Deep plan is $65 a month or $540 a year. Paywalls are only a partial limitation. Through LibKey, Consensus sends you to the full text your university library already pays for, and for part of the paywalled literature it has full text through publisher partnerships. Where neither applies, it works from metadata and abstracts.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Elicit&lt;/strong&gt; is the extraction specialist. It searches over 138 million papers and 545,000 clinical trials, and its strength is pulling structured data out of papers into tables. The pricing page lists a free Basic tier, Plus at $11 per user per month billed annually, Pro at $39 per user per month billed annually with a systematic-review workflow that can screen 5,000 papers, and Scale at $89 per user per month billed annually for collaboration features. Enterprise goes up to 40,000 screened papers with SSO, SAML, and 2FA. The weakness that matters most: in independent evaluations, Elicit's search finds only about 40% of the relevant studies, against roughly 95% for a classic database search. Where the evidence supports it is as a second reviewer during screening and data extraction, not as the main search.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;SciSpace&lt;/strong&gt; is for reading heavy papers, the ones with dense methods sections. &lt;strong&gt;Undermind&lt;/strong&gt; pitches itself for exhaustive searches, when missing one relevant study is the failure you care about. Keep in mind that the evidence for that claim so far comes from the company that makes it, not from independent testing.&lt;/p&gt;

&lt;h3&gt;
  
  
  Free discovery and citation mapping
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Semantic Scholar&lt;/strong&gt; is free, run by the nonprofit Allen Institute for AI, and indexes over 200 million papers with TLDR summaries and citation signals. It is where a literature search should start when the budget is zero.&lt;/p&gt;

&lt;p&gt;For seeing the shape of a field, the mapping tools earn their keep. &lt;strong&gt;ResearchRabbit&lt;/strong&gt; moved from fully free to freemium in 2026, capping the free tier at 50 seed articles before a subscription of around $10 a month. &lt;strong&gt;Connected Papers&lt;/strong&gt; caps free use at five graphs a month, with paid tiers around $4 to $8 a month. &lt;strong&gt;Litmaps&lt;/strong&gt; gives you two maps and 100 articles per map free, with Pro around $10 a month. These tools find papers that keyword search misses, because they follow citation graphs instead of words.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Scite&lt;/strong&gt; does something no one else does at this scale: its Smart Citations analyze over 1.6 billion citation statements and classify each as supporting, contrasting, or merely mentioning a paper. Its Assistant draws on 317 million full-text articles from more than 44 publisher partners, and its Reference Check lets you upload a manuscript and see whether any of your references have been retracted or contradicted. There is no permanent free tier, only a 7-day trial; individual access runs around $20 a month or $12 billed annually. For &lt;strong&gt;citation verification&lt;/strong&gt;, this is the serious instrument.&lt;/p&gt;

&lt;h3&gt;
  
  
  Your own PDFs, your bibliography, and formal reviews
&lt;/h3&gt;

&lt;p&gt;For working with your own documents, &lt;strong&gt;Gemini Notebook&lt;/strong&gt; handles personal PDF collections. A warning applies to the whole category of "chat with PDF" apps: do not assume the tool actually read every page. Check whether figures, tables, footnotes, and scanned pages made it into the model's view before trusting a summary.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Zotero&lt;/strong&gt; remains the source of truth for your bibliography. AI tools can suggest, summarize, and format references. They should not own your reference library, because they hallucinate, and your library is the thing you check them against.&lt;/p&gt;

&lt;p&gt;For formal systematic reviews, where screening protocols, audit trails, and reproducibility matter more than a friendly chat interface, the specialist platforms are &lt;strong&gt;DistillerSR&lt;/strong&gt; and &lt;strong&gt;EPPI-Reviewer&lt;/strong&gt;. Elicit's structured workflow can help as a second reviewer at screening and data extraction, but with roughly 40% recall it should not run the search itself. A general chatbot is not a systematic review tool, whatever its marketing says.&lt;/p&gt;

&lt;h2&gt;
  
  
  The uncomfortable part: where these tools fail
&lt;/h2&gt;

&lt;p&gt;Now the section the vendors would rather you skipped. The failure modes are not edge cases. They are measured, published, and getting more visible.&lt;/p&gt;

&lt;h3&gt;
  
  
  Invented links, dead links, and confident nonsense
&lt;/h3&gt;

&lt;p&gt;Deep-research agents invent links at more than double the rate of ordinary search-augmented chatbots, per the arXiv study above. They answer confidently when they should refuse. And paid versions can be worse in a specific way: they refuse less often, which means more answers, which means more confident errors. You are, in a sense, paying for reduced humility.&lt;/p&gt;

&lt;p&gt;There is also a verification ceiling. The study's UNKNOWN category, 10% to 20% of URLs across models, exists because bot-blocking, paywalls, and ambiguous server responses prevent automated checking. A browser audit found 89% of sampled UNKNOWN URLs were actually live or blocked but operational. So even a good verification tool cannot settle everything, and access-controlled articles remain a blind spot: if the tool cannot read the full text, it is reasoning from an abstract, a snippet, or a guess.&lt;/p&gt;

&lt;h3&gt;
  
  
  The rot is already in the published record
&lt;/h3&gt;

&lt;p&gt;This is not hypothetical. A Nature report estimated that 2.6% of papers at three computer-science conferences in 2025 contained at least one potentially hallucinated citation, up from about 0.3% in 2024. A separate analysis reported by STAT examined over 2 million papers and 97 million citations and found roughly 4,000 fabricated citations across 2,800 papers; the estimated frequency went from one paper in 2,828 in 2023 to one in 458 in 2025, and one in 277 in the first seven weeks of 2026. Fabricated references are compounding, because each generation of papers can cite the previous generation's inventions.&lt;/p&gt;

&lt;p&gt;Add the mundane risks on top: personal accounts used for work data, prices and limits that change mid-project, and the simple fact that asking the same question twice can return a different source set and a different conclusion. Reproducibility, one of the eight criteria, is where these tools are quietly weakest.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which AI research tool should you use?
&lt;/h2&gt;

&lt;p&gt;Depends who you are. That is not a dodge; it is the actual finding. Match the stack to the job:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Student on a budget:&lt;/strong&gt; Semantic Scholar and Google Scholar for discovery, a free Consensus or Elicit account for AI-assisted search, Zotero for references. Total cost: zero.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Thesis or article researcher:&lt;/strong&gt; academic search first (Consensus, Semantic Scholar, a mapping tool), Scite to check how those papers are cited, Zotero as the reference library, and only then Claude or ChatGPT to help you write from the sources you have already chosen. Reverse that order and invented citations end up in the foundation of the thesis.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Systematic review team:&lt;/strong&gt; DistillerSR or EPPI-Reviewer, with a classic database search. Elicit only as a second screener or for data extraction. Not a chatbot.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Professional analyst:&lt;/strong&gt; ChatGPT for structured reports, Gemini for breadth, Perplexity for fast source-linked orientation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Author who wants nuance:&lt;/strong&gt; Claude for reasoning and prose, with every source verified by hand.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Quick facts:&lt;/strong&gt; Perplexity. Two to four minutes, clickable citations.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Breaking news:&lt;/strong&gt; Grok as a real-time signal, verified elsewhere before you repeat anything.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;💡 Tip:&lt;/strong&gt; Whatever stack you pick, run the verification habit: open every citation, check that it says what the report claims it says, and spot-check the metadata. The arXiv study's own fix, an open-source URL checker, cut non-resolving citations from 16.0% to 0.6% in one experiment. Verification works. It just has to actually happen.&lt;/p&gt;

&lt;h2&gt;
  
  
  People Also Ask
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Can AI deep-research tools be trusted to provide accurate citations?
&lt;/h3&gt;

&lt;p&gt;Not blindly. Measured hallucinated-URL rates run from 3% to 13% depending on the model, and on average deep-research agents invent links about twice as often as search-augmented models (10.7% versus 4.8%), despite producing more citations. The spread is wide: OpenAI Deep Research was at 3.5%, Gemini Deep Research at 13.3%. Even Perplexity, the best performer in the Tow Center news-attribution test, failed 37% of the time there. Every citation needs to be opened and checked before it goes into anything you publish.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why do AI research agents invent or break links?
&lt;/h3&gt;

&lt;p&gt;Two different problems get lumped together. Some URLs are genuine link rot: the page existed and is now gone. Others are fabrications: the model generated a plausible-looking address that never existed, which the Wayback Machine test exposes. Retrieval architecture matters more than citation volume, and agents that generate over a hundred URLs per query tend to pad reports with marginal or mismatched sources.&lt;/p&gt;

&lt;h3&gt;
  
  
  What is the difference between web search and an academic research index?
&lt;/h3&gt;

&lt;p&gt;Open-web tools crawl the live internet: news, company pages, government sites, anything current, with highly variable source quality. Academic indexes like those behind Consensus, Elicit, and Semantic Scholar are structured databases of scholarly papers with metadata, abstracts, and citation graphs. They are more reliable for peer-reviewed work but miss unpublished material and often cannot read paywalled full text.&lt;/p&gt;

&lt;h3&gt;
  
  
  What is the best free AI research tool for students?
&lt;/h3&gt;

&lt;p&gt;Semantic Scholar is completely free and indexes over 200 million papers with AI-generated summaries. Pair it with the free tiers of Consensus (3 Deep Searches a month) or Elicit, and manage references in Zotero, which is also free. That combination covers discovery, evidence checking, and bibliography without spending anything.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bottom line
&lt;/h2&gt;

&lt;p&gt;If you must pick one general application: ChatGPT, for the strongest overall results and the most caution about fabrication. If breadth or price matters more: Gemini. If you must pick one academic workflow: Consensus for search, Zotero underneath it as the reference library that never hallucinates.&lt;/p&gt;

&lt;p&gt;But the real answer is the three rules, because no tool choice protects you without them:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Open every citation. Not some of them. Every one that ends up in your work.&lt;/li&gt;
&lt;li&gt;Never treat an AI search as complete. The corpus is partial, the ranking is opaque, and the same question asked twice gives a different answer.&lt;/li&gt;
&lt;li&gt;A human answers for everything that gets published. The tool does not sign the paper. You do.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  🎯 What you now know about AI research tools in 2026
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;No single app wins everything: general deep-research agents and academic indexes do different jobs, and a reliable workflow combines both.&lt;/li&gt;
&lt;li&gt;Citation fabrication is measurable everywhere: 3% to 13% hallucinated URLs across models, and deep-research agents, on average, invent links about twice as often as simpler search-augmented ones.&lt;/li&gt;
&lt;li&gt;ChatGPT leads the general agents overall, Perplexity is fastest with the most checkable citations, Claude reasons best, Gemini covers the most ground, Grok is for breaking news only.&lt;/li&gt;
&lt;li&gt;Consensus plus Zotero is the default academic stack; Scite verifies how papers are actually cited; Elicit helps with extraction and as a second screener, but finds only about 40% of relevant studies, so it is not a main search.&lt;/li&gt;
&lt;li&gt;Fabricated citations are already appearing in published papers at a rising rate, so verification is not paranoia, it is hygiene.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Additional Resources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/html/2604.03173" rel="noopener noreferrer"&gt;Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://suprmind.ai/hub/perplexity/vs-other-ai/" rel="noopener noreferrer"&gt;Perplexity vs ChatGPT, Claude, Gemini and Grok: A 2026 Honest Comparison&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.firecrawl.dev/blog/best-ai-for-research" rel="noopener noreferrer"&gt;Best AI Tools for Research in 2026: From Answer Engines to Research Agents&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://cypris.ai/insights/11-best-ai-tools-for-scientific-literature-review-in-2026" rel="noopener noreferrer"&gt;11 Best AI Tools for Scientific Literature Review in 2026&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://elicit.com/pricing" rel="noopener noreferrer"&gt;Pricing | Elicit: The AI Research Assistant&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://papersflow.ai/blog/best-ai-for-researchers-2026" rel="noopener noreferrer"&gt;Best AI for Researchers in 2026: 10 Tools Compared by Category – PapersFlow&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>programming</category>
      <category>ai</category>
      <category>deepresearch</category>
      <category>aideepresearchtools</category>
    </item>
    <item>
      <title>Claude for Financial Services: Open-Source Finance Agents</title>
      <dc:creator>DiFlowrin</dc:creator>
      <pubDate>Wed, 23 Sep 2026 18:00:18 +0000</pubDate>
      <link>https://dev.to/diflowrin/claude-for-financial-services-open-source-finance-agents-31ep</link>
      <guid>https://dev.to/diflowrin/claude-for-financial-services-open-source-finance-agents-31ep</guid>
      <description>&lt;p&gt;Anthropic has published &lt;strong&gt;Claude for Financial Services&lt;/strong&gt;, an open-source repository of reference agents, skills, and data connectors for the financial-services workflows it sees most often: investment banking, equity research, private equity, and wealth management. The project lives at &lt;a href="https://github.com/anthropics/financial-services" rel="noopener noreferrer"&gt;github.com/anthropics/financial-services&lt;/a&gt;, is written primarily in Python, and ships under the Apache License 2.0. What makes it unusual is that everything in it is available two ways from one source: you can install the same agent as a Claude Cowork plugin, or deploy it through the Claude Managed Agents API behind your own workflow engine. Same system prompt, same skills, you choose where it runs.&lt;/p&gt;

&lt;p&gt;If you work in finance and you have been waiting for a concrete, inspectable example of what financial services AI looks like in production shape, this is the most useful thing Anthropic has released for you. It is not a demo and it is not a product you have to buy. It is a file-based reference implementation, markdown and JSON with no build step, that you can read cover to cover, fork, and tune to how your firm actually works.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this repository matters right now
&lt;/h2&gt;

&lt;p&gt;Financial analysis is full of work that is structured, repetitive, and expensive: building comparable company analyses, updating models after earnings calls, reconciling general ledgers, screening KYC documents. These tasks follow known methods, which makes them good candidates for AI finance workflows, but they also sit inside regulated firms where a black-box answer is not acceptable. As a result, the question for most teams has not been whether AI can do the work, but whether it can do the work in a way a risk committee will sign off on.&lt;/p&gt;

&lt;p&gt;Anthropic's answer with this repository is twofold. First, every agent drafts analyst work product, models, memos, research notes, reconciliations, for review by a qualified professional. The repository states plainly that the agents do not make investment recommendations, execute transactions, bind risk, post to a ledger, or approve onboarding; every output is staged for human sign-off. Second, the whole thing is plain files. There is no hidden service doing something you cannot inspect. A compliance officer can read the system prompt of the KYC Screener the same way they would read a procedure document.&lt;/p&gt;

&lt;p&gt;The timing also reflects a broader push. Anthropic's own &lt;a href="https://www.anthropic.com/news/claude-for-financial-services" rel="noopener noreferrer"&gt;financial services announcement&lt;/a&gt; reports that Claude Opus 4 passed 5 out of 7 levels of the Financial Modeling World Cup competition with 83% accuracy on complex Excel tasks, and that Claude 4 models outperform other frontier models as research agents across financial tasks in Vals AI's Finance Agent benchmark. Those are Anthropic's claims about its own models, so treat them as vendor-reported figures, but they explain why the company felt confident publishing reference workflows rather than just marketing copy.&lt;/p&gt;

&lt;h2&gt;
  
  
  How the repository is organized
&lt;/h2&gt;

&lt;p&gt;The repository splits its content into a few clear layers, and understanding the split is the key to using it well. At the top sit ten named agents, each one owning a workflow end to end. Underneath them are vertical plugins that bundle skills, slash commands, and data connectors by line of business. Alongside both are cookbooks for headless deployment and admin tooling for Microsoft 365.&lt;/p&gt;

&lt;h3&gt;
  
  
  Ten named agents for financial analysis workflows
&lt;/h3&gt;

&lt;p&gt;Each agent is named for the workflow it runs, and each ships as a self-contained plugin that bundles the skills it uses, so installing the agent is all you need. The ten agents group into four functions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Coverage and advisory:&lt;/strong&gt; the Pitch Agent runs comps, precedents, and LBO analysis through to a branded pitch deck, and the Meeting Prep Agent assembles a briefing pack before every client meeting.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Research and modeling:&lt;/strong&gt; the Market Researcher turns a sector or theme into an industry overview, competitive landscape, peer comps, and an ideas shortlist; the Earnings Reviewer takes an earnings call plus filings through a model update to a note draft; the Model Builder produces DCF, LBO, three-statement, and comps models live in Excel.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fund admin and finance ops:&lt;/strong&gt; the Valuation Reviewer ingests GP packages, runs the valuation template, and stages LP reporting; the GL Reconciler finds breaks, traces root cause, and routes items for sign-off; the Month-End Closer handles accruals, roll-forwards, and variance commentary; the Statement Auditor audits LP statements before distribution.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Operations and onboarding:&lt;/strong&gt; the KYC Screener parses onboarding documents, runs the rules engine, and flags gaps.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Because each agent is a starting point rather than a finished product, the expectation is that you install the ones matching your work and then tune the prompts, skills, and connectors to your firm's conventions.&lt;/p&gt;

&lt;h3&gt;
  
  
  Vertical plugins: skills and commands by business line
&lt;/h3&gt;

&lt;p&gt;If you do not want a full agent, you can install the underlying vertical plugins on their own and just get the slash commands and connectors. The core plugin is &lt;strong&gt;financial-analysis&lt;/strong&gt;, which carries the shared modeling skills and the data connectors, and the README says to install it first. On top of that sit investment-banking (CIMs, teasers, process letters, buyer lists, merger models, deal tracking), equity-research (earnings notes, initiations, model updates, thesis and catalyst tracking), private-equity (sourcing, screening, diligence checklists, IC memos, portfolio monitoring), fund-admin (GL recon, break tracing, accruals, roll-forwards, variance commentary, NAV tie-out), and operations (KYC document parsing and rules-grid evaluation). There is also a claude-for-financial-advisors plugin covering advisor workflows such as meeting prep, compliance pre-check, prospect intake, and rebalance review.&lt;/p&gt;

&lt;p&gt;The skill-level detail is where the repository gets concrete. The financial-analysis plugin alone includes skills for comparable company analysis with trading multiples (&lt;code&gt;/comps&lt;/code&gt;), discounted cash flow valuation with WACC and sensitivity analysis (&lt;code&gt;/dcf&lt;/code&gt;), leveraged buyout modeling (&lt;code&gt;/lbo&lt;/code&gt;), populating three-statement financial model templates (&lt;code&gt;/3-statement-model&lt;/code&gt;), and an Excel model audit that does formula tracing, hardcode detection, and balance checks (&lt;code&gt;/debug-model&lt;/code&gt;). The equity-research plugin adds earnings call analysis through &lt;code&gt;/earnings&lt;/code&gt; and &lt;code&gt;/earnings-preview&lt;/code&gt;, plus initiation reports, morning notes, and a catalyst calendar. For private equity AI work, the private-equity plugin covers deal sourcing, screening, diligence checklists, unit economics, IRR/MOIC sensitivity tables, and investment committee memo drafting through &lt;code&gt;/ic-memo&lt;/code&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Financial services MCP data connectors
&lt;/h3&gt;

&lt;p&gt;All of the data connectors are centralized in the financial-analysis core plugin and shared across the rest. They are MCP servers that wire Claude to terminals, research platforms, and document stores. The README's connector table lists twelve providers: Daloopa, Morningstar, S&amp;amp;P Global, FactSet, Moody's, MT Newswires, Aiera, LSEG, PitchBook, Chronograph, Egnyte, and Box. One wrinkle worth knowing: the README describes the core plugin as carrying "all 11 data connectors" while its own table lists twelve, so the documentation is not internally consistent on the count. In practice, what matters is the list itself, and the caveat the README attaches to it: MCP access may require a subscription or API key from the provider. These are enterprise data services, and the connectors being pre-configured does not mean the data is free.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two ways to run the same agent
&lt;/h2&gt;

&lt;p&gt;The dual-track design is the architectural idea at the heart of the repository, and it is worth understanding before you install anything.&lt;/p&gt;

&lt;h3&gt;
  
  
  Installing as Claude Cowork plugins
&lt;/h3&gt;

&lt;p&gt;The interactive track is Cowork. In Cowork, you open Settings, then Plugins, then Add plugin, and either paste the repository URL (&lt;code&gt;https://github.com/anthropics/financial-services&lt;/code&gt;) and pick agents and verticals from the marketplace list, or zip any directory under &lt;code&gt;plugins/&lt;/code&gt; and upload it directly. Once installed, agents appear in Cowork dispatch, skills fire automatically when relevant, and slash commands such as &lt;code&gt;/comps&lt;/code&gt;, &lt;code&gt;/dcf&lt;/code&gt;, &lt;code&gt;/earnings&lt;/code&gt;, and &lt;code&gt;/ic-memo&lt;/code&gt; become available in your session.&lt;/p&gt;

&lt;h3&gt;
  
  
  Installing through Claude Code
&lt;/h3&gt;

&lt;p&gt;For Claude Code, the README gives the exact commands:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Add the marketplace&lt;/span&gt;
claude plugin marketplace add anthropics/financial-services

&lt;span class="c"&gt;# Core skills + connectors (install first)&lt;/span&gt;
claude plugin &lt;span class="nb"&gt;install &lt;/span&gt;financial-analysis@claude-for-financial-services

&lt;span class="c"&gt;# Named agents — pick the ones you want&lt;/span&gt;
claude plugin &lt;span class="nb"&gt;install &lt;/span&gt;pitch-agent@claude-for-financial-services
claude plugin &lt;span class="nb"&gt;install &lt;/span&gt;gl-reconciler@claude-for-financial-services
claude plugin &lt;span class="nb"&gt;install &lt;/span&gt;market-researcher@claude-for-financial-services

&lt;span class="c"&gt;# Vertical skill bundles&lt;/span&gt;
claude plugin &lt;span class="nb"&gt;install &lt;/span&gt;investment-banking@claude-for-financial-services
claude plugin &lt;span class="nb"&gt;install &lt;/span&gt;equity-research@claude-for-financial-services
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Note that the marketplace name after the &lt;code&gt;@&lt;/code&gt; is &lt;code&gt;claude-for-financial-services&lt;/code&gt;, which is not the same as the repository name. Getting that string wrong is the most common install mistake.&lt;/p&gt;

&lt;h3&gt;
  
  
  Deploying through the Claude Managed Agents API
&lt;/h3&gt;

&lt;p&gt;The headless track uses the &lt;a href="https://docs.claude.com/en/api/managed-agents" rel="noopener noreferrer"&gt;Claude Managed Agents API&lt;/a&gt;. Each template under &lt;code&gt;managed-agent-cookbooks/&lt;/code&gt; references the same system prompt and skills as its plugin counterpart, and deployment is a two-line affair:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;ANTHROPIC_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;sk-ant-...
scripts/deploy-managed-agent.sh gl-reconciler
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The deploy script resolves file references, uploads skills, creates leaf-worker subagents, and POSTs the orchestrator to &lt;code&gt;/v1/agents&lt;/code&gt;. The repository also includes &lt;code&gt;scripts/orchestrate.py&lt;/code&gt;, a reference event loop that routes &lt;code&gt;handoff_request&lt;/code&gt; events between agents through your own orchestration layer. However, one capability here is explicitly flagged as early: subagent delegation (&lt;code&gt;callable_agents&lt;/code&gt;) is a research preview, and the per-agent READMEs carry security and handoff guidance for it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the agents actually do in practice
&lt;/h2&gt;

&lt;p&gt;It helps to walk through a couple of workflows end to end, because the value is in the chaining rather than in any single skill.&lt;/p&gt;

&lt;h3&gt;
  
  
  Claude agents for investment banking workflows
&lt;/h3&gt;

&lt;p&gt;Take the Pitch Agent as the clearest example of investment banking automation. It runs comparable company analysis, pulls precedent transactions, builds an LBO, and assembles the results into a branded pitch deck, end to end. The investment-banking vertical underneath it supplies the individual pieces: &lt;code&gt;/one-pager&lt;/code&gt; for strip profiles, &lt;code&gt;/cim&lt;/code&gt; for drafting Confidential Information Memorandums, &lt;code&gt;/teaser&lt;/code&gt; for anonymous one-pagers, &lt;code&gt;/buyer-list&lt;/code&gt; for the strategic and financial buyer universe, &lt;code&gt;/merger-model&lt;/code&gt; for accretion/dilution analysis, and &lt;code&gt;/deal-tracker&lt;/code&gt; for live deal milestones. Similarly, the &lt;code&gt;/ppt-template&lt;/code&gt; command teaches Claude your firm's branded PowerPoint layouts, so the output deck looks like your deck rather than a generic one.&lt;/p&gt;

&lt;h3&gt;
  
  
  Equity research AI: from earnings call to published note
&lt;/h3&gt;

&lt;p&gt;The Earnings Reviewer shows the research loop. It ingests the earnings call and filings, updates the model, and drafts the note. The equity-research plugin's skills cover the surrounding cadence of a coverage desk: pre-earnings scenario analysis, post-earnings quarterly updates, initiations, morning notes, thesis tracking, and a catalyst calendar. Meanwhile, the Model Builder handles the AI-powered DCF and LBO financial modeling side directly in Excel, which matters because Excel is where these models actually live in most firms.&lt;/p&gt;

&lt;h3&gt;
  
  
  Month-end close and KYC screening
&lt;/h3&gt;

&lt;p&gt;On the operations side, the GL Reconciler finds breaks, traces the root cause, and routes items for sign-off, while the Month-End Closer produces accruals, roll-forwards, and variance commentary. The KYC Screener parses onboarding documents and runs a rules-grid evaluation, flagging gaps for a human to resolve. In each case the pattern is the same: the agent does the assembly and the first pass, and a person makes the decision.&lt;/p&gt;

&lt;h2&gt;
  
  
  Partner plugins and the Microsoft 365 add-in
&lt;/h2&gt;

&lt;p&gt;Two partner-built plugins extend the core set. The LSEG plugin covers bond relative value, swap curves, FX carry, options vol, and macro-rates monitoring on LSEG data, and the S&amp;amp;P Global plugin produces tear sheets, earnings previews, and funding digests on S&amp;amp;P Capital IQ. Both live under &lt;code&gt;plugins/partner-built/&lt;/code&gt; and follow the same file-based structure as the rest.&lt;/p&gt;

&lt;p&gt;Separately, the repository includes &lt;code&gt;claude-for-msft-365-install/&lt;/code&gt;, admin tooling for firms that run Claude inside Excel, PowerPoint, Word, and Outlook through the Microsoft 365 add-in. It is a Claude Code plugin, not a Cowork plugin, and it walks an IT admin through generating the customized add-in manifest, granting Azure admin consent, and writing per-user routing config via Microsoft Graph. Notably, it provisions the add-in against your own cloud, Vertex AI, Bedrock, or an internal LLM gateway, instead of Anthropic's API. Install it with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;claude plugin &lt;span class="nb"&gt;install &lt;/span&gt;claude-for-msft-365-install@claude-for-financial-services
/claude-for-msft-365-install:setup
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Making the templates yours
&lt;/h2&gt;

&lt;p&gt;Anthropic is explicit that these are reference templates that get better when you tune them. The suggested customization paths are practical: swap connectors by pointing &lt;code&gt;.mcp.json&lt;/code&gt; at your own data providers and internal systems, drop your firm's terminology and formatting standards into skill files, teach Claude your branded PowerPoint layouts with &lt;code&gt;/ppt-template&lt;/code&gt;, and edit &lt;code&gt;agents/&amp;lt;slug&amp;gt;.md&lt;/code&gt; to match how your team actually runs each workflow.&lt;/p&gt;

&lt;p&gt;Contributing back follows the same file-based logic. New skills go under the relevant vertical's &lt;code&gt;skills/&lt;/code&gt; directory, and &lt;code&gt;python3 scripts/sync-agent-skills.py&lt;/code&gt; propagates them to any agent that bundles them. Before pushing, &lt;code&gt;python3 scripts/check.py&lt;/code&gt; lints every manifest, verifies that all cross-file references resolve, and fails if any bundled skill has drifted from its vertical source. That drift check is a small detail, but it is the kind of thing that keeps a multi-plugin repository coherent as it grows.&lt;/p&gt;

&lt;h2&gt;
  
  
  Limitations and what to watch
&lt;/h2&gt;

&lt;p&gt;A few honest caveats. First, the legal framing is not decorative: nothing in the repository is investment, legal, tax, or accounting advice, and your firm is responsible for verifying outputs and for regulatory compliance. Second, the connectors are only as useful as your subscriptions; a team without FactSet or PitchBook contracts will need to swap in what it actually has. Third, subagent delegation is a research preview, so headless multi-agent deployments deserve extra scrutiny before they touch real client work.&lt;/p&gt;

&lt;p&gt;It is also worth keeping the benchmark numbers in perspective. Anthropic reports strong results for its models on financial tasks, and its &lt;a href="https://claude.com/solutions/financial-services" rel="noopener noreferrer"&gt;financial services page&lt;/a&gt; cites compliance postures such as SOC 2 and FedRAMP, but neither the repository nor those pages publish pricing for the plugins, and there is no independent measurement yet of productivity impact across adopting firms. The fairest reading is that the repository gives you the machinery; the business case still has to be made inside your own workflow.&lt;/p&gt;

&lt;h2&gt;
  
  
  People Also Ask
&lt;/h2&gt;

&lt;h3&gt;
  
  
  How do I install the financial-services plugins in Claude Cowork?
&lt;/h3&gt;

&lt;p&gt;From the Cowork interface, go to Settings, then Plugins, then Add plugin. You can either paste the repository URL, &lt;a href="https://github.com/anthropics/financial-services" rel="noopener noreferrer"&gt;https://github.com/anthropics/financial-services&lt;/a&gt;, and select the agents and verticals you want from the marketplace list, or zip any directory under plugins/ and upload it directly. Install the financial-analysis core plugin first, since it carries the shared skills and all the data connectors.&lt;/p&gt;

&lt;h3&gt;
  
  
  How do I install the financial-services marketplace in Claude Code?
&lt;/h3&gt;

&lt;p&gt;Run &lt;code&gt;claude plugin marketplace add anthropics/financial-services&lt;/code&gt;, then install plugins with commands like &lt;code&gt;claude plugin install financial-analysis@claude-for-financial-services&lt;/code&gt;. The marketplace name after the @ is claude-for-financial-services, spelled exactly like that, and the slash commands appear in a new session once installed.&lt;/p&gt;

&lt;h3&gt;
  
  
  What is the difference between a Claude Cowork plugin and a Managed Agent?
&lt;/h3&gt;

&lt;p&gt;They are two deployment tracks for the same agent. The Cowork plugin runs interactively in your session, while the Managed Agent template deploys headlessly through the /v1/agents API behind your own workflow engine, using agent.yaml, leaf-worker subagents, and steering events. Both reference the same system prompt and skills from the same directory.&lt;/p&gt;

&lt;h3&gt;
  
  
  Do the MCP data connectors include access to the data providers?
&lt;/h3&gt;

&lt;p&gt;No. The connectors for providers such as FactSet, Moody's, LSEG, and PitchBook are pre-configured in the financial-analysis plugin, but the repository notes that MCP access may require a subscription or API key from the provider. You can also point .mcp.json at your own data sources instead.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bottom line
&lt;/h2&gt;

&lt;p&gt;Claude for Financial Services is best read as a reference architecture rather than a product. Its ten agents, six vertical plugins, partner integrations, and Managed Agent cookbooks show, in inspectable files, how Anthropic thinks AI finance workflows should be built: human sign-off on every output, schema-disciplined handoffs, and one source that runs both interactively and headlessly. For a finance team evaluating financial services AI, the fastest way to form an opinion is to clone the repository, install the financial-analysis core, and run &lt;code&gt;/comps&lt;/code&gt; on a company you know well. What you see in that output will tell you more than any benchmark table.&lt;/p&gt;

&lt;h2&gt;
  
  
  Additional Resources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/anthropics/financial-services" rel="noopener noreferrer"&gt;Anthropic Financial Services GitHub Repository&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://claude.com/solutions/financial-services" rel="noopener noreferrer"&gt;Claude Financial Services Solution&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.anthropic.com/news/claude-for-financial-services" rel="noopener noreferrer"&gt;Claude for Financial Services Announcement&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/anthropics/financial-services/blob/main/README.md" rel="noopener noreferrer"&gt;README.md - anthropics/financial-services - GitHub&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/anthropics/financial-services/activity" rel="noopener noreferrer"&gt;Activity · anthropics/financial-services - GitHub&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.strategicmarketresearch.com/market-report/financial-services-market" rel="noopener noreferrer"&gt;Financial Services Market Report (2026): Must-Know Insights &amp;amp; Updates&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;This article includes content created with AI.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>github</category>
      <category>git</category>
      <category>typescript</category>
    </item>
    <item>
      <title>BrowserSkill: AI Agents in Your Real Browser</title>
      <dc:creator>DiFlowrin</dc:creator>
      <pubDate>Thu, 17 Sep 2026 21:20:53 +0000</pubDate>
      <link>https://dev.to/diflowrin/browserskill-ai-agents-in-your-real-browser-3pla</link>
      <guid>https://dev.to/diflowrin/browserskill-ai-agents-in-your-real-browser-3pla</guid>
      <description>&lt;p&gt;Here is the thing nobody tells you about AI agents and browsers. The agent is smart. The browser is yours. And the two do not meet, because your browser holds your logins, your sessions, your accounts, and the agent holds none of that. So every browser automation tool before now made you choose: either hand the agent a sterile, logged-out browser it can barely do anything with, or let it loose in your real one and watch it hijack the tab you were reading. &lt;strong&gt;BrowserSkill&lt;/strong&gt;, an open-source project from Tencent, refuses the choice. It connects agents like Cursor, Claude Code, Codex, OpenClaw, CodeBuddy, WorkBuddy, Pi, Hermes Agent and DeepSeek Harness to your already logged-in browser, and it does the work in a separate window so you keep working. How? That is what the rest of this piece is about.&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem: agents are locked out of the web you actually use
&lt;/h2&gt;

&lt;p&gt;Think about what an agent can and cannot reach. It can write code, run shell commands, read files. But the moment a task touches a website that needs your account, it hits a wall. Your email, your internal dashboards, your admin panels, all of it sits behind a login the agent does not have.&lt;/p&gt;

&lt;h3&gt;
  
  
  📹 Video: Free Tool Gives AI Agents Full Browser Access
&lt;/h3&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/9tmd703wkq0" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Video credit: The Stack&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The usual answers are all bad. Spin up a fresh automation browser: now nothing is logged in, and you are managing separate test accounts. Give the agent your cookies or passwords: a security mess. Let it drive your actual browser window: now it is moving your mouse, stealing your focus, closing your tabs. You become a spectator at your own desk.&lt;/p&gt;

&lt;p&gt;So the real question is not "can an agent use a browser". It can. The question is: can it use &lt;em&gt;your&lt;/em&gt; browser, with &lt;em&gt;your&lt;/em&gt; sessions, without taking the machine away from you. BrowserSkill's answer is yes, and the mechanism is worth understanding because it is genuinely simple.&lt;/p&gt;

&lt;h2&gt;
  
  
  How BrowserSkill actually works
&lt;/h2&gt;

&lt;p&gt;Two local pieces. That is the whole runtime. A command-line tool called &lt;code&gt;bsk&lt;/code&gt;, which runs a small local daemon, and a browser extension. Nothing in the cloud, nothing routed through someone else's servers.&lt;/p&gt;

&lt;p&gt;The chain goes like this. The agent never talks to the browser directly. It calls the &lt;code&gt;bsk&lt;/code&gt; CLI through the shell, the same way it would call any other tool. The CLI passes the request to the local daemon over local IPC. The daemon talks to the extension over a WebSocket on 127.0.0.1. And the extension does the actual browser work, inside a dedicated, visible &lt;strong&gt;Agent Window&lt;/strong&gt; that is separate from your normal windows.&lt;/p&gt;

&lt;p&gt;Why does this architecture matter? Two reasons.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Any agent can use it.&lt;/strong&gt; Anything that can call a shell can call bsk. There is no lock-in to a specific model, agent framework or harness. Cursor, Claude Code, Codex, OpenClaw, CodeBuddy, WorkBuddy, Pi, Hermes Agent, all of them connect the same way.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Your browser stays yours.&lt;/strong&gt; Tasks run in the Agent Window. If the agent genuinely needs a tab you already have open, it must borrow that tab explicitly, return it when the task is done, and leave the rest of your browser alone.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That borrowing rule is the interesting part. The default posture is: hands off the user's stuff. The agent asks, you approve, it returns the tab. Not the other way around.&lt;/p&gt;

&lt;h3&gt;
  
  
  Reuse your login state, skip the test accounts
&lt;/h3&gt;

&lt;p&gt;Because the extension lives in your real browser profile, the agent works with sites you are already signed into. No separate test accounts, no credential handoff. The session you built by logging in like a normal person is the session the agent uses.&lt;/p&gt;

&lt;h3&gt;
  
  
  Human in the loop, built in
&lt;/h3&gt;

&lt;p&gt;And when the task hits something only a human can do, a captcha, a login screen, a confirmation dialog, the agent can ask you to take over, then continue afterwards. This is not a hack. It is a designed feature, and as we will see, it is configurable down to the last switch.&lt;/p&gt;

&lt;h2&gt;
  
  
  Installing it: one line if you have an agent, four steps if you do not
&lt;/h2&gt;

&lt;p&gt;The recommended path is almost funny in how little it asks of you. Already using Cursor, Claude Code, Codex or another shell-capable agent? Copy one line and send it to your agent. It installs the CLI and the skill, then walks you through loading the extension:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Set up browser-skill on this machine by following https://raw.githubusercontent.com/Tencent/BrowserSkill/main/AGENT_INSTALL.md
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is it. The agent does the setup. Which is fitting: a tool for agents, installed by an agent.&lt;/p&gt;

&lt;p&gt;The manual path is not much harder. Four steps.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 1: install the bsk CLI
&lt;/h3&gt;

&lt;p&gt;On macOS or Linux, the recommended install goes to &lt;code&gt;~/.local/bin&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-fsSL&lt;/span&gt; https://raw.githubusercontent.com/Tencent/BrowserSkill/main/install.sh | sh
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;PATH&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;BSK_INSTALL_DIR&lt;/span&gt;&lt;span class="k"&gt;:-&lt;/span&gt;&lt;span class="nv"&gt;$HOME&lt;/span&gt;&lt;span class="p"&gt;/.local/bin&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;:&lt;/span&gt;&lt;span class="nv"&gt;$PATH&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On Windows, from PowerShell, also installing to &lt;code&gt;~/.local/bin&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight powershell"&gt;&lt;code&gt;&lt;span class="n"&gt;irm&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;https://raw.githubusercontent.com/Tencent/BrowserSkill/main/install.ps1&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;iex&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One detail that bites people: the export makes the CLI available in the current Unix shell. A running agent may need the same PATH setting in each shell call, or the installed binary's absolute path. If the agent retains an old PATH after installation, restart it. Then verify the binary in the terminal or agent environment that will actually use it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;bsk &lt;span class="nt"&gt;--version&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Step 2: install the browser extension
&lt;/h3&gt;

&lt;p&gt;Chrome and Microsoft Edge are supported, and the extension is in both stores: the Chrome Web Store and Edge Add-ons. On other Chromium-based browsers, install the Chrome Web Store build; they are expected to work when they support unpacked Chromium extensions. Firefox is planned, not here yet.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 3: install the skill
&lt;/h3&gt;

&lt;p&gt;BrowserSkill ships a skill that teaches your agent harness how to use &lt;code&gt;bsk&lt;/code&gt;. For the supported harnesses, one command:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;bsk install-skill
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Press Space to select the harness, Enter to install. For non-interactive installation, name the harness explicitly, for example &lt;code&gt;bsk install-skill --harness cursor --json&lt;/code&gt;, which also works when the harness is not detected. &lt;code&gt;--yes&lt;/code&gt; alone installs into every detected harness and fails when none are detected. Run &lt;code&gt;bsk install-skill --list&lt;/code&gt; to see internal variants and install paths.&lt;/p&gt;

&lt;p&gt;Want your own instructions instead of the bundled ones? &lt;code&gt;bsk install-skill --harness cursor --source ./SKILL.md&lt;/code&gt;. An explicit &lt;code&gt;--source&lt;/code&gt; stays custom even if its contents match the bundled skill, and existing installations are skipped unless you add &lt;code&gt;--force&lt;/code&gt;. Other shell-capable harnesses work too: copy &lt;code&gt;skill/SKILL.md&lt;/code&gt; into the harness's skills directory as &lt;code&gt;browser-skill/SKILL.md&lt;/code&gt;. DeepSeek Harness is the exception, it uses a dedicated plugin instead, more on that below.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 4: verify the connection
&lt;/h3&gt;

&lt;p&gt;Run &lt;code&gt;bsk doctor&lt;/code&gt; and follow its hints. Open the extension popup and confirm it is connected. Resolve failures before testing browser use. One caveat worth knowing: doctor can pass with no skill installed (it reports &lt;code&gt;N/A&lt;/code&gt;), so verify skill discovery separately.&lt;/p&gt;

&lt;p&gt;Then the first real test. Start a new agent session, confirm &lt;code&gt;browser-skill&lt;/code&gt; is available, and ask it to open &lt;code&gt;https://example.com&lt;/code&gt; and summarize the page. For harnesses with slash-command invocation:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;/browser-skill open example.com and summarize what is on the page.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A successful first run reads the page and stops its BrowserSkill session. If the skill is missing, check the target harness and install path before retrying.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it runs
&lt;/h2&gt;

&lt;p&gt;The platform matrix is broad. Operating systems: macOS on Apple Silicon and Intel, Linux on x64 and ARM64, Windows x64. Browsers: Chrome and Edge supported, other Chromium browsers expected to work, Firefox planned.&lt;/p&gt;

&lt;p&gt;Running inside an agent sandbox that reaps background processes after each command? There is a documented setup for that: keep the daemon in a persistent host environment and connect with a shared &lt;code&gt;BSK_HOME&lt;/code&gt; plus &lt;code&gt;BSK_AUTO_START=0&lt;/code&gt;. Ordinary local use keeps automatic startup by default. And if you want the agent on a server while the browser stays on your desk, you can pair them using the built-in authentication service or a compatible gateway, covered in the remote browser connections documentation.&lt;/p&gt;

&lt;h2&gt;
  
  
  The automation settings: who approves what
&lt;/h2&gt;

&lt;p&gt;This is the part that changed most recently, and the part that decides how much you trust the machine. The extension popup has two independent &lt;strong&gt;Automation settings&lt;/strong&gt;, both enabled by default: "Confirm before borrowing tabs" and "Allow requests for human help". Your saved browser settings are authoritative for every session. Not the CLI flags. The browser settings.&lt;/p&gt;

&lt;p&gt;The four combinations behave exactly as you would expect:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Both on: borrowing requires approval, help requests show the existing UI.&lt;/li&gt;
&lt;li&gt;Confirm on, help off: borrowing requires approval, help requests return disabled.&lt;/li&gt;
&lt;li&gt;Confirm off, help on: borrowing skips confirmation, help requests show the UI.&lt;/li&gt;
&lt;li&gt;Both off: borrowing skips confirmation, help requests return disabled. Fully unattended.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Settings save automatically for the browser profile and apply to existing and new sessions. Turning confirmation off releases pending borrow confirmations; turning help off finishes pending help requests as &lt;code&gt;disabled&lt;/code&gt;. Turning either back on restores its behavior for subsequent operations. Completed borrows are not undone, and finished help requests are not reopened.&lt;/p&gt;

&lt;h3&gt;
  
  
  What changed in 0.3.0
&lt;/h3&gt;

&lt;p&gt;Here is the upgrade note that matters. In version 0.3.0, &lt;code&gt;--unattended&lt;/code&gt;, &lt;code&gt;tab borrow --no-confirm&lt;/code&gt; and &lt;code&gt;BSK_REQUEST_HELP=off&lt;/code&gt; no longer bypass confirmation or disable help. They remain accepted for compatibility, but they are deprecated and cannot override the browser switches. The CLI logs a notice when they are used. Scripts that relied on these inputs alone to avoid waiting must now use the extension settings. &lt;code&gt;session start --json&lt;/code&gt; and &lt;code&gt;session list --json&lt;/code&gt; report the browser's effective &lt;code&gt;interaction&lt;/code&gt; policy, so scripts can read the truth instead of guessing.&lt;/p&gt;

&lt;p&gt;Why the change? Because a command-line flag is a terrible place for a consent decision. The user sitting in front of the browser should own that switch, and now they do.&lt;/p&gt;

&lt;h3&gt;
  
  
  When help is disabled
&lt;/h3&gt;

&lt;p&gt;When help is off, &lt;code&gt;request-help&lt;/code&gt; returns &lt;code&gt;disabled&lt;/code&gt; without confirming any human action. The skill then directs the agent to re-observe and make reasonable efforts to complete authorized steps using existing login state, authorized inputs and available tools. Where task authorization and host rules allow, models with image understanding may attempt graphical verification. But some walls stay walls: phone-only QR scans, face verification, unavailable SMS codes, and image-only captchas for text-only models may remain blocked. A &lt;code&gt;disabled&lt;/code&gt; result neither completes the task nor grants additional permission. Good. A blocked agent should stay blocked.&lt;/p&gt;

&lt;h3&gt;
  
  
  Protocol versions and mixed installations
&lt;/h3&gt;

&lt;p&gt;A few details for people running staggered upgrades. Protocol 1.3 retains connection compatibility with protocols 1.0 through 1.2, and ordinary sessions and default tab borrowing keep working during upgrades. Custom borrowing waits require both daemon and extension at protocol 1.2 or later. The current CLI requires daemon protocol 1.3 for &lt;code&gt;request-help&lt;/code&gt;, because older daemons can answer locally without consulting the browser, which would defeat the whole point. The popup identifies older daemons and &lt;code&gt;bsk status&lt;/code&gt; reports protocol differences. The clean move is to update the CLI, the running daemon and the extension together.&lt;/p&gt;

&lt;h2&gt;
  
  
  Everyday workflows
&lt;/h2&gt;

&lt;p&gt;What does a normal day with this look like? Start tasks with &lt;code&gt;bsk session start&lt;/code&gt;; add &lt;code&gt;--no-focus&lt;/code&gt; if you do not want the Agent Window stealing focus. For unattended operation, turn off the corresponding settings in the extension, not on the command line.&lt;/p&gt;

&lt;p&gt;Need a full-page capture? Two ways. From the extension: Quick actions, then Full-page screenshot. From the agent: &lt;code&gt;bsk screenshot --session &amp;lt;id&amp;gt; --full-page --out page.png&lt;/code&gt;. A full-page screenshot guide covers page support, cancellation and export. Note that new features like full-page screenshots need matching builds of the CLI, daemon and extension, so check versions with &lt;code&gt;bsk --version&lt;/code&gt; and &lt;code&gt;bsk status&lt;/code&gt; if something is missing.&lt;/p&gt;

&lt;p&gt;Updating is one command for the default local setup, once active browser tasks finish:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;bsk update &lt;span class="nt"&gt;--yes&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It restarts a running daemon with default startup settings. If you replaced the binary with the installer instead, restart the existing daemon with &lt;code&gt;bsk daemon restart&lt;/code&gt;. On Windows, if a staged update is reported, wait for the replacement to finish before checking &lt;code&gt;bsk --version&lt;/code&gt;. For a custom port, a host-managed sandbox daemon or a remote server, stop the daemon in its owning host, run &lt;code&gt;bsk update --yes --no-restart-daemon&lt;/code&gt;, and start it there with its original flags and &lt;code&gt;BSK_HOME&lt;/code&gt;. The extension updates through its browser store, and store availability may lag the CLI release.&lt;/p&gt;

&lt;h2&gt;
  
  
  The DeepSeek Harness plugin
&lt;/h2&gt;

&lt;p&gt;DeepSeek Harness users get a first-class path. BrowserSkill ships a dsh plugin on npm as &lt;code&gt;@wxg-prc-cpg/browser-skill-dsh-plugin&lt;/code&gt;. It gives the agent native &lt;code&gt;browser_*&lt;/code&gt; tools and a live view of its browser sessions in the Web UI, and the plugin runs &lt;code&gt;bsk&lt;/code&gt; on the agent's behalf. Same chain as everything else, just with the plugin doing the calling.&lt;/p&gt;

&lt;p&gt;Install the &lt;code&gt;bsk&lt;/code&gt; CLI and connect the extension first. Then add the plugin to a dsh profile and start it (replace &lt;code&gt;web&lt;/code&gt; with your profile name):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;dsh plugin &lt;span class="nt"&gt;--profile&lt;/span&gt; web add @wxg-prc-cpg/browser-skill-dsh-plugin
dsh &lt;span class="nt"&gt;--profile&lt;/span&gt; web
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The plugin includes the &lt;code&gt;browser-skill&lt;/code&gt; skill, so &lt;code&gt;bsk install-skill&lt;/code&gt; is not needed for dsh. Installed plugins do not update automatically; to upgrade, run &lt;code&gt;dsh plugin --profile web update @wxg-prc-cpg/browser-skill-dsh-plugin --latest&lt;/code&gt; and restart the profile afterwards.&lt;/p&gt;

&lt;h2&gt;
  
  
  Under the hood, for developers
&lt;/h2&gt;

&lt;p&gt;The repository is a Cargo plus pnpm workspace, written primarily in TypeScript with the CLI and daemon in Rust, and licensed under MIT. The layout:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;crates/bsk-cli: the bsk CLI and local daemon&lt;/li&gt;
&lt;li&gt;crates/bsk-protocol: shared wire types and JSON schemas&lt;/li&gt;
&lt;li&gt;apps/extension: the browser extension&lt;/li&gt;
&lt;li&gt;packages/ui and packages/i18n: shared extension UI support, including English, Simplified Chinese and Korean localization&lt;/li&gt;
&lt;li&gt;packages/dsh-plugin-browserskill: the DeepSeek Harness plugin&lt;/li&gt;
&lt;li&gt;evals/browser: deterministic local pages and agent-neutral browser capability evaluation&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last piece deserves a sentence. The project ships its own evaluation setup with deterministic local pages, so browser capabilities can be tested without depending on the live web. There is also a scroll-to primitive reference covering its CLI, protocol and plugin entry points, visible bounds and interruption behavior.&lt;/p&gt;

&lt;h2&gt;
  
  
  Limits and what is next
&lt;/h2&gt;

&lt;p&gt;What is not there yet? Firefox, which is planned but not shipped. Some human-only barriers stay human-only when help is disabled: face verification, phone-only QR scans, SMS codes you cannot reach. And the consent model, while much cleaner in 0.3.0, asks mixed-version installations to update all three pieces before the settings are fully enforced.&lt;/p&gt;

&lt;p&gt;None of that changes the core bet. The agent ecosystem is fragmenting into harnesses, frameworks and models, and BrowserSkill's bet is that the browser connection should not fragment with it. One CLI, one extension, any agent that can call a shell. The web you are already logged into, borrowed politely and returned when done.&lt;/p&gt;

&lt;h2&gt;
  
  
  People Also Ask
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Does BrowserSkill work with agents other than Cursor and Claude Code?
&lt;/h3&gt;

&lt;p&gt;Yes. Any agent that can call a shell can use BrowserSkill through the &lt;code&gt;bsk&lt;/code&gt; CLI, with no lock-in to a specific model, framework or harness. The README names Codex, OpenClaw, CodeBuddy, WorkBuddy, Pi and Hermes Agent alongside Cursor and Claude Code, and DeepSeek Harness connects through a dedicated npm plugin instead.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can BrowserSkill run browser tasks without asking me for confirmation?
&lt;/h3&gt;

&lt;p&gt;Yes, but the switch lives in the browser, not the CLI. Turn off "Confirm before borrowing tabs" and "Allow requests for human help" in the extension popup's Automation settings, and tasks run unattended. Since version 0.3.0, the old &lt;code&gt;--unattended&lt;/code&gt; flag and &lt;code&gt;BSK_REQUEST_HELP=off&lt;/code&gt; environment variable are deprecated and cannot override those browser settings.&lt;/p&gt;

&lt;h3&gt;
  
  
  What happens when a task hits a captcha or login screen?
&lt;/h3&gt;

&lt;p&gt;It depends on the help setting. With help requests allowed, the agent pauses and hands the step to you, then picks the task back up once you are done. With help turned off, the &lt;code&gt;request-help&lt;/code&gt; call comes back &lt;code&gt;disabled&lt;/code&gt;, and the agent is told to push on with the login state and inputs it already has. Even then, some barriers do not move: a QR code that only a phone can scan, a face check, or an image-only captcha facing a text-only model can still stop the task cold.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can I run the agent on a server but use my local browser?
&lt;/h3&gt;

&lt;p&gt;Yes. You can pair an agent running on a server with your local browser using the built-in authentication service or a compatible gateway. The setup is covered in the project's remote browser connections documentation.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bottom line
&lt;/h2&gt;

&lt;p&gt;BrowserSkill solves a specific, annoying problem with a specific, clean mechanism: a local CLI and daemon, a browser extension, a separate Agent Window, and a consent model that keeps the human in charge of their own tabs. It reuses the login state you already have, it works with any shell-capable agent, and it is MIT-licensed. If your agents keep bouncing off the logged-in web, this is the bridge.&lt;/p&gt;

&lt;h2&gt;
  
  
  Additional Resources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/Tencent/BrowserSkill" rel="noopener noreferrer"&gt;Tencent/BrowserSkill — Let AI agents use your real, logged-in browser without interrupting your work. CLI + extension for browser automation&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>typescript</category>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
    </item>
    <item>
      <title>Ruflo: Multi-Agent Swarms for Claude Code</title>
      <dc:creator>DiFlowrin</dc:creator>
      <pubDate>Sun, 06 Sep 2026 08:44:13 +0000</pubDate>
      <link>https://dev.to/diflowrin/ruflo-multi-agent-swarms-for-claude-code-5g5k</link>
      <guid>https://dev.to/diflowrin/ruflo-multi-agent-swarms-for-claude-code-5g5k</guid>
      <description>&lt;h2&gt;
  
  
  📋 What You Need Before Ruflo Can Do Anything
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Node.js with npm installed, because every install path below runs through npx or npm&lt;/li&gt;
&lt;li&gt;Claude Code or Codex already working on your machine. Ruflo is a harness around them, not a replacement for them&lt;/li&gt;
&lt;li&gt;A terminal: PowerShell or cmd on Windows, any POSIX shell on macOS or Linux&lt;/li&gt;
&lt;li&gt;A project folder you are comfortable letting Ruflo write configuration into (the full install adds .claude/, CLAUDE.md and helper files)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Here is what most people miss about the &lt;strong&gt;Ruflo agent meta-harness&lt;/strong&gt;: they think it is another AI coding tool. It is not. Ruflo is the thing that wraps around your AI coding tool and turns it from one chatbot into a coordinated team. The project, formerly known as Claude Flow, has a one-line formula for this: Agent = Model + Harness. The model writes. The harness gives it tools, memory, loops, sandboxes and controls so it can actually work. Ruflo is the harness.&lt;/p&gt;

&lt;p&gt;Why should you care? Because a single agent, however good the model, hits a wall fast. One context window. One session of memory. One pair of hands. Ruflo's answer is to put an execution layer around Claude Code and Codex that adds 100+ specialized agents, coordinated swarms, self-learning memory, federated communication across machines and security guardrails. So agents do not just run, they collaborate. Let us go through how it actually works, piece by piece, because the mechanism is the interesting part.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the Ruflo Agent Meta-Harness Actually Is
&lt;/h2&gt;

&lt;p&gt;The word meta-harness sounds like marketing. It is not, it is a precise description. Claude Code is already a harness: it gives a model tools and a loop. Ruflo sits one level up and orchestrates the harness itself. One &lt;code&gt;npx ruflo init&lt;/code&gt; gives Claude Code, in the project's own words, a nervous system: agents self-organize into swarms, learn from every task, remember across sessions and, with federation, talk to agents on other machines without leaking data.&lt;/p&gt;

&lt;p&gt;A few facts about the project itself, because they matter for trust. Ruflo is the renamed Claude Flow, built by ruvnet. The name comes from rUv (ruv.io): the "Ru" is the rUv, the "flo" is, as the README puts it, working until 3am. Underneath it runs on Cognitum.One agentic architecture with a Rust-based engine, embeddings, memory and a plugin system. The repository is TypeScript, licensed MIT, and the license file credits RuvNet.&lt;/p&gt;

&lt;h3&gt;
  
  
  A harness, not a competitor
&lt;/h3&gt;

&lt;p&gt;Pay attention to this part, because people get it wrong constantly. Ruflo is not an alternative to Claude Code or Codex. It depends on them. It is the layer that coordinates what those tools do: routing tasks, spawning specialized agents, keeping memory, enforcing security. You keep writing code in the tool you already use. Ruflo handles the coordination in the background. Comparing Ruflo to Claude Code is like comparing a dispatch system to a truck. The relationship is the point.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why One Agent Is Not Enough Anymore
&lt;/h2&gt;

&lt;p&gt;The thing is, single-agent setups break in predictable ways. State and memory get fragmented across agents, sometimes lost entirely, and infrastructure costs creep up when you run several agents at once, a pattern &lt;a href="https://www.scrapingbee.com/blog/ruflo-ai-agent-orchestration/" rel="noopener noreferrer"&gt;ScrapingBee's write-up&lt;/a&gt; describes well. The time you saved not writing code by hand starts going into wiring agents together instead.&lt;/p&gt;

&lt;p&gt;The README carries a blunt comparison table of Claude Code with and without Ruflo. Read it as the project's own claims, because that is what it is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Collaboration:&lt;/strong&gt; isolated agents with no shared context, versus swarms with shared memory and consensus&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Coordination:&lt;/strong&gt; manual orchestration, versus a queen-led hierarchy using Raft, Byzantine and Gossip consensus&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Memory:&lt;/strong&gt; session-only, versus HNSW vector memory with sub-millisecond retrieval&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Learning:&lt;/strong&gt; static behavior, versus SONA self-learning with pattern matching&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Routing:&lt;/strong&gt; you decide, versus intelligent routing the project states at 89% accuracy&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Background work:&lt;/strong&gt; none, versus 12 auto-triggered workers (audit, optimize, testgaps and others)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Providers:&lt;/strong&gt; Anthropic only, versus 5 providers with failover&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Security:&lt;/strong&gt; standard, versus CVE-hardened with AIDefence&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So the pitch is not "a better model". The pitch is: the model you already pay for, plus memory that persists, plus agents that specialize, plus a coordination layer that learns. Different claim entirely.&lt;/p&gt;

&lt;h2&gt;
  
  
  How Ruflo Works Under the Hood
&lt;/h2&gt;

&lt;p&gt;The architecture is a stack, and each layer has one job. The README lays it out like this:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Orchestration layer:&lt;/strong&gt; the MCP server, a router and 27 hooks. This is where your instructions enter the system.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Swarm coordination:&lt;/strong&gt; a queen agent, a topology and a consensus mechanism. Hierarchical, mesh and adaptive topologies are all supported.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The agents themselves:&lt;/strong&gt; 100+ specialized roles, coder, tester, reviewer, architect, security and so on.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Memory and learning:&lt;/strong&gt; AgentDB, HNSW indexing, SONA and ReasoningBank.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;LLM providers:&lt;/strong&gt; Claude, GPT, Gemini, Cohere and Ollama, with smart routing between them.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;On top of that sits a learning loop. The router sends work to the swarm, the agents execute, results flow into memory, and the system feeds what it learned back into future routing decisions. The README's own diagram is: User to Ruflo (CLI/MCP) to Router to Swarm to Agents to Memory to LLM Providers, with a learning loop closing the circuit back to the router.&lt;/p&gt;

&lt;h3&gt;
  
  
  You do not need to learn the machine to use it
&lt;/h3&gt;

&lt;p&gt;Here is the part that surprised me, in a good way. After &lt;code&gt;init&lt;/code&gt;, you just use Claude Code normally. The hooks system routes tasks, learns from successful patterns and coordinates agents in the background. The project is explicit that you do not need to learn its 314 MCP tools or 26 CLI commands to get value. Obviously you can go deeper, the tools are all there, but the default path is: install, then work as usual, and the swarm wakes up when the task calls for it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Self-learning: SONA and ReasoningBank
&lt;/h3&gt;

&lt;p&gt;Two names worth knowing. SONA is the neural pattern layer: it matches incoming tasks against patterns that worked before. ReasoningBank stores reasoning trajectories, so a strategy that succeeded on one task can be retrieved and reused by a future agent. Add trajectory learning on top and you get the practical effect: a Ruflo setup that has been on your project for a month behaves differently from a fresh one, because it carries your project's history in its memory. A plain Claude Code session starts every morning as sharp as it was on day one, which is to say, exactly as sharp and no sharper.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to Install Ruflo (CLI, Plugins and the Windows Case)
&lt;/h2&gt;

&lt;p&gt;There are two install paths and they are not the same product. Pick wrong and you will think half the features are broken. They are not broken, you installed the lite version. So read this first.&lt;/p&gt;

&lt;h3&gt;
  
  
  Path A: Claude Code plugins (lite)
&lt;/h3&gt;

&lt;p&gt;The plugin path gives you slash commands, a few skills and agent definitions per plugin. Zero files land in your workspace. No hooks are installed, and only &lt;code&gt;ruflo-core&lt;/code&gt; registers its own MCP server. It is for trying a single plugin without committing:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Add the marketplace&lt;/span&gt;
/plugin marketplace add ruvnet/ruflo

&lt;span class="c"&gt;# Install core + any plugins you need&lt;/span&gt;
/plugin &lt;span class="nb"&gt;install &lt;/span&gt;ruflo-core@ruflo
/plugin &lt;span class="nb"&gt;install &lt;/span&gt;ruflo-swarm@ruflo
/plugin &lt;span class="nb"&gt;install &lt;/span&gt;ruflo-rag-memory@ruflo
/plugin &lt;span class="nb"&gt;install &lt;/span&gt;ruflo-neural-trader@ruflo
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;ℹ️ Note:&lt;/strong&gt; Tools from the plugin install of &lt;code&gt;ruflo-core&lt;/code&gt; are callable under names like &lt;code&gt;mcp__plugin_ruflo-core_ruflo__memory_store&lt;/code&gt;, not the bare &lt;code&gt;memory_store&lt;/code&gt; or &lt;code&gt;swarm_init&lt;/code&gt; names the CLI install uses. If a guide references the bare names, it is describing the CLI path.&lt;/p&gt;

&lt;h3&gt;
  
  
  Path B: the full CLI install
&lt;/h3&gt;

&lt;p&gt;The CLI path is the full Ruflo loop: 98 agents, 60+ commands, 30 skills, the MCP server, hooks and the daemon. It writes &lt;code&gt;.claude/&lt;/code&gt;, &lt;code&gt;.claude-flow/&lt;/code&gt;, &lt;code&gt;CLAUDE.md&lt;/code&gt;, helpers and settings into your workspace. This is the path the documentation assumes. The steps:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Open a terminal in your project folder.&lt;/li&gt;
&lt;li&gt;Run the interactive wizard, which works identically on every platform including native Windows PowerShell and cmd: npx ruflo@latest init wizard&lt;/li&gt;
&lt;li&gt;Or, if you want the non-interactive version: npx ruflo@latest init. For a global install: npm install -g &lt;a href="mailto:ruflo@latest"&gt;ruflo@latest&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;On macOS, Linux, WSL or Git-Bash there is also a one-line installer: curl -fsSL &lt;a href="https://cdn.jsdelivr.net/gh/ruvnet/ruflo@main/scripts/install.sh" rel="noopener noreferrer"&gt;https://cdn.jsdelivr.net/gh/ruvnet/ruflo@main/scripts/install.sh&lt;/a&gt; | bash&lt;/li&gt;
&lt;li&gt;Register Ruflo as an MCP server in Claude Code: claude mcp add claude-flow -- npx ruflo@latest mcp start&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;⚠️ Warning:&lt;/strong&gt; The &lt;code&gt;curl ... | bash&lt;/code&gt; form needs a POSIX shell (Git-Bash, WSL, MSYS). On native Windows it fails with &lt;code&gt;'bash' is not recognized&lt;/code&gt;. Use the wizard line from step 2 instead; both end up running the same init flow.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Ruflo Plugin Marketplace for Testing, Security and Architecture
&lt;/h2&gt;

&lt;p&gt;The plugin system is where Ruflo stops being one product and becomes a platform. The README's own index lists 35 plugins, grouped by job. A quick tour of the ones that matter most:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Orchestration:&lt;/strong&gt; ruflo-core (server, health checks, plugin discovery), ruflo-swarm (coordinate agents as a team), ruflo-autopilot (agents running autonomously in a loop), ruflo-loop-workers (background tasks on a timer), ruflo-workflows (reusable multi-step templates)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Memory and knowledge:&lt;/strong&gt; ruflo-agentdb (fast vector database for agent memory), ruflo-rag-memory (hybrid search, graph hops, diversity ranking), ruflo-rvf (save and restore memory across sessions), ruflo-ruvector (GPU-accelerated search and Graph RAG with 103 tools), ruflo-knowledge-graph&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Intelligence:&lt;/strong&gt; ruflo-intelligence (agents learn from past successes), ruflo-ruvllm (run local LLMs like Ollama with smart routing), ruflo-goals (break big goals into plans), ruflo-daa&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Code quality:&lt;/strong&gt; ruflo-testgen (find missing tests and generate them), ruflo-browser (browser testing with Playwright), ruflo-jujutsu (analyze git diffs, score risk, suggest reviewers), ruflo-docs&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Security:&lt;/strong&gt; ruflo-security-audit (scan for vulnerabilities and CVEs), ruflo-aidefence (block prompt injection, detect PII, safety scanning)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Architecture and methodology:&lt;/strong&gt; ruflo-adr (living architecture decision records), ruflo-ddd (scaffold domain-driven design: contexts, aggregates, events), ruflo-sparc (a guided 5-phase development methodology with quality gates), ruflo-metaharness, ruflo-arena (pit agent strategies against each other in tournaments)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;DevOps:&lt;/strong&gt; ruflo-migrations, ruflo-observability (structured logs, traces, metrics), ruflo-cost-tracker (token usage, budgets, cost alerts)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Domain-specific:&lt;/strong&gt; ruflo-neural-trader (AI trading with 4 agents, backtesting, 112+ tools), ruflo-market-data, ruflo-iot-cognitum&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One honesty note on numbers. The plugin index in the README lists 35 plugins, while the capability table in the same README describes the marketplace as 33 native Claude Code plugins plus 21 npm plugins. Both figures come from the project itself, just from different groupings, so do not try to reconcile them into one magic number. The point stands either way: there is a lot, and it is modular.&lt;/p&gt;

&lt;h2&gt;
  
  
  Ruflo AgentDB: HNSW Vector Memory Performance, Measured Honestly
&lt;/h2&gt;

&lt;p&gt;Memory is where Ruflo either proves itself or does not, so let us look at actual numbers. AgentDB is the built-in vector store, indexed with HNSW (Hierarchical Navigable Small World, a graph index for approximate nearest-neighbor search). This is what gives agents persistent RAG memory across sessions: they retain context, learn from previous runs and share knowledge through semantic search.&lt;/p&gt;

&lt;p&gt;The repository's own audit reports measured figures, and they are specific: about 1.9x faster than brute force at N=20k, and about 3.2x to 4.7x faster at N=5k, with recall@10 around 0.99. And here is the part I respect: the same audit admits the approximate index ties or loses at small N, and only wins above the crossover point. The project publishes both the audit document and the benchmark script (&lt;code&gt;scripts/benchmark-intelligence.mjs&lt;/code&gt;) so you can reproduce the numbers yourself.&lt;/p&gt;

&lt;h3&gt;
  
  
  About those bigger claims you will see elsewhere
&lt;/h3&gt;

&lt;p&gt;Third-party coverage quotes far more aggressive figures. One comparison, &lt;a href="https://aisuccesslabjuliangoldie.com/blog/claude-ruflo/" rel="noopener noreferrer"&gt;AI Success Lab&lt;/a&gt;, states HNSW search "up to 12,500 times faster than standard vector lookups". The project's own measured audit says 1.9x to 4.7x against brute force with near-perfect recall. These are not the same claim and they are not measured the same way, so do not average them in your head: the conservative, reproducible number is the one in the repo, and the five-digit one is a blog's framing. When a vendor and a fan disagree about the vendor's product, believe the vendor's benchmark script, because you can run it.&lt;/p&gt;

&lt;p&gt;On the cost side, &lt;a href="https://brightdata.com/blog/ai/ruflo-with-bright-data" rel="noopener noreferrer"&gt;Bright Data's write-up&lt;/a&gt; reports that Ruflo's multi-tier routing (WASM plus LLMs) cuts API costs by up to roughly 75%, and that a default local setup exposes 118 Ruflo skills inside Claude Code. Treat those as that publication's reported figures, not as guarantees for your workload. Your mileage depends on which tasks route to WASM and which need a frontier model.&lt;/p&gt;

&lt;h2&gt;
  
  
  Using Ruflo Federation for Zero-Trust Multi-Agent Collaboration
&lt;/h2&gt;

&lt;p&gt;Now the part that sounds like science fiction and is actually plumbing. Federation lets agents on different machines, different organizations, different cloud regions discover each other, prove who they are and exchange work. The README calls it Slack for agents, and the analogy holds: shared workspaces across trust boundaries, except some channels are trusted and some are not, and the system handles the difference automatically.&lt;/p&gt;

&lt;p&gt;The pipeline, step by step:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Before anything leaves your node, a PII pipeline scans the outbound message. It is a 14-type detection system, and per trust level it can BLOCK, REDACT, HASH or PASS each finding. Emails, keys, personal data: stripped before transit.&lt;/li&gt;
&lt;li&gt;The message is signed. Identity is proven with mTLS plus ed25519 challenge-response. No API keys, no shared secrets.&lt;/li&gt;
&lt;li&gt;The message travels over an encrypted channel. Nobody reads it in transit.&lt;/li&gt;
&lt;li&gt;On the receiving side, identity is checked (forgeries rejected) and prompt injection attempts are blocked.&lt;/li&gt;
&lt;li&gt;Both sides write an audit trail. Every federation event produces a structured, searchable record, with HIPAA, SOC2 and GDPR audit trails available as compliance modes.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Trust is behavioral, not binary. The scoring formula is published: 0.4 × success + 0.2 × uptime + 0.2 × threat + 0.2 × integrity. New agents start untrusted and see discovery info only, not your memory. Upgrades require history. Downgrades are instant, no human in the loop. See the design? Misbehave once and you are demoted on the spot; earn your way up slowly. That is what zero-trust means here in practice.&lt;/p&gt;

&lt;p&gt;The lifecycle is exposed through 9 MCP tools and 10 CLI commands. The README's example (note it invokes the &lt;code&gt;claude-flow&lt;/code&gt; package name, the project's former identity, which the commands still use):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Team A: initialize federation and generate keypair&lt;/span&gt;
npx claude-flow@latest federation init

&lt;span class="c"&gt;# Team A: join Team B's federation endpoint&lt;/span&gt;
npx claude-flow@latest federation &lt;span class="nb"&gt;join &lt;/span&gt;wss://team-b.example.com:8443

&lt;span class="c"&gt;# Team A: send a task — PII is stripped automatically before it leaves&lt;/span&gt;
npx claude-flow@latest federation send &lt;span class="nt"&gt;--to&lt;/span&gt; team-b &lt;span class="nt"&gt;--type&lt;/span&gt; task-request &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--message&lt;/span&gt; &lt;span class="s2"&gt;"Analyze transaction patterns for account anomalies"&lt;/span&gt;

&lt;span class="c"&gt;# Team A: check peer trust levels and session health&lt;/span&gt;
npx claude-flow@latest federation status
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;💡 Tip:&lt;/strong&gt; The worked example in the docs is two teams sharing fraud signals without sharing customer data. The PII stripping is what makes that sentence possible instead of a compliance incident. There is also an opt-in WireGuard mesh layer documented under &lt;code&gt;docs/federation/&lt;/code&gt; if you need packet-layer reachability tied to federation trust.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Web UI and the GOAP A* Planner at goal.ruv.io
&lt;/h2&gt;

&lt;p&gt;Two hosted front-ends ship with the project, and both are self-hostable. They are worth knowing about even if you live in the terminal, because they show what the harness can do when you drive it from a chat window or a goal statement.&lt;/p&gt;

&lt;h3&gt;
  
  
  flo.ruv.io: multi-model chat with MCP tool calling
&lt;/h3&gt;

&lt;p&gt;The web UI is a multi-model AI chat with built-in Model Context Protocol (MCP) tool calling. Six curated frontier models come out of the box: Qwen 3.6 Max (the default), Claude Sonnet 4.6, Claude Haiku 4.5, Gemini 2.5 Pro, Gemini 2.5 Flash and OpenAI, all via OpenRouter. You can add any OpenAI-compatible endpoint: vLLM, Ollama, LM Studio, Together, Groq, self-hosted. There is native support for ruvLLM, the project's self-improving local model layer (it lives in &lt;code&gt;ruvnet/RuVector/examples/ruvLLM&lt;/code&gt;), which routes to MicroLoRA adapters and learns from your trajectories via SONA, fully offline if you want.&lt;/p&gt;

&lt;p&gt;The tool story is the interesting bit. About 210 tools are ready to call, in 5 server groups (Core, Intelligence, Agents, Memory, DevTools), plus an 18-tool gallery that runs entirely in your browser via WASM and works offline. One model response can fire 4 to 6 or more tools in parallel, shown as cards so you can see exactly what ran. You can also paste in your own MCP servers (HTTP, SSE or stdio) and they join the same flow. Memory is backed by AgentDB plus HNSW, so "remember my favorite color is indigo" actually survives for weeks. Self-hosting is a first-class option: the source lives in &lt;code&gt;ruflo/src/ruvocal/&lt;/code&gt; with a multi-stage Dockerfile (&lt;code&gt;INCLUDE_DB=true&lt;/code&gt; builds in MongoDB) and a &lt;code&gt;cloudbuild.yaml&lt;/code&gt; for Google Cloud Run.&lt;/p&gt;

&lt;h3&gt;
  
  
  goal.ruv.io: plain English in, executable plan out
&lt;/h3&gt;

&lt;p&gt;The second front-end is a GOAP (Goal-Oriented Action Planning) planner. You type something like "ship the auth refactor with tests and a PR", and the system extracts success criteria, constraints and implicit preconditions, then runs an A* search through the state space to find the shortest viable path of actions. This is classic game-AI planning ported to software work.&lt;/p&gt;

&lt;p&gt;Three details make it more than a demo. First, adaptive replanning: when an action fails or new information arrives, the planner re-runs A* from the current state instead of starting over. Second, the live dashboard at &lt;code&gt;/agents&lt;/code&gt; shows every spawned agent with role, current step, memory namespace, token budget and status, and you can kill runaway workers or reassign tasks. Third, every action node maps to a real tool call, Ruflo's MCP tools, your custom servers or shell, scheduled in parallel where the dependency graph allows. Plans, trajectories and outcomes flow back into AgentDB, so future plans retrieve past solutions. The planner gets smarter with every run. The source is in &lt;code&gt;v3/goal_ui/&lt;/code&gt; (Vite plus Supabase), and you can run your own with &lt;code&gt;cd v3/goal_ui &amp;amp;&amp;amp; npm install &amp;amp;&amp;amp; npm run dev&lt;/code&gt; from the &lt;code&gt;goal&lt;/code&gt; branch.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Ruflo Claims About Speed, and How to Check It
&lt;/h2&gt;

&lt;p&gt;The project publishes a benchmark matrix comparing itself to LangGraph, AutoGen and CrewAI on darwin-arm64 and linux-x64, and it claims wins on cold start, single turn and RSS by margins from 1.3x up to 1953x. Now, pay attention to how you read that. These are the project's own numbers, run on the project's own workload spec. To its credit, the repo links the workload specification, the progress log and the raw matrix JSON for both platforms, so the claim is auditable rather than vibes. But it is still the vendor's benchmark of the vendor's product. Report it as that, run it yourself if the numbers matter to you, and do not treat the 1953x end of the range as the typical case, because ranges like that never are.&lt;/p&gt;

&lt;h3&gt;
  
  
  MetaHarness and ruflo verify: auditing your own setup
&lt;/h3&gt;

&lt;p&gt;Two quieter tools deserve a mention. MetaHarness grades your agent setup from 1 to 100, scans tool configurations for security issues, snapshots the project so you can catch regressions between runs, and finds templates matching your repo. The &lt;code&gt;ruflo eject&lt;/code&gt; command turns a Ruflo project into a standalone agent toolkit with its own name. Separately, &lt;code&gt;ruflo verify&lt;/code&gt; lets you cryptographically check that your installed bytes match the signed witness manifest. For a tool whose whole job is running autonomous agents on your machine, that verify step is not a luxury.&lt;/p&gt;

&lt;p&gt;The documentation is organized by audience, which is a small mercy: a Status doc (what currently works), a User Guide (every command and flag), the MetaHarness guide, a verification doc and a team gateway checklist for safer multi-person workflows.&lt;/p&gt;

&lt;h2&gt;
  
  
  When Ruflo Is the Wrong Tool
&lt;/h2&gt;

&lt;p&gt;Every tool article skips this part, and every reader pays for it later. So, the honest list:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Simple, single-agent tasks.&lt;/strong&gt; Swarm coordination has overhead. For a one-step task, plain Claude Code is faster because there is nothing to coordinate. The &lt;a href="https://aisuccesslabjuliangoldie.com/blog/claude-ruflo/" rel="noopener noreferrer"&gt;AI Success Lab comparison&lt;/a&gt; puts the crossover at roughly when a task has more than three logical steps that do not depend on each other. Below that, you are paying coordination tax for nothing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tiny vector workloads.&lt;/strong&gt; The project's own audit says the HNSW index ties or loses to brute force at small N. If your agent memory is a few hundred entries, the fancy index buys you nothing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The plugin path, if you expect the full loop.&lt;/strong&gt; No hooks, no daemon, and most plugins do not register an MCP server. If you installed via /plugin install and the documented behavior is missing, that is why.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;If you do not have the underlying tools.&lt;/strong&gt; Ruflo is an execution layer around Claude Code and Codex. No Claude Code or Codex, nothing to harness. It also does not replace your model provider access; it routes to providers, it does not include them.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Do you see the pattern? Ruflo's costs are all coordination costs, and coordination only pays when there is something worth coordinating. One agent, one step, one session: skip it. A swarm writing, testing and reviewing across a large codebase for months: that is the case it was built for.&lt;/p&gt;

&lt;h2&gt;
  
  
  People Also Ask
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Is Ruflo free, and what does it actually cost to run?
&lt;/h3&gt;

&lt;p&gt;Ruflo itself is free and open source under the MIT license, and it runs locally. According to &lt;a href="https://aisuccesslabjuliangoldie.com/blog/claude-ruflo/" rel="noopener noreferrer"&gt;AI Success Lab&lt;/a&gt;, it adds nothing to your monthly bill on top of Claude Code, while your Claude Code subscription tier remains the non-negotiable foundation cost. Your real variable cost is model API usage, which is exactly what the multi-tier routing and the &lt;code&gt;ruflo-cost-tracker&lt;/code&gt; plugin (budgets and cost alerts) exist to manage.&lt;/p&gt;

&lt;h3&gt;
  
  
  Does Ruflo work with models other than Claude?
&lt;/h3&gt;

&lt;p&gt;Yes. At the core level, the router is not tied to Anthropic: it can send work to GPT, Gemini, Cohere or a local Ollama instance, and fail over between providers when one is unavailable. The chat front-end widens that further through OpenRouter, where the curated list defaults to Qwen 3.6 Max and also includes recent Claude and Gemini releases, and any OpenAI-compatible endpoint, vLLM, LM Studio, Together, Groq or self-hosted, can be plugged in. If you want nothing leaving your machine at all, the ruvLLM layer runs fully offline.&lt;/p&gt;

&lt;h3&gt;
  
  
  How many plugins does Ruflo have?
&lt;/h3&gt;

&lt;p&gt;The README's plugin index lists 35 plugins across categories like orchestration, memory, intelligence, code quality, security, architecture, DevOps and domain-specific tools. Elsewhere in the same README, the marketplace is described as 33 native Claude Code plugins plus 21 npm plugins. Both figures come from the project itself and reflect different groupings, so treat them as two views of the same ecosystem rather than one definitive count.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can Ruflo agents on different machines share data safely?
&lt;/h3&gt;

&lt;p&gt;That is exactly what federation is for. Identity is proven with mTLS plus ed25519 challenge-response (no API keys or shared secrets), a 14-type PII pipeline scans every outbound message with BLOCK, REDACT, HASH or PASS policies per trust level, and behavioral trust scoring upgrades reliable peers slowly while downgrading misbehaving ones instantly. Every event lands in a structured audit trail, with HIPAA, SOC2 and GDPR compliance modes.&lt;/p&gt;

&lt;h2&gt;
  
  
  So, Do You Actually Need a Harness?
&lt;/h2&gt;

&lt;p&gt;Back to the opening formula: Agent = Model + Harness. The industry spent two years obsessing over the model part and treating the harness as an afterthought, then wondered why agents forgot everything between sessions and could not coordinate two tasks without a human playing dispatcher. Ruflo's bet is that the harness is where the leverage was hiding all along: memory that persists, agents that specialize, coordination that learns, security that assumes nobody is trusted by default.&lt;/p&gt;

&lt;p&gt;The project is young, the README is honest enough to publish its own audit caveats, and the install takes one npx command, so the cost of finding out whether it fits your workflow is an afternoon, not a procurement process. Try the plugin path if you are curious, the full CLI path if you are serious, and the hosted demos if you do not want to install anything at all. Anyway. Your move.&lt;/p&gt;

&lt;h2&gt;
  
  
  🎯 What You Now Know About Ruflo
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Ruflo (formerly Claude Flow) is an agent meta-harness: an execution layer around Claude Code and Codex, not a competitor to them. Agent = Model + Harness.&lt;/li&gt;
&lt;li&gt;One npx ruflo@latest init wizard installs the full loop on any platform, including native Windows; the plugin path is a lite version with slash commands only and no hooks.&lt;/li&gt;
&lt;li&gt;The architecture stacks an orchestration layer (MCP server, router, 27 hooks) over swarm coordination, 100+ specialized agents, AgentDB/HNSW memory with SONA and ReasoningBank learning, and five LLM providers.&lt;/li&gt;
&lt;li&gt;The repo's own audit measures AgentDB at roughly 1.9x to 4.7x faster than brute force with recall@10 near 0.99, and admits the index loses at small N; bigger third-party speed claims exist but are not the project's measured figures.&lt;/li&gt;
&lt;li&gt;Federation gives zero-trust collaboration across machines: mTLS plus ed25519 identity, a 14-type PII pipeline, behavioral trust scoring and compliance-grade audit trails.&lt;/li&gt;
&lt;li&gt;The plugin ecosystem (35 indexed plugins) covers testing, security audits, DDD and SPARC methodologies, observability and cost tracking; flo.ruv.io and goal.ruv.io add a multi-model chat UI and a GOAP A* goal planner, both self-hostable.&lt;/li&gt;
&lt;li&gt;Ruflo is the wrong tool for simple single-step tasks: coordination overhead only pays off when work has multiple independent steps worth parallelizing.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Additional Resources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/ruvnet/ruflo" rel="noopener noreferrer"&gt;ruvnet/ruflo on GitHub&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://brightdata.com/blog/ai/ruflo-with-bright-data" rel="noopener noreferrer"&gt;Ruflo + Bright Data for Enterprise Agentic Coding&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://aisuccesslabjuliangoldie.com/blog/claude-ruflo/" rel="noopener noreferrer"&gt;Claude Ruflo Vs Claude Code (Honest 2026 Comparison)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.scrapingbee.com/blog/ruflo-ai-agent-orchestration/" rel="noopener noreferrer"&gt;Ruflo: Multi-Agent AI Orchestration for Claude Code &amp;amp; Codex&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/ruvnet/ruflo/discussions/851" rel="noopener noreferrer"&gt;Multiple agents for the same prompt · ruvnet/ruflo · Discussion #851 · GitHub&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.linkedin.com/posts/faizan170_aiorchestration-multiagentsystems-ruflo-activity-7433913131678142464-Z17_" rel="noopener noreferrer"&gt;Ruflo's Multi-Agent Claude Swarm for Autonomous Dev | Faizan Amin posted on the topic | LinkedIn&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;This article includes content created with AI.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>performance</category>
      <category>database</category>
      <category>security</category>
      <category>programming</category>
    </item>
    <item>
      <title>Free Claude Code and the Identity Question</title>
      <dc:creator>DiFlowrin</dc:creator>
      <pubDate>Thu, 03 Sep 2026 19:34:15 +0000</pubDate>
      <link>https://dev.to/diflowrin/free-claude-code-and-the-identity-question-8f7</link>
      <guid>https://dev.to/diflowrin/free-claude-code-and-the-identity-question-8f7</guid>
      <description>&lt;h2&gt;
  
  
  What Free Claude Code Actually Proves About Coding Agents
&lt;/h2&gt;

&lt;p&gt;The story going around is that Free Claude Code works because an extension accepts a fake identity. It does not. It works because Claude Code was built to read two environment variables, &lt;strong&gt;ANTHROPIC_BASE_URL&lt;/strong&gt; and &lt;strong&gt;ANTHROPIC_AUTH_TOKEN&lt;/strong&gt;, and to send its requests wherever they point. That is a documented feature for corporate gateways, and Free Claude Code, an MIT-licensed Python project that runs a local AI proxy on your own machine, uses exactly that door. The real question is not whether the endpoint can be swapped. It is who is standing at the other end of it, and what your agent hands over when it gets there.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the Free Claude Code Debate Started at All
&lt;/h2&gt;

&lt;p&gt;Coding agents got popular fast. Everyone has one open in a terminal or an editor, everyone burns through tokens, and everyone eventually looks at the bill. Then a project appears that says: run Claude Code, Codex, Pi, OpenCode and more, for free, from your terminal, app, IDE or phone. Of course people notice.&lt;/p&gt;

&lt;h3&gt;
  
  
  📹 Video: The BEST Free Claude Code Setup (Ollama + Proxy Tutorial)
&lt;/h3&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/bMdoyQAcSEU" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Video credit: SelfTaughtDev&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;And then the second reaction comes. If it works without the vendor's own API, something must be broken. Right?&lt;/p&gt;

&lt;p&gt;Wrong. We have to understand the plumbing. A coding agent is a client. It speaks a wire format, it sends an Authorization header, it reads a base URL from configuration. The vendor's own documentation describes an LLM gateway as a proxy your organisation runs between the agent and the model provider, and it tells you which variables to set for it. Enterprises asked for that. It shipped. Free Claude Code is a single developer using the same switch that a bank's platform team uses, on localhost instead of behind a corporate firewall.&lt;/p&gt;

&lt;p&gt;So the interesting part is not the trick. There is no trick. The interesting part is what a configurable endpoint means for the person who configures it, and for the company whose laptops they configure it on.&lt;/p&gt;

&lt;h2&gt;
  
  
  Free Claude Code as a Local AI Proxy: How It Works
&lt;/h2&gt;

&lt;p&gt;The repository is &lt;a href="https://github.com/Alishahryar1/free-claude-code" rel="noopener noreferrer"&gt;Alishahryar1/free-claude-code&lt;/a&gt;. Python, MIT licence, and its README says the thing plainly at the top: independent open-source project, not affiliated with or endorsed by Anthropic, Claude and Claude Code are trademarks of Anthropic. Credit where it is due, that disclaimer is doing more honest work than most of the commentary about it.&lt;/p&gt;

&lt;p&gt;The mechanism, as the project describes it: you install with a shell or PowerShell one-liner, you start &lt;code&gt;fcc-server&lt;/code&gt; (or the desktop launcher on Windows and macOS), and an Admin UI opens. You paste a provider key there, pick a model from a searchable dropdown, click &lt;strong&gt;Apply&lt;/strong&gt;. Then you run your agent through a wrapper command: &lt;code&gt;fcc-claude&lt;/code&gt;, &lt;code&gt;fcc-codex&lt;/code&gt;, &lt;code&gt;fcc-pi&lt;/code&gt;, &lt;code&gt;fcc-opencode&lt;/code&gt;, &lt;code&gt;fcc-cline&lt;/code&gt;, &lt;code&gt;fcc-hermes&lt;/code&gt;, &lt;code&gt;fcc-dsh&lt;/code&gt;, &lt;code&gt;fcc-grok&lt;/code&gt;, &lt;code&gt;fcc-muse&lt;/code&gt;, &lt;code&gt;fcc-aider&lt;/code&gt;. Ten agents, one model catalog.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Four Settings That Redirect the Agent
&lt;/h3&gt;

&lt;p&gt;For the VS Code integration, the README asks you to add environment variables to your user settings JSON. This is the whole "identity" question, in four lines:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Setting&lt;/th&gt;
&lt;th&gt;Value from the project's setup&lt;/th&gt;
&lt;th&gt;What it does&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;ANTHROPIC_BASE_URL&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;&lt;a href="http://localhost:8082" rel="noopener noreferrer"&gt;http://localhost:8082&lt;/a&gt;&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Sends agent requests to the local proxy instead of the vendor endpoint&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;ANTHROPIC_AUTH_TOKEN&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;freecc&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Value used for the Authorization header; matched to the Admin UI&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;claudeCode.disableLoginPrompt&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;true&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Stops the extension asking you to sign in&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;CLAUDE_CODE_ENABLE_GATEWAY_MODEL_DISCOVERY&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;1&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Lets the client discover models offered by the gateway&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Notice what the token is. The word &lt;code&gt;freecc&lt;/code&gt;. Not a signed credential, not a key with a checked prefix. And that is not a bug someone found: the vendor's environment-variable reference describes &lt;code&gt;ANTHROPIC_AUTH_TOKEN&lt;/code&gt; as a custom value for the Authorization header, prefixed with &lt;code&gt;Bearer&lt;/code&gt;, precisely because corporate gateways use their own token formats. Any string. By design.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the Trust Actually Sits in a Multi-Provider Model Router
&lt;/h2&gt;

&lt;p&gt;So if there is no exploit, is there no risk? There is plenty. It just lives somewhere else than the headlines put it.&lt;/p&gt;

&lt;p&gt;Think about what a coding agent sends. Your file contents. Your error strings. Your prompts, which are often a description of exactly where your system is weak. Point the base URL at a proxy and the proxy sees all of it, in cleartext if the URL is &lt;code&gt;http://&lt;/code&gt;, which on &lt;code&gt;localhost&lt;/code&gt; is fine, and on someone else's host is not fine at all. That is the actual lesson of a project like this, and it applies to every OpenAI-compatible gateway, not to one repository. Cline's own provider documentation walks users through setting a base URL and a key for any compatible endpoint, and the guidance is the same: it will not be the official vendor URL, so you had better know whose it is (&lt;a href="https://docs.cline.bot/provider-config/openai-compatible" rel="noopener noreferrer"&gt;Cline docs&lt;/a&gt;).&lt;/p&gt;

&lt;h3&gt;
  
  
  Bearer Token Authentication and Proxy Authentication in Practice
&lt;/h3&gt;

&lt;p&gt;To its credit, the project takes the local surface seriously. The README says you can protect the local proxy with a bearer token by enabling &lt;strong&gt;Proxy Authentication&lt;/strong&gt; in the Admin UI, and the Codex integration does not even ask you to paste a secret: it runs &lt;code&gt;fcc-codex --print-proxy-auth-token&lt;/code&gt; as an auth command so the client reads the current token itself. That is a better pattern than a hardcoded string in a settings file, and it is worth copying.&lt;/p&gt;

&lt;p&gt;The gap is not in the token. The gap is that a base URL is a trust decision, and nothing in a text field tells you how much trust you just spent.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which AI Providers the Project Routes To, and Why That Matters
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;50 ToS-friendly providers. 1.3B+ free tokens every month.&lt;/strong&gt; The project also states that it follows provider terms and removes integrations if they stop being allowed.Free Claude Code, project README&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The provider catalog in the README is long and specific: NVIDIA NIM, OpenRouter, Groq, ClinePass, xAI, QwenCloud in two separate plan flavours, Together AI, DeepInfra, SiliconFlow, Nebius, Chutes, Featherless, ZenMux, W&amp;amp;B Inference, Azure OpenAI, Google AI Studio, Vertex AI, DeepSeek, Mistral and Codestral, OpenCode Zen and Go, Vercel AI Gateway, Amazon Bedrock, Hugging Face, Cohere, GitHub Models, Kimi, MiniMax, Cerebras, SambaNova, Kilo.ai, Fireworks, Novita, Cloudflare Workers AI, Z.ai, TokenRouter, NaraRoute, Poolside, LLM7.io, Ollama Cloud, plus local LM Studio, llama.cpp and Ollama. Each one has its own Admin UI setting and its own model slug format, like &lt;code&gt;nvidia_nim/nvidia/nemotron-3-super-120b-a12b&lt;/code&gt; or &lt;code&gt;groq/llama-3.3-70b-versatile&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The repository's own structure matches that claim rather than dressing it up. Under &lt;code&gt;src/free_claude_code/providers/&lt;/code&gt; the tree carries seventeen directories, including &lt;code&gt;cloudflare&lt;/code&gt;, &lt;code&gt;deepseek&lt;/code&gt;, &lt;code&gt;gemini&lt;/code&gt;, &lt;code&gt;github_models&lt;/code&gt;, &lt;code&gt;groq&lt;/code&gt;, &lt;code&gt;kilo&lt;/code&gt;, &lt;code&gt;lmstudio&lt;/code&gt;, &lt;code&gt;mistral&lt;/code&gt;, &lt;code&gt;nvidia_nim&lt;/code&gt;, &lt;code&gt;open_router&lt;/code&gt;, &lt;code&gt;vertex&lt;/code&gt;. There is also a &lt;code&gt;tests/providers/&lt;/code&gt; directory, a &lt;code&gt;tests/contracts/&lt;/code&gt; directory, an &lt;code&gt;ARCHITECTURE.md&lt;/code&gt;, a &lt;code&gt;CONTRIBUTING.md&lt;/code&gt; and a &lt;code&gt;uv.lock&lt;/code&gt;. This is a maintained codebase, not a gist.&lt;/p&gt;

&lt;p&gt;And the honest caveat is the project's own: free-tier availability and limits are controlled by each provider and may change. Also, a fallback list can cost you twice. The README warns that a failed request may reach and consume usage from more than one provider before it succeeds.&lt;/p&gt;

&lt;h2&gt;
  
  
  Model Fallback Routing, Reasoning Control and the Rest of the Machinery
&lt;/h2&gt;

&lt;p&gt;The features worth knowing about, as the project describes them:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Fallback models.&lt;/strong&gt; An ordered list under Model Config. After retries are exhausted, the proxy tries your next configured model without making you restart the turn, across every connected client.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tier routing.&lt;/strong&gt; MODEL is the fallback for every request, while MODEL_FABLE, MODEL_OPUS, MODEL_SONNET and MODEL_HAIKU each override one Claude Code tier. Route the heavy tier to a hosted model, the cheap tier to a local one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reasoning control.&lt;/strong&gt; Admin UI, Model Config, Reasoning: take the effort the client sent, turn it off, or override with Low through Max. Providers that do not support a control keep their own behaviour.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Token savings on terminal output.&lt;/strong&gt; The project claims up to &lt;strong&gt;90%&lt;/strong&gt; fewer terminal-output tokens with the optional RTK filter, plus five in-proxy optimisations that answer quota probes, command-prefix detection, titles, suggestions and filepaths without calling a provider at all.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Surfaces.&lt;/strong&gt; Native launchers, VS Code, the Codex App, JetBrains ACP, Discord, Telegram, and a local Chat Sessions view in Admin with persisted history, streaming, fallback and compaction.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Voice notes.&lt;/strong&gt; Re-run the installer with --voice-nim, --voice-local or --voice-all for NVIDIA NIM or local Whisper transcription.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Compare that with a managed product, where the router is the vendor's business logic and you mostly get a picker. Cursor, for instance, documents a router that spans a fixed set of models and hides which one answered unless a team admin turns visibility on (&lt;a href="https://cursor.com/help/models-and-usage/available-models" rel="noopener noreferrer"&gt;Cursor Docs&lt;/a&gt;). Here the routing table is yours, and so is the responsibility for it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Terms of Service, Not Certificates: Where the Argument Should Go
&lt;/h2&gt;

&lt;p&gt;People reach for the word vulnerability because it sounds decisive. It is the wrong word here, and using it wrongly costs you the argument in the room where it matters.&lt;/p&gt;

&lt;p&gt;Configurable endpoint&lt;/p&gt;

&lt;p&gt;A documented client setting that sends agent traffic to a chosen host. Intended for gateways. Not an exploit.&lt;/p&gt;

&lt;p&gt;Local proxy&lt;/p&gt;

&lt;p&gt;A service on your own machine that receives agent requests and forwards them to a provider, translating formats along the way.&lt;/p&gt;

&lt;p&gt;Provider terms&lt;/p&gt;

&lt;p&gt;The contract that actually governs whether a given key may be used from a coding agent. The project flags specific plans as personal, interactive use only, and links each provider's own rules.&lt;/p&gt;

&lt;p&gt;Read that last one again, because it is where a company gets hurt. The README notes that Kimi Code subscription keys and QwenCloud Coding Plan keys are for local, personal, interactive coding-agent use, with the keys and endpoints not interchangeable. A developer wiring a personal-use key into a team workflow is not a TLS problem. It is a compliance problem, and no certificate check would have caught it.&lt;/p&gt;

&lt;p&gt;The client side is not defenceless either. The vendor's documentation states that when the base URL points at a non-first-party host, MCP tool search is disabled by default, and that as of v2.1.196 Remote Control is disabled when the base URL points anywhere other than &lt;code&gt;api.anthropic.com&lt;/code&gt;. Features degrade on purpose once you leave the vendor's endpoint. That is a client that knows where it is.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a Security Team Should Actually Do About It
&lt;/h2&gt;

&lt;p&gt;Not panic. Not ban a repository. Treat the endpoint as the control point, because it is.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Manage the variables. If the base URL and the auth token can be set from a project file or an editor setting, then they are configuration you should be shipping and auditing, not discovering.&lt;/li&gt;
&lt;li&gt;Log the hosts. You do not need to read prompts to know that agent traffic left for an address nobody approved.&lt;/li&gt;
&lt;li&gt;Write the key policy down. Which provider plans may be used from a corporate machine, and which are personal-only. Then enforce it on keys, not on tools.&lt;/li&gt;
&lt;li&gt;Prefer token helpers to pasted strings. The Codex setup in this project reads the proxy token through a command instead of hardcoding it. Do the same internally.&lt;/li&gt;
&lt;li&gt;Separate the machines. A router that can reach forty-odd providers has no business sitting on a laptop with production credentials.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The ecosystem gap is real, but it is a standards gap, not a hole: agents authenticate to gateways with bearer strings and trust whatever answers in the right wire format, because that is what makes a gateway possible at all. Response signing, endpoint pinning by policy, and an agent that tells you loudly which host it is talking to would all help. None of them exist as a shared standard today. Until they do, the base URL field is the policy.&lt;/p&gt;

&lt;h2&gt;
  
  
  People Also Ask
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Does Free Claude Code require Anthropic API access?
&lt;/h3&gt;

&lt;p&gt;No. It routes agent requests to providers you configure yourself, from NVIDIA NIM and OpenRouter to local LM Studio, llama.cpp or Ollama. The project is independent and states it is not affiliated with or endorsed by Anthropic. What it does need is a valid key or local server for whichever provider you pick, and it asks you to respect that provider's terms, including plans it flags as personal, interactive use only.&lt;/p&gt;

&lt;h3&gt;
  
  
  How do you install it on Windows, macOS or Linux?
&lt;/h3&gt;

&lt;p&gt;On macOS and Linux the README gives a single command: &lt;code&gt;curl -fsSL "https://raw.githubusercontent.com/Alishahryar1/free-claude-code/main/scripts/install.sh" | sh&lt;/code&gt;. On Windows it is the PowerShell equivalent using &lt;code&gt;install.ps1&lt;/code&gt;. Re-running the same command updates the install, and the README points you to &lt;code&gt;scripts/install.sh&lt;/code&gt; and &lt;code&gt;scripts/install.ps1&lt;/code&gt; if you want to read them first, which you should. You then start it from the desktop launcher on Windows and macOS, or with &lt;code&gt;fcc-server&lt;/code&gt; on Linux.&lt;/p&gt;

&lt;h3&gt;
  
  
  What happens if Claude Code still asks me to log in?
&lt;/h3&gt;

&lt;p&gt;The README covers this case. Open the state file at &lt;code&gt;%USERPROFILE%\.claude.json&lt;/code&gt; on Windows or &lt;code&gt;~/.claude.json&lt;/code&gt; on macOS, Linux and WSL, and merge in &lt;code&gt;"hasCompletedOnboarding": true&lt;/code&gt; without removing the file's other fields. If the file does not exist, create it with that single property inside a complete JSON object, then restart the agent or the IDE.&lt;/p&gt;

&lt;h3&gt;
  
  
  Which version am I running, and how do I remove it?
&lt;/h3&gt;

&lt;p&gt;Run &lt;code&gt;fcc-server --version&lt;/code&gt; to check the installed version without starting the server. Uninstalling uses the matching &lt;code&gt;uninstall.sh&lt;/code&gt; or &lt;code&gt;uninstall.ps1&lt;/code&gt; script, and the README is specific about scope: it removes Free Claude Code, its desktop launcher and commands, and &lt;code&gt;~/.fcc/&lt;/code&gt;, while keeping uv, Python, shared PATH entries and the coding agents themselves. Stop every running FCC command first.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Part Everyone Skips
&lt;/h2&gt;

&lt;p&gt;Free Claude Code is a well-built local router with a long provider list, a real Admin UI, an MIT licence and a proxy token you can turn on. It is also a demonstration that the thing standing between your codebase and a stranger's server is a text field in a settings file. Both of those are true at the same time, and only one of them is the project's fault.&lt;/p&gt;

&lt;p&gt;So use it as an argument, by all means. Just make the right one. Not "the extension accepts a fake identity", which will get you corrected by anyone who has read the gateway documentation. Say instead: our agents authenticate to whatever answers in the right format, we do not currently control or log where they point, and the credentials people wire in may not be licensed for the work they are doing. That version survives a meeting.&lt;/p&gt;

&lt;h2&gt;
  
  
  📊 What Engineering Leaders Should Take From This
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;No exploit, a documented door:&lt;/strong&gt; Free Claude Code works through the same base-URL and auth-token variables enterprise gateways use.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The token was never the identity:&lt;/strong&gt; the setup uses the literal string freecc, and the vendor's own reference describes that variable as accepting any string for the Authorization header.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Breadth is real and self-consistent:&lt;/strong&gt; the project claims &lt;strong&gt;50&lt;/strong&gt; providers and &lt;strong&gt;1.3B+&lt;/strong&gt; free tokens monthly, and its providers/ folder holds &lt;strong&gt;17&lt;/strong&gt; provider directories with tests and contract tests alongside.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fallback has a cost:&lt;/strong&gt; the project warns that one failed request may consume usage from more than one provider before it succeeds.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Your real exposure is licensing and egress:&lt;/strong&gt; personal-use provider plans on corporate machines, and unlogged agent traffic to unapproved hosts. Control the endpoint and the keys, not the repository.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Additional Resources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/Alishahryar1/free-claude-code" rel="noopener noreferrer"&gt;Free Claude Code - GitHub&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.cline.bot/provider-config/openai-compatible" rel="noopener noreferrer"&gt;Cline OpenAI Compatible provider docs&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.cline.bot/api/overview" rel="noopener noreferrer"&gt;Cline API overview&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://cursor.com/help/models-and-usage/available-models" rel="noopener noreferrer"&gt;Available models | Cursor Docs&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://code.claude.com/docs/en/env-vars" rel="noopener noreferrer"&gt;Environment variables - Claude Code Docs&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://code.claude.com/docs/en/authentication" rel="noopener noreferrer"&gt;Authentication - Claude Code Docs&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;This article includes content created with AI.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>github</category>
      <category>programming</category>
      <category>git</category>
    </item>
    <item>
      <title>Clip Architect: MoneyPrinterTurbo as a Windows Desktop App</title>
      <dc:creator>DiFlowrin</dc:creator>
      <pubDate>Thu, 27 Aug 2026 09:40:19 +0000</pubDate>
      <link>https://dev.to/diflowrin/clip-architect-moneyprinterturbo-as-a-windows-desktop-app-1k0e</link>
      <guid>https://dev.to/diflowrin/clip-architect-moneyprinterturbo-as-a-windows-desktop-app-1k0e</guid>
      <description>&lt;h2&gt;
  
  
  What Clip Architect Actually Changes About Local AI Video Generation
&lt;/h2&gt;

&lt;p&gt;Here's what people get wrong about a tool like this. The hard part was never really the AI writing the script. It's the plumbing around it, the part nobody photographs for the landing page. &lt;strong&gt;Clip Architect&lt;/strong&gt; is a Windows desktop application that wraps the open-source &lt;a href="https://github.com/harry0703/MoneyPrinterTurbo" rel="noopener noreferrer"&gt;MoneyPrinterTurbo&lt;/a&gt; pipeline (the one that turns a topic into a scripted, narrated, subtitled short video) inside a Tauri 2 shell, with a React 19 interface and a Python backend running underneath as a private local service. You give it a topic, you get an MP4 sized for TikTok, Reels or Shorts, and nothing in between gets uploaded anywhere except to whichever provider you configured, with the key you supplied yourself. No account, no subscription, no cloud render queue. Once you get that one distinction, wrapper versus engine, the rest of this holds together on its own.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the Terminal Step Was the Real Barrier
&lt;/h2&gt;

&lt;p&gt;Let's look at where the friction actually sat. Upstream MoneyPrinterTurbo is a Python web app built on FastAPI with a Streamlit interface: &lt;a href="https://github.com/diflowrin/Clip-Architect" rel="noopener noreferrer"&gt;you start it from a terminal and use it in a browser&lt;/a&gt;. Fine for a developer. It stops being fine the moment the person who wants the video has never opened a terminal in their life, and most people who want a video have never opened a terminal in their life. Closing that gap is the whole reason Clip Architect exists: a Tauri shell owns the window and the process lifecycle, a React frontend replaces Streamlit, and the Python backend starts and stops with the app itself, quietly, in the background. You install it, you open it, and a command line never comes up.&lt;/p&gt;

&lt;p&gt;The chain underneath doesn't change. Give it a subject, an LLM writes the script and the search keywords, stock footage or your own files supply the picture, a text-to-speech engine speaks the narration, and FFmpeg cuts the clips to the voice track, burns in subtitles, mixes background music and writes the final MP4. Every one of those stages already existed in MoneyPrinterTurbo. What changed is who can reach them.&lt;/p&gt;

&lt;h2&gt;
  
  
  How the Pieces Talk to Each Other
&lt;/h2&gt;

&lt;h3&gt;
  
  
  The Shell, the Backend, and the Port Nobody Hardcodes
&lt;/h3&gt;

&lt;p&gt;The Rust shell reserves a free loopback port, starts the backend with that port, waits for it to accept connections, and kills it on exit. The frontend never assumes a fixed port, it gets the base URL from the shell over a status event, each time, fresh. The child process sits inside a Windows Job Object with kill-on-close, so a crash or a Task Manager kill can't leave a render running in the background, unattended, still burning through somebody's API quota. That's not a small detail. A process nobody can see, still working, still spending money, is exactly the kind of thing that makes people stop trusting desktop software.&lt;/p&gt;

&lt;p&gt;First startup is slow: around 70 seconds cold, while Windows Defender scans the roughly 290 MB bundle, and around 20 seconds warm after that. The timeout is set generously on purpose, because a false failure on someone's first launch is worse than a slow one.&lt;/p&gt;

&lt;h3&gt;
  
  
  Two Directories, Split for a Reason
&lt;/h3&gt;

&lt;p&gt;Upstream derives every file path from its own location on disk, which breaks the moment the code runs from a read-only install directory. Clip Architect resolves two separate directories instead. One holds &lt;code&gt;config.toml&lt;/code&gt; and storage (renders, downloaded footage, uploads), and it's writable, per-user app data. The other holds fonts, songs, public assets and the bundled &lt;code&gt;ffmpeg.exe&lt;/code&gt;, and it isn't writable, it sits inside the install itself. That split is what makes a Microsoft Store build possible at all: MSIX installs under Program Files, and the app simply cannot write there.&lt;/p&gt;

&lt;h2&gt;
  
  
  What You Can Actually Configure
&lt;/h2&gt;

&lt;p&gt;Nothing ships with credentials. The app sits inert until you supply at least an LLM provider key, and, unless you're rendering entirely from local files, a Pexels or Pixabay key on top of that. That's the honest way to build this: it doesn't touch your money before you've told it to.&lt;/p&gt;

&lt;h3&gt;
  
  
  Script and Keyword Generation
&lt;/h3&gt;

&lt;p&gt;Scripts and stock-footage search keywords can come from any of &lt;strong&gt;23 LLM providers&lt;/strong&gt;: OpenAI, Anthropic Claude, Google Gemini, DeepSeek, Qwen, Moonshot/Kimi, xAI Grok, Groq, OpenRouter, Ollama, and more besides. Adding a new one is described as a registry entry, not a new code path, which tells you something: the list was built to keep growing, not to sit fixed where it is now.&lt;/p&gt;

&lt;h3&gt;
  
  
  Narration and Subtitles
&lt;/h3&gt;

&lt;p&gt;Narration comes from Edge TTS (free, no account needed), Azure Speech, or ElevenLabs. Voices are listed per engine, and the app tells you up front when a key is missing instead of letting the render fail silently later, once you've already waited for it. Subtitles are burned in with a choice of bundled font, size, position, colour and outline, and you see the result before you commit to rendering it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Footage and Music
&lt;/h3&gt;

&lt;p&gt;Footage comes from Pexels, Pixabay, or your own local files. Background music comes from 8 bundled tracks, or whatever you upload yourself. Neither Pexels nor Pixabay is hosted by the app: it reaches out to both at render time, using the free API key you provided.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Component&lt;/th&gt;
&lt;th&gt;Role&lt;/th&gt;
&lt;th&gt;Notes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Tauri 2 / Rust shell&lt;/td&gt;
&lt;td&gt;Window, process supervision&lt;/td&gt;
&lt;td&gt;Kills backend on exit via Windows Job Object&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;React 19 frontend&lt;/td&gt;
&lt;td&gt;Generate, Library, Settings views&lt;/td&gt;
&lt;td&gt;Replaces upstream Streamlit UI&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Python backend&lt;/td&gt;
&lt;td&gt;FastAPI + moviepy + FFmpeg + TTS&lt;/td&gt;
&lt;td&gt;Runs as PyInstaller sidecar, &lt;code&gt;ca-desktop-backend.exe&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;FFmpeg&lt;/td&gt;
&lt;td&gt;Cutting, subtitles, music mix, export&lt;/td&gt;
&lt;td&gt;Bundled LGPL build, roughly 110 MB&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  The Review Pass, and Why It Exists
&lt;/h2&gt;

&lt;p&gt;Before you spend real API quota and real render minutes on a full video, Clip Architect offers an optional review pass. It downloads the clips first and lets you reorder, remove or replace any of them before anything else happens. Finished tasks can be reopened for another pass later, too. There's also a voice-over-first option: you record and listen to the narration on its own, with its duration, before committing to anything downstream of it. None of this is decoration. Footage matching by keyword misses sometimes, and paying twice for a clip you were always going to swap out, once in provider quota and once in render time, is exactly the waste a review step is there to stop.&lt;/p&gt;

&lt;p&gt;Saved projects, a library of past renders, and a shared cache of downloaded clips round the workflow out. That cache matters more than it sounds like it should: &lt;a href="https://diflowrin.com/clip-architect/" rel="noopener noreferrer"&gt;footage downloaded for one video stays available to the next one&lt;/a&gt;, so two videos on neighbouring subjects don't pay twice for the same shot, and you can browse the cache and delete anything you'd rather never see again.&lt;/p&gt;

&lt;h2&gt;
  
  
  Installing It and What It Needs
&lt;/h2&gt;

&lt;p&gt;The straightforward path is the Microsoft Store listing: one click, no build steps, updates handled by Windows from then on. That path needs Windows 10 build 1809 or newer with the WebView2 runtime, which most current Windows 10 or 11 machines already carry without you doing anything. Building it from source is a different matter, and needs Rust with the MSVC toolchain, Node.js 20 or newer, and uv, since the backend pins Python 3.11.&lt;/p&gt;

&lt;p&gt;The credential-free path through the whole thing is local footage plus Edge TTS, and it needs no account of any kind. That path earns its place twice over: it's the honest first run for someone who just installed the app and hasn't typed in a single key yet, and it's also what a Microsoft Store certification reviewer, testing cold with nothing configured, needs in order to see the software actually do something.&lt;/p&gt;

&lt;h2&gt;
  
  
  What It Doesn't Do Yet
&lt;/h2&gt;

&lt;p&gt;Being fair about the limits is part of describing a tool honestly, so here they are. Clip Architect is Windows-only: the shell, the MSIX packaging and the Windows Media Foundation encoder it leans on are all specific to that platform, full stop. Subtitle fonts cover Latin, Vietnamese and Thai; there's no CJK, Cyrillic or Greek font shipped, so other scripts come out as blank glyphs instead of crashing anything. H.264 output is Constrained Baseline, because it runs through Windows' own encoder rather than a bundled one, a patent decision made on purpose rather than an oversight, and the files come out larger than a libx264 build would produce. And a handful of upstream integrations, AI-generated video among them, plus extra music generation and cross-posting to other platforms, are sitting vendored in the codebase but not yet wired into the interface.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Licence Question, Answered Plainly
&lt;/h2&gt;

&lt;p&gt;Clip Architect is dual-licensed. GPL-3.0 is the default, and it costs nothing: use it, study it, modify it, sell the videos you make with it. What GPL-3.0 asks in return only kicks in if you distribute the software itself, or something built from it, in which case that too has to ship under GPL-3.0, with source attached. A separate commercial licence, granted by written agreement, covers what GPL doesn't: embedding the app in a closed-source product, redistributing it without the source, or shipping it under your own name.&lt;/p&gt;

&lt;p&gt;Neither licence reaches into what you actually produce. Whatever you render is yours, to publish, sell or monetise as you like, and your own footage and scripts were always yours to begin with. The upstream MoneyPrinterTurbo code, credited to its original author Harry, is MIT-licensed and keeps that notice intact; MIT is one-way compatible with GPL, so the combined work ships under GPL-3.0 by choice, not by obligation, since nothing bundled forces that choice on anyone. And that choice is exactly what leaves room for the second licence to exist at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  People Also Ask
&lt;/h2&gt;

&lt;h3&gt;
  
  
  How does Clip Architect turn a topic into a finished short video automatically?
&lt;/h3&gt;

&lt;p&gt;It runs the same pipeline start to finish without you touching a terminal at all: the LLM you configured drafts the script and the footage keywords, the clips get matched and downloaded, the voice track gets recorded, and FFmpeg assembles the result, subtitles and music included, into an MP4 ready to post.&lt;/p&gt;

&lt;h3&gt;
  
  
  Which LLM providers and text-to-speech engines does Clip Architect support?
&lt;/h3&gt;

&lt;p&gt;Scripts and footage keywords can be generated by any of 23 LLM providers, including OpenAI, Anthropic Claude, Google Gemini, DeepSeek, Qwen, Moonshot/Kimi, xAI Grok, Groq, OpenRouter and Ollama. Narration can use Edge TTS at no cost and with no account, or Azure Speech and ElevenLabs if you supply their keys.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can Clip Architect run entirely on my local Windows machine without cloud rendering?
&lt;/h3&gt;

&lt;p&gt;Yes. Rendering, cutting and encoding all happen on your own machine through the bundled FFmpeg build; nothing goes to a cloud render queue. The only traffic leaving your machine goes to whichever provider you configured, using the key you supplied yourself, for the script, the voice, or the footage.&lt;/p&gt;

&lt;h3&gt;
  
  
  How do I install Clip Architect from the Microsoft Store and what are the system requirements?
&lt;/h3&gt;

&lt;p&gt;It installs from the Microsoft Store in one click, with updates handled by Windows afterward. It needs Windows 10 build 1809 or newer with the WebView2 runtime; building it from source instead requires Rust with the MSVC toolchain, Node.js 20+, and uv for the Python 3.11 backend.&lt;/p&gt;

&lt;h2&gt;
  
  
  What This Means If You're Choosing a Local Video Pipeline
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;It's a shell, not a rival engine:&lt;/strong&gt; Clip Architect doesn't compete with MoneyPrinterTurbo, it packages it, so the generation logic and its limits come along unchanged, not replaced.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Local-first is the whole pitch:&lt;/strong&gt; your footage, your scripts, and your rendered videos never leave your machine except toward the one provider you configured yourself.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The review pass saves money before it saves time:&lt;/strong&gt; checking clips before a full render is what stops you burning quota and minutes on footage you were going to swap out anyway.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The licence lets you monetise freely:&lt;/strong&gt; GPL-3.0 or the commercial licence, neither one claims any right over the videos that come out of the app.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Windows-only, for now:&lt;/strong&gt; the shell, the MSIX packaging and the H.264 encoder are all tied to that one platform, so this isn't a cross-platform answer yet.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Additional Resources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/diflowrin/Clip-Architect" rel="noopener noreferrer"&gt;Clip Architect GitHub repository&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/harry0703/MoneyPrinterTurbo" rel="noopener noreferrer"&gt;MoneyPrinterTurbo main GitHub project page&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/harry0703/MoneyPrinterTurbo/blob/main/README-en.md" rel="noopener noreferrer"&gt;MoneyPrinterTurbo English README on GitHub&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://diflowrin.com/clip-architect/" rel="noopener noreferrer"&gt;Clip Architect project page&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.verdent.ai/guides/moneyprinterturbo-github" rel="noopener noreferrer"&gt;MoneyPrinterTurbo builder guide and architecture overview&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://pyshine.com/MoneyPrinterTurbo-AI-Powered-Video-Generation/" rel="noopener noreferrer"&gt;MoneyPrinterTurbo AI-powered one-click short video generation article&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;This article includes content created with AI.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>programming</category>
      <category>python</category>
      <category>ai</category>
      <category>react</category>
    </item>
    <item>
      <title>DeepSeek Harness: Plugin-First Agent Runtime</title>
      <dc:creator>DiFlowrin</dc:creator>
      <pubDate>Thu, 20 Aug 2026 11:43:52 +0000</pubDate>
      <link>https://dev.to/diflowrin/deepseek-harness-plugin-first-agent-runtime-2djh</link>
      <guid>https://dev.to/diflowrin/deepseek-harness-plugin-first-agent-runtime-2djh</guid>
      <description>&lt;h2&gt;
  
  
  Where DeepSeek Harness Stands Right Now
&lt;/h2&gt;

&lt;p&gt;DeepSeek Harness (&lt;strong&gt;dsh&lt;/strong&gt;) is an open-source agent harness from DeepSeek AI, built on one idea: &lt;strong&gt;everything is a plugin&lt;/strong&gt;. It launched on &lt;strong&gt;August 13, 2026&lt;/strong&gt; under the MIT license, the same day DeepSeek raised V4 Pro API prices — a free harness paired with a monetization move. It is in developer preview, and the project says in capitals that there will be compatibility-breaking changes. You can run it from npm in one command or build it from source with pnpm. It ships with no model, no key, no provider — you bring those yourself.&lt;/p&gt;

&lt;h2&gt;
  
  
  What DeepSeek Harness Actually Is
&lt;/h2&gt;

&lt;p&gt;DeepSeek Harness is an open-source agent harness developed by &lt;a href="https://deepseek.com/harness" rel="noopener noreferrer"&gt;DeepSeek AI&lt;/a&gt;. Not a model. Not a coding assistant you talk to. A harness — the thing that wraps around a model and gives it hands: tools, sessions, sandboxes, scheduling, a UI.&lt;/p&gt;

&lt;p&gt;And here is the part that matters. The architecture is &lt;strong&gt;plugin-based to the core&lt;/strong&gt;. Models, tools, skills, sessions, sandboxes, storage, loops, scheduling, the UI itself — all of it is plugins. The kernel underneath is &lt;a href="https://github.com/cordiverse/cordis" rel="noopener noreferrer"&gt;Cordis&lt;/a&gt;, a meta-framework whose design is laid out in the paper &lt;em&gt;A Programming Paradigm for Spatiotemporal Composability&lt;/em&gt;. Cordis handles loading, unloading and dependency relationships between plugins. It does not carry any agent capability itself. The capabilities come from the plugins, and they cooperate through Cordis services and events.&lt;/p&gt;

&lt;p&gt;So what does that buy you? You can select, replace or extend any capability at the configuration layer — without touching the source code. That is the claim, and it is the whole design philosophy: a thin kernel plus a bag of composable parts, not a monolithic agent tool.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why It Matters Now
&lt;/h2&gt;

&lt;p&gt;Because of the timing. &lt;a href="https://venturebeat.com/technology/deepseek-harness-launches-as-open-source-rival-to-claude-code-alongside-v4-pro-on-api-with-higher-prices" rel="noopener noreferrer"&gt;VentureBeat&lt;/a&gt; reports that DeepSeek launched Harness v0.1 on &lt;strong&gt;August 13, 2026&lt;/strong&gt;, alongside the official DeepSeek-V4-Pro model — and alongside a price increase. V4-Pro peak input moved to &lt;strong&gt;$1.32&lt;/strong&gt; per million tokens and peak output to &lt;strong&gt;$3.96&lt;/strong&gt;; off-peak sits at &lt;strong&gt;$0.66&lt;/strong&gt; input and &lt;strong&gt;$1.98&lt;/strong&gt; output.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;DeepSeek open-sourced its agent harness under MIT on the same day it raised V4 Pro API prices — free orchestration, paid inference.&lt;/strong&gt;VentureBeat, August 13, 2026&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Read that structure. The harness is free, permissively licensed, yours to modify. The model inference is not free — it is billed by whatever API provider you configure. The giveaway and the monetization are two halves of one move. And the positioning is explicit: an alternative to integrated coding-agent environments such as Anthropic's Claude Code.&lt;/p&gt;

&lt;p&gt;What can it do already? Per VentureBeat's reporting: inspect repositories, edit files, execute shell commands, search files and the web, maintain plans, invoke skills, delegate work to subagents, and enforce approval policies. That is a full agentic coding loop, not a demo.&lt;/p&gt;

&lt;h2&gt;
  
  
  How the Plugin Architecture Works
&lt;/h2&gt;

&lt;p&gt;The repository tells the story before you read a word of documentation. It is a TypeScript &lt;strong&gt;monorepo&lt;/strong&gt; managed with pnpm, and the &lt;code&gt;packages/&lt;/code&gt; directory holds dozens of plugin packages — session, sandbox, mcp, llm, plan, schedule, skill, subagent, terminal, workflow, and many more. Each capability is its own package. There are also &lt;code&gt;apps/cli&lt;/code&gt; and &lt;code&gt;apps/web&lt;/code&gt;, a Python SDK under &lt;code&gt;python/&lt;/code&gt;, vendored dependencies under &lt;code&gt;vendor/&lt;/code&gt; (including Cordis itself), and a VitePress site under &lt;code&gt;website/&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Cordis kernel&lt;/p&gt;

&lt;p&gt;The plugin framework underneath DeepSeek Harness. Responsible only for plugin loading, unloading and dependency relationships — no agent capabilities live in the kernel.&lt;/p&gt;

&lt;p&gt;Capabilities as plugins&lt;/p&gt;

&lt;p&gt;Models, tools, skills, sessions, sandboxes, storage, loops, scheduling and UI are all provided by plugins that cooperate through Cordis services and events.&lt;/p&gt;

&lt;p&gt;Configuration-layer composition&lt;/p&gt;

&lt;p&gt;Developers can select, replace or extend any capability in configuration, without modifying source code.&lt;/p&gt;

&lt;p&gt;Append-only session log&lt;/p&gt;

&lt;p&gt;According to &lt;a href="https://www.eigent.ai/blog/deepseek-harness-agent-runtime" rel="noopener noreferrer"&gt;Eigent's write-up&lt;/a&gt;, every agent run is recorded in an append-only log: system prompts, reasoning, tool calls and results, subagent scheduling, every context injection. Full traceability of what the model saw and did.&lt;/p&gt;

&lt;h3&gt;
  
  
  It can orchestrate its own competitors
&lt;/h3&gt;

&lt;p&gt;One detail worth sitting with. &lt;a href="https://www.mindstudio.ai/blog/deepseek-harness-agentic-coding" rel="noopener noreferrer"&gt;MindStudio&lt;/a&gt; reports that the plugin composability extends to entire other agent harnesses — Claude Code or Codex can be called as sub-agents inside a DeepSeek-orchestrated workflow. Rival and wrapper at the same time. The line between the two gets blurry fast when everything is a plugin.&lt;/p&gt;

&lt;h2&gt;
  
  
  Running It: npm and Source
&lt;/h2&gt;

&lt;h3&gt;
  
  
  How to run DeepSeek Harness from npm
&lt;/h3&gt;

&lt;p&gt;Install Node.js, then one command:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;npx @deepseek-ai/dsh web&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That starts the Web UI at &lt;code&gt;http://127.0.0.1:3080&lt;/code&gt; by default and opens it in your default browser on a local launch. Over SSH it only prints the host URL, because the SSH client or editor owns the forwarded address. Pass &lt;code&gt;--no-open&lt;/code&gt; to run the server without opening a browser.&lt;/p&gt;

&lt;p&gt;One caveat the README does not spell out but installers should know: &lt;a href="https://www.atlascloud.ai/blog/tips/how-to-install-deepseek-harness" rel="noopener noreferrer"&gt;AtlasCloud's setup guide&lt;/a&gt; states the repository requires Node.js &lt;strong&gt;^22.19.0 or &amp;gt;=24.0.0&lt;/strong&gt; — nothing on the 23.x line qualifies at any patch level. An older Node may not satisfy the engine requirement even if npx itself runs.&lt;/p&gt;

&lt;h3&gt;
  
  
  DeepSeek Harness run from source with pnpm
&lt;/h3&gt;

&lt;p&gt;From a repository checkout:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;git clone &lt;a href="https://github.com/deepseek-ai/deepseek-harness.git" rel="noopener noreferrer"&gt;https://github.com/deepseek-ai/deepseek-harness.git&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;cd deepseek-harness&lt;/li&gt;
&lt;li&gt;pnpm install&lt;/li&gt;
&lt;li&gt;pnpm run build&lt;/li&gt;
&lt;li&gt;pnpm dsh web&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;code&gt;pnpm run build&lt;/code&gt; prepares the repository artifacts. &lt;code&gt;pnpm dsh web&lt;/code&gt; then uses those built artifacts without rebuilding.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Fine Print: Configuration, Credentials, Context
&lt;/h2&gt;

&lt;p&gt;Here is where early adopters need to pay attention. The harness ships with &lt;strong&gt;no credentials, no default provider, and no bundled model&lt;/strong&gt;, per AtlasCloud. It will not do a single useful thing until you give it an OpenAI-compatible base URL, an API key, and at least one model ID. Three routes work: DeepSeek's own API, an OpenAI-compatible gateway, or a local model via Ollama.&lt;/p&gt;

&lt;p&gt;Two more details from the same guide deserve a table, because they will bite you if you miss them:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Setting&lt;/th&gt;
&lt;th&gt;Default&lt;/th&gt;
&lt;th&gt;Implication&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Hand-declared model context window&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;262,144 tokens&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;V4's full 1,048,576-token window stays off until configured&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Default max tokens&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;32,768&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Output limits are far below the model's ceiling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Credential storage&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Plain text&lt;/strong&gt; at ~/.dsh/.credentials.yaml&lt;/td&gt;
&lt;td&gt;Keys sit unencrypted on disk; no source documents mitigations&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Settings&lt;/td&gt;
&lt;td&gt;~/.dsh/settings.yaml&lt;/td&gt;
&lt;td&gt;Plugins and models configured via YAML&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;And a warning about a comfortable assumption: a local Web UI is not local inference. &lt;a href="https://atoms.dev/blog/deepseek-harness" rel="noopener noreferrer"&gt;Atoms.dev&lt;/a&gt; puts it plainly — the browser interface runs locally, but the configured model may still call a paid remote API. The same source notes the standard permission preset is workspace-write with approval prompts, and that a danger-full-access mode exists which deliberately bypasses filesystem confinement. Do not reach for that one casually.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Real Usage Costs
&lt;/h2&gt;

&lt;p&gt;Agentic coding at scale burns tokens. MindStudio's real-world test — building a real-time ISS tracker with a live API feed and a 3D globe visualization — consumed roughly &lt;strong&gt;20 million tokens&lt;/strong&gt; across two turns and about &lt;strong&gt;35 minutes&lt;/strong&gt;, with output alone near &lt;strong&gt;240,000 tokens&lt;/strong&gt;. Cache hit rates landed in the &lt;strong&gt;95–100%&lt;/strong&gt; range, unusually high for this category of tool.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;One moderately complex task: ~20 million tokens, two turns, 35 minutes.&lt;/strong&gt;MindStudio, August 2026&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;So budget accordingly. High cache hit rates soften the bill, but 20 million tokens is 20 million tokens.&lt;/p&gt;

&lt;p&gt;On benchmarks, be careful. &lt;a href="https://composio.dev/content/best-agent-harness-deepseek-v4-flash" rel="noopener noreferrer"&gt;Composio&lt;/a&gt; tested DeepSeek v4 Flash across several harnesses — 240 runs, a &lt;strong&gt;53.8%&lt;/strong&gt; overall pass rate, DeepAgents at &lt;strong&gt;53.3%&lt;/strong&gt; with a 187.1-second median and $0.045 per successful task. But DeepSeek Harness itself was not in that test. Nobody has published a head-to-head benchmark of it against rival harnesses. Anyone claiming otherwise is guessing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Reports Disagree
&lt;/h2&gt;

&lt;p&gt;Because the project is days old and moving fast, third-party write-ups contradict each other. Two examples worth knowing:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Default port.&lt;/strong&gt; The README states &lt;a href="http://127.0.0.1:3080" rel="noopener noreferrer"&gt;http://127.0.0.1:3080&lt;/a&gt;, and Atoms.dev agrees. MindStudio's two setup articles both say port &lt;strong&gt;3018&lt;/strong&gt;. Whether the default changed between versions or one report is simply wrong — no source settles it. Trust the README.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Whether a DeepSeek API key is mandatory.&lt;/strong&gt; MindStudio says you need a DeepSeek API key before the harness will let you do anything. AtlasCloud says the opposite: no default provider, any OpenAI-compatible endpoint works, including local Ollama. The architecture — model provider as plugin — supports AtlasCloud's reading.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Preset modes.&lt;/strong&gt; Atoms.dev lists Standard, PTC, Minimal, Creator. CometAPI lists Standard, Minimal, Code, Creator. MindStudio says full, code, minimal, creator. Four modes, three namings — check the current docs before scripting against them.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The Developer Preview Caveat
&lt;/h2&gt;

&lt;p&gt;The project is unusually honest about its own status. The README states it is in &lt;strong&gt;developer preview&lt;/strong&gt;, iterating rapidly, and — in capitals — &lt;strong&gt;THERE WILL BE COMPATIBILITY-BREAKING CHANGES&lt;/strong&gt;. No source provides a roadmap or a date for a stable release. Pin a package version and verify current documentation before adopting it for anything serious, as Atoms.dev advises.&lt;/p&gt;

&lt;p&gt;The licence, at least, is settled: &lt;strong&gt;MIT&lt;/strong&gt;, copyright DeepSeek, with third-party dependencies and their licenses disclosed in THIRD_PARTY_NOTICES.md. Commercial use, modification, redistribution — all permitted.&lt;/p&gt;

&lt;h2&gt;
  
  
  Community and Contributing
&lt;/h2&gt;

&lt;p&gt;Feedback and bug reports go through GitHub Discussions. Plugin authors are asked to add the &lt;code&gt;dsh-plugin&lt;/code&gt; topic to their repositories for discoverability — a small signal that third-party plugin development is an expected part of the model, not an afterthought. There is also a Discord community. For contributors: CONTRIBUTING.md, a development guide, and architecture documentation. For agents working in the repo: AGENTS.md. Documentation ships in English and Chinese.&lt;/p&gt;

&lt;h2&gt;
  
  
  People Also Ask
&lt;/h2&gt;

&lt;h3&gt;
  
  
  What is the default Web UI port for DeepSeek Harness?
&lt;/h3&gt;

&lt;p&gt;The README states the Web UI runs at &lt;code&gt;http://127.0.0.1:3080&lt;/code&gt; by default. Some third-party guides report port 3018 instead, and no source clarifies whether the default changed between versions — the repository documentation is the authoritative reference.&lt;/p&gt;

&lt;h3&gt;
  
  
  Do I need a DeepSeek API key to use DeepSeek Harness?
&lt;/h3&gt;

&lt;p&gt;Sources disagree. MindStudio reports a DeepSeek API key is required before anything works, while AtlasCloud reports the harness ships with no default provider and accepts any OpenAI-compatible base URL, key and model ID — including local models via Ollama. The plugin-based design supports the latter reading.&lt;/p&gt;

&lt;h3&gt;
  
  
  What Node.js version does DeepSeek Harness require?
&lt;/h3&gt;

&lt;p&gt;AtlasCloud's installation guide states the engine requirement is Node.js &lt;strong&gt;^22.19.0 or &amp;gt;=24.0.0&lt;/strong&gt;. The 23.x line does not qualify at any patch level, and an older Node may fail the engine check even when npx itself runs.&lt;/p&gt;

&lt;h3&gt;
  
  
  Is DeepSeek Harness free for commercial use?
&lt;/h3&gt;

&lt;p&gt;Yes — the code is MIT licensed, which permits commercial use, modification and redistribution. Note that model inference is still billed by whichever API provider you configure, and the developer-preview status means breaking changes are expected.&lt;/p&gt;

&lt;h2&gt;
  
  
  📊 What This Means for Teams Evaluating DeepSeek Harness
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Architecture is the product:&lt;/strong&gt; a thin Cordis kernel with every capability — models, tools, sessions, sandboxes, UI — as a swappable plugin, composable in configuration without source changes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Free harness, paid inference:&lt;/strong&gt; MIT-licensed code launched the same day V4 Pro API prices rose; the browser UI is local but the model bill is not.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Bring your own everything:&lt;/strong&gt; no bundled model, no default provider, no credentials — you supply an OpenAI-compatible endpoint, key and model ID.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Watch the defaults:&lt;/strong&gt; hand-declared models get a &lt;strong&gt;262,144&lt;/strong&gt;-token context window and &lt;strong&gt;32,768&lt;/strong&gt; max tokens, and credentials sit in plain text under ~/.dsh.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Preview means preview:&lt;/strong&gt; breaking changes are promised in writing, no stable-release roadmap exists, and third-party guides already contradict each other on ports, modes and requirements. Pin versions; verify docs.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The Bottom Line
&lt;/h2&gt;

&lt;p&gt;DeepSeek Harness is a serious piece of infrastructure from a major lab, released with unusual openness and unusual candor about its own instability. The plugin-first design is not marketing decoration — the monorepo structure, the Cordis kernel, the dsh-plugin ecosystem push all point the same direction. What it is not yet is finished. Treat it as what it says it is: a developer preview worth building on carefully, with pinned versions, isolated workspaces, and eyes open about where the tokens — and the money — go.&lt;/p&gt;

&lt;h2&gt;
  
  
  Additional Resources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/deepseek-ai/deepseek-harness" rel="noopener noreferrer"&gt;DeepSeek Harness GitHub repository&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://deepseek.com/harness" rel="noopener noreferrer"&gt;DeepSeek Harness official site&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/deepseek-ai/deepseek-harness/blob/master/README.md" rel="noopener noreferrer"&gt;DeepSeek Harness README on GitHub&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/deepseek-ai/deepseek-harness/blob/master/CONTRIBUTING.md" rel="noopener noreferrer"&gt;DeepSeek Harness contributing guide&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/deepseek-ai/deepseek-harness/blob/master/LICENSE" rel="noopener noreferrer"&gt;DeepSeek Harness license&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/deepseek-ai/deepseek-harness/blob/master/THIRD_PARTY_NOTICES.md" rel="noopener noreferrer"&gt;DeepSeek Harness third-party notices&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;This article includes content created with AI.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>typescript</category>
      <category>github</category>
      <category>git</category>
      <category>ai</category>
    </item>
    <item>
      <title>I Ran 4 Free OpenRouter Models Through a Full Article Pipeline. Half of Them Broke.</title>
      <dc:creator>DiFlowrin</dc:creator>
      <pubDate>Fri, 07 Aug 2026 16:17:12 +0000</pubDate>
      <link>https://dev.to/diflowrin/i-ran-4-free-openrouter-models-through-a-full-article-pipeline-half-of-them-broke-34cp</link>
      <guid>https://dev.to/diflowrin/i-ran-4-free-openrouter-models-through-a-full-article-pipeline-half-of-them-broke-34cp</guid>
      <description>&lt;p&gt;Most model comparisons test one prompt and score the output. That tells you almost nothing about whether a model survives a real pipeline.&lt;/p&gt;

&lt;p&gt;A real pipeline has steps. Draft, then a second pass that reads the draft back and rewrites it. Maybe structured output somewhere. Maybe a tool call. Each step is a new chance for something to break, and the breakages don't correlate neatly with how good the prose is.&lt;/p&gt;

&lt;p&gt;So I tested four free models on OpenRouter through an actual two-step content workflow instead of a single call. Same prompts, same setup, no paid tokens anywhere.&lt;/p&gt;

&lt;p&gt;Here's what broke and what didn't.&lt;/p&gt;

&lt;h2&gt;
  
  
  The setup
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3b1qfi5v1bq43b6xbwjj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3b1qfi5v1bq43b6xbwjj.png" alt=" " width="800" height="427"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The pipeline is minimal on purpose:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Generate&lt;/strong&gt; — take a topic, produce a full article&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Improve&lt;/strong&gt; — feed the draft back with an SEO-optimization instruction, get a revised version&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Score&lt;/strong&gt; — the tool computes a content score on both versions&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Step 2 is the interesting one. It's a longer context, it requires the model to actually read and revise rather than generate fresh, and it's where things fell over.&lt;/p&gt;

&lt;p&gt;Everything runs through OpenRouter's OpenAI-compatible endpoint, so swapping models is a one-line change:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;https://openrouter.ai/api/v1/chat/completions&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;method&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;POST&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Authorization&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;`Bearer &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;OPENROUTER_API_KEY&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Content-Type&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;application/json&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="na"&gt;body&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
    &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;nvidia/nemotron-3-ultra-550b-a55b:free&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;// ← the only thing that changes&lt;/span&gt;
    &lt;span class="na"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt; &lt;span class="na"&gt;role&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;user&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;prompt&lt;/span&gt; &lt;span class="p"&gt;}],&lt;/span&gt;
  &lt;span class="p"&gt;}),&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's the whole reason this test is cheap to run. If you want to reproduce it, you're changing a string in a loop.&lt;/p&gt;

&lt;h2&gt;
  
  
  The results
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Draft quality&lt;/th&gt;
&lt;th&gt;Improve pass&lt;/th&gt;
&lt;th&gt;Verdict&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;inclusionai/ling-3.0-flash:free&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Good, fast&lt;/td&gt;
&lt;td&gt;❌ Failed every run&lt;/td&gt;
&lt;td&gt;One-shot only&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;nvidia/nemotron-3-ultra-550b-a55b:free&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Good, fast&lt;/td&gt;
&lt;td&gt;✅ Clean&lt;/td&gt;
&lt;td&gt;Best of four&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;nvidia/nemotron-nano-9b-v2:free&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Weak (score 52)&lt;/td&gt;
&lt;td&gt;✅ → 67&lt;/td&gt;
&lt;td&gt;Viable with two passes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;nvidia/nemotron-3-super-120b-a12b:free&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;❌ Errored&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;Broken&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwyx1bcso3vmmsyovfhdo.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwyx1bcso3vmmsyovfhdo.png" alt=" " width="800" height="429"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fceg05t5gcbrtcivslqfj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fceg05t5gcbrtcivslqfj.png" alt=" " width="800" height="428"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;inclusionai/ling-3.0-flash:free&lt;/code&gt; — good writer
&lt;/h3&gt;

&lt;p&gt;Fast, and the prose was genuinely better than I expected from a free tier. Clean structure, no obvious LLM tells.&lt;/p&gt;

&lt;p&gt;Then step 2 failed. Every single run, not intermittently. I don't have a root cause — it could be a context limit on the free tier, a provider-side timeout on longer inputs, or something in how the revision prompt is shaped. Consistent failure usually points at a hard constraint rather than flakiness, but I didn't dig further.&lt;/p&gt;

&lt;p&gt;Result: you get a good first draft and no way to iterate. Whether that's fatal depends entirely on whether your workflow has a step 2.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;nvidia/nemotron-3-ultra-550b-a55b:free&lt;/code&gt; — the only one that just worked
&lt;/h3&gt;

&lt;p&gt;Fast, solid text, improve pass completed without complaint. Nothing to write up, which is the point. If you test one model from this list, test this one.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;nvidia/nemotron-nano-9b-v2:free&lt;/code&gt; — the interesting failure
&lt;/h3&gt;

&lt;p&gt;Two things went wrong here, and only one of them matters.&lt;/p&gt;

&lt;p&gt;First, instruction-following: I set length to "short" and got 1,557 words. Small models are notoriously loose about constraints expressed in prose. If length actually matters to you, enforce it downstream — &lt;code&gt;max_tokens&lt;/code&gt;, a post-generation truncation step, or a structured output schema — rather than asking politely in the prompt.&lt;/p&gt;

&lt;p&gt;Second, quality: content score of 52 on the first draft. That's below what I'd publish.&lt;/p&gt;

&lt;p&gt;But the improve pass took it to &lt;strong&gt;67&lt;/strong&gt;. That's a real jump, and it reframes the model entirely. A 9B model isn't "too small to use" — it's a model that needs two passes to reach where a bigger model lands in one.&lt;/p&gt;

&lt;p&gt;That trade-off is worth doing the math on. Two calls to a small fast model versus one call to a large one isn't obviously worse on latency or cost, and on free tiers it's not worse on either. The question is whether your pipeline can tolerate the extra step.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;nvidia/nemotron-3-super-120b-a12b:free&lt;/code&gt; — nothing
&lt;/h3&gt;

&lt;p&gt;Errored out. Never produced an article. Not much to analyse.&lt;/p&gt;

&lt;p&gt;Free-tier endpoints get rate-limited, deprioritized, and occasionally just fall over under load. This might work fine next week. That's the deal with free tiers.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd actually take from this
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Test the chain, not the prompt.&lt;/strong&gt; The most useful finding wasn't a quality ranking. It was that a model with good output couldn't complete the workflow. Single-prompt benchmarks would have ranked Ling highly and told you nothing about the thing that actually disqualified it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Small models are two-pass models.&lt;/strong&gt; Nemotron Nano's 52 → 67 is the clearest signal in the whole test. Don't evaluate a small model on its first output. Evaluate it on its second.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Consistent failure ≠ flaky failure.&lt;/strong&gt; Ling failing every time is diagnostically different from failing sometimes. One suggests a constraint you can find and work around, the other suggests infrastructure you can't control. Log which one you're seeing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The model is half the cost surface.&lt;/strong&gt; Images in this pipeline came from Pexels, pulled in automatically. Full run — draft, revision, images — cost nothing end to end. When people price out an AI content workflow they price the tokens and forget everything around them. Sometimes the stock photo bill is the bigger line item.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvl7yf9xwjwfmn5ff251u.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvl7yf9xwjwfmn5ff251u.png" alt=" " width="800" height="429"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Caveats
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq6saq4ri77n6nne8jixz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq6saq4ri77n6nne8jixz.png" alt=" " width="800" height="340"&gt;&lt;/a&gt;&lt;br&gt;
Free tiers change constantly. Models get added, throttled, and deprecated, and providers rotate what's behind a given slug. Everything above is a snapshot from a single run each — not a statistically meaningful benchmark. Content scores come from the tool's own scorer, so treat them as directional and internally comparable, not absolute.&lt;/p&gt;

&lt;p&gt;I also haven't tested any of these on non-English content, and I'd bet that's where the free/paid gap widens most.&lt;/p&gt;

&lt;p&gt;If you've run free models in a multi-step pipeline, I'm curious where yours broke — especially if you got Ling's revision step working, because I'd like to know what I did wrong.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>openrouter</category>
      <category>productivity</category>
      <category>llm</category>
    </item>
    <item>
      <title>Arroyo: Discover Real-Time Data Processing in Rust</title>
      <dc:creator>DiFlowrin</dc:creator>
      <pubDate>Sat, 16 May 2026 19:21:10 +0000</pubDate>
      <link>https://dev.to/diflowrin/arroyo-discover-real-time-data-processing-in-rust-19p7</link>
      <guid>https://dev.to/diflowrin/arroyo-discover-real-time-data-processing-in-rust-19p7</guid>
      <description>&lt;h2&gt;
  
  
  Why Arroyo Matters in Stream Processing
&lt;/h2&gt;

&lt;p&gt;The demand for real-time analytics solutions has surged, rendering traditional data processing capabilities insufficient for modern needs. &lt;strong&gt;Arroyo&lt;/strong&gt; answers this call by providing a &lt;strong&gt;distributed stream processing engine&lt;/strong&gt; that enables developers to harness data from both bounded and unbounded sources efficiently. Built in &lt;strong&gt;Rust&lt;/strong&gt;, Arroyo combines performance with safety, making it an excellent choice for applications that require low-latency and high-throughput data processing.&lt;/p&gt;

&lt;h2&gt;
  
  
  How Arroyo Works: A Deep Dive
&lt;/h2&gt;

&lt;p&gt;At its core, Arroyo operates on a &lt;strong&gt;dataflow model&lt;/strong&gt;, allowing it to process streams of data fluidly. By supporting stateful computations, Arroyo empowers developers to build more complex streaming applications. These applications can handle tasks such as joining streams or applying time-windows—all defined using a familiar &lt;strong&gt;SQL-like syntax&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;This engine scales to millions of events per second, making it suitable for applications where performance is paramount. It employs a distributed architecture, allowing workloads to be balanced across multiple nodes, which dramatically enhances performance and reliability.&lt;/p&gt;

&lt;h3&gt;
  
  
  Stateful Stream Processing in Arroyo
&lt;/h3&gt;

&lt;p&gt;A standout feature of Arroyo is its capability for &lt;strong&gt;stateful stream processing&lt;/strong&gt;. Unlike traditional stream processing engines, Arroyo can maintain state information across multiple processing operations. This means you can analyze trends, averages, or any long-term insights without losing context. By supporting complex operations like windowing and joining, users can design intricate analytics pipelines that return meaningful insights in real-time.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Benefits of Using Arroyo
&lt;/h2&gt;

&lt;p&gt;Adopting Arroyo offers numerous advantages, particularly for developers and teams focused on real-time analytics.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Low Latency:&lt;/strong&gt; Arroyo is designed for speed. The ability to process millions of events per second makes it ideal for applications that cannot afford delays.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Fault Tolerance:&lt;/strong&gt; With built-in state &lt;strong&gt;checkpointing&lt;/strong&gt;, Arroyo ensures that data processing can recover from failures without losing valuable data.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;SQL Support:&lt;/strong&gt; Using SQL for defining pipelines greatly lowers the hurdle for data teams, making Arroyo more accessible for users familiar with traditional databases.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Flexible Integrations:&lt;/strong&gt; Arroyo easily integrates with popular data systems like &lt;strong&gt;Kafka&lt;/strong&gt; and &lt;strong&gt;Iceberg&lt;/strong&gt;, enabling seamless data ingestion and output.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Open Source:&lt;/strong&gt; As a project within the open-source community, Arroyo benefits from contributions from developers around the world, enhancing its capabilities over time.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Real-Time Analytics Capability
&lt;/h3&gt;

&lt;p&gt;With Arroyo, businesses can perform &lt;strong&gt;real-time analytics&lt;/strong&gt; on their data streams. For example, financial services can monitor transactions for fraud as they occur, while e-commerce platforms can track customer behaviors and adapt inventory dynamically. The ability to analyze data in real-time means faster decision-making and improved user experiences.&lt;/p&gt;

&lt;h2&gt;
  
  
  Practical Examples of Arroyo in Use
&lt;/h2&gt;

&lt;p&gt;Understanding how Arroyo fits into real-world scenarios can illuminate its power and flexibility. Here are practical workflows that leverage Arroyo’s features:&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Event-Driven Data Processing in E-Commerce
&lt;/h3&gt;

&lt;p&gt;Imagine a scenario where an online retailer is running a flash sale. Arroyo can process streams of user interactions in real-time, tracking clicks, views, and cart additions. By analyzing this data, the retailer can identify trends, adjusting prices dynamically and recommending related products, enhancing the buying experience significantly.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Anomaly Detection in Financial Services
&lt;/h3&gt;

&lt;p&gt;In banking, Arroyo can be employed to analyze transaction streams to detect anomalies that may indicate fraud. Real-time monitoring allows banks to respond swiftly, alerting customers and freezing suspicious activities before greater damage occurs. This enhances security and builds customer trust.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fcajysjz5ces8oteua7s7.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fcajysjz5ces8oteua7s7.webp" alt="Arroyo: Discover Real-Time Data Processing in Rust" width="800" height="457"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Smart Urban Planning
&lt;/h3&gt;

&lt;p&gt;Municipalities are increasingly using live data from sensor networks (such as traffic cameras and IoT devices). Arroyo can consolidate these streams, providing insights into urban dynamics—like traffic patterns and pollution levels—allowing city planners to make informed decisions that improve living conditions.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Path Ahead for Arroyo
&lt;/h2&gt;

&lt;p&gt;The future looks bright for Arroyo and its community of developers. Being based in &lt;strong&gt;Rust&lt;/strong&gt;, a language known for its safety and performance, positions Arroyo strongly in the realm of high-efficiency applications. Its focus on &lt;strong&gt;stateful stream processing&lt;/strong&gt; and SQL integration mean that it can catch the attention of enterprise-level developers seeking ways to cope with rising demands for streaming analytics.&lt;/p&gt;

&lt;p&gt;However, like any technology, Arroyo has areas for improvement. Better documentation and example use cases could further ease the onboarding process for new users. As Arroyo grows, maintaining a healthy plugin ecosystem to support more connectors and data sinks will be essential for its adoption.&lt;/p&gt;

&lt;h2&gt;
  
  
  People Also Ask
&lt;/h2&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;### What is Arroyo used for?

Arroyo is primarily used for **distributed stream processing**, enabling real-time analytics on large volumes of data from both bounded and unbounded sources. Its capabilities are beneficial in various domains like e-commerce, finance, and smart city applications.



### Is Arroyo written in Rust?

Yes, Arroyo is implemented in **Rust**, which brings performance and memory safety to the engine, making it suitable for high-throughput applications.



### How do I get started with Arroyo?

To get started with Arroyo, you can install it via **Homebrew**, a shell installation script, or Docker. Comprehensive documentation and a tutorial are available on the official website.



### Does Arroyo support Kafka streams?

Yes, Arroyo includes support for **Kafka streaming**, allowing users to integrate and process streaming data efficiently from Kafka sources.



### Where can I find Arroyo documentation?

Documentation for Arroyo, including installation instructions and tutorials, can be found at the [official documentation site](https://doc.arroyo.dev/).
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;
&lt;h2&gt;
  
  
  Sources &amp;amp; References
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Original Source:&lt;/strong&gt; &lt;a href="https://github.com/ArroyoSystems/arroyo" rel="noopener noreferrer"&gt;https://github.com/ArroyoSystems/arroyo&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;### Additional Resources&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;- [Official GitHub Repository](https://github.com/ArroyoSystems/arroyo)

- [Arroyo Documentation](https://doc.arroyo.dev/)

- [Developer Setup Guide](https://doc.arroyo.dev/developing/dev-setup/)

- [Arroyo Real-time Analytics Tutorial](https://github.com/ArroyoSystems/analytics-tutorial)

- [ArroyoSystems GitHub Organization](https://github.com/ArroyoSystems)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

</description>
      <category>rust</category>
      <category>webdev</category>
      <category>programming</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>Get Shit Done: Streamlined AI Development Workflow</title>
      <dc:creator>DiFlowrin</dc:creator>
      <pubDate>Sat, 16 May 2026 18:06:15 +0000</pubDate>
      <link>https://dev.to/diflowrin/get-shit-done-streamlined-ai-development-workflow-4c07</link>
      <guid>https://dev.to/diflowrin/get-shit-done-streamlined-ai-development-workflow-4c07</guid>
      <description>&lt;h2&gt;
  
  
  Introduction to Get Shit Done
&lt;/h2&gt;

&lt;p&gt;The realm of software development has evolved dramatically over the years, particularly with the advent of Artificial Intelligence (AI) coding assistants. One such innovative framework designed to enhance productivity and streamline workflows is &lt;strong&gt;Get Shit Done&lt;/strong&gt; (GSD). This meta-prompting and context-engineering framework is tailored for AI coding assistants, including &lt;strong&gt;Claude Code&lt;/strong&gt;, &lt;strong&gt;OpenCode&lt;/strong&gt;, &lt;strong&gt;Gemini CLI&lt;/strong&gt;, and &lt;strong&gt;Codex&lt;/strong&gt;. Its primary objective is to support &lt;strong&gt;spec-driven development&lt;/strong&gt; workflows, ensuring that developers can efficiently manage their tasks from inception to deployment.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Essence of Get Shit Done
&lt;/h2&gt;

&lt;p&gt;At its core, GSD is designed to tackle some of the most pressing challenges faced by developers today, particularly the issue of &lt;strong&gt;context rot&lt;/strong&gt;. This phenomenon refers to the degradation of context quality that occurs as AI assistants fill their context windows. By addressing this issue, GSD empowers developers to maintain clarity and focus throughout the development process.&lt;/p&gt;

&lt;h3&gt;
  
  
  📹 Video: This Is How AI Will Change Your Workflow
&lt;/h3&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/FXkzsqk3a0w"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Video credit: Alex Hormozi&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;GSD introduces a structured approach through a series of well-defined phases: discuss, plan, execute, verify, and ship. This framework not only organizes tasks effectively but also emphasizes the importance of &lt;strong&gt;fresh context&lt;/strong&gt; for each stage of development. As a result, developers can leverage AI to enhance their productivity significantly.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why GSD Was Developed
&lt;/h3&gt;

&lt;p&gt;The inception of GSD was driven by the experiences of solo developers who needed a streamlined tool that diverged from the complexities associated with larger engineering organizations. Traditional tools often incorporate extensive processes such as sprint ceremonies, story points, and rigorous stakeholder syncs, which can hinder creativity and flexibility. GSD, on the other hand, minimizes complexity and maximizes efficiency, allowing developers to focus on building great products.&lt;/p&gt;

&lt;h2&gt;
  
  
  Understanding the GSD Framework
&lt;/h2&gt;

&lt;p&gt;The &lt;strong&gt;Get Shit Done&lt;/strong&gt; framework operates through a series of six commands that guide users through the development process. Each command is designed to accomplish a specific task, ensuring clarity and efficiency at every step.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Initialize: Setting the Foundation
&lt;/h3&gt;

&lt;p&gt;The first step in the GSD workflow is to initialize the project using the command &lt;code&gt;/gsd-new-project&lt;/code&gt;. This command prompts developers to answer questions that lead to research, requirements gathering, and the creation of a roadmap. Once approved, this roadmap serves as the foundation for subsequent phases.&lt;/p&gt;

&lt;p&gt;If developers already have existing code, they can run the command &lt;code&gt;/gsd-map-codebase&lt;/code&gt; to analyze their stack, architecture, and conventions. This preliminary analysis ensures that the questions posed during the initialization phase are relevant and tailored to the specific project.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Discuss: Capturing Vision and Decisions
&lt;/h3&gt;

&lt;p&gt;The second command, &lt;code&gt;/gsd-discuss-phase 1&lt;/code&gt;, allows developers to elaborate on their project vision. A roadmap typically provides a high-level overview, but discussions can uncover the finer details necessary for successful implementation. This phase captures decisions regarding layouts, API structures, error handling, and data management, addressing any uncertainties before planning begins.&lt;/p&gt;

&lt;p&gt;The insights gleaned from this discussion phase feed directly into the planning stage, ensuring that developers have a comprehensive understanding of their project requirements.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Plan: Strategizing Development
&lt;/h3&gt;

&lt;p&gt;After discussion, the next step is planning. The command &lt;code&gt;/gsd-plan-phase 1&lt;/code&gt; initiates a loop of research, planning, and verification until the plans are deemed satisfactory. Each plan is crafted to be manageable, allowing execution to occur within a fresh context window. This ensures that developers have the necessary clarity and focus when moving forward.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Execute: Bringing Plans to Life
&lt;/h3&gt;

&lt;p&gt;Execution is where ideas transform into tangible outcomes. Using the command &lt;code&gt;/gsd-execute-phase 1&lt;/code&gt;, developers can run plans in parallel waves, utilizing subagents that operate with fresh context windows. Each task is assigned an atomic commit, allowing for a clean git history that reflects the completed work.&lt;/p&gt;

&lt;p&gt;With a primary context window maintained at 30-40%, developers can step away from the workflow and return to find completed tasks ready for review.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Verify: Ensuring Quality and Functionality
&lt;/h3&gt;

&lt;p&gt;Verification is a crucial step in maintaining the integrity of the development process. The command &lt;code&gt;/gsd-verify-work 1&lt;/code&gt; allows developers to review the work that has been built. If any issues are identified, they can generate a diagnosis and a fix plan that is ready for immediate re-execution. This automated feedback loop eliminates the need for manual debugging, simplifying the verification process.&lt;/p&gt;

&lt;h3&gt;
  
  
  6. Repeat and Ship: Finalizing Development
&lt;/h3&gt;

&lt;p&gt;Once the verification phase is complete, the workflow proceeds to the final commands: &lt;code&gt;/gsd-ship 1&lt;/code&gt;, &lt;code&gt;/gsd-complete-milestone&lt;/code&gt;, and &lt;code&gt;/gsd-new-milestone&lt;/code&gt;. These commands facilitate the transition from development to deployment, allowing developers to loop through the discuss, plan, execute, verify, and ship phases until a milestone is achieved. Once completed, developers can archive, tag, and start a new milestone with a clean slate.&lt;/p&gt;

&lt;h2&gt;
  
  
  Benefits of Using GSD Framework in Software Development
&lt;/h2&gt;

&lt;p&gt;The implementation of the &lt;strong&gt;Get Shit Done&lt;/strong&gt; framework offers several advantages that can significantly enhance the software development experience:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Increased Efficiency:&lt;/strong&gt; By breaking down the development process into manageable phases and utilizing AI for task execution, developers can complete projects faster and with reduced effort.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Enhanced Clarity:&lt;/strong&gt; The structured approach minimizes confusion and ensures that all team members have a clear understanding of project objectives and requirements.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Improved Quality:&lt;/strong&gt; The verification phase allows for immediate identification and resolution of issues, leading to higher-quality outputs.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Streamlined Collaboration:&lt;/strong&gt; GSD is designed for both solo developers and larger teams, facilitating collaborative efforts without the overhead of cumbersome project management tools.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Flexibility:&lt;/strong&gt; The framework can be adapted to various development environments and workflows, making it a versatile tool for developers across different industries.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Who Can Benefit from the GSD Framework?
&lt;/h2&gt;

&lt;p&gt;The &lt;strong&gt;Get Shit Done&lt;/strong&gt; framework is particularly beneficial for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Solo Developers:&lt;/strong&gt; Individual developers looking for an efficient way to manage their projects can leverage GSD to streamline their workflows.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Small to Medium-Sized Teams:&lt;/strong&gt; Teams with limited resources can benefit from the automation and structure provided by GSD, allowing them to maximize their productivity without the need for extensive project management infrastructure.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Freelancers:&lt;/strong&gt; Freelancers working on multiple projects can use GSD to keep track of their tasks, ensuring that they meet deadlines and maintain quality.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Startups:&lt;/strong&gt; New ventures can utilize GSD to establish a solid development foundation, enabling them to build and launch products quickly and effectively.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Getting Started with Get Shit Done
&lt;/h2&gt;

&lt;p&gt;To begin using the &lt;strong&gt;Get Shit Done&lt;/strong&gt; framework, developers can install it via the command line with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx get-shit-done-cc@latest
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This command prompts users to select their runtime environment, whether it be &lt;strong&gt;Claude Code&lt;/strong&gt;, &lt;strong&gt;OpenCode&lt;/strong&gt;, &lt;strong&gt;Gemini CLI&lt;/strong&gt;, or others, and offers the choice of a global or local installation. Users can also configure their installation to include only the skills they require, fostering a more personalized experience.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion: The Future of AI-Powered Development
&lt;/h2&gt;

&lt;p&gt;The &lt;strong&gt;Get Shit Done&lt;/strong&gt; framework represents a significant step forward in the evolution of software development, particularly in the context of AI-powered tools. By addressing the challenges of context rot and providing a streamlined workflow, GSD empowers developers to focus on what truly matters: building high-quality products efficiently. As the landscape of technology continues to evolve, frameworks like GSD will play a crucial role in shaping the future of software development.&lt;/p&gt;

&lt;h2&gt;
  
  
  People Also Ask
&lt;/h2&gt;

&lt;h3&gt;
  
  
  What is Get Shit Done?
&lt;/h3&gt;

&lt;p&gt;Get Shit Done is a lightweight meta-prompting and context-engineering framework designed for AI coding assistants, aimed at enhancing productivity through a structured development process.&lt;/p&gt;

&lt;h3&gt;
  
  
  How does the GSD framework work?
&lt;/h3&gt;

&lt;p&gt;The GSD framework operates through a series of six commands that guide developers through the phases of initialization, discussion, planning, execution, verification, and shipping.&lt;/p&gt;

&lt;h3&gt;
  
  
  Which AI coding tools does Get Shit Done support?
&lt;/h3&gt;

&lt;p&gt;GSD supports various AI coding tools, including Claude Code, OpenCode, Gemini CLI, Codex, Copilot, and others.&lt;/p&gt;

&lt;h3&gt;
  
  
  What is spec-driven development?
&lt;/h3&gt;

&lt;p&gt;Spec-driven development is an approach that emphasizes the importance of specifications in guiding the development process, ensuring that the final product aligns with the original vision and requirements.&lt;/p&gt;

&lt;h3&gt;
  
  
  How does the discuss-plan-execute-verify-ship loop work?
&lt;/h3&gt;

&lt;p&gt;This loop involves discussing project details, planning tasks, executing them in parallel, verifying the results, and eventually shipping the completed work, allowing for iterative development and continuous improvement.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources &amp;amp; References
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Original Source:&lt;/strong&gt; &lt;a href="https://github.com/gsd-build/get-shit-done" rel="noopener noreferrer"&gt;https://github.com/gsd-build/get-shit-done&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;### Additional Resources&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;- [GitHub Repository](https://github.com/gsd-build/get-shit-done)

- [README](https://github.com/gsd-build/get-shit-done/blob/main/README.md)

- [User Guide](https://github.com/gsd-build/get-shit-done/blob/main/docs/USER-GUIDE.md)

- [Official Documentation](https://gsd-build-get-shit-done.mintlify.app)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

</description>
      <category>programming</category>
      <category>ai</category>
      <category>typescript</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>OpenSSL: Essential Toolkit for Secure Communications</title>
      <dc:creator>DiFlowrin</dc:creator>
      <pubDate>Sun, 03 May 2026 15:26:08 +0000</pubDate>
      <link>https://dev.to/diflowrin/openssl-essential-toolkit-for-secure-communications-3n40</link>
      <guid>https://dev.to/diflowrin/openssl-essential-toolkit-for-secure-communications-3n40</guid>
      <description>&lt;h2&gt;
  
  
  Why OpenSSL Matters Now in Open Source Security
&lt;/h2&gt;

&lt;p&gt;The importance of &lt;strong&gt;OpenSSL&lt;/strong&gt; cannot be overstated, especially as the landscape of web security continues to evolve rapidly. With the increasing number of cyber threats, understanding how to implement secure communication protocols is crucial for developers and organizations alike. OpenSSL serves as a vital tool in this arena, providing a comprehensive &lt;strong&gt;SSL/TLS toolkit&lt;/strong&gt; that facilitates encrypted communications across the internet.&lt;/p&gt;

&lt;p&gt;As the backbone of secure traffic on the web, OpenSSL enables the establishment of encrypted channels that protect sensitive data from eavesdropping and tampering. This is increasingly vital as more businesses shift to online operations. Misunderstandings about OpenSSL often arise; many developers know it exists but are unaware of its extensive capabilities and the nuances of its usage. Understanding OpenSSL can empower developers to build safer applications and ensure privacy for their users.&lt;/p&gt;

&lt;h3&gt;
  
  
  📹 Video: Install OpenSSL on Windows 10/11: Easy Tutorial
&lt;/h3&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/6zpBKVLox34"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Video credit: OurTechRoom&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  How OpenSSL Works: Mechanisms Behind the Magic
&lt;/h2&gt;

&lt;p&gt;At its core, OpenSSL is a robust library that provides implementations of the &lt;strong&gt;TLS&lt;/strong&gt; and &lt;strong&gt;SSL&lt;/strong&gt; protocols. It encapsulates a wide range of cryptographic functions, making it an essential &lt;strong&gt;encryption library&lt;/strong&gt; for developers. The library supports various cryptographic algorithms, including symmetric and asymmetric cryptography, hashing functions, and digital signature verification.&lt;/p&gt;

&lt;p&gt;OpenSSL operates through a command-line interface and offers a rich set of APIs. When you generate an &lt;strong&gt;SSL certificate&lt;/strong&gt; using OpenSSL, it involves creating a private key, generating a certificate signing request (CSR), and obtaining a signed certificate from a &lt;strong&gt;certificate authority&lt;/strong&gt; (CA). This process ensures that the identity of the communicating parties is authenticated and that the data being exchanged is encrypted.&lt;/p&gt;

&lt;p&gt;Moreover, OpenSSL also implements the &lt;strong&gt;DTLS protocol&lt;/strong&gt; for datagram-based applications and supports &lt;strong&gt;QUIC&lt;/strong&gt;, a modern transport layer network protocol designed for performance and security. This versatility makes OpenSSL not just a library for web servers but a critical component for a multitude of applications that require secure communications.&lt;/p&gt;

&lt;h2&gt;
  
  
  Real Benefits of Using OpenSSL
&lt;/h2&gt;

&lt;p&gt;One of the biggest advantages of OpenSSL is its open-source nature, allowing developers to contribute to its improvement and adapt it for their specific needs. This community-driven aspect fosters innovation and ensures that OpenSSL remains up-to-date with the latest security standards.&lt;/p&gt;

&lt;p&gt;OpenSSL provides a straightforward method for implementing &lt;strong&gt;secure communication protocols&lt;/strong&gt;. By utilizing OpenSSL, developers can easily set up HTTPS, ensuring that data transmitted between clients and servers is encrypted. This is particularly beneficial for applications handling sensitive user information, such as financial applications or personal data.&lt;/p&gt;

&lt;p&gt;Additionally, OpenSSL’s extensive documentation and community support make it accessible for both novice and experienced developers. Whether you’re looking for an &lt;strong&gt;OpenSSL installation and setup&lt;/strong&gt; guide or need help with specific cryptographic functions, chances are someone in the community has tackled a similar issue.&lt;/p&gt;

&lt;h2&gt;
  
  
  Practical Examples: Workflows with OpenSSL
&lt;/h2&gt;

&lt;p&gt;Let’s look at some practical examples to illustrate how you can work with OpenSSL effectively. First, installing OpenSSL can vary depending on your operating system:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Windows:&lt;/strong&gt; You can download the latest version of OpenSSL from a trusted source and follow the installer.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Mac:&lt;/strong&gt; Installing through Homebrew is straightforward: brew install openssl.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Linux:&lt;/strong&gt; Use your package manager, e.g., sudo apt-get install openssl for Debian-based systems.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Once installed, generating an SSL certificate is a common task. Here’s a brief &lt;strong&gt;OpenSSL certificate generation tutorial&lt;/strong&gt;:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;Generate a private key:&lt;br&gt;
openssl genrsa -out mydomain.key 2048&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Create a certificate signing request (CSR):&lt;br&gt;
openssl req -new -key mydomain.key -out mydomain.csr&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Self-sign the certificate (for testing purposes):&lt;br&gt;
openssl x509 -req -days 365 -in mydomain.csr -signkey mydomain.key -out mydomain.crt&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This workflow illustrates how you can quickly generate a self-signed certificate, which is useful for development and testing. For production, you would send the CSR to a CA for signing.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F5pkngj7urj7t9ljxkmmo.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F5pkngj7urj7t9ljxkmmo.webp" alt="OpenSSL: Essential Toolkit for Secure Communications" width="800" height="457"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Another important aspect is using the OpenSSL command line for various cryptographic tasks. For instance, to encrypt a file, you can use:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;openssl enc &lt;span class="nt"&gt;-aes-256-cbc&lt;/span&gt; &lt;span class="nt"&gt;-salt&lt;/span&gt; &lt;span class="nt"&gt;-in&lt;/span&gt; file.txt &lt;span class="nt"&gt;-out&lt;/span&gt; file.enc
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And to decrypt it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;openssl enc &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="nt"&gt;-aes-256-cbc&lt;/span&gt; &lt;span class="nt"&gt;-in&lt;/span&gt; file.enc &lt;span class="nt"&gt;-out&lt;/span&gt; file.txt
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  What’s Next for OpenSSL: Future Directions and Limitations
&lt;/h2&gt;

&lt;p&gt;The future of OpenSSL is bright, especially with the growing emphasis on security in the tech landscape. As threats become more sophisticated, OpenSSL will continue to evolve to meet the demands of modern security protocols. Developers should keep an eye on updates from the OpenSSL community, particularly regarding support for new cryptographic standards and protocols.&lt;/p&gt;

&lt;p&gt;However, there are limitations to be aware of. While OpenSSL is incredibly powerful, it’s essential to stay updated with its security patches. Vulnerabilities can arise, as seen in past incidents, which underscores the need for regular updates and adherence to security best practices.&lt;/p&gt;

&lt;p&gt;Additionally, understanding the fundamentals of cryptography is crucial when using OpenSSL. Misconfigurations can lead to severe security flaws, so developers must invest time in learning about cryptographic principles, certificate management, and secure setup practices.&lt;/p&gt;

&lt;h2&gt;
  
  
  People Also Ask
&lt;/h2&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;### What is OpenSSL and what does it do?

OpenSSL is an open-source software library that provides tools and protocols for implementing secure communications over networks. It supports various cryptographic functions and is widely used for generating SSL/TLS certificates.



### How do I install OpenSSL on Windows, Mac, or Linux?

Installation varies by operating system. For Windows, download the installer; for Mac, use Homebrew; for Linux, employ your package manager, e.g., `sudo apt-get install openssl`.



### How do I generate an SSL certificate with OpenSSL?

You can generate an SSL certificate using OpenSSL by creating a private key, generating a CSR, and then self-signing the certificate or having it signed by a CA.



### What is the difference between SSL and TLS?

SSL (Secure Sockets Layer) is the predecessor to TLS (Transport Layer Security). TLS is a more secure and updated version of SSL, and it is recommended for secure communications.



### How do I clone the OpenSSL GitHub repository?

You can clone the OpenSSL GitHub repository by running the command `git clone https://github.com/openssl/openssl.git` in your terminal.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;
&lt;h2&gt;
  
  
  Sources &amp;amp; References
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Original Source:&lt;/strong&gt; &lt;a href="https://github.com/openssl/openssl" rel="noopener noreferrer"&gt;https://github.com/openssl/openssl&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;### Additional Resources&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;- [OpenSSL Official Website](https://www.openssl.org/)

- [OpenSSL GitHub Repository](https://github.com/openssl/openssl)

- [OpenSSL Documentation and Wiki](https://wiki.openssl.org/)

- [OpenSSL Source Downloads](https://openssl-library.org/source/)

- [OpenSSL Git Repository Information](https://wiki.openssl.org/index.php/Use_of_Git)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

</description>
      <category>webdev</category>
      <category>programming</category>
      <category>tutorial</category>
      <category>beginners</category>
    </item>
    <item>
      <title>MoltUI: Innovative Terminal Molecular Visualization Tool</title>
      <dc:creator>DiFlowrin</dc:creator>
      <pubDate>Sat, 02 May 2026 15:26:10 +0000</pubDate>
      <link>https://dev.to/diflowrin/moltui-innovative-terminal-molecular-visualization-tool-4df3</link>
      <guid>https://dev.to/diflowrin/moltui-innovative-terminal-molecular-visualization-tool-4df3</guid>
      <description>&lt;h2&gt;
  
  
  Executive Summary
&lt;/h2&gt;

&lt;p&gt;MoltUI is an open-source terminal molecular viewer that leverages Unicode to provide a unique way to visualize molecular structures directly in the command line. This tool stands out for its ability to render complex molecules in a straightforward manner, making it accessible for both experienced chemists and developers alike. As molecular visualization becomes increasingly critical in scientific research, tools like MoltUI are paving the way for efficient and effective data representation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why MoltUI Matters Now
&lt;/h2&gt;

&lt;p&gt;The rise of computational chemistry and molecular biology has created a pressing need for tools that can visualize molecular structures quickly and effectively. Traditional graphical interfaces often require significant resources and can be cumbersome to use, especially in a terminal-based workflow. MoltUI fills this gap by providing a &lt;strong&gt;terminal molecular viewer&lt;/strong&gt; based on Unicode, enabling users to visualize molecular structures without the overhead of a full graphical application.&lt;/p&gt;

&lt;p&gt;Moreover, the growth of remote work and server-based computations has made terminal applications more relevant than ever. Researchers and developers can now run &lt;strong&gt;molecular visualization terminal&lt;/strong&gt; tools like MoltUI on remote servers, allowing for quick and easy access to molecular data across various platforms.&lt;/p&gt;

&lt;h2&gt;
  
  
  How MoltUI Works
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Understanding the Mechanism of MoltUI
&lt;/h3&gt;

&lt;p&gt;MoltUI operates by converting molecular data from various formats, such as SMILES or MOL files, into a visual representation that can be displayed in the terminal. It uses Unicode characters to represent atoms and bonds, which allows for a surprisingly detailed depiction of molecular structures with minimal resource usage.&lt;/p&gt;

&lt;p&gt;When you run MoltUI, it parses the input file, extracts the necessary structural information, and then formats it into ASCII art. The choice of Unicode makes it possible to depict complex molecules in a visually appealing way, while still being text-based. This approach not only makes it lightweight but also ensures that it can be used in environments where graphical output is not feasible.&lt;/p&gt;

&lt;h3&gt;
  
  
  Installation and Setup
&lt;/h3&gt;

&lt;p&gt;Installing MoltUI is straightforward. It can be done via Python's package manager, &lt;strong&gt;pip&lt;/strong&gt;, which is a familiar tool for many developers:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;moltui
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After installation, users can initiate the viewer by simply running:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;moltui path/to/your/molecule.mol
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This command will launch the viewer and display the molecular structure right in your terminal.&lt;/p&gt;

&lt;h2&gt;
  
  
  Real Benefits of Using MoltUI
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Accessibility and Efficiency
&lt;/h3&gt;

&lt;p&gt;The primary benefit of using MoltUI is its accessibility. By providing a &lt;strong&gt;Unicode-based molecular visualization&lt;/strong&gt;, it allows scientists and developers who may not have access to high-end graphical tools to visualize molecular data effectively. This can be particularly beneficial in educational settings, where students can learn about molecular structures without needing extensive software installations.&lt;/p&gt;

&lt;p&gt;Additionally, the efficiency of working within a terminal cannot be overstated. Users can quickly visualize structures as part of their workflows without switching contexts or waiting for graphical interfaces to load. This is crucial for researchers who often work with large datasets and need to iterate rapidly.&lt;/p&gt;

&lt;h3&gt;
  
  
  Integration with Other Tools
&lt;/h3&gt;

&lt;p&gt;MoltUI is designed to integrate well with other command-line tools, making it a flexible part of a larger data analysis workflow. For example, it can be easily combined with other Python-based tools for data analysis or molecular simulations. This integration capability enhances its usability and allows for the automation of tasks, which is a significant advantage for developers looking to streamline their processes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Practical Examples of Using MoltUI
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Visualizing Different Molecules
&lt;/h3&gt;

&lt;p&gt;One of the most straightforward applications of MoltUI is its use in educational settings. For instance, a chemistry student can visualize the structure of common molecules like water (H2O) or glucose (C6H12O6) by simply inputting the corresponding MOL file:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F7ptwd2gkak9skmllh429.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F7ptwd2gkak9skmllh429.webp" alt="MoltUI: Innovative Terminal Molecular Visualization Tool" width="800" height="457"&gt;&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;moltui glucose.mol
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This command will yield a clear representation of the glucose molecule, allowing the student to explore its structure interactively.&lt;/p&gt;

&lt;h3&gt;
  
  
  Integrating MoltUI into a Workflow
&lt;/h3&gt;

&lt;p&gt;Consider a scenario in which a researcher is analyzing the outputs of a molecular dynamics simulation. The researcher can use MoltUI to visualize each frame of the simulation in real-time:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;frame&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;simulation_frames&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;moltui&lt;/span&gt; &lt;span class="n"&gt;frame&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;mol&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This approach enables the researcher to quickly identify structural changes over time, facilitating a more in-depth analysis of the molecular behavior.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's Next for MoltUI?
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Future Developments and Limitations
&lt;/h3&gt;

&lt;p&gt;As with any tool, there are always opportunities for improvement. While MoltUI is excellent for basic visualizations, advanced features such as interactive viewing or dynamic updates while manipulating molecular structures could enhance its functionality. The developer community can contribute to these enhancements, making it a collaborative effort that aligns with the spirit of open-source development.&lt;/p&gt;

&lt;p&gt;Furthermore, while MoltUI currently supports a variety of molecular formats, expanding its compatibility to include more complex formats and 3D visualizations could greatly enhance its usability. This would enable users to work with a wider range of molecular data, catering to more advanced research needs.&lt;/p&gt;

&lt;h2&gt;
  
  
  People Also Ask
&lt;/h2&gt;

&lt;h3&gt;
  
  
  What is MoltUI?
&lt;/h3&gt;

&lt;p&gt;MoltUI is an open-source terminal molecular viewer that allows users to visualize molecular structures using Unicode characters directly in the command line. It simplifies the process of molecular visualization, making it accessible for both researchers and developers.&lt;/p&gt;

&lt;h3&gt;
  
  
  How to install MoltUI terminal viewer?
&lt;/h3&gt;

&lt;p&gt;To install MoltUI, use the Python package manager pip. Simply run the command &lt;code&gt;pip install moltui&lt;/code&gt; in your terminal, and you will be able to visualize molecular structures by executing &lt;code&gt;moltui path/to/your/molecule.mol&lt;/code&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  What is a terminal molecular viewer?
&lt;/h3&gt;

&lt;p&gt;A terminal molecular viewer is a tool that allows users to visualize molecular structures directly in a command-line interface, using text and Unicode characters instead of graphical displays. This approach is lightweight and efficient, especially in environments where graphical output is not available.&lt;/p&gt;

&lt;h3&gt;
  
  
  Is MoltUI based on Unicode?
&lt;/h3&gt;

&lt;p&gt;Yes, MoltUI utilizes Unicode characters to represent atoms and bonds, enabling it to create detailed molecular visualizations that can be displayed in a terminal.&lt;/p&gt;

&lt;h3&gt;
  
  
  How does MoltUI visualize molecules?
&lt;/h3&gt;

&lt;p&gt;MoltUI visualizes molecules by parsing molecular data files and converting that data into a text-based representation using ASCII art and Unicode characters, allowing for effective representation of molecular structures in the terminal.&lt;/p&gt;

&lt;h2&gt;
  
  
  📊 Key Findings &amp;amp; Takeaways
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Accessibility:&lt;/strong&gt; MoltUI democratizes molecular visualization, making it accessible to users without high-end graphical tools.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Efficiency:&lt;/strong&gt; The terminal-based approach allows for rapid visualizations, which is crucial for iterative workflows.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Integration Potential:&lt;/strong&gt; Its compatibility with other command-line tools enhances its utility in data analysis pipelines.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Future Opportunities:&lt;/strong&gt; Expanding features and format compatibility could significantly enhance MoltUI's capabilities.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Sources &amp;amp; References
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Original Source:&lt;/strong&gt; &lt;a href="https://github.com/kszenes/moltui" rel="noopener noreferrer"&gt;https://github.com/kszenes/moltui&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;### Additional Resources&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;- [MoltUI GitHub Repository](https://github.com/kszenes/moltui)

- [MoltUI - Terminal Molecular Viewer](https://github.com/kszenes/moltui)

- [Molty GitHub Topic](https://github.com/topics/molty)

- [Molt GitHub Topics](https://github.com/topics/molt?l=javascript)

- [Uranusjr Molt Python Manager](https://github.com/uranusjr/molt)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

</description>
      <category>webdev</category>
      <category>programming</category>
      <category>tutorial</category>
      <category>beginners</category>
    </item>
  </channel>
</rss>
