<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Samod Alex</title>
    <description>The latest articles on DEV Community by Samod Alex (@samod_alex).</description>
    <link>https://dev.to/samod_alex</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4097888%2Fcdac303a-3b80-4735-9e1b-cf49458cd949.png</url>
      <title>DEV Community: Samod Alex</title>
      <link>https://dev.to/samod_alex</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/samod_alex"/>
    <language>en</language>
    <item>
      <title>ChatGPT and Grok Just Joined the Pentagon's AI Platform. Claude Didn't.</title>
      <dc:creator>Samod Alex</dc:creator>
      <pubDate>Thu, 03 Sep 2026 04:37:58 +0000</pubDate>
      <link>https://dev.to/samod_alex/chatgpt-and-grok-just-joined-the-pentagons-ai-platform-claude-didnt-21p1</link>
      <guid>https://dev.to/samod_alex/chatgpt-and-grok-just-joined-the-pentagons-ai-platform-claude-didnt-21p1</guid>
      <description>&lt;p&gt;The Department of War (the renamed Department of Defense) &lt;a href="https://techcrunch.com/2026/08/31/the-pentagon-now-has-its-own-version-of-chatgpt-and-grok/" rel="noopener noreferrer"&gt;rolled out custom versions of ChatGPT and Grok&lt;/a&gt; on its internal AI platform this week. More than three million civilian and military personnel now have access to tools built for what the department calls “warfighter needs.”&lt;/p&gt;

&lt;p&gt;Google's Gemini was already part of that ecosystem. Anthropic's Claude isn't — and given the Pentagon's very public fight with Anthropic over how AI can be used, that absence is difficult to treat as incidental.&lt;/p&gt;

&lt;h2&gt;
  
  
  What just launched
&lt;/h2&gt;

&lt;p&gt;GenAI.mil is the Pentagon's secure enterprise AI platform, built to give personnel access to frontier models inside authorized government cloud infrastructure rather than sending government work through ordinary consumer services. Google launched Gemini for Government on the platform in December 2025, making it the first enterprise AI tool available through GenAI.mil.&lt;/p&gt;

&lt;p&gt;On August 31, two more tools joined it: ChatGPT Mil, built through &lt;a href="https://openai.com/index/bringing-chatgpt-to-genaimil/" rel="noopener noreferrer"&gt;OpenAI's government program&lt;/a&gt;, and Grok for Government, delivered under SpaceX's Starshield AI brand. The Pentagon says more than 1.7 million unique users have signed up since GenAI.mil launched nine months ago, out of more than three million personnel — roughly 57% of the department.&lt;/p&gt;

&lt;p&gt;That's a striking adoption rate for a platform operating across one of the world's largest organizations.&lt;/p&gt;

&lt;h2&gt;
  
  
  ChatGPT Mil: the familiar option
&lt;/h2&gt;

&lt;p&gt;ChatGPT Mil looks deliberately close to the product most developers already know: chat, file uploads, projects, and custom GPTs. The Pentagon says it is intended for document-heavy unclassified work including planning, policy, logistics, administration, procurement, research, and mission support.&lt;/p&gt;

&lt;p&gt;OpenAI says the deployment runs in authorized government cloud infrastructure, with model- and platform-level safeguards, and that data processed on GenAI.mil remains isolated from OpenAI's public and commercial model-training systems.&lt;/p&gt;

&lt;p&gt;More capabilities are scheduled to arrive over time. For now, the pitch is relatively straightforward: take the familiar ChatGPT experience and put it inside a government-controlled environment.&lt;/p&gt;

&lt;h2&gt;
  
  
  Grok for Government: a different pitch
&lt;/h2&gt;

&lt;p&gt;Grok's positioning is more operational. The Pentagon says Starshield AI's Grok for Government is accredited for Controlled Unclassified Information at Impact Level 5, and includes deep-thinking inference, three reasoning modes — Auto, Fast, and Expert — along with customizable workspaces, persistent projects, and reusable “playbooks” designed to preserve institutional knowledge as personnel rotate.&lt;/p&gt;

&lt;p&gt;Its stated use cases range from market research for acquisition professionals to supply-chain management for logisticians. In other words, the department is pitching Grok not simply as a chatbot, but as an enterprise reasoning and workflow tool that can be embedded into recurring military and administrative processes.&lt;/p&gt;

&lt;h2&gt;
  
  
  The tool that's missing
&lt;/h2&gt;

&lt;p&gt;Here's what makes this more interesting than another government AI launch: Claude isn't on GenAI.mil.&lt;/p&gt;

&lt;p&gt;Gemini, ChatGPT, and Grok are now part of the platform's frontier-model lineup. Anthropic isn't.&lt;/p&gt;

&lt;p&gt;The Pentagon has not publicly said that its dispute with Anthropic is the specific reason Claude is absent from GenAI.mil. But the timing matters: the department spent much of 2026 fighting Anthropic over precisely the restrictions that the company wanted attached to government use of Claude.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Anthropic said no
&lt;/h2&gt;

&lt;p&gt;Back in February 2026, Anthropic and Defense Secretary Pete Hegseth were in a public standoff. Anthropic said negotiations had reached an impasse over two narrow restrictions: using Claude for mass domestic surveillance of Americans and using it in fully autonomous weapons. Anthropic argued that today's frontier models were not reliable enough for fully autonomous lethal weapons, while also saying it supported lawful national-security uses outside those two exceptions.&lt;/p&gt;

&lt;p&gt;Hegseth's position was fundamentally different. The Pentagon argued that the department, not a private AI company, should determine what lawful military applications its technology could support. On February 27, Hegseth directed the department to designate Anthropic a supply-chain risk to national security, following President Trump's order that federal agencies stop using Anthropic's technology.&lt;/p&gt;

&lt;p&gt;Anthropic sued in California and Washington, D.C.&lt;/p&gt;

&lt;p&gt;The litigation then split into two very different tracks. In California, Judge Rita Lin issued a preliminary injunction on March 26 blocking the government from implementing the presidential directive and the supply-chain designation while the case proceeded. The order explicitly said it did not prevent the Department of War from moving to other AI providers, provided those actions complied with applicable law.&lt;/p&gt;

&lt;p&gt;The parallel D.C. litigation initially went the other way: the D.C. Circuit declined to block the Pentagon's action while that case continued.&lt;/p&gt;

&lt;p&gt;Then, on August 27, the California case produced the much bigger development. Judge Lin ruled on the merits of Anthropic's challenge, granting Anthropic summary judgment on key claims against the Department of War and finding the challenged designation unlawful. The court also rejected the government's request to pause the effect of the ruling.&lt;/p&gt;

&lt;p&gt;That makes the August decision Anthropic's biggest court victory yet in the dispute — but not its first court victory.&lt;/p&gt;

&lt;p&gt;The legal backdrop is also more nuanced than “the Pentagon allows AI to kill without humans.” A 2023 DoD directive permits autonomous and semi-autonomous weapon systems, including systems with automated functions for acquiring, tracking, identifying, prioritizing, and engaging targets. But the directive also requires appropriate levels of human judgment over the use of force, and its definition of semi-autonomous systems specifically preserves operator control over the decision to select individual targets or target groups for engagement.&lt;/p&gt;

&lt;p&gt;Anthropic's objection was narrower. It wanted its models excluded from fully autonomous weapons and mass domestic surveillance; it was not arguing that the United States should never develop or deploy autonomous weapons of any kind.&lt;/p&gt;

&lt;p&gt;The disagreement wasn't purely technical, either. At a January appearance at SpaceX, Hegseth said the department was building “war-ready weapons and systems, not chatbots for an Ivy League faculty lounge.”&lt;/p&gt;

&lt;h2&gt;
  
  
  The bet that paid off
&lt;/h2&gt;

&lt;p&gt;xAI made the opposite calculation.&lt;/p&gt;

&lt;p&gt;As the Pentagon-Anthropic fight escalated in February, TechCrunch reported that xAI was preparing to become classified-ready and quoted defense-focused investor Sachin Seth of Trousdale Ventures as estimating that competitors could need six to 12 months to catch up with Anthropic's capabilities. TechCrunch also reported that, given Elon Musk's public rhetoric, xAI appeared unlikely to object to giving the Pentagon broad control over how its technology was used.&lt;/p&gt;

&lt;p&gt;Six months later, Grok for Government is on GenAI.mil.&lt;/p&gt;

&lt;p&gt;That doesn't prove the Anthropic dispute directly caused Grok's deployment. But the timeline makes the strategic contrast difficult to miss: Anthropic drew a line over two uses; xAI positioned itself as willing to meet the department's requirements; and the Pentagon now has Grok inside its secure enterprise AI platform.&lt;/p&gt;

&lt;h2&gt;
  
  
  Spreading the risk on purpose
&lt;/h2&gt;

&lt;p&gt;None of this is limited to chatbots.&lt;/p&gt;

&lt;p&gt;In May, the Pentagon announced agreements with AWS, Microsoft, Nvidia, and Reflection AI to deploy AI technologies and models on classified networks. The department described the broader strategy in terms of building an architecture that prevents vendor lock-in and gives the Joint Force access to a diverse suite of AI capabilities.&lt;/p&gt;

&lt;p&gt;That context matters because the Pentagon's approach increasingly looks less like choosing one “official AI” and more like building a portfolio.&lt;/p&gt;

&lt;p&gt;Gemini covers one part of the platform. ChatGPT covers another. Grok adds another. Other vendors are being wired into higher-classification environments.&lt;/p&gt;

&lt;p&gt;The department's own language is revealing: “eliminating vendor lock” and promoting a “vibrant American AI ecosystem.”&lt;/p&gt;

&lt;p&gt;A vendor disagreement therefore doesn't have to become an infrastructure crisis. The department can add another model.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this matters beyond the Pentagon
&lt;/h2&gt;

&lt;p&gt;For developers watching from outside government contracting, the important signal isn't really the feature list.&lt;/p&gt;

&lt;p&gt;It's that &lt;strong&gt;AI safety policies can become procurement issues when a government customer wants broader control over lawful uses than a model provider is willing to grant.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The Anthropic dispute turned that question into an actual legal and commercial fight: a company imposed usage boundaries, the government rejected them, the company was designated a supply-chain risk, and a federal court later ruled that the government's action was unlawful.&lt;/p&gt;

&lt;p&gt;At the same time, the Pentagon has responded by broadening its AI supplier base rather than building its strategy around a single model provider.&lt;/p&gt;

&lt;p&gt;That's the part worth watching.&lt;/p&gt;

&lt;p&gt;The long-term question isn't whether Claude, ChatGPT, Gemini, or Grok is individually “the Pentagon's AI.”&lt;/p&gt;

&lt;p&gt;It's whether government AI procurement is moving toward a model where &lt;strong&gt;no single frontier AI company is allowed to become indispensable — and where the ability to accept or reject a vendor's safety boundaries becomes part of the competitive equation.&lt;/strong&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://zyvop.com/chatgpt-and-grok-just-joined-the-pentagon-s-ai-platform-claude-didn-t-y3wc3" rel="noopener noreferrer"&gt;ZyVOP&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;💡 For more articles like this, &lt;a href="https://zyvop.com/newsletter" rel="noopener noreferrer"&gt;subscribe to the ZyVOP newsletter&lt;/a&gt;!&lt;/p&gt;

</description>
      <category>anthropic</category>
      <category>aipolicy</category>
      <category>autonomousweapons</category>
      <category>pentagon</category>
    </item>
    <item>
      <title>Hy4 Preview: Inside Tencent's 770-Billion-Parameter Open-Weight Flagship</title>
      <dc:creator>Samod Alex</dc:creator>
      <pubDate>Tue, 01 Sep 2026 06:55:31 +0000</pubDate>
      <link>https://dev.to/samod_alex/hy4-preview-inside-tencents-770-billion-parameter-open-weight-flagship-5ga4</link>
      <guid>https://dev.to/samod_alex/hy4-preview-inside-tencents-770-billion-parameter-open-weight-flagship-5ga4</guid>
      <description>&lt;p&gt;&lt;em&gt;A sharper attention mechanism, four residual streams, and aggressive API pricing — but a model this large still puts self-hosting firmly in data-center territory.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;On August 28, Tencent's Hy team released &lt;strong&gt;Hy4 Preview&lt;/strong&gt;, a 770-billion-parameter mixture-of-experts language model under the Apache 2.0 license. It is more than twice the size of Hy3, expands the advertised context window from 256K to &lt;strong&gt;1 million tokens&lt;/strong&gt;, and ships as a roughly &lt;strong&gt;1.56TB BF16 checkpoint&lt;/strong&gt;. &lt;a href="https://www.tencent.com/tencent-releases-and-open-sources-tencent-hy4-preview/?utm_source=chatgpt.com" rel="noopener noreferrer"&gt;Tencent's official Hy4 announcement&lt;/a&gt; &lt;a href="https://huggingface.co/tencent/Hy4-preview?utm_source=chatgpt.com" rel="noopener noreferrer"&gt;Hy4 Preview on Hugging Face&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The headline is 770 billion parameters. The more interesting story is how Tencent is trying to make those parameters practical: by activating only a fraction of them per token, making long-context attention sparse, expanding the residual pathway, and building speculative decoding directly into the checkpoint.&lt;/p&gt;

&lt;p&gt;It is also a useful case study in a distinction that is getting harder to ignore: &lt;strong&gt;open weights do not necessarily mean accessible hardware&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The generation-over-generation jump
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Specification&lt;/th&gt;
&lt;th&gt;Hy3&lt;/th&gt;
&lt;th&gt;Hy4 Preview&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Backbone parameters&lt;/td&gt;
&lt;td&gt;295B&lt;/td&gt;
&lt;td&gt;770B&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Active backbone parameters&lt;/td&gt;
&lt;td&gt;21B&lt;/td&gt;
&lt;td&gt;49B&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MTP parameters&lt;/td&gt;
&lt;td&gt;3.8B&lt;/td&gt;
&lt;td&gt;10B&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Context window&lt;/td&gt;
&lt;td&gt;256K&lt;/td&gt;
&lt;td&gt;1M&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model weights&lt;/td&gt;
&lt;td&gt;~598GB BF16&lt;/td&gt;
&lt;td&gt;~1.56TB BF16&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Tencent's Hy3 release lists 295B total parameters, 21B active parameters, a 3.8B MTP layer and a 256K context window. Hy4 moves to a 770B backbone with 49B active parameters, a 10B MTP layer and a 1M-token context. The Hy4 specification table explicitly excludes the MTP layer from the 770B headline figure, so the auxiliary prediction module should not be silently mixed into that backbone comparison. &lt;a href="https://github.com/Tencent-Hunyuan/Hy3?utm_source=chatgpt.com" rel="noopener noreferrer"&gt;Hy3 official repository&lt;/a&gt; &lt;a href="https://github.com/Tencent-Hunyuan/Hy4-preview?utm_source=chatgpt.com" rel="noopener noreferrer"&gt;Hy4 official repository&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Hy4's 770B parameters are spread across &lt;strong&gt;78 layers&lt;/strong&gt;. The first layer uses a dense feed-forward network; the remaining 77 use MoE blocks with &lt;strong&gt;256 routed experts and one shared expert&lt;/strong&gt;. Each token activates &lt;strong&gt;eight routed experts plus the shared expert&lt;/strong&gt;, producing the 49B active-parameter figure. &lt;a href="https://github.com/Tencent-Hunyuan/Hy4-preview?utm_source=chatgpt.com#model-introduction" rel="noopener noreferrer"&gt;Hy4 architecture and model specifications&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;So Hy4 is not simply "a 770B model." It is a model with an enormous capacity pool whose &lt;strong&gt;per-token computation is deliberately much smaller than the headline parameter count suggests&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What changed under the hood
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Gated sparse attention for million-token context
&lt;/h3&gt;

&lt;p&gt;The most consequential architectural change is the attention mechanism.&lt;/p&gt;

&lt;p&gt;In conventional full attention, the amount of pairwise interaction grows rapidly as the context grows. DeepSeek Sparse Attention, or &lt;strong&gt;DSA&lt;/strong&gt;, reduces that burden by using a lightweight indexer to identify relevant positions and then applying the expensive attention calculation to a selected subset rather than the entire context. DeepSeek introduced DSA in the experimental V3.2-Exp release in September 2025 specifically to improve long-context training and inference efficiency. &lt;a href="https://deepseek.com/news/v3-2-exp/?utm_source=chatgpt.com" rel="noopener noreferrer"&gt;DeepSeek V3.2-Exp announcement&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Hy4 uses a &lt;strong&gt;Gated DSA&lt;/strong&gt; variant and combines it with &lt;strong&gt;IndexCache&lt;/strong&gt;, which Tencent describes as enabling cross-layer sparse-index reuse. Its published configuration sets the sparse index top-k to &lt;strong&gt;2,048&lt;/strong&gt;. &lt;a href="https://github.com/Tencent-Hunyuan/Hy4-preview?utm_source=chatgpt.com#model-specifications" rel="noopener noreferrer"&gt;Hy4 technical configuration&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The distinction matters because DSA does &lt;strong&gt;not&lt;/strong&gt; simply turn quadratic attention into linear attention. The sparse selection mechanism still has computational cost. What changes is the expensive full-attention stage: instead of computing attention over the entire context, the model operates on a much smaller selected set.&lt;/p&gt;

&lt;p&gt;The IndexCache work tackles another part of the problem. Its research paper argues that sparse indices produced by neighboring layers are sufficiently similar that rebuilding them independently is wasteful; reusing those indices can remove a large share of indexer computation while retaining model quality in their experiments. &lt;a href="https://arxiv.org/abs/2603.12201?utm_source=chatgpt.com" rel="noopener noreferrer"&gt;IndexCache research paper&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Tencent explicitly says Hy4 uses IndexCache for &lt;strong&gt;cross-layer sparse-index reuse&lt;/strong&gt;. It has not, however, published a full technical report explaining every detail of its Gated DSA implementation. That means any more specific explanation of what the gate learns should be treated as an inference, not a documented fact.&lt;/p&gt;

&lt;p&gt;The practical takeaway is simpler: &lt;strong&gt;Hy4's 1M-token context is backed by an attention architecture designed to avoid paying the full dense-attention cost across that entire window.&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Four residual streams
&lt;/h3&gt;

&lt;p&gt;Hy4 also changes the residual pathway.&lt;/p&gt;

&lt;p&gt;Standard transformers maintain a single residual stream through the network. Hyper-Connections expand that into multiple streams and learn how information is routed between them. The original Hyper-Connections work proposed the approach as a way of addressing limitations associated with conventional residual connections, including a trade-off between gradient behavior and representation collapse. &lt;a href="https://arxiv.org/abs/2409.19606?utm_source=chatgpt.com" rel="noopener noreferrer"&gt;Hyper-Connections paper&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Hy4 uses &lt;strong&gt;four residual streams&lt;/strong&gt; in an implementation Tencent calls &lt;strong&gt;identity Hyper-Connections (iHC)&lt;/strong&gt;. &lt;a href="https://github.com/Tencent-Hunyuan/Hy4-preview?utm_source=chatgpt.com#model-specifications" rel="noopener noreferrer"&gt;Hy4 model specifications&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That gives the network more pathways through which information can move between layers, but it also introduces a new optimization problem: having multiple streams does not guarantee that a trained model will use all of them equally.&lt;/p&gt;

&lt;p&gt;Recent work has investigated this kind of &lt;strong&gt;stream dominance&lt;/strong&gt;, making Hy4's identity-constrained implementation interesting without implying that the broader stability question has been settled. In other words, Tencent is deploying a research direction that is still being actively understood rather than dropping in a universally established replacement for the residual stream. &lt;a href="https://arxiv.org/abs/2409.19606?utm_source=chatgpt.com" rel="noopener noreferrer"&gt;Hyper-Connections research&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  A built-in draft model for speculative decoding
&lt;/h3&gt;

&lt;p&gt;Hy4 includes a native &lt;strong&gt;multi-token prediction (MTP)&lt;/strong&gt; layer specifically intended to support speculative decoding. Tencent says the MTP layer contains &lt;strong&gt;10B parameters in total, with 0.7B activated&lt;/strong&gt;. &lt;a href="https://github.com/Tencent-Hunyuan/Hy4-preview?utm_source=chatgpt.com#model-introduction" rel="noopener noreferrer"&gt;Hy4 official architecture details&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Speculative decoding works by letting a cheaper predictive component propose several future tokens, after which the main model verifies them. Accepted predictions allow the system to produce more output per expensive main-model step.&lt;/p&gt;

&lt;p&gt;The interesting part of Hy4 is that the MTP machinery is part of the released model rather than requiring users to supply an external draft model. Tencent's deployment examples explicitly enable MTP in both &lt;strong&gt;vLLM and SGLang&lt;/strong&gt;, using the Hy4 FP8 checkpoint. &lt;a href="https://github.com/Tencent-Hunyuan/Hy4-preview?utm_source=chatgpt.com#deployment" rel="noopener noreferrer"&gt;Hy4 deployment recipes&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That does not make 770B inference inexpensive. It does show that Tencent is treating serving efficiency as part of the model design itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  The benchmark result is interesting — but it is Tencent's
&lt;/h2&gt;

&lt;p&gt;Tencent's main comparative result comes from a &lt;strong&gt;blind side-by-side evaluation&lt;/strong&gt;, not an independently administered leaderboard.&lt;/p&gt;

&lt;p&gt;The company says &lt;strong&gt;163 internal experts&lt;/strong&gt; rated outputs on &lt;strong&gt;203 engineering tasks&lt;/strong&gt;. Hy4 averaged &lt;strong&gt;2.99/4&lt;/strong&gt;, compared with &lt;strong&gt;2.92 for GLM 5.3&lt;/strong&gt; and &lt;strong&gt;2.94 for Kimi K3&lt;/strong&gt;. Tencent reports a 46.8% win rate against GLM 5.3 and a 51.2% win rate against Kimi K3, with the remaining results split between ties and losses. &lt;a href="https://github.com/Tencent-Hunyuan/Hy4-preview?utm_source=chatgpt.com#built-for-productivity" rel="noopener noreferrer"&gt;Hy4 benchmark details&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That is evidence that Tencent considers Hy4 competitive with those models. It is not, by itself, proof of a definitive ranking.&lt;/p&gt;

&lt;p&gt;The difference is important because the score gaps are small and the evaluation is designed by the model's developer. A different task distribution, evaluator pool or scoring methodology could produce a different ordering.&lt;/p&gt;

&lt;p&gt;For now, the fairest description is that &lt;strong&gt;Hy4 has a strong self-reported showing at the open-model frontier&lt;/strong&gt;, while the independent benchmark picture is still developing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Price is a more tangible advantage
&lt;/h2&gt;

&lt;p&gt;Hy4's pricing is unusually aggressive.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Input / 1M tokens&lt;/th&gt;
&lt;th&gt;Cached input / 1M&lt;/th&gt;
&lt;th&gt;Output / 1M&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Hy4 Preview&lt;/td&gt;
&lt;td&gt;$0.834&lt;/td&gt;
&lt;td&gt;$0.042&lt;/td&gt;
&lt;td&gt;$2.501&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GLM 5.3&lt;/td&gt;
&lt;td&gt;$1.40&lt;/td&gt;
&lt;td&gt;$0.26&lt;/td&gt;
&lt;td&gt;$4.40&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kimi K3&lt;/td&gt;
&lt;td&gt;$3.00&lt;/td&gt;
&lt;td&gt;$0.30&lt;/td&gt;
&lt;td&gt;$15.00&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Tencent lists Hy4 at &lt;strong&gt;$0.834 per million input tokens, $0.042 for cache hits and $2.501 per million output tokens&lt;/strong&gt;. &lt;a href="https://www.tencent.com/tencent-releases-and-open-sources-tencent-hy4-preview/?utm_source=chatgpt.com" rel="noopener noreferrer"&gt;Tencent Hy4 announcement and pricing&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Kimi's official K3 pricing page lists &lt;strong&gt;$0.30 cached input, $3.00 uncached input and $15 output&lt;/strong&gt;, and confirms a 1,048,576-token context window. &lt;a href="https://www.kimi.ai/resources/kimi-k3-pricing?utm_source=chatgpt.com" rel="noopener noreferrer"&gt;Kimi K3 official pricing&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;For GLM 5.3, Z.ai's current listed rates are &lt;strong&gt;$1.40 input, $0.26 cached input and $4.40 output&lt;/strong&gt;. The vendor pricing record does not itself state a context length, so that detail should not be presented as vendor-confirmed here. &lt;a href="https://www.solvency.dev/models/glm-5.3/?utm_source=chatgpt.com" rel="noopener noreferrer"&gt;GLM 5.3 pricing record with Z.ai source&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;At those list prices, Hy4 is about &lt;strong&gt;3.6× cheaper than Kimi K3 on input&lt;/strong&gt; and about &lt;strong&gt;1.7× cheaper than GLM 5.3&lt;/strong&gt;. Its output price is also much lower than either comparison.&lt;/p&gt;

&lt;p&gt;The cache price is arguably the more interesting number. Long-running agents and document-heavy applications often resend the same context repeatedly. Dropping the price from $0.834 to $0.042 per million tokens after a cache hit makes repeated reads of large contexts dramatically cheaper.&lt;/p&gt;

&lt;p&gt;That turns the million-token window from a headline specification into something closer to a &lt;strong&gt;pricing strategy&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Open weights, expensive hardware
&lt;/h2&gt;

&lt;p&gt;Hy4 is released under the &lt;strong&gt;Apache 2.0 license&lt;/strong&gt;, and Tencent has published the weights alongside deployment instructions for vLLM and SGLang. &lt;a href="https://github.com/Tencent-Hunyuan/Hy4-preview?utm_source=chatgpt.com" rel="noopener noreferrer"&gt;Hy4 license and repository&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;But the physical size of the model changes what "open" means in practice.&lt;/p&gt;

&lt;p&gt;The BF16 checkpoint is about &lt;strong&gt;1.56TB&lt;/strong&gt; on Hugging Face. Tencent also provides an FP8 version and an eight-way tensor-parallel deployment example for it. &lt;a href="https://huggingface.co/tencent/Hy4-preview?utm_source=chatgpt.com" rel="noopener noreferrer"&gt;Hy4 Hugging Face files&lt;/a&gt; &lt;a href="https://github.com/Tencent-Hunyuan/Hy4-preview?utm_source=chatgpt.com#vllm" rel="noopener noreferrer"&gt;Hy4 FP8 deployment recipe&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;So there are really two different kinds of accessibility here.&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;model artifacts are open&lt;/strong&gt;: you can download them, inspect them, modify them and deploy them under the published license.&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;compute is not ordinary&lt;/strong&gt;: the official serving path assumes a multi-GPU system, and the full BF16 checkpoint alone requires more memory than a conventional workstation can provide.&lt;/p&gt;

&lt;p&gt;For the typical developer, that makes a hosted endpoint far more realistic than buying enough hardware to serve the full model locally.&lt;/p&gt;

&lt;p&gt;The weights may be open. The infrastructure is still specialized.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Tencent actually said about Hy4 before launch
&lt;/h2&gt;

&lt;p&gt;There is a useful distinction between what Tencent announced about Hunyuan 4 and what eventually appeared in Hy4 Preview.&lt;/p&gt;

&lt;p&gt;During Tencent's August 12 earnings call, management said the company was training a larger &lt;strong&gt;Hunyuan 4&lt;/strong&gt; and expected to release it later in 2026. In the same broader discussion, executives said Tencent was also upgrading its multimodal capabilities. They did &lt;strong&gt;not&lt;/strong&gt;, in the material available from the call, explicitly say that the forthcoming Hunyuan 4 language model itself would launch as a multimodal checkpoint. &lt;a href="https://earningscalls.dev/transcripts/tencent-holdings-limited_700_earnings_call_transcript_2026-08-12?utm_source=chatgpt.com" rel="noopener noreferrer"&gt;August 12, 2026 Tencent earnings-call transcript&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That distinction matters because the released Hy4 Preview is &lt;strong&gt;text-only&lt;/strong&gt;. Simon Willison independently described it as a text-input model with no vision support after inspecting the release. &lt;a href="https://simonwillison.net/2026/Aug/29/hy4/?utm_source=chatgpt.com" rel="noopener noreferrer"&gt;Simon Willison's Hy4 analysis&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;So there is no need to frame this as a broken multimodal promise. The more defensible conclusion is simpler: &lt;strong&gt;Tencent discussed Hunyuan's multimodal work and a future Hunyuan 4, while the model that shipped on August 28 is a text model.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The chat template reveals a surprisingly simple reasoning control
&lt;/h2&gt;

&lt;p&gt;Some of the most interesting details are buried not in the model card but in the chat template.&lt;/p&gt;

&lt;p&gt;Simon Willison inspected Hy4's &lt;code&gt;chat_template.jinja&lt;/code&gt; and found that its &lt;code&gt;reasoning_effort&lt;/code&gt; parameter accepts only two values: &lt;code&gt;high&lt;/code&gt; and &lt;code&gt;no_think&lt;/code&gt;, with &lt;code&gt;high&lt;/code&gt; as the default. &lt;a href="https://simonwillison.net/2026/Aug/29/hy4/?utm_source=chatgpt.com" rel="noopener noreferrer"&gt;Hy4 chat-template analysis by Simon Willison&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That makes the interface essentially a switch rather than a graduated reasoning dial.&lt;/p&gt;

&lt;p&gt;Other current reasoning models expose multiple effort levels or explicit reasoning budgets. Hy4's released template keeps the control much simpler: deep reasoning on, or reasoning off.&lt;/p&gt;

&lt;p&gt;Willison also ran his recurring test asking a model to generate an SVG of a pelican riding a bicycle. Hy4 produced a competent vector illustration, but the reasoning trace contained clipped fragments such as:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Maybe add sunglasses? no.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Willison's observation was that the internal reasoning was somewhat truncated and ungrammatical. That does not demonstrate a capability failure. It is a small but interesting example of a broader possibility: &lt;strong&gt;hidden reasoning tokens do not necessarily need to read like polished human language if their job is simply to help the model solve the task.&lt;/strong&gt; &lt;a href="https://simonwillison.net/2026/Aug/29/hy4/?utm_source=chatgpt.com" rel="noopener noreferrer"&gt;Simon Willison's pelican test and reasoning trace&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The bigger story
&lt;/h2&gt;

&lt;p&gt;Hy4 is easy to describe as "Tencent's 770B model."&lt;/p&gt;

&lt;p&gt;That misses the more interesting point.&lt;/p&gt;

&lt;p&gt;The architecture is built around a collection of techniques that all attack the same underlying problem: &lt;strong&gt;how do you increase model capability without making every token proportionally more expensive to compute?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;MoE keeps only part of the 770B parameter pool active for each token. Gated DSA reduces the amount of expensive full attention required over long contexts. IndexCache reuses sparse-selection information across layers. Four residual streams expand the network's information pathways. Native MTP gives the serving stack another route to higher effective decoding throughput. &lt;a href="https://github.com/Tencent-Hunyuan/Hy4-preview?utm_source=chatgpt.com" rel="noopener noreferrer"&gt;Hy4 architecture overview&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The result is not a fundamentally new type of model. It is something more practical: &lt;strong&gt;a concentrated bundle of scaling and inference techniques that are increasingly becoming part of the same design problem.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;There is a second story, too.&lt;/p&gt;

&lt;p&gt;Tencent's own earnings-call commentary makes clear that it is thinking about models, products and compute as one system. Management described Hunyuan as something to be co-designed with products such as WorkBuddy and CodeBuddy, while also emphasizing investment in training larger models and providing inference capacity. &lt;a href="https://earningscalls.dev/transcripts/tencent-holdings-limited_700_earnings_call_transcript_2026-08-12?utm_source=chatgpt.com" rel="noopener noreferrer"&gt;Tencent Q2 2026 earnings-call transcript&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That helps explain the economics of Hy4. Aggressive API pricing is not separate from the architecture story. It is part of it.&lt;/p&gt;

&lt;p&gt;The weights are open. The model is huge. The context is enormous. The serving stack is heavily optimized.&lt;/p&gt;

&lt;p&gt;And the most important number may not be &lt;strong&gt;770B&lt;/strong&gt; at all.&lt;/p&gt;

&lt;p&gt;It may be the distance between &lt;strong&gt;total model capacity, active computation and the cost of turning that capacity into useful tokens&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Hy4 Preview is Tencent's latest attempt to shrink that distance.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The weights may be open. The compute bill is not.&lt;/strong&gt;&lt;/p&gt;




&lt;h3&gt;
  
  
  Sources
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Primary&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;a href="https://www.tencent.com/tencent-releases-and-open-sources-tencent-hy4-preview/?utm_source=chatgpt.com" rel="noopener noreferrer"&gt;Tencent — Hy4 Preview announcement&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;a href="https://github.com/Tencent-Hunyuan/Hy4-preview?utm_source=chatgpt.com" rel="noopener noreferrer"&gt;Tencent Hy4 Preview GitHub repository&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;a href="https://huggingface.co/tencent/Hy4-preview?utm_source=chatgpt.com" rel="noopener noreferrer"&gt;Hy4 Preview — Hugging Face model card&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;a href="https://github.com/Tencent-Hunyuan/Hy3?utm_source=chatgpt.com" rel="noopener noreferrer"&gt;Tencent Hy3 GitHub repository&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;a href="https://earningscalls.dev/transcripts/tencent-holdings-limited_700_earnings_call_transcript_2026-08-12?utm_source=chatgpt.com" rel="noopener noreferrer"&gt;Tencent August 12, 2026 earnings-call transcript&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Technical background&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;a href="https://deepseek.com/news/v3-2-exp/?utm_source=chatgpt.com" rel="noopener noreferrer"&gt;DeepSeek V3.2-Exp / DSA announcement&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;a href="https://arxiv.org/abs/2603.12201?utm_source=chatgpt.com" rel="noopener noreferrer"&gt;IndexCache research paper&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;a href="https://arxiv.org/abs/2409.19606?utm_source=chatgpt.com" rel="noopener noreferrer"&gt;Hyper-Connections research paper&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Independent analysis and pricing&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;a href="https://simonwillison.net/2026/Aug/29/hy4/?utm_source=chatgpt.com" rel="noopener noreferrer"&gt;Simon Willison — Introducing Hy4 Preview&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;a href="https://www.kimi.ai/resources/kimi-k3-pricing?utm_source=chatgpt.com" rel="noopener noreferrer"&gt;Kimi K3 official pricing&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;a href="https://www.solvency.dev/models/glm-5.3/?utm_source=chatgpt.com" rel="noopener noreferrer"&gt;GLM 5.3 pricing record, with Z.ai source&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://zyvop.com/hy4-preview-inside-tencent-s-770-billion-parameter-open-weight-flagship-6z1j4" rel="noopener noreferrer"&gt;ZyVOP&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;💡 For more articles like this, &lt;a href="https://zyvop.com/newsletter" rel="noopener noreferrer"&gt;subscribe to the ZyVOP newsletter&lt;/a&gt;!&lt;/p&gt;

</description>
      <category>largelanguagemodels</category>
      <category>hy4</category>
      <category>tencent</category>
      <category>openweightai</category>
    </item>
    <item>
      <title>The Developers Who Got Fired Built an AI CEO to Replace Their Bosses</title>
      <dc:creator>Samod Alex</dc:creator>
      <pubDate>Fri, 28 Aug 2026 12:40:06 +0000</pubDate>
      <link>https://dev.to/samod_alex/the-developers-who-got-fired-built-an-ai-ceo-to-replace-their-bosses-a10</link>
      <guid>https://dev.to/samod_alex/the-developers-who-got-fired-built-an-ai-ceo-to-replace-their-bosses-a10</guid>
      <description>&lt;p&gt;&lt;em&gt;Open Executive shot up the Hacker News front page with a premise that sounds like a joke: developers who say they lost their jobs during an “AI transformation” built an open-source system designed to do the work of a CEO and the executives around them.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The joke is easy to understand. The engineering is more interesting.&lt;/em&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The short version:&lt;/strong&gt; Open Executive is an open-source multi-agent system from SenteLabs. It presents one executive persona to the user, while routing work behind the scenes to eight specialist agents covering strategy, finance, HR, legal, operations, marketing, product, and board communications. It combines Claude models, ChromaDB-based retrieval, episodic memory, a scheduler, company-specific documents, and an evaluation suite. The project quickly attracted attention on Hacker News, where its premise triggered an argument that went well beyond the software itself.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The origin story needs a little caution.&lt;/p&gt;

&lt;p&gt;A post promoting the project says some developers were let go as part of an “AI Transformation” and subsequently built Open Executive as a response. That claim became part of the project's appeal online, but it is still best treated as the project's own account rather than as an independently established fact.&lt;/p&gt;

&lt;p&gt;The reaction, however, is very real.&lt;/p&gt;

&lt;p&gt;When checked on August 27, the Hacker News submission had climbed to hundreds of points and hundreds of comments, making it one of the site's more visible discussions that day. The comments quickly moved from the joke itself to a much harder question:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How much of executive work is actually impossible to automate?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;And once you look at what Open Executive has built, that question becomes more interesting.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Open Executive Actually Does
&lt;/h2&gt;

&lt;p&gt;Open Executive is not a chatbot with a CEO prompt pasted on top.&lt;/p&gt;

&lt;p&gt;It is a small virtual organization hiding behind one interface.&lt;/p&gt;

&lt;p&gt;The system exposes a single executive voice to the user, while eight specialist agents handle different domains:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Chief Strategy Officer&lt;/strong&gt; — competitive analysis, M&amp;amp;A, positioning, and OKRs&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Chief Financial Officer&lt;/strong&gt; — financial modeling, fundraising, unit economics, and cash flow&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Chief HR/People Officer&lt;/strong&gt; — hiring, compensation, performance, and culture&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;General Counsel&lt;/strong&gt; — contracts, IP basics, employment law, and compliance&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Chief Operating Officer&lt;/strong&gt; — process design, vendors, and operational scaling&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Chief Marketing Officer&lt;/strong&gt; — go-to-market, brand, communications, and PR&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Chief Product Officer&lt;/strong&gt; — roadmap, prioritization, and product strategy&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Board Communications Director&lt;/strong&gt; — board decks, investor updates, and governance&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The user never has to decide which specialist to call.&lt;/p&gt;

&lt;p&gt;Send a question to the Executive, and the orchestrator determines which specialists should be involved. Those agents retrieve relevant information, produce their own analysis, and the system combines the results into one response.&lt;/p&gt;

&lt;p&gt;That distinction matters.&lt;/p&gt;

&lt;p&gt;The product is not really trying to make one model pretend to be exceptionally knowledgeable. It is trying to make several specialized models behave like one coherent organization.&lt;/p&gt;

&lt;p&gt;And the system goes beyond a chat interface.&lt;/p&gt;

&lt;p&gt;The repository supports a web UI, CLI, Slack, email, Telegram, Google Chat, and Discord integrations. It also maintains episodic memory across sessions and includes a scheduler capable of surfacing follow-ups and time-sensitive actions.&lt;/p&gt;

&lt;p&gt;In other words, the goal is not simply to answer questions.&lt;/p&gt;

&lt;p&gt;It is to stay involved.&lt;/p&gt;

&lt;h2&gt;
  
  
  Under the Hood
&lt;/h2&gt;

&lt;p&gt;The architecture is surprisingly straightforward.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmermaid.ink%2Fimg%2Fpako%3AeNpVkc1u2zAMx1-F49nekKIYhmAoENtxO2xBh3jbxeqBkRlbqCwZktw0DfJcu-_JBinA2t348SP__DihtB3jEvfaHuRALsCPShgAgFUrcMPeU88CHyDPb6BoBa6fWc5BPTHcOzmwD46CdZ937sNNqWnuGBprDAe4fv9R4MOlV5HKy5PAhjXLAM3EUpFWPniBZ2EuWJmwqhX46Q0Bq55N8EmiiXLcH-HPb6iVISM5mt-4Jx2Nu23C7id2FJQ1PgY35B45KNNH57uz3SxDNAtLrkszXuSrJL9uBZaDsyNVBWxXt_-WWKd0vWgFFrPSIVcGvhp70NylE_1HXcUudpzIHKGych7jCm-06kXibluBzdGEgb164Q5er7tlP1njXxvXV5cKzHBkN5LqcHnCMPAY_9fxnmYdMLtEfpFTtNPsI7O3JtQ0Kn3EJeY0TZpzf_SBxwwKrczjhmST_NqakEH8Um8Zfn4RmMHW7mywGdyxfuKgJGWwcop0Bp6Mzz07tccsiTTqJc6yuJ6e8XzOcNeXVluHS3x3GFRgPP8F_MfN1Q" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmermaid.ink%2Fimg%2Fpako%3AeNpVkc1u2zAMx1-F49nekKIYhmAoENtxO2xBh3jbxeqBkRlbqCwZktw0DfJcu-_JBinA2t348SP__DihtB3jEvfaHuRALsCPShgAgFUrcMPeU88CHyDPb6BoBa6fWc5BPTHcOzmwD46CdZ937sNNqWnuGBprDAe4fv9R4MOlV5HKy5PAhjXLAM3EUpFWPniBZ2EuWJmwqhX46Q0Bq55N8EmiiXLcH-HPb6iVISM5mt-4Jx2Nu23C7id2FJQ1PgY35B45KNNH57uz3SxDNAtLrkszXuSrJL9uBZaDsyNVBWxXt_-WWKd0vWgFFrPSIVcGvhp70NylE_1HXcUudpzIHKGych7jCm-06kXibluBzdGEgb164Q5er7tlP1njXxvXV5cKzHBkN5LqcHnCMPAY_9fxnmYdMLtEfpFTtNPsI7O3JtQ0Kn3EJeY0TZpzf_SBxwwKrczjhmST_NqakEH8Um8Zfn4RmMHW7mywGdyxfuKgJGWwcop0Bp6Mzz07tccsiTTqJc6yuJ6e8XzOcNeXVluHS3x3GFRgPP8F_MfN1Q" alt="Mermaid Diagram" width="437" height="888"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The stack is deliberately conventional:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Choice&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;LLM backbone&lt;/td&gt;
&lt;td&gt;Anthropic Claude API&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Default model&lt;/td&gt;
&lt;td&gt;Claude Sonnet 4.6&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Deep reasoning&lt;/td&gt;
&lt;td&gt;Claude Opus 4.7&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Backend&lt;/td&gt;
&lt;td&gt;Python 3.11 + FastAPI&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Frontend&lt;/td&gt;
&lt;td&gt;Next.js 15 + Tailwind&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vector store&lt;/td&gt;
&lt;td&gt;ChromaDB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Episodic memory&lt;/td&gt;
&lt;td&gt;SQLite&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;License&lt;/td&gt;
&lt;td&gt;Apache 2.0&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The interesting part is what the project does with those components.&lt;/p&gt;

&lt;h2&gt;
  
  
  Retrieval Is Split Into Two Layers
&lt;/h2&gt;

&lt;p&gt;Each specialist can retrieve from two different sources.&lt;/p&gt;

&lt;p&gt;The first is the project's built-in knowledge base: a git-tracked collection of Markdown files intended to provide general business and MBA-style reference material.&lt;/p&gt;

&lt;p&gt;The second is the company's private knowledge.&lt;/p&gt;

&lt;p&gt;Users can upload things such as pitch decks, financial models, strategic documents, and other company material. Those documents are chunked and stored separately from the built-in knowledge.&lt;/p&gt;

&lt;p&gt;That separation is important because the system does not treat company-specific information as part of the permanent system prompt. Instead, retrieved context is injected into the user turn.&lt;/p&gt;

&lt;p&gt;That is a relatively clean design for an application that mixes shared knowledge with private business information.&lt;/p&gt;

&lt;p&gt;It also gives the project an obvious extension point: the more useful company data you provide, the more context the specialist agents can work with.&lt;/p&gt;

&lt;h2&gt;
  
  
  Different Jobs Get Different Models
&lt;/h2&gt;

&lt;p&gt;Open Executive does not send every task to its most expensive model.&lt;/p&gt;

&lt;p&gt;The Executive orchestrator and most specialists use Claude Sonnet 4.6. Four areas — strategy, finance, legal, and board communications — are assigned Claude Opus 4.7 with extended thinking.&lt;/p&gt;

&lt;p&gt;That is a sensible cost/quality tradeoff.&lt;/p&gt;

&lt;p&gt;A routing system that sends every trivial question to the strongest available model would be expensive. A system that treats financial modeling and “what should I put in this Slack message?” as equivalent would be wasteful.&lt;/p&gt;

&lt;p&gt;The interesting engineering decision is therefore not simply which model is strongest.&lt;/p&gt;

&lt;p&gt;It is deciding &lt;strong&gt;where stronger reasoning is worth paying for.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Anthropic describes Opus 4.7 as its generally available successor to Opus 4.6, with improvements in complex reasoning, software engineering, and long-running agentic work. Sonnet 4.6 is positioned as the more capable, lower-cost member of the family.&lt;/p&gt;

&lt;h2&gt;
  
  
  It Has Memory, But It Doesn't Just Keep Everything
&lt;/h2&gt;

&lt;p&gt;One of the better design choices is the memory system.&lt;/p&gt;

&lt;p&gt;After a response, a separate lightweight model pass extracts things such as decisions, initiatives, and advice and stores them in SQLite.&lt;/p&gt;

&lt;p&gt;The next session can then start with a compact record of what happened previously.&lt;/p&gt;

&lt;p&gt;That's different from simply replaying the entire conversation every time.&lt;/p&gt;

&lt;p&gt;For an executive assistant, that matters. A real executive relationship depends on continuity:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;What did we decide last month?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Why did we decide it?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;What were we supposed to follow up on?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Loading an entire historical transcript would be expensive and noisy. Compressing the useful state into structured memory is a much more practical pattern.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Scheduler Is a Small Detail That Reveals a Lot
&lt;/h2&gt;

&lt;p&gt;The scheduler might be the least glamorous feature in the repository.&lt;/p&gt;

&lt;p&gt;It is also one of the easiest places for an otherwise good agent system to fail.&lt;/p&gt;

&lt;p&gt;Open Executive can proactively surface scheduled actions and follow-ups, which means background jobs need to be claimed safely.&lt;/p&gt;

&lt;p&gt;The project handles that with an atomic database update using &lt;code&gt;UPDATE … RETURNING&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;There is also an explicit architectural constraint: the API runs as a single instance because running multiple scheduler instances without additional coordination could cause scheduled actions to fire twice.&lt;/p&gt;

&lt;p&gt;That's not a glamorous feature.&lt;/p&gt;

&lt;p&gt;It is good engineering.&lt;/p&gt;

&lt;p&gt;The repository is unusually explicit about this limitation rather than pretending horizontal scaling is already solved.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Evaluation System Is More Interesting Than the Demo
&lt;/h2&gt;

&lt;p&gt;This is probably the part of Open Executive I would pay the most attention to as an engineer.&lt;/p&gt;

&lt;p&gt;Adding another agent isn't treated as simply adding another prompt.&lt;/p&gt;

&lt;p&gt;The repository's contribution guidance requires new agent work to be wired into the architecture, accompanied by tests, and covered by evaluation scenarios.&lt;/p&gt;

&lt;p&gt;The project also has a dedicated evaluation suite that uses an LLM judge to score executive responses across dimensions including persona coherence, domain accuracy, company-context usage, routing quality, and actionability.&lt;/p&gt;

&lt;p&gt;There are currently 29 scored scenarios covering the eight domains, with regression thresholds intended to prevent quality from quietly deteriorating as the system changes.&lt;/p&gt;

&lt;p&gt;That is important because agent systems have an uncomfortable property that ordinary unit tests don't capture well:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;the code can be correct while the behavior gets worse.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A prompt change can improve one workflow and damage another. A routing change can make one specialist more useful while causing the wrong specialist to be selected elsewhere.&lt;/p&gt;

&lt;p&gt;For that reason, behavioral evaluation is not just a nice extra. It is part of the engineering surface.&lt;/p&gt;

&lt;h2&gt;
  
  
  It Doesn't Completely Lock Itself to Anthropic
&lt;/h2&gt;

&lt;p&gt;The repository is built around the Anthropic API, but it also includes a provider abstraction that allows other OpenAI-compatible endpoints and local inference setups.&lt;/p&gt;

&lt;p&gt;That means a deployment can, in principle, mix hosted and local models rather than forcing every task through one provider.&lt;/p&gt;

&lt;p&gt;The tradeoff is straightforward: local models do not automatically provide all of the same capabilities as Anthropic's hosted models, particularly around prompt caching, extended thinking, web search, and reliable tool use.&lt;/p&gt;

&lt;p&gt;That's actually a healthier way to present the feature.&lt;/p&gt;

&lt;p&gt;“Supports local models” is not the same thing as “local models are equivalent.”&lt;/p&gt;

&lt;p&gt;They're not.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Privacy Story Is Better Than the Marketing Usually Sounds
&lt;/h2&gt;

&lt;p&gt;Company profiles, uploaded documents, and vector data are intended to remain in the deployment's own storage rather than being copied into a third-party SaaS database.&lt;/p&gt;

&lt;p&gt;But there is an important qualification.&lt;/p&gt;

&lt;p&gt;The relevant context still has to be sent to the model provider when inference happens.&lt;/p&gt;

&lt;p&gt;So this is not “your data never leaves your infrastructure” in the absolute sense.&lt;/p&gt;

&lt;p&gt;It is closer to:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;your application state stays under your control, while the model calls still send the required inference context to the provider you configured.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That's a much more useful way to think about privacy for an LLM application.&lt;/p&gt;

&lt;p&gt;The repository also states that Anthropic does not train on API inputs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Then Hacker News Got Involved
&lt;/h2&gt;

&lt;p&gt;This is where the project becomes bigger than its architecture.&lt;/p&gt;

&lt;p&gt;The Hacker News submission used the title:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“CEO fired developers to make room for AI. Developers create open source AI CEO.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That framing was almost guaranteed to generate an argument.&lt;/p&gt;

&lt;p&gt;And it did.&lt;/p&gt;

&lt;p&gt;The discussion split between people treating the project as clever satire, people questioning whether executive work is really automatable, and people pointing out the legal and organizational responsibilities that don't disappear just because an AI can produce recommendations.&lt;/p&gt;

&lt;p&gt;One recurring objection was straightforward:&lt;/p&gt;

&lt;p&gt;An AI can recommend a decision.&lt;/p&gt;

&lt;p&gt;It cannot simply absorb the legal liability, fiduciary responsibility, accountability, or political consequences that come with making that decision on behalf of a real company.&lt;/p&gt;

&lt;p&gt;That objection is hard to dismiss.&lt;/p&gt;

&lt;p&gt;But it also misses part of what makes the project interesting.&lt;/p&gt;

&lt;p&gt;A CEO does not spend every minute of the day exercising some mysterious form of irreplaceable judgment.&lt;/p&gt;

&lt;p&gt;A large portion of executive work involves gathering information, comparing options, preparing reports, reviewing metrics, communicating decisions, tracking initiatives, and producing documents.&lt;/p&gt;

&lt;p&gt;Those are exactly the sorts of activities language models and agentic systems are getting better at.&lt;/p&gt;

&lt;p&gt;The harder question is where the boundary sits.&lt;/p&gt;

&lt;h2&gt;
  
  
  Is an AI CEO Actually the Point?
&lt;/h2&gt;

&lt;p&gt;Probably not.&lt;/p&gt;

&lt;p&gt;The more interesting interpretation is that Open Executive is a demonstration of how many supposedly distinct white-collar roles can be decomposed into information retrieval, analysis, synthesis, communication, and workflow execution.&lt;/p&gt;

&lt;p&gt;That's also why the project's architecture matters more than its joke.&lt;/p&gt;

&lt;p&gt;If you remove the word “CEO,” what remains is a multi-agent business system with:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Domain-specialized models&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Retrieval over business knowledge&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Company-specific context&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Persistent memory&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Scheduled actions&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Integrations&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Evaluation&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Model routing&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Provider abstraction&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those components have obvious applications far beyond replacing an executive.&lt;/p&gt;

&lt;p&gt;The “AI CEO” framing simply makes the underlying idea impossible to ignore.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Real Experiment
&lt;/h2&gt;

&lt;p&gt;Open Executive is not going to eliminate the need for a human CEO tomorrow.&lt;/p&gt;

&lt;p&gt;Its value is somewhere else.&lt;/p&gt;

&lt;p&gt;It is an open-source experiment in asking how much organizational work can be turned into software.&lt;/p&gt;

&lt;p&gt;And that question applies just as much to developers as it does to executives.&lt;/p&gt;

&lt;p&gt;If writing code, analyzing requirements, preparing reports, maintaining documentation, and operating systems can increasingly be delegated to software, then there is no obvious reason to assume the automation boundary stops at the engineering department.&lt;/p&gt;

&lt;p&gt;That is the uncomfortable part of the joke.&lt;/p&gt;

&lt;p&gt;The developers didn't build an AI CEO because an AI CEO is obviously practical.&lt;/p&gt;

&lt;p&gt;They built one because the idea exposes a symmetry that companies often ignore:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Once you start measuring jobs by which tasks can be automated, every layer of the organization eventually has to answer the same question.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;For now, Open Executive is still a software project, not a corporate officer.&lt;/p&gt;

&lt;p&gt;But as a piece of open-source engineering — and as a provocation — it is considerably more interesting than the headline suggests.&lt;/p&gt;

&lt;h3&gt;
  
  
  Try It
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/SenteLabsAI/OpenExecutive.git
&lt;span class="nb"&gt;cd &lt;/span&gt;OpenExecutive

&lt;span class="nb"&gt;cp&lt;/span&gt; .env.example .env
&lt;span class="c"&gt;# Add your ANTHROPIC_API_KEY&lt;/span&gt;

make dev
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The repository's current quick-start documentation lists Python 3.11+, Node 22+, an API on port &lt;code&gt;8000&lt;/code&gt;, and a Next.js UI on port &lt;code&gt;3000&lt;/code&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Sources
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;a href="https://github.com/SenteLabsAI/OpenExecutive" rel="noopener noreferrer"&gt;Open Executive — GitHub&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;a href="https://github.com/SenteLabsAI/OpenExecutive/blob/main/docs/architecture.md" rel="noopener noreferrer"&gt;Open Executive — Technical Architecture&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;a href="https://news.ycombinator.com/item?id=49458418" rel="noopener noreferrer"&gt;Hacker News discussion&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;a href="https://www.anthropic.com/news/claude-sonnet-4-6" rel="noopener noreferrer"&gt;Claude Sonnet 4.6 — Anthropic&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;a href="https://www.anthropic.com/news/claude-opus-4-7" rel="noopener noreferrer"&gt;Claude Opus 4.7 — Anthropic&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://zyvop.com/the-developers-who-got-fired-built-an-ai-ceo-to-replace-their-bosses-rpovs" rel="noopener noreferrer"&gt;ZyVOP&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;💡 For more articles like this, &lt;a href="https://zyvop.com/newsletter" rel="noopener noreferrer"&gt;subscribe to the ZyVOP newsletter&lt;/a&gt;!&lt;/p&gt;

</description>
      <category>opensource</category>
      <category>ai</category>
      <category>developertools</category>
      <category>claude</category>
    </item>
  </channel>
</rss>
