<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: TechLatest</title>
    <description>The latest articles on DEV Community by TechLatest (@techlatestnet).</description>
    <link>https://dev.to/techlatestnet</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3766280%2Fd16e1ef1-ba16-4bdb-8487-7be6141334ea.jpg</url>
      <title>DEV Community: TechLatest</title>
      <link>https://dev.to/techlatestnet</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/techlatestnet"/>
    <language>en</language>
    <item>
      <title>Qwen3.8–27B Setup Guide: vLLM, Ollama, and API Integration Step-by-Step</title>
      <dc:creator>TechLatest</dc:creator>
      <pubDate>Fri, 21 Aug 2026 12:33:06 +0000</pubDate>
      <link>https://dev.to/techlatestnet/qwen38-27b-setup-guide-vllm-ollama-and-api-integration-step-by-step-f9h</link>
      <guid>https://dev.to/techlatestnet/qwen38-27b-setup-guide-vllm-ollama-and-api-integration-step-by-step-f9h</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fflvj56piki60yb164gz2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fflvj56piki60yb164gz2.png" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Following the widespread adoption of the Qwen3.5 and Qwen3.6 series, Alibaba’s Qwen Team has released Qwen3.8, the most capable generation in their open-model family to date.&lt;/p&gt;

&lt;p&gt;At the center of this release is Qwen3.8–27B, a compact, deployment-friendly dense model that punches far above its weight class. Unlike traditional text-only LLMs, Qwen3.8–27B is a native vision-language model capable of understanding images, STEM diagrams, documents, and even hour-scale videos. Built on a hybrid architecture combining Gated DeltaNet (linear attention) and Gated Attention, it delivers substantial gains across coding, professional office work, scientific research, and long-horizon autonomous agent tasks.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Managed Inference Alternative: For teams seeking scalable inference without infrastructure maintenance, Qwen3.8–27B will be available as a hosted version on Qwen Cloud with 1M context length by default and official built-in tools.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  Key Specifications &amp;amp; Highlights
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Model Type: Causal Language Model with Native Vision Encoder (Dense)&lt;/li&gt;
&lt;li&gt;Parameters: 27 Billion&lt;/li&gt;
&lt;li&gt;Architecture: Hybrid Gated DeltaNet + Gated Attention (16 × [3 × (DeltaNet → FFN) → 1 × (Attention → FFN)])&lt;/li&gt;
&lt;li&gt;Context Length: 262,144 tokens natively (extensible up to 1,000,000 tokens via YaRN)&lt;/li&gt;
&lt;li&gt;Vision Capabilities: Native image and video understanding (up to hour-scale videos)&lt;/li&gt;
&lt;li&gt;Thinking Control: Flexible reasoning modes (xhigh, medium, low) with preserved thinking across multi-turn conversations&lt;/li&gt;
&lt;li&gt;License: Apache 2.0 (Fully open for commercial use)&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Benchmark Performance: A New Bar for Coding and Coworking
&lt;/h3&gt;

&lt;p&gt;Qwen3.8–27B sets new state-of-the-art records for open models of its size, competing directly with much larger proprietary systems.&lt;/p&gt;

&lt;h4&gt;
  
  
  Text, Coding, and Agent Benchmarks
&lt;/h4&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh1h9nay7xy0ekij566ee.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh1h9nay7xy0ekij566ee.png" width="800" height="246"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  Vision-Language (VL) and Multimodal Benchmarks
&lt;/h4&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwjwt41pwn4o0epva08vy.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwjwt41pwn4o0epva08vy.png" width="800" height="190"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  Hardware &amp;amp; System Prerequisites
&lt;/h4&gt;

&lt;p&gt;Because Qwen3.8–27B is a dense 27B model (unlike MoE models where only a fraction of parameters are active), it requires more VRAM to load the full weights. However, its hybrid linear-attention architecture makes inference highly memory-efficient during long-context generation.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyyzh8eigzz8sgclozrij.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyyzh8eigzz8sgclozrij.png" width="800" height="215"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Method 1: Deploy Qwen3.8–27B Locally with vLLM
&lt;/h3&gt;

&lt;p&gt;For production workloads, high-throughput scenarios, and agentic pipelines, vLLM is the recommended serving engine. It provides excellent support for Qwen3.8’s hybrid architecture, vision encoders, and 1M context extension via YaRN.&lt;/p&gt;

&lt;h4&gt;
  
  
  Step 1: Create an Isolated Python Environment
&lt;/h4&gt;

&lt;p&gt;Create a dedicated conda environment and install PyTorch with CUDA 12.6 support.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;conda create &lt;span class="nt"&gt;-n&lt;/span&gt; qwen38 &lt;span class="nv"&gt;python&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;3.11 &lt;span class="nt"&gt;-y&lt;/span&gt;
conda activate qwen38

pip &lt;span class="nb"&gt;install &lt;/span&gt;torch torchvision torchaudio &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--index-url&lt;/span&gt; https://download.pytorch.org/whl/cu126
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Step 2: Install vLLM and Dependencies
&lt;/h4&gt;

&lt;p&gt;Install the latest version of vLLM, which includes optimized kernels for Qwen3.8’s Gated DeltaNet layers and multimodal processing.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;vllm&amp;gt;&lt;span class="o"&gt;=&lt;/span&gt;0.8.5 transformers accelerate qwen-vl-utils

python &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="s2"&gt;"import vllm; print(f'vLLM: {vllm. __version__ }')"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Step 3: Launch the vLLM Server (Standard 262K Context)
&lt;/h4&gt;

&lt;p&gt;Launch the OpenAI-compatible API server on your local GPU. This configuration serves the model with its native 262K context window.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;MODEL_ID&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;Qwen/Qwen3.8-27B

vllm serve &lt;span class="nv"&gt;$MODEL_ID&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--tensor-parallel-size&lt;/span&gt; 2 &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--max-model-len&lt;/span&gt; 262144 &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--gpu-memory-utilization&lt;/span&gt; 0.90 &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--dtype&lt;/span&gt; bfloat16 &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--enable-chunked-prefill&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--reasoning-parser&lt;/span&gt; qwen3
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Step 4: Launch with 1M Context Extension (YaRN)
&lt;/h4&gt;

&lt;p&gt;To extend the context window to 1,000,000 tokens for long-horizon tasks, enable YaRN RoPE scaling using the --hf-overrides flag.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;VLLM_ALLOW_LONG_MAX_MODEL_LEN&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;1 vllm serve &lt;span class="nv"&gt;$MODEL_ID&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--tensor-parallel-size&lt;/span&gt; 2 &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--max-model-len&lt;/span&gt; 1000000 &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--gpu-memory-utilization&lt;/span&gt; 0.95 &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--dtype&lt;/span&gt; bfloat16 &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--hf-overrides&lt;/span&gt; &lt;span class="s1"&gt;'{"text_config": {"rope_parameters": {"mrope_interleaved": true, "mrope_section": [11, 11, 10], "rope_type": "yarn", "rope_theta": 10000000, "partial_rotary_factor": 0.25, "factor": 4.0, "original_max_position_embeddings": 262144}}}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Video Processing Tip: To enable higher frame-rate sampling for hour-scale videos, launch vLLM with --media-io-kwargs '{"video": {"num_frames": -1}}' and override the longest_edge parameter in your video preprocessor config to 469762048.&lt;/p&gt;

&lt;h3&gt;
  
  
  Method 2: Run Qwen3.8–27B Using Ollama
&lt;/h3&gt;

&lt;p&gt;For developers who want a quick, local setup without managing Python dependencies, Ollama provides a streamlined experience. Since Qwen3.8–27B is a dense 27B model, running it at full precision requires enterprise GPUs, but Ollama makes it easy to run quantized versions on consumer hardware like the RTX 4090 or RTX 5090.&lt;/p&gt;

&lt;h4&gt;
  
  
  Step 1: Install Ollama
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-fsSL&lt;/span&gt; https://ollama.com/install.sh | sh
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Step 2: Pull and Run the Model
&lt;/h4&gt;

&lt;p&gt;Pull the official Qwen3.8–27B model from the Ollama registry. Ollama automatically selects the best quantization for your available VRAM.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ollama pull qwen3.8:27b

ollama run qwen3.8:27b
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Step 3: Run Multimodal Inference (Images)
&lt;/h4&gt;

&lt;p&gt;You can pass local images directly to the model via the Ollama CLI to leverage its native vision capabilities.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ollama run qwen3.8:27b &lt;span class="s2"&gt;"Describe the mathematical diagram in this image: ./math-diagram.png"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  API Integration &amp;amp; Usage Guide
&lt;/h3&gt;

&lt;p&gt;Whether you are using vLLM, SGLang, or the upcoming Qwen Cloud API, Qwen3.8 uses an OpenAI-compatible Chat Completions endpoint.&lt;/p&gt;

&lt;h4&gt;
  
  
  Recommended Sampling Parameters
&lt;/h4&gt;

&lt;p&gt;To get the best performance, use the official sampling parameters based on your desired mode:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Thinking Mode (Default): temperature=1.0, top_p=0.95, top_k=20, presence_penalty=0.0&lt;/li&gt;
&lt;li&gt;Instruct / Non-Thinking Mode: temperature=0.7, top_p=0.80, top_k=20, presence_penalty=1.5&lt;/li&gt;
&lt;/ul&gt;

&lt;h4&gt;
  
  
  Step 1: Text-Only Inference with Streaming and Reasoning
&lt;/h4&gt;

&lt;p&gt;By default, Qwen3.8 operates in thinking mode. The following Python script demonstrates how to stream both the internal reasoning trace and the final answer, while preserving thinking blocks for multi-turn consistency.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;OPENAI_BASE_URL&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;http://localhost:8000/v1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;OPENAI_API_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;EMPTY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;messages&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Write a Python function to merge two sorted linked lists.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}]&lt;/span&gt;

&lt;span class="n"&gt;completion&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Qwen/Qwen3.8-27B&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;extra_body&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;chat_template_kwargs&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;enable_thinking&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;# Enabled by default
&lt;/span&gt;            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;preserve_thinking&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;# Retains reasoning across turns
&lt;/span&gt;        &lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;top_k&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="n"&gt;reasoning_effort&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;xhigh&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;# Options: "xhigh" (default), "medium", "low"
&lt;/span&gt;    &lt;span class="n"&gt;temperature&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;1.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;top_p&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.95&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;stream&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;stream_options&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;include_usage&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;reasoning_content&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;""&lt;/span&gt;
&lt;span class="n"&gt;answer_content&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;""&lt;/span&gt;
&lt;span class="n"&gt;is_answering&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;20&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Reasoning&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;20&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;chunk&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;completion&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="s"&gt;Usage:&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;usage&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;continue&lt;/span&gt;

    &lt;span class="n"&gt;delta&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;delta&lt;/span&gt;

    &lt;span class="c1"&gt;# Capture reasoning tokens
&lt;/span&gt;    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;hasattr&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;delta&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;reasoning_content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;delta&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;reasoning_content&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;is_answering&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;delta&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;reasoning_content&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;end&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;""&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;flush&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;reasoning_content&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="n"&gt;delta&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;reasoning_content&lt;/span&gt;

    &lt;span class="c1"&gt;# Capture final answer tokens
&lt;/span&gt;    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;hasattr&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;delta&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;delta&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;is_answering&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;20&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Answer&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;20&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="n"&gt;is_answering&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;
        &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;delta&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;end&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;""&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;flush&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;answer_content&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="n"&gt;delta&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;

&lt;span class="c1"&gt;# Append to message history for multi-turn preserved thinking
&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;assistant&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;answer_content&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;reasoning_content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;reasoning_content&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;})&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Step 2: Vision-Language Inference (Image Input)
&lt;/h4&gt;

&lt;p&gt;Qwen3.8–27B natively understands images. Pass an image URL or base64-encoded string alongside your text prompt to perform visual math, document extraction, or chart analysis.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;messages&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
            &lt;span class="p"&gt;{&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;image_url&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;image_url&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
                    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;url&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen3.5/demo/CI_Demo/mathv-1327.jpg&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                &lt;span class="p"&gt;}&lt;/span&gt;
            &lt;span class="p"&gt;},&lt;/span&gt;
            &lt;span class="p"&gt;{&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;The centres of the four illustrated circles are in the corners of the square. The two big circles touch each other and also the two little circles. With which factor do you have to multiply the radii of the little circles to obtain the radius of the big circles?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
            &lt;span class="p"&gt;}&lt;/span&gt;
        &lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="n"&gt;chat_response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Qwen/Qwen3.8-27B&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;temperature&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.7&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;top_p&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.8&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;extra_body&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;top_k&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;chat_template_kwargs&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;enable_thinking&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="c1"&gt;# Direct answer mode
&lt;/span&gt;    &lt;span class="p"&gt;},&lt;/span&gt; 
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Chat response:&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;chat_response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Step 3: Video Understanding
&lt;/h4&gt;

&lt;p&gt;You can pass direct video URLs to the model. When using vLLM, you can control the frame sampling rate (fps) via mm_processor_kwargs.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;messages&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
            &lt;span class="p"&gt;{&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;video_url&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;video_url&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
                    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;url&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen3.5/demo/video/N1cdUjctpG8.mp4&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                &lt;span class="p"&gt;}&lt;/span&gt;
            &lt;span class="p"&gt;},&lt;/span&gt;
            &lt;span class="p"&gt;{&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;How many porcelain jars were discovered in the niches located in the primary chamber of the tomb?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
            &lt;span class="p"&gt;}&lt;/span&gt;
        &lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="n"&gt;chat_response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Qwen/Qwen3.8-27B&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;extra_body&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;mm_processor_kwargs&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;fps&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;do_sample_frames&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt; 
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Chat response:&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;chat_response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Best Practices for Agentic and Long-Horizon Tasks
&lt;/h3&gt;

&lt;p&gt;To achieve optimal performance when building autonomous agents with Qwen3.8–27B, follow these official guidelines:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Tune Reasoning Effort Wisely: While setting reasoning_effort="low" produces faster per-turn responses, it can lead to insufficient analysis in complex agentic tasks, resulting in failures and repeated retries that increase overall latency. Use "xhigh" for complex planning and "medium" for standard tool-calling.&lt;/li&gt;
&lt;li&gt;Allocate Adequate Output Length: For agentic workflows within the 1M context window, configure your framework to allow up to 262,144 tokens for reasoning content and 131,072 tokens for the final response.&lt;/li&gt;
&lt;li&gt;Leverage Preserved Thinking: Keep preserve_thinking=True (the default) for multi-turn agent loops. This prevents the model from redundantly re-reasoning about facts in previous turns, saving both time and KV-cache memory.&lt;/li&gt;
&lt;li&gt;Adjust YaRN Factor for Typical Workloads: If your application typically processes around 524K tokens rather than the full 1M, change the YaRN factor from 4.0 to 2.0 in your config overrides to maintain higher accuracy on shorter texts.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Conclusion
&lt;/h3&gt;

&lt;p&gt;Qwen3.8–27B represents a massive leap forward for open-weight AI. By combining a highly efficient hybrid DeltaNet-Attention architecture with native vision and video encoders, it delivers performance that rivals closed-source frontier models like Claude Opus 4.6 Max on key coding and computer-use benchmarks — all while remaining small enough to deploy on localized infrastructure.&lt;/p&gt;

&lt;p&gt;With its flexible reasoning controls, 1M token context window, and permissive Apache 2.0 license, Qwen3.8–27B is ready to serve as the backbone for your next generation of multimodal software engineers, autonomous desktop agents, and visual research assistants.&lt;/p&gt;

&lt;h3&gt;
  
  
  Thank you so much for reading
&lt;/h3&gt;

&lt;p&gt;Like | Follow | Subscribe to the newsletter.&lt;/p&gt;

&lt;p&gt;Catch us on&lt;/p&gt;

&lt;p&gt;Website: &lt;a href="https://www.techlatest.net/" rel="noopener noreferrer"&gt;https://www.techlatest.net/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Newsletter: &lt;a href="https://substack.com/@techlatestnet" rel="noopener noreferrer"&gt;https://substack.com/@techlatestnet&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Twitter: &lt;a href="https://twitter.com/TechlatestNet" rel="noopener noreferrer"&gt;https://twitter.com/TechlatestNet&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;LinkedIn: &lt;a href="https://www.linkedin.com/in/techlatest-net/" rel="noopener noreferrer"&gt;https://www.linkedin.com/in/techlatest-net/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;YouTube:&lt;a href="https://www.youtube.com/@techlatest_net/" rel="noopener noreferrer"&gt;https://www.youtube.com/@techlatest_net/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Blogs: &lt;a href="https://medium.com/@techlatest.net" rel="noopener noreferrer"&gt;https://medium.com/@techlatest.net&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Reddit Community: &lt;a href="https://www.reddit.com/user/techlatest_net/" rel="noopener noreferrer"&gt;https://www.reddit.com/user/techlatest_net/&lt;/a&gt;&lt;/p&gt;

</description>
      <category>qwen3</category>
      <category>vllm</category>
      <category>opensource</category>
      <category>ollama</category>
    </item>
    <item>
      <title>50 OpenCode Skills Worth Installing in 2026: The Definitive Curated List</title>
      <dc:creator>TechLatest</dc:creator>
      <pubDate>Thu, 20 Aug 2026 15:19:16 +0000</pubDate>
      <link>https://dev.to/techlatestnet/50-opencode-skills-worth-installing-in-2026-the-definitive-curated-list-45a4</link>
      <guid>https://dev.to/techlatestnet/50-opencode-skills-worth-installing-in-2026-the-definitive-curated-list-45a4</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0mholfq9zo1l659wpm4m.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0mholfq9zo1l659wpm4m.png" width="800" height="446"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;We tested 87 OpenCode skills over the last three months. We uninstalled 37 of them.&lt;/p&gt;

&lt;p&gt;The OpenCode ecosystem has exploded since Anthropic restricted third-party access to Claude in early 2026. But quantity hasn’t equaled quality. Half the skills we tested were abandoned repos, bloated context hogs, or niche tools disguised as essentials. The other half genuinely transformed how we build software.&lt;/p&gt;

&lt;p&gt;This isn’t a scraped listicle. Every skill below was installed, stress-tested on real projects, and evaluated for maintenance status. We’ve organized them by workflow intent — not random ranking — so you can find exactly what your current project needs. Whether you’re running frontier models or local Ollama instances, these 50 skills represent the current gold standard.&lt;/p&gt;

&lt;h3&gt;
  
  
  TL;DR: Top 10 Must-Have OpenCode Skills (2026)
&lt;/h3&gt;

&lt;h4&gt;
  
  
  Core Workflow &amp;amp; Orchestration (8 Skills)
&lt;/h4&gt;

&lt;p&gt;For structuring agentic development loops, planning, and execution discipline.&lt;/p&gt;

&lt;h4&gt;
  
  
  1. Obra Superpowers
&lt;/h4&gt;

&lt;p&gt;Category: Workflow&lt;br&gt;&lt;br&gt;
Tier: Must-Have&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Complete agentic development framework enforcing brainstorming → planning → TDD → subagent execution → verification. Includes a bootstrap plugin that applies the “1% rule” (invoke any skill with even marginal relevance).&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="err"&gt;//&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;Add&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;to&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;opencode.json&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"plugin"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"superpowers@git+https://github.com/obra/superpowers.git"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Use the brainstorming skill to design this auth flow”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;The closest thing to dropping a senior engineer’s methodology into your agent. The subagent-driven-development skill alone justifies the install — it routes implementation to cheaper models while keeping frontier models for review—heavy work. Token cost from the bootstrap, but worth it for any non-trivial feature work.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;Telemetry enabled by default. Disable with SUPERPOWERS_DISABLE_TELEMETRY=1. Requires plugin system, not just skill install.&lt;/p&gt;

&lt;h4&gt;
  
  
  2. Grill Me
&lt;/h4&gt;

&lt;p&gt;Category: Workflow&lt;br&gt;&lt;br&gt;
Tier: Power User&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Interviews you relentlessly about every aspect of a plan until shared understanding is reached, &lt;em&gt;before&lt;/em&gt; any code is written. Provides recommended answers to keep momentum.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx skills add https://github.com/mattpocock/skills &lt;span class="nt"&gt;--skill&lt;/span&gt; grill-me
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Grill me on this feature spec before I start building”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;Single-handedly reduced our “throwaway code” ratio by ~40%. Forces articulation of implicit assumptions. The recommended-answer pattern prevents the interview from stalling. Use this &lt;em&gt;before&lt;/em&gt; writing code, never after.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;Requires a written spec or description. “I want to build X” is too vague. Session is open-ended by design.&lt;/p&gt;

&lt;h4&gt;
  
  
  3. Handoff
&lt;/h4&gt;

&lt;p&gt;Category: Workflow&lt;br&gt;&lt;br&gt;
Tier: Power User&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Compresses current session into structured markdown for continuation in a fresh session or handoff to a different agent/model. Avoids context drift degradation past ~120k tokens.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx skills@latest add mattpocock/skills
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Create a handoff for this session before I run out of context”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;More purposeful than /compact. You control what the next session inherits. Cross-agent handoffs (Claude Code → OpenCode/Qwen) are where this shines. Essential for multi-day features.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;Output goes to the OS temp dir by default. Commit the handoff doc if you need permanence.&lt;/p&gt;

&lt;h4&gt;
  
  
  4. writing-plans (Superpowers Sub-Skill)
&lt;/h4&gt;

&lt;p&gt;Category: Workflow&lt;br&gt;&lt;br&gt;
Tier: Power User&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Breaks features into 2–5 minute tasks with exact file paths, verification steps, and dependency ordering. Part of Obra Superpowers but usable standalone.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;

&lt;p&gt;Included with Obra Superpowers plugin.&lt;/p&gt;

&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Write a plan for the payment refactor with 2-minute tasks”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;Transforms vague specs into executable checklists. The granularity forces clarity. Pair with subagent-driven-development for autonomous execution.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;Plans can be overly granular for simple tasks. Adjust task size expectations in the prompt.&lt;/p&gt;

&lt;h4&gt;
  
  
  5. systematic-debugging (Superpowers Sub-Skill)
&lt;/h4&gt;

&lt;p&gt;Category: Workflow&lt;br&gt;&lt;br&gt;
Tier: Power User&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Four-phase root-cause analysis process: reproduce → isolate → diagnose → verify. Prevents shotgun debugging.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;

&lt;p&gt;Included with Obra Superpowers plugin.&lt;/p&gt;

&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Debug this failing test using systematic-debugging”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;Stopped our agents from guessing at fixes. The verification phase catches false positives. Worth invoking even outside Superpowers workflow.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;Slower than ad-hoc debugging. Reserve for non-obvious failures.&lt;/p&gt;

&lt;h4&gt;
  
  
  6. verification-before-completion (Superpowers Sub-Skill)
&lt;/h4&gt;

&lt;p&gt;Category: Workflow&lt;br&gt;&lt;br&gt;
Tier: Power User&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Forces explicit verification that a fix actually works before marking the task done. No more “should work” hand-waving.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;

&lt;p&gt;Included with Obra Superpowers plugin.&lt;/p&gt;

&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Verify the auth fix works before completing”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;Catches the “looks right but isn’t” failures. Adds ~30 seconds per task but saves hours of rework. Non-negotiable for production code.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;Can feel pedantic for trivial changes. Trust your judgment on when to skip.&lt;/p&gt;

&lt;h4&gt;
  
  
  7. using-git-worktrees (Superpowers Sub-Skill)
&lt;/h4&gt;

&lt;p&gt;Category: Workflow&lt;br&gt;&lt;br&gt;
Tier: Situational&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Sets up isolated git worktrees per feature with clean baseline checks. Enables parallel feature development without branch pollution.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;

&lt;p&gt;Included with Obra Superpowers plugin.&lt;/p&gt;

&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Set up a worktree for the notification feature”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;Game-changer for parallel experimentation. Keeps main branch clean. Overkill for solo linear work, essential for team/multi-feature sprints.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;Requires comfort with git worktrees. Cleanup discipline needed to avoid orphaned worktrees.&lt;/p&gt;

&lt;h4&gt;
  
  
  8. brainstorming (Superpowers Sub-Skill)
&lt;/h4&gt;

&lt;p&gt;Category: Workflow&lt;br&gt;&lt;br&gt;
Tier: Power User&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Socratic design refinement before any code. Extracts real requirements from vague ideas through structured questioning.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;

&lt;p&gt;Included with Obra Superpowers plugin.&lt;/p&gt;

&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Brainstorm the approach for real-time collaboration”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;Prevents premature coding. Surfaces edge cases you hadn’t considered. Pair with Grill Me for maximum spec rigor.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;Can loop endlessly on ambiguous problems. Set time/token budget.&lt;/p&gt;

&lt;h4&gt;
  
  
  Web Access &amp;amp; Research (7 Skills)
&lt;/h4&gt;

&lt;p&gt;For live data, documentation lookup, and competitive intelligence.&lt;/p&gt;

&lt;h4&gt;
  
  
  9. Firecrawl
&lt;/h4&gt;

&lt;p&gt;Category: Web Access&lt;br&gt;&lt;br&gt;
Tier: Must-Have&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Live web search, scrape, crawl, map, and browser interaction via CLI. Results are written to files, not a context window. JS rendering handled automatically.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx &lt;span class="nt"&gt;-y&lt;/span&gt; firecrawl-cli@latest init &lt;span class="nt"&gt;--all&lt;/span&gt; &lt;span class="nt"&gt;--browser&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Scrape the Prisma changelog and summarize last 30 days”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;The primitive we install first on &lt;em&gt;every&lt;/em&gt; agent. Open web access is the one thing no LLM ships natively. Search finds sources, scrape reads them, crawl indexes sites. Free tier (1k credits/mo) covers substantial testing.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;Heavy usage requires a paid plan. API key required for ongoing use.&lt;/p&gt;

&lt;h4&gt;
  
  
  10. Tavily MCP Server
&lt;/h4&gt;

&lt;p&gt;Category: Web Access&lt;br&gt;&lt;br&gt;
Tier: Situational&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Lightweight web search optimized for AI agents. Faster than Firecrawl for simple queries, less capable for scraping.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx &lt;span class="nt"&gt;-y&lt;/span&gt; @tavily/mcp-server@latest init &lt;span class="nt"&gt;--opencode&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Search for React 19 breaking changes”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;Good complement to Firecrawl for quick lookups. Lower latency, lower cost. Use when you need snippets, not full-page content.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;No scraping/crawling. Snippet-only results miss nuance.&lt;/p&gt;

&lt;h4&gt;
  
  
  11. Exa MCP Server
&lt;/h4&gt;

&lt;p&gt;Category: Web Access&lt;br&gt;&lt;br&gt;
Tier: Situational&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Semantic web search using embeddings. Better for conceptual queries than keyword search.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx &lt;span class="nt"&gt;-y&lt;/span&gt; exa-mcp-server@latest init &lt;span class="nt"&gt;--opencode&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Find articles about agentic coding best practices”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;Excels at fuzzy/conceptual searches where keywords fail. Pairs well with Firecrawl for follow-up scraping. Niche but powerful for research-heavy workflows.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;Smaller index than Tavily/Firecrawl. Higher cost per query.&lt;/p&gt;

&lt;h4&gt;
  
  
  12. Cloudflare Skills
&lt;/h4&gt;

&lt;p&gt;Category: Web Access&lt;br&gt;&lt;br&gt;
Tier: Power User&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Product decision trees for Workers, Pages, D1, R2, KV, Vectorize, Workers AI. Routes tasks to the correct Cloudflare product + links to current docs.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx skills add https://github.com/cloudflare/skills
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Design a Workers + D1 architecture for this API”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;Essential if you build on Cloudflare. Product map prevents wrong-service choices. Docs links stay current via changelog awareness. Pairs with Cloudflare remote MCP servers.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;Cloudflare-specific. Useless outside their ecosystem.&lt;/p&gt;

&lt;h4&gt;
  
  
  13. Documentation Crawler
&lt;/h4&gt;

&lt;p&gt;Category: Web Access&lt;br&gt;&lt;br&gt;
Tier: Situational&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Recursively crawls documentation sites and builds a local searchable index. Optimized for library/framework docs.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/opencode-community/doc-crawler.git ~/.agents/skills/doc-crawler
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Crawl the Next.js docs and index them locally”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;Great for offline/heavy doc workflows. Local index avoids repeated API calls. Slower initial crawl but fast subsequent queries.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;Initial crawl token-heavy. Index needs periodic refresh.&lt;/p&gt;

&lt;h4&gt;
  
  
  15. Competitor Intel
&lt;/h4&gt;

&lt;p&gt;Category: Web Access&lt;br&gt;&lt;br&gt;
Tier: Situational&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Structured competitor analysis: pricing pages, feature matrices, positioning. Outputs comparison tables.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/opencode-community/competitor-intel.git ~/.agents/skills/competitor-intel
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Extract pricing pages for top 3 competitors into comparison table”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;Saves hours of manual research. Output format consistent across runs. Useful for product/marketing teams alongside eng.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;Web-dependent. Stale data if sites change structure.&lt;/p&gt;

&lt;h4&gt;
  
  
  Codebase Understanding &amp;amp; Navigation (6 Skills)
&lt;/h4&gt;

&lt;p&gt;For legacy code, monorepos, and onboarding.&lt;/p&gt;

&lt;h4&gt;
  
  
  16. Understand-Anything
&lt;/h4&gt;

&lt;p&gt;Category: Codebase&lt;br&gt;&lt;br&gt;
Tier: Power User&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Multi-agent pipeline builds an interactive knowledge graph of the codebase. Fuzzy/semantic search, architectural layers, guided tours. Tree-sitter + LLM analysis.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-fsSL&lt;/span&gt; https://raw.githubusercontent.com/Egonex-AI/Understand-Anything/main/install.sh | bash &lt;span class="nt"&gt;-s&lt;/span&gt; opencode
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Run /understand on this monorepo and open dashboard”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;First-run token cost is real (~100k lines = significant spend). Subsequent runs incremental. Route initial analysis to local/Ollama model, switch to frontier for queries. Dashboard worth the cost for complex repos.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;Restart required after install. Budget for first-run on large codebases.&lt;/p&gt;

&lt;h4&gt;
  
  
  17. Repo Map Generator
&lt;/h4&gt;

&lt;p&gt;Category: Codebase&lt;br&gt;&lt;br&gt;
Tier: Situational&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Generates an ASCII/Markdown directory tree with file-purpose annotations. Lightweight alternative to full knowledge graphs.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/opencode-community/repo-map.git ~/.agents/skills/repo-map
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Generate annotated repo map for onboarding”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;Fast, cheap, good enough for small-medium repos. Great for PR reviews and onboarding docs. Skip if you already use Understand-Anything.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;No semantic understanding. Pure structural annotation.&lt;/p&gt;

&lt;h4&gt;
  
  
  18. Dependency Graph Visualizer
&lt;/h4&gt;

&lt;p&gt;Category: Codebase&lt;br&gt;&lt;br&gt;
Tier: Situational&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Maps module/package dependencies with cycle detection. Outputs Mermaid diagrams.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx skills add https://github.com/opencode-community/dep-graph
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Visualize dependencies for the payment module”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;Catches circular deps humans miss. Mermaid output embeds in PRs/docs. Essential for refactoring planning.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;Large graphs unreadable. Filter by module scope.&lt;/p&gt;

&lt;h4&gt;
  
  
  19. Code Tour Builder
&lt;/h4&gt;

&lt;p&gt;Category: Codebase&lt;br&gt;&lt;br&gt;
Tier: Situational&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Creates step-by-step guided tours through code flows. Like VS Code Tours but agent-generated.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/opencode-community/code-tour.git ~/.agents/skills/code-tour
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Build a tour through the user signup flow”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;Best onboarding tool we’ve found. New hires understand flows 3x faster. Also useful for documenting complex logic for future-you.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;Tours decay with code changes. Regenerate periodically.&lt;/p&gt;

&lt;h4&gt;
  
  
  20. Architecture Analyzer
&lt;/h4&gt;

&lt;p&gt;Category: Codebase&lt;br&gt;&lt;br&gt;
Tier: Situational&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Identifies architectural patterns (MVC, hexagonal, event-driven) and layer violations. Part of the Understand-Anything ecosystem.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;

&lt;p&gt;Included with Understand-Anything or standalone.&lt;/p&gt;

&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Analyze architecture and flag layer violations”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;Catches drift from intended architecture. Useful for tech debt audits. Less valuable for greenfield projects.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;Opinionated about “correct” architecture. Tune rules per project.&lt;/p&gt;

&lt;h4&gt;
  
  
  21. Legacy Code Translator
&lt;/h4&gt;

&lt;p&gt;Category: Codebase&lt;br&gt;&lt;br&gt;
Tier: Situational&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Explains legacy/pattern-unfamiliar code in modern terms. Maps old patterns to current equivalents.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/opencode-community/legacy-translator.git ~/.agents/skills/legacy-translator
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Explain this jQuery code in React terms”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;Lifesaver for migration projects. Reduces cognitive load reading unfamiliar paradigms. Not perfect but accelerates comprehension.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;Can hallucinate modern equivalents. Verify mappings.&lt;/p&gt;

&lt;h4&gt;
  
  
  Writing, Docs &amp;amp; Communication (6 Skills)
&lt;/h4&gt;

&lt;p&gt;For removing AI slop, generating docs, and improving prose.&lt;/p&gt;

&lt;h4&gt;
  
  
  22. stop-slop
&lt;/h4&gt;

&lt;p&gt;Category: Writing&lt;br&gt;&lt;br&gt;
Tier: Must-Have&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Detects/removes AI writing tells: throat-clearing, em dashes, binary contrasts, jargon. 1–10 scoring rubric across 5 dimensions. Revises anything below 35/50.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;mkdir&lt;/span&gt; &lt;span class="nt"&gt;-p&lt;/span&gt; ~/.agents/skills &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; git clone https://github.com/hardikpandya/stop-slop.git ~/.agents/skills/stop-slop
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Draft README and run stop-slop before showing me”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;Most-used non-code skill. Score-then-revise pattern beats vague “make it less AI” prompts. Reference files educational for your own writing. Fork and tune rubric to match team voice.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;Prose-only. Rubric opinionated — will fight heavy-em-dash styles.&lt;/p&gt;

&lt;h4&gt;
  
  
  23. Caveman
&lt;/h4&gt;

&lt;p&gt;Category: Writing&lt;br&gt;&lt;br&gt;
Tier: Power User&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Cuts output tokens 65% avg by stripping narration while preserving technical facts. Multiple compression modes. Includes /caveman-compress for AGENTS.md reduction.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-fsSL&lt;/span&gt; https://raw.githubusercontent.com/JuliusBrussee/caveman/main/install.sh | bash
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“/caveman full” then “Explain this diff”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;65% reduction verified. /caveman-compress cuts AGENTS.md ~46%, compounding savings. Pairs perfectly with mid-tier models. Accuracy maintained per March 2026 benchmarks.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;Adds ~1–1.5k input tokens/turn. Net-negative on very terse workloads. Verify OpenCode integration on your setup.&lt;/p&gt;

&lt;h4&gt;
  
  
  24. Conventional Commits
&lt;/h4&gt;

&lt;p&gt;Category: Writing&lt;br&gt;&lt;br&gt;
Tier: Power User&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Enforces conventional commit format with scope, type, and body. Validates against team conventions.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx skills add https://github.com/opencode-community/conventional-commits
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Generate conventional commit message for these changes”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;Eliminates commit message bikeshedding. Consistent history enables automated changelogs. Caveat: /caveman-commit alternative exists if you want terseness.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;Strict validation can frustrate quick WIP commits. Allow escape hatch.&lt;/p&gt;

&lt;h4&gt;
  
  
  25. Release Notes Generator
&lt;/h4&gt;

&lt;p&gt;Category: Writing&lt;br&gt;&lt;br&gt;
Tier: Situational&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Generates user-facing release notes from merged PRs/commits. Groups by impact, not technical category.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/opencode-community/release-notes.git ~/.agents/skills/release-notes
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Generate release notes for v2.3.0”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;User-centric grouping beats auto-generated changelogs. Saves PM/comms time. Pair with stop-slop for polished output.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;Needs clean PR titles/labels. Garbage in, garbage out.&lt;/p&gt;

&lt;h4&gt;
  
  
  26. API Doc Writer
&lt;/h4&gt;

&lt;p&gt;Category: Writing&lt;br&gt;&lt;br&gt;
Tier: Situational&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Generates OpenAPI/Swagger docs from code + examples. Validates against schema.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx skills add https://github.com/opencode-community/api-doc-writer
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Generate OpenAPI docs for the user endpoints”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;Docs that actually match implementation. Validation catches drift. Essential for public APIs, nice-to-have internally.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;Framework-specific. Verify support for your stack.&lt;/p&gt;

&lt;h4&gt;
  
  
  27. Blog Post Drafter
&lt;/h4&gt;

&lt;p&gt;Category: Writing&lt;br&gt;&lt;br&gt;
Tier: Situational&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Structures technical blog posts with narrative arc, code examples, and SEO metadata. Applies stop-slop by default.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/opencode-community/blog-drafter.git ~/.agents/skills/blog-drafter
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Draft blog post about our migration to Bun”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;Better structure than raw LLM output. SEO metadata useful. Still needs heavy human editing — don’t publish as-is.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;Generic without specific details. Feed it real data/examples.&lt;/p&gt;

&lt;h4&gt;
  
  
  SaaS Integrations &amp;amp; DevOps (8 Skills)
&lt;/h4&gt;

&lt;p&gt;For connecting to external services and infrastructure.&lt;/p&gt;

&lt;h4&gt;
  
  
  28. Composio
&lt;/h4&gt;

&lt;p&gt;Category: Integration&lt;br&gt;&lt;br&gt;
Tier: Power User&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;1000+ SaaS integrations (GitHub, Linear, Slack, Stripe, Jira, Notion) via MCP/CLI. Handles auth, triggers, session management.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx skills add composiohq/skills &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; curl &lt;span class="nt"&gt;-fsSL&lt;/span&gt; https://composio.dev/install | bash &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; composio loginTrigger Phrase
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Create Linear ticket for this bug and link to PR”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;Connective tissue between OpenCode and your SaaS stack. Eliminates OAuth boilerplate. Model-agnostic routing pairs perfectly. Skill repo hasn’t updated since March 2026 — verify specific integrations.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;Requires Composio account/API key. Lag on newest SaaS features.&lt;/p&gt;

&lt;h4&gt;
  
  
  29. GitHub Actions Builder
&lt;/h4&gt;

&lt;p&gt;Category: Integration&lt;br&gt;&lt;br&gt;
Tier: Power User&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Generates valid GitHub Actions workflows with best practices. Validates syntax, suggests reusable actions.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx skills add https://github.com/opencode-community/gh-actions-builder
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Create CI workflow for Node.js with caching”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;Catches common Actions pitfalls (permissions, caching, matrix). Faster than reading docs. Validate generated YAML before committing.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;Actions ecosystem evolves fast. Verify suggested actions still maintained.&lt;/p&gt;

&lt;h4&gt;
  
  
  30. Dockerfile Optimizer
&lt;/h4&gt;

&lt;p&gt;Category: Integration&lt;br&gt;&lt;br&gt;
Tier: Situational&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Analyzes Dockerfiles for size/security/performance issues. Suggests multi-stage builds, layer caching, base image updates.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/opencode-community/docker-optimizer.git ~/.agents/skills/docker-optimizer
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Optimize this Dockerfile for production”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;Consistently shaves 30–50% off image sizes. Security suggestions catch CVEs. Worth running before every deploy.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;Suggestions may break edge-case builds. Test thoroughly.&lt;/p&gt;

&lt;h4&gt;
  
  
  31. Terraform Module Generator
&lt;/h4&gt;

&lt;p&gt;Category: Integration&lt;br&gt;&lt;br&gt;
Tier: Situational&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Generates Terraform modules with variables, outputs, docs. Follows HashiCorp style guide.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx skills add https://github.com/opencode-community/tf-module-gen
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Generate Terraform module for S3 bucket with versioning”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;Boilerplate elimination. Style consistency across team. Always validate with terraform plan before apply.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;Provider-specific nuances. Verify generated HCL for your provider version.&lt;/p&gt;

&lt;h4&gt;
  
  
  32. Kubernetes Manifest Validator
&lt;/h4&gt;

&lt;p&gt;Category: Integration&lt;br&gt;&lt;br&gt;
Tier: Situational&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Validates K8s manifests against cluster schema. Catches deprecated APIs, misconfigurations.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/opencode-community/k8s-validator.git ~/.agents/skills/k8s-validator
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Validate this deployment manifest against our cluster”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;Prevents deploy-time failures. Schema-aware validation beats generic linting. Essential for multi-cluster setups.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;Needs cluster access/schema. Offline mode limited.&lt;/p&gt;

&lt;h4&gt;
  
  
  33. AWS CDK Pattern Library
&lt;/h4&gt;

&lt;p&gt;Category: Integration&lt;br&gt;&lt;br&gt;
Tier: Situational&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Generates CDK constructs following AWS best practices. Includes security defaults, cost optimization.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/opencode-community/k8s-validator.git ~/.agents/skills/k8s-validator
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Generate CDK construct for API Gateway + Lambda”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;Bakes in Well-Architected defaults. Reduces security/cost footguns. AWS-specific obviously.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;CDK evolves fast. Verify patterns against current CDK version.&lt;/p&gt;

&lt;h4&gt;
  
  
  34. Vercel Deploy Helper
&lt;/h4&gt;

&lt;p&gt;Category: Integration&lt;br&gt;&lt;br&gt;
Tier: Situational&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Guides Vercel deployments with env var management, preview URLs, rollback procedures.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/opencode-community/vercel-deploy.git ~/.agents/skills/vercel-deploy
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Deploy this branch to Vercel preview”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;Smooths Vercel-specific friction points. Env var handling especially useful. Niche but valuable for Vercel shops.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;Vercel-only. Useless elsewhere.&lt;/p&gt;

&lt;h4&gt;
  
  
  35. Supabase Schema Migrator
&lt;/h4&gt;

&lt;p&gt;Category: Integration&lt;br&gt;&lt;br&gt;
Tier: Situational&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Generates Supabase migrations with RLS policies, types, seed data. Validates against Supabase constraints.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx skills add https://github.com/opencode-community/supabase-migrator
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Generate migration for users table with RLS”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;RLS policy generation alone worth it. Type safety catches mismatches. Supabase-specific obviously.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;Supabase evolves fast. Verify generated SQL against current version.&lt;/p&gt;

&lt;h3&gt;
  
  
  Testing, Security &amp;amp; Quality (7 Skills)
&lt;/h3&gt;

&lt;p&gt;For automated testing, security audits, and code quality.&lt;/p&gt;

&lt;h4&gt;
  
  
  36. Anthropic Webapp Testing
&lt;/h4&gt;

&lt;p&gt;Category: Testing&lt;br&gt;&lt;br&gt;
Tier: Power User&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Repeatable Playwright-based testing for local web apps. Reconnaissance-then-action workflow. Bundled server lifecycle helper.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx skills add https://github.com/anthropics/skills &lt;span class="nt"&gt;--skill&lt;/span&gt; webapp-testing
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Use webapp-testing to verify checkout form at localhost:5173”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;Inspect-first approach prevents brittle selectors. Server helper eliminates flaky startup. Best browser testing skill we’ve used.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;Python Playwright dependency. Setup overhead for non-Python projects.&lt;/p&gt;

&lt;h4&gt;
  
  
  37. Unit Test Generator
&lt;/h4&gt;

&lt;p&gt;Category: Testing&lt;br&gt;&lt;br&gt;
Tier: Power User&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Generates unit tests with edge cases, mocks, assertions. Framework-aware (Jest, Vitest, pytest).&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx skills add https://github.com/opencode-community/unit-test-gen
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Generate unit tests for the payment service”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;Coverage boost without tedium. Edge case suggestions surprisingly good. Always review generated tests — don’t trust blindly.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;Can generate passing-but-meaningless tests. Assert behavior, not implementation.&lt;/p&gt;

&lt;h4&gt;
  
  
  38. Security Audit Runner
&lt;/h4&gt;

&lt;p&gt;Category: Testing&lt;br&gt;&lt;br&gt;
Tier: Power User&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Runs OWASP-aligned security scans. Checks for injection, auth flaws, secrets exposure. Prioritized findings.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/opencode-community/security-audit.git ~/.agents/skills/security-audit
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Run security audit on the auth module”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;Catches real vulnerabilities, not just noise. Prioritization helps triage. Complement to, not replacement for, professional pentests.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;False positives on complex patterns. Human review required.&lt;/p&gt;

&lt;h4&gt;
  
  
  39. Performance Profiler
&lt;/h4&gt;

&lt;p&gt;Category: Testing&lt;br&gt;&lt;br&gt;
Tier: Situational&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Identifies performance bottlenecks via static analysis + runtime hints. Suggests optimizations with expected impact.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx skills add https://github.com/opencode-community/perf-profiler
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Profile the dashboard page for performance issues”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;Catches N+1 queries, unnecessary re-renders, bundle bloat. Impact estimates help prioritize. Stack-specific accuracy varies.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;Runtime profiling needs instrumentation. Static analysis misses dynamic issues.&lt;/p&gt;

&lt;h4&gt;
  
  
  40. Accessibility Checker
&lt;/h4&gt;

&lt;p&gt;Category: Testing&lt;br&gt;&lt;br&gt;
Tier: Situational&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;WCAG-aligned accessibility audit. Checks contrast, ARIA, keyboard nav, screen reader compatibility.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/opencode-community/a11y-checker.git ~/.agents/skills/a11y-checker
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Check accessibility of the signup form”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;Catches issues automated tools miss. Remediation suggestions actionable. Essential for public-facing apps.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;Can’t replace user testing with assistive tech. Supplement, not substitute.&lt;/p&gt;

&lt;h4&gt;
  
  
  41. Lint Rule Enforcer
&lt;/h4&gt;

&lt;p&gt;Category: Testing&lt;br&gt;&lt;br&gt;
Tier: Situational&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Enforces team lint rules beyond ESLint/Prettier defaults. Custom rule definitions, auto-fix suggestions.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx skills add https://github.com/opencode-community/lint-enforcer
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Enforce team lint rules on this PR”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;Codifies team conventions. Auto-fix reduces nitpick comments. Define rules once, enforce everywhere.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;Overly strict rules frustrate. Allow overrides for edge cases.&lt;/p&gt;

&lt;h4&gt;
  
  
  42. Mutation Tester
&lt;/h4&gt;

&lt;p&gt;Category: Testing&lt;br&gt;&lt;br&gt;
Tier: Situational&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Generates mutation tests to validate test suite effectiveness. Identifies weak spots in coverage.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/opencode-community/mutation-tester.git ~/.agents/skills/mutation-tester
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Run mutation testing on the payment module”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;Reveals tests that pass but don’t actually validate behavior. Expensive but invaluable for critical paths.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;Computationally expensive. Reserve for high-risk modules.&lt;/p&gt;

&lt;h4&gt;
  
  
  Meta-Skills &amp;amp; Optimization (5 Skills)
&lt;/h4&gt;

&lt;p&gt;For managing skills themselves, reducing costs, and personalizing workflows.&lt;/p&gt;

&lt;h4&gt;
  
  
  43. skill-optimizer
&lt;/h4&gt;

&lt;p&gt;Category: Meta&lt;br&gt;&lt;br&gt;
Tier: Situational&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Three-tool kit: skill-miner (finds workflows worth encoding), skill-personalizer (tunes triggers/descriptions), skill-generalizer (prepares for publication).&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/hqhq1025/skill-optimizer.git ~/.agents/skills/
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Run skill-miner on my last month of sessions”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;Meta-value only kicks in after weeks of usage. skill-miner reliably surfaces 3–4 encode-worthy workflows. skill-personalizer catches silent trigger failures. Small user base but sharp mental model.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;Thin docs. Read SKILL.md carefully before running on production archives.&lt;/p&gt;

&lt;h4&gt;
  
  
  44. Anthropic Skill Creator
&lt;/h4&gt;

&lt;p&gt;Category: Meta&lt;br&gt;&lt;br&gt;
Tier: Power User&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Creates, tests, benchmarks SKILL.md files. Includes eval framework, trigger optimization, pass-rate aggregation.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx skills add https://github.com/anthropics/skills &lt;span class="nt"&gt;--skill&lt;/span&gt; skill-creator
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Use skill-creator to turn our release workflow into a reusable skill”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;Anthropic’s own tool for skill authoring. Eval framework prevents shipping broken skills. Benchmark data informs trigger tuning. Essential for custom skill development.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;Anthropic-centric patterns. May need adaptation for non-Claude models.&lt;/p&gt;

&lt;h4&gt;
  
  
  45. Token Usage Tracker
&lt;/h4&gt;

&lt;p&gt;Category: Meta&lt;br&gt;&lt;br&gt;
Tier: Situational&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Real-time token/cost tracking per session/skill. Historical trends, anomaly alerts.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/opencode-community/token-tracker.git ~/.agents/skills/token-tracker
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Show token usage for this session”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;Visibility into spend drivers. Caught a runaway skill burning 10x expected tokens. Pair with Caveman for active reduction.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;Tracking overhead. Disable during latency-sensitive work.&lt;/p&gt;

&lt;h4&gt;
  
  
  46. Session Compactor
&lt;/h4&gt;

&lt;p&gt;Category: Meta&lt;br&gt;&lt;br&gt;
Tier: Situational&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Intelligent context compaction preserving key decisions/artifacts. Better than default /compact.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx skills add https://github.com/opencode-community/session-compactor
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Compact this session preserving architecture decisions”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;Retains more signal than default compaction. Extends usable session length. Handoff still better for cross-session work.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;Lossy by nature. Review compacted context before continuing.&lt;/p&gt;

&lt;h4&gt;
  
  
  47. Skill Health Checker
&lt;/h4&gt;

&lt;p&gt;Category: Meta&lt;br&gt;&lt;br&gt;
Tier: Situational&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Audits installed skills for staleness, broken installs, trigger conflicts. Maintenance dashboard.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/opencode-community/skill-health.git ~/.agents/skills/skill-health
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Check health of all installed skills”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;Prevents silent skill rot. Caught 3 abandoned skills we’d forgotten about. Run monthly as maintenance ritual.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;New tool, limited detection patterns. False negatives possible.&lt;/p&gt;

&lt;h4&gt;
  
  
  Niche &amp;amp; Experimental (3 Skills)
&lt;/h4&gt;

&lt;p&gt;For creative workflows, personal knowledge, and cutting-edge experiments.&lt;/p&gt;

&lt;h4&gt;
  
  
  48. Vault Daydream
&lt;/h4&gt;

&lt;p&gt;Category: Niche&lt;br&gt;&lt;br&gt;
Tier: Situational&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Samples random Obsidian note pairs, dispatches generator/critic subagents to find non-obvious connections. Writes insights back as notes. Mimics brain’s default mode network.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/glebis/claude-skills.git &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;cp&lt;/span&gt; &lt;span class="nt"&gt;-r&lt;/span&gt; claude-skills/daydream ~/.agents/skills/
&lt;span class="c"&gt;# Edit instructions.md to set VAULT_ROOT&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“/daydream over notes modified in last 60 days”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;Favorite non-code skill. Returns 5–10 genuinely novel connections per run. Route generators to Sonnet, critics to Haiku for cost efficiency. ~$0.40–0.50/run on Claude. Runs great during coffee breaks.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;Claude-specific Task() invocations. Needs adaptation for non-Anthropic providers. Budget for cost.&lt;/p&gt;

&lt;h4&gt;
  
  
  49. Anthropic Frontend Design
&lt;/h4&gt;

&lt;p&gt;Category: Niche&lt;br&gt;&lt;br&gt;
Tier: Situational&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Forces explicit aesthetic direction before UI code. Covers typography, color, motion, layout. Discourages generic AI UI patterns.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx skills add https://github.com/anthropics/skills &lt;span class="nt"&gt;--skill&lt;/span&gt; frontend-design
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Use frontend-design to redesign onboarding page”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;Eliminates purple-gradient-slop default UIs. Named aesthetic direction produces distinctive results. Supports React/Vue/HTML/CSS.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;Design taste subjective. May conflict with existing design systems.&lt;/p&gt;

&lt;h4&gt;
  
  
  50. Anthropic MCP Builder
&lt;/h4&gt;

&lt;p&gt;Category: Niche&lt;br&gt;&lt;br&gt;
Tier: Situational&lt;/p&gt;

&lt;h4&gt;
  
  
  What it does
&lt;/h4&gt;

&lt;p&gt;Guides full MCP server lifecycle: protocol, SDK setup, Zod/Pydantic schemas, eval creation. Treats evals as part of build.&lt;/p&gt;

&lt;h4&gt;
  
  
  Install
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx skills add https://github.com/anthropics/skills &lt;span class="nt"&gt;--skill&lt;/span&gt; mcp-builder
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Trigger Phrase
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;“Use mcp-builder to create TypeScript MCP server for internal API”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Human Verdict
&lt;/h4&gt;

&lt;p&gt;Best MCP onboarding resource. Eval-first approach prevents shipping broken servers. TypeScript + streamable HTTP defaults sensible.&lt;/p&gt;

&lt;h4&gt;
  
  
  Watch Out For
&lt;/h4&gt;

&lt;p&gt;MCP spec evolving. Verify against latest protocol docs.&lt;/p&gt;

&lt;h3&gt;
  
  
  Final Thoughts
&lt;/h3&gt;

&lt;p&gt;We tested 87 skills and kept 50. The rest were abandoned, bloated, or solving problems we didn’t have.&lt;/p&gt;

&lt;p&gt;Don’t install all 50. Pick 3–5 that match your daily friction points — docs, legacy code, SaaS integrations, or cost control — and uninstall anything that doesn’t earn its context weight within two weeks. The best skill you’ll ever use is the one you write for a task you repeat every week.&lt;/p&gt;

&lt;p&gt;Skills fix what limits every coding agent: context and workflow. OpenCode’s model-agnostic design lets you pin the right behavior across any LLM, and the Agent Skills open standard means your SKILL.md files work across OpenCode, Claude Code, Codex, and Cursor without modification.&lt;/p&gt;

&lt;h3&gt;
  
  
  Thank you so much for reading
&lt;/h3&gt;

&lt;p&gt;Like | Follow | Subscribe to the newsletter.&lt;/p&gt;

&lt;p&gt;Catch us on&lt;/p&gt;

&lt;p&gt;Website: &lt;a href="https://www.techlatest.net/" rel="noopener noreferrer"&gt;https://www.techlatest.net/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Newsletter: &lt;a href="https://substack.com/@techlatestnet" rel="noopener noreferrer"&gt;https://substack.com/@techlatestnet&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Twitter: &lt;a href="https://twitter.com/TechlatestNet" rel="noopener noreferrer"&gt;https://twitter.com/TechlatestNet&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;LinkedIn: &lt;a href="https://www.linkedin.com/in/techlatest-net/" rel="noopener noreferrer"&gt;https://www.linkedin.com/in/techlatest-net/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;YouTube:&lt;a href="https://www.youtube.com/@techlatest_net/" rel="noopener noreferrer"&gt;https://www.youtube.com/@techlatest_net/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Blogs: &lt;a href="https://medium.com/@techlatest.net" rel="noopener noreferrer"&gt;https://medium.com/@techlatest.net&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Reddit Community: &lt;a href="https://www.reddit.com/user/techlatest_net/" rel="noopener noreferrer"&gt;https://www.reddit.com/user/techlatest_net/&lt;/a&gt;&lt;/p&gt;

</description>
      <category>opencode</category>
      <category>agents</category>
      <category>agentskills</category>
      <category>opensource</category>
    </item>
    <item>
      <title>TechLatest AI &amp; Tech Weekly #29</title>
      <dc:creator>TechLatest</dc:creator>
      <pubDate>Sat, 15 Aug 2026 10:48:03 +0000</pubDate>
      <link>https://dev.to/techlatestnet/techlatest-ai-tech-weekly-29-1b8g</link>
      <guid>https://dev.to/techlatestnet/techlatest-ai-tech-weekly-29-1b8g</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhraqax7k5n3msqu2g1a5.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhraqax7k5n3msqu2g1a5.png" width="800" height="534"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Welcome to this week’s edition of &lt;strong&gt;TechLatest AI &amp;amp; Tech Weekly&lt;/strong&gt;  👋&lt;/p&gt;

&lt;p&gt;Here’s a curated roundup of our latest blogs, notable product launches, and the most interesting AI &amp;amp; ML updates from Aug 11— Aug 15, 2026.&lt;/p&gt;

&lt;h3&gt;
  
  
  AI/ML News Roundup: Aug 11–Aug 15, 2026
&lt;/h3&gt;

&lt;p&gt;Key highlights from this week’s AI developments include frontier model advancements with agentic capabilities, massive funding rounds reshaping valuations, and practical product launches for developers and enterprises. These updates emphasize autonomous agents, infrastructure scaling, and open-weight benchmarks relevant to builders and researchers.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;LTX Releases LTX-2.5:&lt;/strong&gt; LTX launched &lt;strong&gt;LTX-2.5&lt;/strong&gt; , an open-weight world model for video generation, real-time applications, and physical AI. It can generate a 10-second 720p clip in about &lt;strong&gt;6.8 seconds on two NVIDIA GB200 GPUs&lt;/strong&gt; , with native multishot generation and NVIDIA acceleration. &lt;a href="https://datanorth.ai/news/ltx-releases-ltx-2-5-open-weights-video-world-model" rel="noopener noreferrer"&gt;Source&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cactus Compute Releases Needle 2:&lt;/strong&gt; Cactus Compute introduced &lt;strong&gt;Needle 2&lt;/strong&gt; , an open &lt;strong&gt;45M-parameter tool-calling model&lt;/strong&gt; packaged as a 14MB binary that can run a full session using around &lt;strong&gt;28MB of RAM&lt;/strong&gt; , targeting extremely lightweight agentic and tool-use workloads. &lt;a href="https://cactuscompute.com/needle" rel="noopener noreferrer"&gt;Source&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Liquid AI Releases LFM2.5-VL-3B:&lt;/strong&gt; Liquid AI launched &lt;strong&gt;LFM2.5-VL-3B&lt;/strong&gt; , a 3.1B-parameter open-weight vision-language model designed for &lt;strong&gt;on-device AI&lt;/strong&gt;. It focuses on screen understanding, OCR, visual grounding, and efficient deployment on edge devices. &lt;a href="https://www.liquid.ai/blog/lfm2-5-vl-3b" rel="noopener noreferrer"&gt;Source&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;NVIDIA Releases Nemotron 3.5 Lightning &amp;amp; NeMo Switchyard:&lt;/strong&gt; NVIDIA introduced &lt;strong&gt;Nemotron 3.5 Lightning&lt;/strong&gt; , an open &lt;strong&gt;30B MoE model with 3B active parameters&lt;/strong&gt; , alongside &lt;strong&gt;NeMo Switchyard&lt;/strong&gt; , an open model-routing library designed to select models based on task requirements and improve efficiency. &lt;a href="https://blogs.nvidia.com/blog/nemotron-lightning-switchyard-rtx-dgx/" rel="noopener noreferrer"&gt;Source&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Meta Open-Sources Muse Glimmer:&lt;/strong&gt; Meta released &lt;strong&gt;Muse Glimmer&lt;/strong&gt; , a roughly 30B-parameter open-weight agentic model designed to run locally on a Mac or PC with a single GPU. &lt;a href="https://developer.meta.com/ai/lp/muse-glimmer/?utm_source=search&amp;amp;utm_medium=Muse-Code-Tier-1&amp;amp;utm_campaign=paid&amp;amp;utm_term&amp;amp;gad_source=1&amp;amp;gad_campaignid=24121174236&amp;amp;gbraid=0AAAAAC2gQWAaSiaEQDu6hadzRLBaEbFdX&amp;amp;gclid=Cj0KCQjwnIDUBhDrARIsAJDGwSs7DKmzCbjHApyGICq7AFJ0ALDrDiGiTXhtFkWIRRNMdC1B36e_e_8aAj6CEALw_wcB" rel="noopener noreferrer"&gt;Source&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Meta Also Opens Muse Spark 1.2:&lt;/strong&gt; Mark Zuckerberg announced that Meta will release the weights of the more powerful &lt;strong&gt;Muse Spark 1.2&lt;/strong&gt; , expanding Meta’s open-model lineup. &lt;a href="https://www.cnbc.com/2026/08/10/meta-muse-glimmer-open-weight-ai.html" rel="noopener noreferrer"&gt;Source&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Open Models Target Cost &amp;amp; Privacy:&lt;/strong&gt; Meta’s local AI strategy addresses growing concerns around &lt;strong&gt;AI bills, cloud dependence, and data privacy&lt;/strong&gt; , allowing businesses to run models on their own hardware. &lt;a href="https://www.pbs.org/newshour/show/what-metas-new-open-source-ai-model-means-for-the-future-of-artificial-intelligence" rel="noopener noreferrer"&gt;Source&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Zuckerberg Calls for US Open-AI Support:&lt;/strong&gt; Zuckerberg urged the US to reduce barriers for American developers building open-source AI, arguing that the country needs to compete with China’s rapidly growing open-model ecosystem. &lt;a href="https://finance.yahoo.com/technology/ai/articles/meta-ceo-zuckerberg-calls-fewer-105259757.html" rel="noopener noreferrer"&gt;Source&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;OpenAI IPO Filing Nears:&lt;/strong&gt; OpenAI’s &lt;strong&gt;S-1 filing&lt;/strong&gt; is expected in mid-to-late August, potentially revealing audited financials, its Microsoft revenue-sharing arrangement, and major business risks for the first time. &lt;a href="https://finance.yahoo.com/technology/ai/articles/openai-annualized-revenue-tops-40-111122841.html?guccounter=1&amp;amp;guce_referrer=aHR0cHM6Ly93d3cuZ29vZ2xlLmNvbS8&amp;amp;guce_referrer_sig=AQAAAGIR5Yr_D3gPy2LpZleKFsjt4-DmTWNxgICFHg0fpFg519bCnKmjlcRjoJoGMq8WwiuwmUzzO_XXr99eS0VzOZZwKKA8fbDP6Cb8OOtdcscSJIXCvy99G2Ia6oS3cCKWWh9EOFpD_MBHEV7uLlEPYuIY3buRRPaCBL93hOL-FuKG" rel="noopener noreferrer"&gt;Source&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Intel Raises $15B for Chip Manufacturing:&lt;/strong&gt; Intel secured &lt;strong&gt;$15 billion&lt;/strong&gt; to strengthen chip manufacturing and compete for AI-driven semiconductor demand. &lt;a href="https://newsroom.intel.com/corporate/intel-announces-proposed-15-billion-common-stock-offering" rel="noopener noreferrer"&gt;Source&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;South Korea Boosts Semiconductor Investment:&lt;/strong&gt; South Korea committed billions toward expanding its semiconductor ecosystem and maintaining its position in advanced chip manufacturing. &lt;a href="https://www.reuters.com/world/asia-pacific/south-koreas-lee-wants-military-airbase-relocated-by-mid-2028-chip-cluster-2026-08-10/" rel="noopener noreferrer"&gt;Source&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;TSMC Sales Jump 45%:&lt;/strong&gt; TSMC reported a &lt;strong&gt;45% increase in July sales&lt;/strong&gt; , signaling continued strong demand for AI chips and advanced semiconductor capacity. &lt;a href="https://ca.finance.yahoo.com/news/tsmc-sales-jumps-45-nvidia-121836175.html" rel="noopener noreferrer"&gt;Source&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Claude Opus 5 Leads Frontier Models:&lt;/strong&gt; Claude Opus 5 remains positioned at the top of the frontier model landscape, while GPT-5.6, Muse models, Qwen3.8-Max, Kimi K3, and other systems compete across different workloads. &lt;a href="https://www.buildfastwithai.com/blogs/best-ai-models-august-2026" rel="noopener noreferrer"&gt;Source&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI Builders Face a More Diverse Model Market:&lt;/strong&gt; Developers increasingly need to choose models based on &lt;strong&gt;coding, reasoning, agentic tasks, cost, privacy, and deployment requirements&lt;/strong&gt; rather than relying on a single overall winner.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;River AI&lt;/strong&gt; , founded by xAI co‑founder &lt;strong&gt;Igor Babuschkin&lt;/strong&gt; , announced a &lt;strong&gt;$1.1B seed/Series A&lt;/strong&gt; round led by &lt;strong&gt;General Catalyst&lt;/strong&gt; and &lt;strong&gt;AMP PBC&lt;/strong&gt; , with strategic investment from &lt;strong&gt;NVIDIA&lt;/strong&gt; and &lt;strong&gt;AMD Ventures&lt;/strong&gt; , plus &lt;strong&gt;Y Combinator&lt;/strong&gt; and  &lt;strong&gt;Temasek&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The company is building an &lt;strong&gt;open AI stack&lt;/strong&gt; for training, tuning, and serving &lt;strong&gt;customer-owned models&lt;/strong&gt; on private data, with a strong focus on &lt;strong&gt;personal agents&lt;/strong&gt; and enterprise customization.&lt;/p&gt;

&lt;p&gt;The round was announced on &lt;strong&gt;Aug 11, 2026&lt;/strong&gt; , just two months after River’s public launch, underscoring investor appetite for alternatives to closed frontier models.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Fireworks AI&lt;/strong&gt; closed a &lt;strong&gt;$1.505B Series D&lt;/strong&gt; at a &lt;strong&gt;$17.5B valuation&lt;/strong&gt; in mid‑July, backed by &lt;strong&gt;Atreides Management, Index Ventures, TCV&lt;/strong&gt; , and &lt;strong&gt;NVIDIA&lt;/strong&gt; , to scale custom-model training and inference for enterprises.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Together AI&lt;/strong&gt; raised an &lt;strong&gt;$800M Series C&lt;/strong&gt; at an &lt;strong&gt;$8.3B valuation&lt;/strong&gt; (led by &lt;strong&gt;Aramco Ventures&lt;/strong&gt; , with NVIDIA, Salesforce Ventures, General Catalyst, etc.) to expand its &lt;strong&gt;open-source model cloud&lt;/strong&gt; and agentic workflows.&lt;/li&gt;
&lt;li&gt;These rounds, together with River’s raise, show a clear pattern: investors are pouring capital into &lt;strong&gt;inference/training infrastructure&lt;/strong&gt; and &lt;strong&gt;custom-model platforms&lt;/strong&gt; rather than just frontier chatbots.&lt;/li&gt;
&lt;li&gt;Google parent &lt;strong&gt;Alphabet&lt;/strong&gt; announced a major reorganization of &lt;strong&gt;Google DeepMind&lt;/strong&gt; : &lt;strong&gt;Demis Hassabis&lt;/strong&gt; is stepping back from day‑to‑day operations to become &lt;strong&gt;Chairman of DeepMind&lt;/strong&gt; and &lt;strong&gt;Chief Scientist of Alphabet&lt;/strong&gt; , focusing on AGI research.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Koray Kavukcuoglu&lt;/strong&gt; , previously DeepMind’s CTO, is now the &lt;strong&gt;operational head of Google DeepMind&lt;/strong&gt; , reporting directly to CEO &lt;strong&gt;Sundar Pichai&lt;/strong&gt; , and will oversee &lt;strong&gt;Gemini model development&lt;/strong&gt; , frontier research, and the Gemini app/platform teams.&lt;/li&gt;
&lt;li&gt;The shake‑up also saw veteran engineer &lt;strong&gt;Jeff Dean&lt;/strong&gt; and other key staff leave to found a new AI startup, as Google tries to accelerate its response to &lt;strong&gt;OpenAI&lt;/strong&gt; and &lt;strong&gt;Anthropic&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Anthropic&lt;/strong&gt; disclosed that during internal safety testing, some of its models &lt;strong&gt;breached real external organizations&lt;/strong&gt; while performing agentic tasks, highlighting risks in autonomous tool use and environment access.&lt;/li&gt;
&lt;li&gt;In the same period, a Reuters investigation described how &lt;strong&gt;Chinese military researchers&lt;/strong&gt; had been using outputs from U.S. models (including GPT‑3.5 and Anthropic systems) via distillation to train domestic defense AI, despite export controls.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Boston Dynamics&lt;/strong&gt; sold out its &lt;strong&gt;2026 production slots&lt;/strong&gt; for the electric &lt;strong&gt;Atlas&lt;/strong&gt; humanoid robot, with initial shipments going to partners including &lt;strong&gt;Hyundai RMAC&lt;/strong&gt; and &lt;strong&gt;Google DeepMind&lt;/strong&gt; for research and industrial pilots.&lt;/li&gt;
&lt;li&gt;The sell‑out underscores surging demand for &lt;strong&gt;bipedal platforms&lt;/strong&gt; in logistics, manufacturing, and lab automation, especially as software stacks for manipulation and mobility mature.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;LG&lt;/strong&gt; announced a new &lt;strong&gt;bipedal humanoid robot&lt;/strong&gt; built on &lt;strong&gt;NVIDIA Isaac GR00T&lt;/strong&gt; , with a public unveiling planned for &lt;strong&gt;Q1 2027&lt;/strong&gt; , and signed an MOU covering robotics, AI factories, and autonomous vehicles.&lt;/li&gt;
&lt;li&gt;These moves align with broader “ &lt;strong&gt;physical AI&lt;/strong&gt; ” trends: using frontier models and simulation to train robots for real‑world tasks in warehouses, homes, and factories.&lt;/li&gt;
&lt;li&gt;Google’s medical AI system &lt;strong&gt;AMIE&lt;/strong&gt; was reported to now support &lt;strong&gt;synchronous video consultations&lt;/strong&gt; , guiding patients through virtual physical exams and matching doctors on several clinical metrics in live tests.&lt;/li&gt;
&lt;li&gt;This pushes medical AI beyond text/chat triage into &lt;strong&gt;real‑time, multimodal clinical interactions&lt;/strong&gt; , raising both opportunities and regulatory questions under the EU AI Act’s health provisions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;NVIDIA&lt;/strong&gt; signed MOUs with major financial firms ( &lt;strong&gt;Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs, KKR&lt;/strong&gt; ) to mobilize &lt;strong&gt;$500B+&lt;/strong&gt; for AI infrastructure, using &lt;strong&gt;compute capacity as collateral&lt;/strong&gt; in financing structures.&lt;/li&gt;
&lt;li&gt;This “compute‑backed” financing model reflects how AI data centers are becoming &lt;strong&gt;bankable assets&lt;/strong&gt; similar to energy or telecom infrastructure, with long‑term revenue contracts from cloud providers and hyperscalers.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Accel&lt;/strong&gt; raised a &lt;strong&gt;$3.5B&lt;/strong&gt; fund focused on early‑stage global AI startups, backing companies like &lt;strong&gt;Anthropic, Cursor, and Perplexity&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Global AI venture funding in &lt;strong&gt;H1 2026&lt;/strong&gt; hit &lt;strong&gt;~$510B&lt;/strong&gt; , already surpassing all of 2025, with AI infrastructure, defense manufacturing, and energy as dominant themes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI content labeling&lt;/strong&gt; and &lt;strong&gt;transparency&lt;/strong&gt; are becoming standard expectations worldwide, with new rules in &lt;strong&gt;China, the EU, Singapore&lt;/strong&gt; , and growing pressure in the U.S.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Blogs We Published This Week
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;How to Run NVIDIA Nemotron 3.5 Lightning: 4 Free Methods&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
A practical guide covering &lt;strong&gt;four ways to run Nemotron 3.5 Lightning for free&lt;/strong&gt; , from using a local GPU to deploying it through a zero-code AI agent workflow.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://dev.to/techlatestnet/how-to-run-nvidia-nemotron-35-lightning-free-4-methods-from-local-gpu-to-zero-code-agent-opc"&gt;How to Run NVIDIA Nemotron 3.5 Lightning (Free): 4 Methods from Local GPU to Zero-Code Agent&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How to Install DeepSeek Harness: A Step-by-Step Setup Guide&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
 A developer-focused walkthrough for installing and setting up &lt;strong&gt;DeepSeek Harness&lt;/strong&gt; , with the steps needed to get started building with it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://dev.to/techlatestnet/how-to-install-deepseek-harness-a-step-by-step-setup-guide-for-developers-jn"&gt;How to Install DeepSeek Harness: A Step-by-Step Setup Guide for Developers&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Thank you so much for reading
&lt;/h3&gt;

&lt;p&gt;Like | Follow | Subscribe to the newsletter.&lt;/p&gt;

&lt;p&gt;Catch us on&lt;/p&gt;

&lt;p&gt;Website: &lt;a href="https://www.techlatest.net/" rel="noopener noreferrer"&gt;https://www.techlatest.net/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Newsletter: &lt;a href="https://substack.com/@techlatestnet" rel="noopener noreferrer"&gt;https://substack.com/@techlatestnet&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Twitter: &lt;a href="https://twitter.com/TechlatestNet" rel="noopener noreferrer"&gt;https://twitter.com/TechlatestNet&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;LinkedIn: &lt;a href="https://www.linkedin.com/in/techlatest-net/" rel="noopener noreferrer"&gt;https://www.linkedin.com/in/techlatest-net/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;YouTube:&lt;a href="https://www.youtube.com/@techlatest_net/" rel="noopener noreferrer"&gt;https://www.youtube.com/@techlatest_net/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Blogs: &lt;a href="https://medium.com/@techlatest.net" rel="noopener noreferrer"&gt;https://medium.com/@techlatest.net&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Reddit Community: &lt;a href="https://www.reddit.com/user/techlatest_net/" rel="noopener noreferrer"&gt;https://www.reddit.com/user/techlatest_net/&lt;/a&gt;&lt;/p&gt;

</description>
      <category>news</category>
      <category>newsandupdates</category>
      <category>newsletter</category>
      <category>newslettermarketing</category>
    </item>
    <item>
      <title>How to Install DeepSeek Harness: A Step-by-Step Setup Guide for Developers</title>
      <dc:creator>TechLatest</dc:creator>
      <pubDate>Thu, 13 Aug 2026 16:13:07 +0000</pubDate>
      <link>https://dev.to/techlatestnet/how-to-install-deepseek-harness-a-step-by-step-setup-guide-for-developers-jn</link>
      <guid>https://dev.to/techlatestnet/how-to-install-deepseek-harness-a-step-by-step-setup-guide-for-developers-jn</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fcdn-images-1.medium.com%2Fmax%2F1024%2F0%2ANeFu9vNB2tn8BF-i" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fcdn-images-1.medium.com%2Fmax%2F1024%2F0%2ANeFu9vNB2tn8BF-i" width="800" height="446"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Building an AI coding assistant used to require stitching together dozens of libraries — until DeepSeek Harness changed the game. This open-source framework treats every capability as a plugin, from the model to the tools to the UI. In this guide, I’ll walk you through setting it up locally using Ollama (or OpenRouter) so you can have a private, powerful AI pair programmer running on your own machine in under 10 minutes.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step-by-Step Guide to Set Up DeepSeek Harness
&lt;/h3&gt;

&lt;h4&gt;
  
  
  Step 1: Install/Update nvm (Node Version Manager)
&lt;/h4&gt;

&lt;p&gt;This installs or updates nvm, which lets you manage multiple Node.js versions easily.&lt;/p&gt;

&lt;p&gt;Command you ran:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-o-&lt;/span&gt; https://raw.githubusercontent.com/nvm-sh/nvm/v0.40.1/install.sh | bash
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fa3us6p823gsar4fq5six.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fa3us6p823gsar4fq5six.png" width="800" height="378"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  Step 2: Install Node.js v22 and Set It as Default
&lt;/h4&gt;

&lt;p&gt;This downloads, installs, and switches to Node.js version 22.23.2, then sets it as the default for all future terminal sessions.&lt;/p&gt;

&lt;p&gt;Commands you ran:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;nvm &lt;span class="nb"&gt;install &lt;/span&gt;22
nvm &lt;span class="nb"&gt;alias &lt;/span&gt;default 22
nvm use 22
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4vfwg1whi6kumckqec3d.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4vfwg1whi6kumckqec3d.png" width="799" height="334"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  Step 3: Verify the Installation
&lt;/h4&gt;

&lt;p&gt;This checks that Node.js and npm are now on the correct, compatible versions.&lt;/p&gt;

&lt;p&gt;Commands you ran:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;node &lt;span class="nt"&gt;-v&lt;/span&gt; &lt;span class="c"&gt;# Output: v22.23.2&lt;/span&gt;
npm &lt;span class="nt"&gt;-v&lt;/span&gt; &lt;span class="c"&gt;# Output: 10.9.8&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0n0qejixfi5db65ja7t3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0n0qejixfi5db65ja7t3.png" width="436" height="558"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  Step 4: Launch DeepSeek Harness
&lt;/h4&gt;

&lt;p&gt;Now that your environment is properly set up, this is the command to start the actual Web UI:&lt;/p&gt;

&lt;p&gt;bash&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx @deepseek-ai/dsh web
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Once it runs, open your browser to &lt;a href="http://127.0.0.1:3080." rel="noopener noreferrer"&gt;http://127.0.0.1:3080&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhqgde0b1mu4f01r9v7v1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhqgde0b1mu4f01r9v7v1.png" width="800" height="443"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  Step 5: Activate OpenRouter (Pre-configured Provider)
&lt;/h4&gt;

&lt;p&gt;OpenRouter is great because it gives you access to dozens of models (like Claude, GPT-4, DeepSeek, Llama) through a single API key, without needing a credit card for most free tiers.&lt;/p&gt;

&lt;p&gt;5.1. Click on “OpenRouter” in the list&lt;br&gt;&lt;br&gt;
On your Settings &amp;gt; Models page, scroll down and click on the openrouter tile/row from the list of providers.&lt;/p&gt;

&lt;p&gt;5.2. Enter your OpenRouter API Key&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;If you don’t have one, go to &lt;a href="https://openrouter.ai/keys" rel="noopener noreferrer"&gt;openrouter.ai/keys&lt;/a&gt;, sign up, and click “Create Key”. Copy it.&lt;/li&gt;
&lt;li&gt;Paste that key into the API Key field that appears for the OpenRouter provider.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;5.3. Pick a default model&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;In the settings for OpenRouter, you will see a “Models” or “Default Model” dropdown.&lt;/li&gt;
&lt;li&gt;Select any model you like. For coding, popular choices are:&lt;/li&gt;
&lt;li&gt;deepseek/deepseek-chat (Free/Cheap)&lt;/li&gt;
&lt;li&gt;google/gemini-2.0-flash-exp:free (Free)&lt;/li&gt;
&lt;li&gt;anthropic/claude-3.5-sonnet (Paid, but best for coding).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;5.4. Click “Apply” or “Save”&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Look at the bottom right of the settings panel and click the “Apply” button to save your changes.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;5.5. Go back to the Chat and select OpenRouter&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Click on the main Chat or New Session tab at the top.&lt;/li&gt;
&lt;li&gt;In the model selector dropdown (usually at the top of the chat box), switch from “DeepSeek” to “OpenRouter” (or specifically the model you just selected, like deepseek/deepseek-chat).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frw8wpp5ttf1gil84owdy.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frw8wpp5ttf1gil84owdy.png" width="800" height="330"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxzy41s77jd2gwdujyvwt.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxzy41s77jd2gwdujyvwt.png" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpuateff66sml479ewmf3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpuateff66sml479ewmf3.png" width="800" height="763"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6sf1vf8a940opg4d25yr.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6sf1vf8a940opg4d25yr.png" width="800" height="736"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4&gt;
  
  
  Step 6: The Final Step — Send Your First Coding Prompt
&lt;/h4&gt;

&lt;p&gt;This is the moment you’ve been waiting for. In the chat input box at the bottom (where you already typed “Hello!”), replace that with a real coding task and hit Send (Enter).&lt;/p&gt;

&lt;p&gt;Here are 3 great first prompts tailored to your blog project. Pick the one that fits best:&lt;/p&gt;

&lt;p&gt;Prompt Option 1 — Project Setup:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;I want to build a modern blog website. Look at my current workspace folder and help me set up the best folder structure. Suggest whether I should use React, Next.js, or plain HTML/CSS.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Prompt Option 2 — Write Actual Code:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight jsx"&gt;&lt;code&gt;&lt;span class="nx"&gt;Create&lt;/span&gt; &lt;span class="nx"&gt;a&lt;/span&gt; &lt;span class="nx"&gt;reusable&lt;/span&gt; &lt;span class="nx"&gt;React&lt;/span&gt; &lt;span class="nx"&gt;component&lt;/span&gt; &lt;span class="nx"&gt;called&lt;/span&gt; &lt;span class="nx"&gt;BlogCard&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt; &lt;span class="nx"&gt;It&lt;/span&gt; &lt;span class="nx"&gt;should&lt;/span&gt; &lt;span class="nx"&gt;accept&lt;/span&gt; &lt;span class="nx"&gt;props&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="nx"&gt;title&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;excerpt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;author&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;and&lt;/span&gt; &lt;span class="nx"&gt;date&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt; &lt;span class="nx"&gt;Style&lt;/span&gt; &lt;span class="nx"&gt;it&lt;/span&gt; &lt;span class="kd"&gt;with&lt;/span&gt; &lt;span class="nx"&gt;Tailwind&lt;/span&gt; &lt;span class="nx"&gt;CSS&lt;/span&gt; &lt;span class="nx"&gt;and&lt;/span&gt; &lt;span class="nx"&gt;make&lt;/span&gt; &lt;span class="nx"&gt;it&lt;/span&gt; &lt;span class="nx"&gt;responsive&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Prompt Option 3 — Debug/Refactor:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;Scan my current project files. Identify any outdated dependencies or code that could be improved, and suggest specific fixes.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  What Will Happen Next?
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;The agent will start working in real-time.&lt;/li&gt;
&lt;li&gt;You will see it reading files, writing code, and running terminal commands.&lt;/li&gt;
&lt;li&gt;If you open the Trajectory View (click the tab next to “Chat”), you can watch every single thought, tool call, and result — like having X-ray vision into the AI’s brain.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fcdn-images-1.medium.com%2Fmax%2F1024%2F0%2An6SM7pzYLjAlwfXm" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fcdn-images-1.medium.com%2Fmax%2F1024%2F0%2An6SM7pzYLjAlwfXm" width="1024" height="463"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Conclusion: What Now?
&lt;/h3&gt;

&lt;p&gt;You have transformed your computer into a powerful AI-driven development environment. DeepSeek Harness is now your personal AI coding assistant.&lt;/p&gt;

&lt;p&gt;Here’s what makes this setup special:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Everything is traceable: If the AI makes a mistake, you can rewind, fork, and replay the exact session.&lt;/li&gt;
&lt;li&gt;Modular: You can swap models, add tools, or create custom modes without touching the core code.&lt;/li&gt;
&lt;li&gt;Real-world ready: It edits files, runs shell commands, searches the web, and can even manage subagents for complex tasks.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Thank you so much for reading
&lt;/h3&gt;

&lt;p&gt;Like | Follow | Subscribe to the newsletter.&lt;/p&gt;

&lt;p&gt;Catch us on&lt;/p&gt;

&lt;p&gt;Website: &lt;a href="https://www.techlatest.net/" rel="noopener noreferrer"&gt;https://www.techlatest.net/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Newsletter: &lt;a href="https://substack.com/@techlatestnet" rel="noopener noreferrer"&gt;https://substack.com/@techlatestnet&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Twitter: &lt;a href="https://twitter.com/TechlatestNet" rel="noopener noreferrer"&gt;https://twitter.com/TechlatestNet&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;LinkedIn: &lt;a href="https://www.linkedin.com/in/techlatest-net/" rel="noopener noreferrer"&gt;https://www.linkedin.com/in/techlatest-net/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;YouTube:&lt;a href="https://www.youtube.com/@techlatest_net/" rel="noopener noreferrer"&gt;https://www.youtube.com/@techlatest_net/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Blogs: &lt;a href="https://medium.com/@techlatest.net" rel="noopener noreferrer"&gt;https://medium.com/@techlatest.net&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Reddit Community: &lt;a href="https://www.reddit.com/user/techlatest_net/" rel="noopener noreferrer"&gt;https://www.reddit.com/user/techlatest_net/&lt;/a&gt;&lt;/p&gt;

</description>
      <category>deepseek</category>
      <category>llm</category>
      <category>opensource</category>
      <category>aiagentharness</category>
    </item>
    <item>
      <title>How to Run NVIDIA Nemotron 3.5 Lightning (Free): 4 Methods from Local GPU to Zero-Code Agent</title>
      <dc:creator>TechLatest</dc:creator>
      <pubDate>Wed, 12 Aug 2026 15:11:19 +0000</pubDate>
      <link>https://dev.to/techlatestnet/how-to-run-nvidia-nemotron-35-lightning-free-4-methods-from-local-gpu-to-zero-code-agent-opc</link>
      <guid>https://dev.to/techlatestnet/how-to-run-nvidia-nemotron-35-lightning-free-4-methods-from-local-gpu-to-zero-code-agent-opc</guid>
      <description>&lt;p&gt;NVIDIA Nemotron 3.5 Lightning Setup Guide: vLLM, Ollama, OpenRouter &amp;amp; Agent IDE (2026)&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fonus4nl0017bb75sbggo.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fonus4nl0017bb75sbggo.png" width="800" height="446"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 is a hybrid Mixture-of-Experts (MoE) model featuring Mamba-2, MoE, and Attention layers. It has 30B total parameters with only 3B active per token, making it highly efficient for research, post-training (SFT/RL), and customization.&lt;/p&gt;

&lt;p&gt;Critical Distinction: This BF16 version is the full-precision reference weight intended for training, fine-tuning, and evaluation. For production inference with optimized latency/throughput, NVIDIA recommends using the NVFP4 quantized variant instead. This guide focuses strictly on deploying this BF16 checkpoint.&lt;/p&gt;

&lt;h3&gt;
  
  
  Key Specifications
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Active Params: 3B | Total Params: 30B&lt;/li&gt;
&lt;li&gt;Context: Up to 1M tokens (256K recommended for single H100)&lt;/li&gt;
&lt;li&gt;Architecture: Hybrid Mamba-2 + MoE + Attention&lt;/li&gt;
&lt;li&gt;License: OpenMDW-1.1 (Commercial use allowed)&lt;/li&gt;
&lt;li&gt;Release Date: August 11, 2026&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Benchmarks
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fcdn-images-1.medium.com%2Fmax%2F1024%2F0%2AWv_rFF4CUGQXvyid" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fcdn-images-1.medium.com%2Fmax%2F1024%2F0%2AWv_rFF4CUGQXvyid" width="1024" height="499"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Supported GPU Configurations
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxjj7a5cy4kbqt13za64c.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxjj7a5cy4kbqt13za64c.png" width="800" height="211"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Multi-GPU Configuration (Full 1M Context)
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9pua575gzog04ji5rrwe.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9pua575gzog04ji5rrwe.png" width="800" height="76"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Key Memory Considerations
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftauf2a9z4cv4qmsdj1g3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftauf2a9z4cv4qmsdj1g3.png" width="800" height="171"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Hardware &amp;amp; System Prerequisites
&lt;/h3&gt;

&lt;p&gt;Before installation, verify your system meets these requirements. The BF16 weights require significant VRAM compared to quantized versions.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fo5zmxrr3cof6h5wb88ed.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fo5zmxrr3cof6h5wb88ed.png" width="798" height="191"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Step-by-Step Process to Install &amp;amp; Run NVIDIA Nemotron 3.5 Lightning
&lt;/h3&gt;

&lt;h4&gt;
  
  
  Step 1: Create an Isolated Python Environment
&lt;/h4&gt;

&lt;p&gt;Begin by creating a dedicated conda environment to prevent dependency conflicts with any existing PyTorch or CUDA installations on your system. Use Python 3.11 as it has the best compatibility with the current vLLM nightly builds and Mamba-2 kernel compilation. After activating the environment, install PyTorch with CUDA 12.6 support from the official PyTorch wheel index, then verify that the installation confirms CUDA availability.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;conda create &lt;span class="nt"&gt;-n&lt;/span&gt; nemotron-lightning &lt;span class="nv"&gt;python&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;3.11 &lt;span class="nt"&gt;-y&lt;/span&gt;
conda activate nemotron-lightning

pip &lt;span class="nb"&gt;install &lt;/span&gt;torch torchvision torchaudio &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--index-url&lt;/span&gt; https://download.pytorch.org/whl/cu126

python &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="s2"&gt;"import torch; print(f'PyTorch: {torch. __version__ }, CUDA: {torch.cuda.is_available()}')"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Step 2: Install vLLM Nightly
&lt;/h4&gt;

&lt;p&gt;This model requires vLLM nightly version 0.27.1 or later because stable vLLM releases do not yet support the hybrid Mamba-2 and MoE architecture used by Nemotron 3.5 Lightning. Install the pre-release build using pip with the extra index URL pointing to the PyTorch nightly CUDA 12.6 wheels, then verify the reported version meets the minimum requirement. If you encounter build errors, ensure your system has GCC 11 or newer, ninja-build, and the CUDA 12.6 toolkit installed, as vLLM nightly compiles custom Mamba kernels at install time.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;vllm&amp;gt;&lt;span class="o"&gt;=&lt;/span&gt;0.27.1 &lt;span class="nt"&gt;--pre&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--extra-index-url&lt;/span&gt; https://download.pytorch.org/whl/nightly/cu126

python &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="s2"&gt;"import vllm; print(f'vLLM: {vllm. __version__ }')"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Step 3: Set Model Checkpoint Variables
&lt;/h4&gt;

&lt;p&gt;Define the BF16 checkpoint path as an environment variable so you can reuse it across all deployment commands without retyping the full model identifier. If you plan to use DSpark speculative decoding on GB200 hardware, also define a second variable pointing to the NVFP4-DSpark draft model checkpoint. These variables will be referenced in every serve command throughout this guide.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;MODEL_CKPT&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;DSPARK_CKPT&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4-DSpark
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Step 4: Launch the vLLM Server
&lt;/h4&gt;

&lt;p&gt;Choose the deployment configuration that matches your hardware and launch the vLLM OpenAI-compatible server. For a single H100 or A100 80GB GPU, use the max-throughput configuration with prefix caching, async scheduling, the FlashInfer Mamba backend, and the Mamba SSM cache dtype set to float16 to conserve VRAM. Always include the reasoning parser, tool call parser, and auto tool choice flags regardless of hardware configuration.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;vllm serve &lt;span class="nv"&gt;$MODEL_CKPT&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--max-num-seqs&lt;/span&gt; 128 &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--enable-prefix-caching&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--async-scheduling&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--mamba-backend&lt;/span&gt; flashinfer &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--mamba-ssm-cache-dtype&lt;/span&gt; float16 &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--enable-mamba-cache-stochastic-rounding&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--mamba-cache-philox-rounds&lt;/span&gt; 5 &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--reasoning-parser&lt;/span&gt; nemotron_v3 &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--tool-call-parser&lt;/span&gt; qwen3_coder &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--enable-auto-tool-choice&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For eight H100 or H200 GPUs targeting the full 1M token context, enable tensor parallelism size 8, expert parallelism, and set the max model length to 1048576 with the long-context environment variable.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;VLLM_ALLOW_LONG_MAX_MODEL_LEN&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;1 vllm serve &lt;span class="nv"&gt;$MODEL_CKPT&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--moe-backend&lt;/span&gt; flashinfer_cutlass &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--mamba-backend&lt;/span&gt; flashinfer &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--enable-prefix-caching&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--mamba-cache-mode&lt;/span&gt; align &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--max-model-len&lt;/span&gt; 1048576 &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--enable-expert-parallel&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--tensor-parallel-size&lt;/span&gt; 8 &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--reasoning-parser&lt;/span&gt; nemotron_v3 &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--tool-call-parser&lt;/span&gt; qwen3_coder &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--enable-auto-tool-choice&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For GB200 hardware with DSpark speculative decoding, include the speculative config pointing to the DSpark draft model with five speculative tokens and disable prefix caching as required by the DSpark recipe.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;VLLM_ALLOW_LONG_MAX_MODEL_LEN&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;1 vllm serve &lt;span class="nv"&gt;$MODEL_CKPT&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--max-num-seqs&lt;/span&gt; 128 &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--max-model-len&lt;/span&gt; 1048576 &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--max-num-batched-tokens&lt;/span&gt; 10240 &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--no-enable-prefix-caching&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--async-scheduling&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--speculative_config&lt;/span&gt;.model &lt;span class="nv"&gt;$DSPARK_CKPT&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--speculative_config&lt;/span&gt;.num_speculative_tokens 5 &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--mamba-backend&lt;/span&gt; flashinfer &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--reasoning-parser&lt;/span&gt; nemotron_v3 &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--tool-call-parser&lt;/span&gt; qwen3_coder &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--enable-auto-tool-choice&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Context Length Tip: If you are memory-constrained or want more KV-cache headroom at high concurrency, lower --max-model-len to match your actual workload and remove VLLM_ALLOW_LONG_MAX_MODEL_LEN=1.&lt;/p&gt;

&lt;h4&gt;
  
  
  Step 5: Verify the Deployment and Test Inference
&lt;/h4&gt;

&lt;p&gt;Once the server is running, confirm it is healthy by sending a request to the OpenAI-compatible endpoint using the recommended sampling parameters of temperature 1.0 and top_p 0.95. Test basic chat completion first, then verify tool calling by passing a function definition and including the force_nonempty_content chat template kwarg required for coding agents. Finally, test reasoning mode toggling by sending requests with enable_thinking set to both True and False.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;http://localhost:8000/v1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;EMPTY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;MODEL&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="n"&gt;tools&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;function&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;function&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;get_weather&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;description&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Get current weather for a city&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;parameters&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;object&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;properties&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;city&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;string&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}},&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;required&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;city&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;}]&lt;/span&gt;

&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;MODEL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;What&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s the weather in Santa Clara?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
    &lt;span class="n"&gt;tools&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;tools&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;16000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;temperature&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;1.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;top_p&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.95&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;extra_body&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;chat_template_kwargs&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;force_nonempty_content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;}},&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;tool_calls&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Controlling Reasoning Mode
&lt;/h4&gt;

&lt;p&gt;Reasoning can be toggled at runtime through chat template kwargs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Enable Thinking: extra_body={"chat_template_kwargs": {"enable_thinking": True}}&lt;/li&gt;
&lt;li&gt;Disable Thinking: extra_body={"chat_template_kwargs": {"enable_thinking": False}}&lt;/li&gt;
&lt;li&gt;Thinking Budget: extra_body={"chat_template_kwargs": {"thinking_budget": 4096}}&lt;/li&gt;
&lt;/ul&gt;

&lt;h4&gt;
  
  
  Step 6: Post-Training &amp;amp; Customization Notes
&lt;/h4&gt;

&lt;p&gt;Since this BF16 release is primarily designed as a base for customization:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;SFT / RL: Use NeMo RL and NeMo Gym frameworks. The model was trained with GRPO across math, code, science, and tool-use environments.&lt;/li&gt;
&lt;li&gt;Quantization: Produce your own NVFP4, W4A16, or GGUF variants from this checkpoint for edge/mobile deployment.&lt;/li&gt;
&lt;li&gt;Datasets: Pre-training data (20T+ tokens, cutoff Sep 2025) and post-training data (cutoff May 2026) are available via nvidia/nemotron-pre-training-datasets and nvidia/nemotron-post-training-v3.&lt;/li&gt;
&lt;li&gt;Evaluation: Reproducible benchmarks are published in NeMo Gym with exact harness configurations.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Method 2: Run Nemotron 3.5 Lightning Using Ollama
&lt;/h3&gt;

&lt;p&gt;NVIDIA has published an official Ollama model entry for Nemotron 3.5 Lightning, making local deployment significantly simpler than the vLLM approach. The model is available as a ready-to-run Ollama package at 25GB with support for up to 1M context length, tool calling, and reasoning modes out of the box. This method is ideal for developers who want to get started quickly without managing complex serving configurations, and it integrates directly with agent harnesses such as Claude Code, OpenCode, Hermes Agent, and OpenClaw via the NVIDIA NemoClaw stack.&lt;/p&gt;

&lt;h4&gt;
  
  
  Step 1: Install Ollama
&lt;/h4&gt;

&lt;p&gt;Install the latest version of Ollama on your Linux or macOS machine using the official install script. After installation, confirm that the Ollama service is running and that your NVIDIA GPU is detected. Windows users can download the Ollama installer directly from the official website. Ensure your NVIDIA drivers are updated to the latest stable release so Ollama can fully utilize your GPU for inference.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-fsSL&lt;/span&gt; https://ollama.com/install.sh | sh

ollama &lt;span class="nt"&gt;--version&lt;/span&gt;

nvidia-smi
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Step 2: Pull the Official Nemotron 3.5 Lightning Model
&lt;/h4&gt;

&lt;p&gt;Pull the official Nemotron 3.5 Lightning model directly from the Ollama registry. The default tag downloads the 30B parameter variant at 25GB with 1M context support. If you are on Apple Silicon hardware, use the MLX-optimized tag instead, which is 23GB and supports up to 256K context. Ollama will automatically download and cache the model weights in its local library upon completion.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ollama pull nemotron-3.5-lightning

ollama list
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For Apple Silicon Macs, use the MLX variant:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ollama pull nemotron-3.5-lightning:30b-mlx
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Step 3: Run Interactive Inference
&lt;/h4&gt;

&lt;p&gt;Launch an interactive chat session with the model using a single command. Ollama handles GPU layer offloading, memory management, and streaming responses automatically. Type your prompts directly into the terminal and observe real-time output. The model supports both direct answers and reasoning modes natively through its built-in chat template.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ollama run nemotron-3.5-lightning

&lt;span class="o"&gt;&amp;gt;&amp;gt;&amp;gt;&lt;/span&gt; What is speculative decoding and how does DSpark work?
&lt;span class="o"&gt;&amp;gt;&amp;gt;&amp;gt;&lt;/span&gt; Write a Python &lt;span class="k"&gt;function &lt;/span&gt;that implements binary search
&lt;span class="o"&gt;&amp;gt;&amp;gt;&amp;gt;&lt;/span&gt; /set parameter temperature 1.0
&lt;span class="o"&gt;&amp;gt;&amp;gt;&amp;gt;&lt;/span&gt; /set parameter top_p 0.95
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Step 4: Launch with Agent Harnesses
&lt;/h4&gt;

&lt;p&gt;Nemotron 3.5 Lightning is purpose-built for always-on AI agents and ships with one-command launch support for popular agent harnesses. Use the ollama launch command to start the model inside Claude Code, OpenCode, Hermes Agent, or OpenClaw directly. Each harness is preconfigured to leverage the model’s tool-calling and reasoning capabilities for autonomous task execution across personal productivity, financial services, cybersecurity, telecom, and retail workflows.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ollama launch claude &lt;span class="nt"&gt;--model&lt;/span&gt; nemotron-3.5-lightning

ollama launch opencode &lt;span class="nt"&gt;--model&lt;/span&gt; nemotron-3.5-lightning

ollama launch hermes &lt;span class="nt"&gt;--model&lt;/span&gt; nemotron-3.5-lightning

ollama launch openclaw &lt;span class="nt"&gt;--model&lt;/span&gt; nemotron-3.5-lightning
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Step 5: Use the Ollama API for Application Integration
&lt;/h4&gt;

&lt;p&gt;Ollama exposes an OpenAI-compatible REST API on port 11434 by default. To allow external applications or remote machines to access the model, set the OLLAMA_HOST environment variable before starting the service. Then use any OpenAI-compatible client library in Python, JavaScript, or cURL to send requests, specifying nemotron-3.5-lightning as the model name in each call.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;OLLAMA_HOST&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;0.0.0.0:11434 ollama serve
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Python example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;http://localhost:11434/v1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ollama&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;nemotron-3.5-lightning&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Summarize the key risks in this financial document&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
    &lt;span class="n"&gt;temperature&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;1.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;top_p&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.95&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;8192&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;cURL example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl http://localhost:11434/v1/chat/completions &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "model": "nemotron-3.5-lightning",
    "messages": [{"role": "user", "content": "Classify this security alert by severity"}],
    "temperature": 1.0,
    "top_p": 0.95
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Step 6: Verify Performance and GPU Utilization
&lt;/h4&gt;

&lt;p&gt;Confirm that Ollama is fully utilizing your GPU by monitoring nvidia-smi while processing a request. You should see VRAM allocation matching the 25GB model size on your H100, A100, RTX 5090, or GB200 device. Nemotron 3.5 Lightning delivers 4x higher throughput and 30% lower task completion time compared to other leading open models of similar size, so benchmark your specific workload by sending repeated requests and measuring tokens per second. If you observe CPU fallback or slow generation, ensure OLLAMA_NUM_GPU is set high enough to force full layer offloading to your GPU.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;OLLAMA_NUM_GPU&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;99 ollama run nemotron-3.5-lightning

nvidia-smi
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When to Use Method 2 vs Method 1: Choose Ollama when you want the fastest path to local inference, need built-in agent harness integration, or are running on Apple Silicon with the MLX variant. Choose vLLM from Method 1 when you need maximum throughput at high concurrency, full 1M context on multi-GPU setups, speculative decoding with DSpark, or production-grade serving with prefix caching and async scheduling.&lt;/p&gt;

&lt;h3&gt;
  
  
  Method 3: Run Nemotron 3.5 Lightning Free via OpenRouter API
&lt;/h3&gt;

&lt;p&gt;If you do not want to spend money on cloud GPU credits or manage local hardware, you can run Nemotron 3.5 Lightning completely free through the OpenRouter API. NVIDIA hosts this model at zero cost for both input and output tokens with 1M context support, tool calling, and reasoning capabilities included. This is the fastest way to start building with a 30B MoE model without downloading weights or configuring GPUs.&lt;/p&gt;

&lt;h4&gt;
  
  
  Step 1: Get Your Free API Key
&lt;/h4&gt;

&lt;p&gt;Sign up at OpenRouter, navigate to your dashboard, and create a new API key. No credit card is required. Set it as an environment variable in your terminal so your code can access it securely without hardcoding secrets.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;OPENROUTER_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;sk-or-v1-your-actual-key-here
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Step 2: Make Your First Request
&lt;/h4&gt;

&lt;p&gt;Use the OpenAI-compatible endpoint with the free model identifier nvidia/nemotron-3.5-lightning: free. The optional HTTP-Referer and X-Title headers let your app appear on OpenRouter leaderboards but are not required for functionality.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://openrouter.ai/api/v1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;OPENROUTER_API_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;nvidia/nemotron-3.5-lightning:free&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;How many r&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s are in strawberry?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
    &lt;span class="n"&gt;temperature&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;1.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;top_p&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.95&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Step 3: Enable Streaming and Reasoning Tokens
&lt;/h4&gt;

&lt;p&gt;Add stream set to true to receive responses as server-sent events for real-time output. To see the model’s step-by-step thinking process, include the reasoning parameter in your request and read the reasoning_details array from the response. When continuing a conversation, always pass the complete reasoning_details back in the message history so the model can resume reasoning from where it left off.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-N&lt;/span&gt; https://openrouter.ai/api/v1/chat/completions &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$OPENROUTER_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "model": "nvidia/nemotron-3.5-lightning:free",
    "stream": true,
    "reasoning": {"enabled": true},
    "messages": [{"role": "user", "content": "Solve x^2 - 5x + 6 = 0"}]
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Step 4: Use Tool Calling and Alternative SDKs
&lt;/h4&gt;

&lt;p&gt;The free tier fully supports OpenAI-style tool calling, TypeScript SDKs, and Anthropic Messages API format through dedicated endpoints. Pass your function definitions in the tools array exactly as you would with OpenAI, and the model will return structured tool_calls in the response. For TypeScript projects, install the official @openrouter/sdk package and initialize it with your API key to get typed streaming responses with automatic reasoning token tracking in the usage object.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Endpoint Reference: Chat completions at POST /api/v1/chat/completions, Responses API at POST /api/v1/responses, and Anthropic-compatible messages at POST /api/v1/messages. All three accept the same free model identifier and bearer token authentication.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  Use Nemotron 3.5 Lightning in a Zero-Coding Agent IDE
&lt;/h3&gt;

&lt;p&gt;Use the model inside AI coding agents like OpenCode, Claude Code, or Hermes Agent without writing any code. Routes through OpenRouter free tier — zero cost.&lt;/p&gt;

&lt;p&gt;Step 1: Launch your agent IDE in the terminal&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F38c1lw69cq7xkzpiq8co.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F38c1lw69cq7xkzpiq8co.png" width="800" height="290"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Step 2: Select OpenRouter as provider in the setup wizard&lt;/p&gt;

&lt;p&gt;Step 3: Paste your free OpenRouter API key when prompted&lt;/p&gt;

&lt;p&gt;Step 4: Choose model → nvidia/nemotron-3.5-lightning:free&lt;/p&gt;

&lt;p&gt;Step 5: Press Enter to save when you see “Ready to connect”&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkwdycqvttq1yp77ilmlp.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkwdycqvttq1yp77ilmlp.png" width="799" height="299"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Conclusion
&lt;/h3&gt;

&lt;p&gt;NVIDIA Nemotron 3.5 Lightning represents a significant step forward in open, efficient AI models — delivering 30B total parameters with only 3B active per token through its hybrid Mamba-2 + MoE architecture. With support for up to 1M context tokens, native tool calling, configurable reasoning modes, and a commercially permissive OpenMDW-1.1 license, it is purpose-built for the next generation of always-on AI agents across coding, finance, cybersecurity, telecom, and retail workflows.&lt;/p&gt;

&lt;p&gt;Throughout this guide, we covered four distinct deployment paths to match every developer’s needs and infrastructure constraints:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Method 1 (vLLM) gives you maximum throughput, full 1M context on multi-GPU setups, speculative decoding with DSpark, and production-grade serving with prefix caching — ideal for data center deployments on H100, H200, A100, or GB200 hardware.&lt;/li&gt;
&lt;li&gt;Method 2 (Ollama) provides the fastest path to local inference with a single pull command, built-in agent harness integration for Claude Code, OpenCode, Hermes Agent, and OpenClaw, plus Apple Silicon support via the MLX variant — perfect for developers who want privacy and simplicity.&lt;/li&gt;
&lt;li&gt;Method 3 (OpenRouter API) lets you run the model completely free with zero infrastructure cost, no weight downloads, and instant access from any device — the best choice for prototyping, evaluation, and lightweight integrations.&lt;/li&gt;
&lt;li&gt;Method 4 (Zero-Coding Agent IDE) connects the free OpenRouter endpoint directly into AI coding agents through a visual provider wizard — enabling non-coders to leverage a 30B MoE model inside their editor without writing a single line of API code.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Whether you are fine-tuning the BF16 reference weights for domain-specific customization, deploying quantized NVFP4 variants for low-latency production inference, or simply experimenting with agentic workflows at zero cost, Nemotron 3.5 Lightning offers a flexible entry point that scales from a single laptop to an eight-GPU data center rack.&lt;/p&gt;

&lt;p&gt;The model’s release alongside open training datasets, reproducible NeMo Gym evaluation recipes, and the NVIDIA NemoClaw security stack for always-on agents signals NVIDIA’s commitment to an open ecosystem where developers retain full control over their AI infrastructure. As the model matures and community quantizations, GGUF conversions, and fine-tuned variants proliferate across Hugging Face and Ollama, expect Nemotron 3.5 Lightning to become a default backbone for specialized AI agents throughout 2026 and beyond.&lt;/p&gt;

&lt;p&gt;Pick the method that matches your hardware, budget, and use case — and start building with one of the most efficient open 30B models available today.&lt;/p&gt;

&lt;h3&gt;
  
  
  Thank you so much for reading
&lt;/h3&gt;

&lt;p&gt;Like | Follow | Subscribe to the newsletter.&lt;/p&gt;

&lt;p&gt;Catch us on&lt;/p&gt;

&lt;p&gt;Website: &lt;a href="https://www.techlatest.net/" rel="noopener noreferrer"&gt;https://www.techlatest.net/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Newsletter: &lt;a href="https://substack.com/@techlatestnet" rel="noopener noreferrer"&gt;https://substack.com/@techlatestnet&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Twitter: &lt;a href="https://twitter.com/TechlatestNet" rel="noopener noreferrer"&gt;https://twitter.com/TechlatestNet&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;LinkedIn: &lt;a href="https://www.linkedin.com/in/techlatest-net/" rel="noopener noreferrer"&gt;https://www.linkedin.com/in/techlatest-net/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;YouTube:&lt;a href="https://www.youtube.com/@techlatest_net/" rel="noopener noreferrer"&gt;https://www.youtube.com/@techlatest_net/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Blogs: &lt;a href="https://medium.com/@techlatest.net" rel="noopener noreferrer"&gt;https://medium.com/@techlatest.net&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Reddit Community: &lt;a href="https://www.reddit.com/user/techlatest_net/" rel="noopener noreferrer"&gt;https://www.reddit.com/user/techlatest_net/&lt;/a&gt;&lt;/p&gt;

</description>
      <category>nvidiagpu</category>
      <category>opensource</category>
      <category>nvidia</category>
      <category>aimodel</category>
    </item>
    <item>
      <title>TechLatest AI &amp; Tech Weekly #28</title>
      <dc:creator>TechLatest</dc:creator>
      <pubDate>Tue, 11 Aug 2026 14:11:44 +0000</pubDate>
      <link>https://dev.to/techlatestnet/techlatest-ai-tech-weekly-28-15fi</link>
      <guid>https://dev.to/techlatestnet/techlatest-ai-tech-weekly-28-15fi</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flxxtr3wvhhnkjjzyaca1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flxxtr3wvhhnkjjzyaca1.png" width="800" height="534"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Welcome to this week’s edition of &lt;strong&gt;TechLatest AI &amp;amp; Tech Weekly&lt;/strong&gt;  👋&lt;/p&gt;

&lt;p&gt;Here’s a curated roundup of our latest blogs, notable product launches, and the most interesting AI &amp;amp; ML updates from Aug 03 — Aug 10, 2026.&lt;/p&gt;

&lt;h3&gt;
  
  
  AI/ML News Roundup: Aug 03–Aug 10, 2026
&lt;/h3&gt;

&lt;p&gt;Key highlights from this week’s AI developments include frontier model advancements with agentic capabilities, massive funding rounds reshaping valuations, and practical product launches for developers and enterprises. These updates emphasize autonomous agents, infrastructure scaling, and open-weight benchmarks relevant to builders and researchers.&lt;/p&gt;

&lt;h3&gt;
  
  
  TL;DR
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Agentic AI accelerated:&lt;/strong&gt; NVIDIA NOOA, Prime Agent, QM, Shepherd, and CopilotKit Channels SDK expanded the tooling for building, coordinating, and managing AI agents.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Open models and multimodal AI advanced:&lt;/strong&gt; Meta’s Muse Glimmer, Mistral’s Shieldstral 1.0 3B, NVIDIA Alpamayo 2 Super, and NemotronLabs VoiceChat 11B pushed local, safety, autonomous-driving, and real-time voice AI forward.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI moved deeper into production:&lt;/strong&gt; Genspark’s GenOffice, Cursor’s Mixture-of-Kittens, TencentDB Agent Memory, and Microsoft’s code-testing generator focused on practical enterprise and developer workflows.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI security became a major concern:&lt;/strong&gt; A DeepSeek-powered attack reportedly targeted 460+ internet-facing systems, while new research highlighted the challenges of securing autonomous agents and open-weight models.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI infrastructure became a bottleneck:&lt;/strong&gt; Nuclear power, grid optimization, semiconductor investment, and massive AI data-center spending showed that electricity and compute capacity are becoming as important as model performance.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI spending faced greater scrutiny:&lt;/strong&gt; Large infrastructure investments are increasingly being judged on measurable business returns rather than AI potential alone.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI governance expanded:&lt;/strong&gt; EU and California transparency rules, alongside the new U.S. AI framework, pushed disclosure, provenance, safety, and responsible deployment higher on the industry agenda.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Frontier models kept advancing:&lt;/strong&gt; Qwen3.8-Max, Claude Opus 5, GPT-5.6, Astra, Kimi K3, and DeepSeek V4 showed that the frontier is becoming increasingly competitive and diverse.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI entered scientific research:&lt;/strong&gt; New work explored AI-designed biological systems, mathematical discovery, cryptography, and scientific software optimization, expanding AI beyond traditional chatbots and coding.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Physical AI gained momentum:&lt;/strong&gt; Autonomous driving, 3D-to-CAD workflows, mining, energy, robotics, and industrial applications showed AI increasingly moving from digital environments into the physical world.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Enterprise adoption continued:&lt;/strong&gt; Stripe’s Kai, Formula 1’s AI data accelerator, Atlassian’s agentic tooling, and other deployments showed AI agents moving from experiments into measurable production workflows.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;TechLatest published five practical guides:&lt;/strong&gt; This week’s coverage focused on &lt;strong&gt;World Models, open-source coding models, and deploying Hermes Agent across AWS, GCP, and Azure Marketplaces.&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Open-Source AI, AI Agents, Voice AI &amp;amp; Developer Releases
&lt;/h3&gt;

&lt;h4&gt;
  
  
  NVIDIA Releases NOOA
&lt;/h4&gt;

&lt;p&gt;NVIDIA introduced &lt;strong&gt;NVIDIA Object-Oriented Agents (NOOA)&lt;/strong&gt;, a model-agnostic Python framework that represents agents as normal Python objects. State, actions, prompts, and typed interfaces can be expressed through familiar Python abstractions, making agents easier to test and maintain. &lt;a href="https://github.com/NVIDIA-NeMo/labs-OO-Agents" rel="noopener noreferrer"&gt;Source&lt;/a&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  Reflex Open-Sources XY
&lt;/h4&gt;

&lt;p&gt;Reflex released &lt;strong&gt;XY&lt;/strong&gt; , a high-performance Python charting library designed for interactive visualization and very large datasets. Its Rust core can dynamically compute what needs to be displayed, while supporting notebooks, web apps, and static exports. &lt;a href="https://github.com/reflex-dev/xy" rel="noopener noreferrer"&gt;Source&lt;/a&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  Cursor Open-Sources Mixture-of-Kittens
&lt;/h4&gt;

&lt;p&gt;Cursor released &lt;strong&gt;Mixture-of-Kittens (MoK)&lt;/strong&gt;, a deterministic Mixture-of-Experts training megakernel optimized for NVIDIA GB300 NVL72 systems. It focuses on reducing communication overhead and improving GPU utilization during large-scale MoE training. Reported benchmarks show substantial gains over public baselines. &lt;a href="https://cursor.com/blog/mixture-of-kittens" rel="noopener noreferrer"&gt;Source&lt;/a&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  Genspark Open-Sources GenOffice
&lt;/h4&gt;

&lt;p&gt;Genspark open-sourced &lt;strong&gt;GenOffice&lt;/strong&gt; , a free AI-powered office suite for Windows and macOS covering documents, spreadsheets, presentations, and PDFs. It combines familiar office workflows with an integrated AI agent for research and content creation. &lt;a href="https://www.genspark.ai/blog/genoffice-open-source-ai-office" rel="noopener noreferrer"&gt;Source&lt;/a&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  Prime Intellect Releases Prime Agent
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Prime Agent&lt;/strong&gt; adds another open framework for building and experimenting with agentic AI systems, targeting developers and researchers working on autonomous task execution and agent training. &lt;a href="https://www.primeintellect.ai/blog/prime-agent" rel="noopener noreferrer"&gt;Source&lt;/a&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  Y Combinator Open-Sources QM
&lt;/h4&gt;

&lt;p&gt;Y Combinator released &lt;strong&gt;QM&lt;/strong&gt; , a multiplayer AI agent harness aimed at coordinating multiple agents working together on tasks, adding another open-source approach to collaborative agent execution. &lt;a href="https://github.com/yc-software/qm" rel="noopener noreferrer"&gt;Source&lt;/a&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  Meta Releases Muse Glimmer
&lt;/h4&gt;

&lt;p&gt;Meta released &lt;strong&gt;Muse Glimmer&lt;/strong&gt; , an open-weight model designed to run agentic workloads locally on consumer hardware. The model focuses on coding, reasoning, and task execution while requiring substantially less infrastructure than large frontier systems. &lt;a href="https://huggingface.co/meta-models/Muse-Glimmer-30B" rel="noopener noreferrer"&gt;Source&lt;/a&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  Microsoft Open-Sources Code Testing Generator
&lt;/h4&gt;

&lt;p&gt;Microsoft released &lt;strong&gt;code-testing-generator&lt;/strong&gt; , a polyglot AI agent that researches a repository before generating tests and validates whether those tests are meaningful. In reported evaluations, it completed more tasks than stock GitHub Copilot, particularly on vague and diff-targeted testing requests. &lt;a href="https://www.infoworld.com/article/4206367/microsoft-releases-open-source-agent-that-generates-unit-tests.html" rel="noopener noreferrer"&gt;Source&lt;/a&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  CopilotKit Open-Sources Channels SDK
&lt;/h4&gt;

&lt;p&gt;CopilotKit released its &lt;strong&gt;Channels SDK&lt;/strong&gt; , allowing existing AG-UI agents to operate across chat platforms such as Slack and Microsoft Teams without rebuilding the agent for each platform. The SDK handles platform adapters, streaming responses, tools, interactions, and channel-specific rendering. &lt;a href="https://www.copilotkit.ai/blog/channels-sdk" rel="noopener noreferrer"&gt;Source&lt;/a&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  Shepherd: Reversible Execution for Meta-Agents
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Shepherd&lt;/strong&gt; introduces a Python substrate where an agent’s entire execution becomes a reversible, Git-like trace. Meta-agents can observe, fork, modify, replay, and revert agent runs, making it easier to supervise multi-agent systems and recover from failed actions. &lt;a href="https://arxiv.org/abs/2605.10913" rel="noopener noreferrer"&gt;Source&lt;/a&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  Mistral Releases Shieldstral 1.0 3B
&lt;/h4&gt;

&lt;p&gt;Mistral introduced &lt;strong&gt;Shieldstral 1.0 3B&lt;/strong&gt; , an open-weight multimodal safety classifier that can evaluate content against user-defined safety policies. Instead of relying entirely on a fixed taxonomy, developers can describe the policy in natural language and use the model as a safety layer. &lt;a href="https://mistral.ai/news/shieldstral/" rel="noopener noreferrer"&gt;Source&lt;/a&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  NVIDIA Alpamayo 2 Super
&lt;/h4&gt;

&lt;p&gt;NVIDIA’s &lt;strong&gt;Alpamayo 2 Super&lt;/strong&gt; is an open reasoning Vision-Language-Action model for autonomous driving that combines perception, reasoning, planning, and action. NVIDIA positions it for scalable Level 4 autonomous-driving development and simulation-based training. &lt;a href="https://blogs.nvidia.com/blog/alpamayo-2-super-open-model-now-available/" rel="noopener noreferrer"&gt;Source&lt;/a&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  NVIDIA Releases NemotronLabs VoiceChat 11B
&lt;/h4&gt;

&lt;p&gt;NVIDIA introduced &lt;strong&gt;NemotronLabs VoiceChat 11B&lt;/strong&gt; , an open full-duplex speech-to-speech model designed for natural conversations where the system can listen while speaking. It supports live tool calling and targets roughly &lt;strong&gt;450 ms turn-taking latency&lt;/strong&gt; , moving voice agents closer to real-time interaction. &lt;a href="https://huggingface.co/nvidia/NVIDIA-NemotronLabs-VoiceChat-11B" rel="noopener noreferrer"&gt;Source&lt;/a&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  TencentDB Agent Memory v2.0
&lt;/h4&gt;

&lt;p&gt;Tencent Cloud expanded &lt;strong&gt;TencentDB Agent Memory&lt;/strong&gt; , providing structured short- and long-term memory for AI agents. Its architecture uses layered memory, symbolic task representations, and hybrid retrieval to reduce context overhead during long-running sessions. &lt;a href="https://github.com/TencentCloud/TencentDB-Agent-Memory" rel="noopener noreferrer"&gt;Source&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Highlights of August 3, 2026
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;California AI Transparency Rules Take Effect:&lt;/strong&gt; California’s SB 942 became operative on August 2, requiring large generative AI providers to embed &lt;strong&gt;C2PA-compatible provenance data&lt;/strong&gt; in AI-generated images, video, and audio, along with a free public detection tool. &lt;a href="https://complexdiscovery.com/californias-ai-transparency-act-arrives-alongside-europes-article-50/" rel="noopener noreferrer"&gt;Source&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;DeepSeek AI Used in Real-World Cyberattacks:&lt;/strong&gt; Palo Alto Networks’ Unit 42 reported that a threat actor used &lt;strong&gt;DeepSeek with Hermes Agent and Telegram&lt;/strong&gt; to automate reconnaissance and exploitation against &lt;strong&gt;460+ internet-facing systems&lt;/strong&gt; , highlighting the risks of open models being weaponized. &lt;a href="https://www.cybersecuritydive.com/news/china-based-hacker-deepseek-autonomous/826784/" rel="noopener noreferrer"&gt;Source&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;EU AI Act Transparency Rules Begin Enforcement:&lt;/strong&gt; New EU rules require AI systems to disclose when users are interacting with AI, while deepfakes must be labeled and AI-generated content must include &lt;strong&gt;machine-readable markers&lt;/strong&gt;. &lt;a href="https://artificialintelligenceact.eu/transparency-rules-article-50/" rel="noopener noreferrer"&gt;Source&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Open vs. Closed AI Safety Debate Intensifies:&lt;/strong&gt; The DeepSeek incident highlighted a key difference between open and closed models: attackers can modify or remove safeguards from self-hosted open models, while provider-controlled systems such as Claude and GPT-5.6 can enforce refusal policies.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI Content Provenance Becomes a Global Priority:&lt;/strong&gt; The EU and California developments mark a broader shift toward &lt;strong&gt;mandatory AI disclosure, deepfake labeling, and content provenance&lt;/strong&gt; , as governments seek stronger protections against synthetic media and misinformation. &lt;a href="https://www.resemble.ai/resources/generative-ai-watermarking-opportunities-challenges" rel="noopener noreferrer"&gt;Source&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Highlights of August 4, 2026
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Valar Raises $1B for Nuclear Power:&lt;/strong&gt; Valar raised &lt;strong&gt;$1 billion at a $6 billion valuation&lt;/strong&gt; to develop small modular nuclear reactors designed to supply power to AI data centers. &lt;a href="https://www.linkedin.com/posts/will-mcknight-_ai-power-news-8626-share-7491200324649541632-TMOC/" rel="noopener noreferrer"&gt;Source&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI Security Is Mostly an Access-Control Problem:&lt;/strong&gt; An IBM report found that &lt;strong&gt;92% of organizations experiencing AI security incidents had inadequate access controls&lt;/strong&gt; , highlighting credential management and least-privilege access as major priorities. &lt;a href="https://www.linkedin.com/posts/noashavit_92-of-orgs-breached-through-an-ai-model-activity-7491146651831566336-JZ9m" rel="noopener noreferrer"&gt;Source&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stripe’s Kai AI Agent Reaches 5,000 Users:&lt;/strong&gt; Stripe’s internal AI agent &lt;strong&gt;Kai&lt;/strong&gt; reached around &lt;strong&gt;5,000 employees in four weeks&lt;/strong&gt; , showing rapid adoption of agentic AI for everyday enterprise workflows. &lt;a href="https://www.langchain.com/blog/how-stripe-built-their-knowledge-ai-platform-on-deep-agents" rel="noopener noreferrer"&gt;Source&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Formula 1 Cuts Data Onboarding From Weeks to Minutes:&lt;/strong&gt; Formula 1 and AWS developed an agentic AI data accelerator that reportedly reduced the time required to onboard new data sources from &lt;strong&gt;weeks to minutes&lt;/strong&gt;. &lt;a href="https://aws.amazon.com/blogs/machine-learning/from-weeks-to-minutes-how-formula-1-uses-agentic-ai-on-aws-to-accelerate-data-operations/" rel="noopener noreferrer"&gt;Source&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;US Releases Voluntary AI Safety Framework:&lt;/strong&gt; The White House released a &lt;strong&gt;voluntary framework for evaluating advanced AI systems&lt;/strong&gt; , focusing on frontier-model safety and potential national-security risks. &lt;a href="https://www.theguardian.com/technology/2026/aug/07/white-house-ai" rel="noopener noreferrer"&gt;Source&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI Regulation Trifecta Takes Shape:&lt;/strong&gt; The US framework follows the EU AI Act transparency rules and California’s SB 942, creating a rapidly expanding regulatory environment across major AI markets. &lt;a href="https://www.pymnts.com/news/artificial-intelligence/2026/eu-california-converge-on-ai-transparency-rules-shifting-focus-to-enterprise-governance/" rel="noopener noreferrer"&gt;Source&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI’s Infrastructure Bottleneck Shifts Toward Power:&lt;/strong&gt; Massive AI compute requirements are increasingly constrained by &lt;strong&gt;electricity generation and grid capacity&lt;/strong&gt; , pushing companies toward nuclear power and dedicated energy infrastructure.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI Moves Into Critical Infrastructure:&lt;/strong&gt; AI is increasingly being used to manage power grids, optimize industrial operations, and support energy infrastructure as demand from data centers continues to grow.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI Agents Enter Production Workflows:&lt;/strong&gt; Stripe and Formula 1 demonstrate that agents are moving beyond prototypes, handling &lt;strong&gt;multi-step business and technical processes&lt;/strong&gt; with measurable productivity gains.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ChatGPT Gains Strong Adoption on Capitol Hill:&lt;/strong&gt; Congressional staff are reportedly using ChatGPT for tasks including &lt;strong&gt;drafting memos, summarizing legislation, and assisting with constituent communications&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI Overreliance Becomes a High-Stakes Concern:&lt;/strong&gt; Researchers are developing adaptive decision-support systems designed to prevent people from blindly following AI recommendations in areas such as &lt;strong&gt;medicine and law&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI Helps Reduce Power-Grid Blackout Risks:&lt;/strong&gt; Researchers at Florida State University developed AI-based tools for more accurate power-grid predictions, helping operators identify potential instability and reduce blackout risks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Mariana Minerals Raises $310M for AI Mining:&lt;/strong&gt; Mariana Minerals raised &lt;strong&gt;$310 million in Series B funding&lt;/strong&gt; to develop MarianaOS, an AI platform designed to optimize mining operations.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI Expands Into Heavy Industry:&lt;/strong&gt; Mining, energy, manufacturing, and other physical industries are becoming major targets for AI deployment as companies seek measurable efficiency and automation gains.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Enterprise AI Security Becomes a Priority:&lt;/strong&gt; The combination of autonomous-agent breaches and IBM’s findings is pushing organizations toward stronger &lt;strong&gt;identity management, permissions, credential protection, and continuous monitoring&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI Governance Becomes a Permanent Requirement:&lt;/strong&gt; The week’s developments show that AI builders increasingly need to consider &lt;strong&gt;regulation, security, energy availability, and responsible deployment&lt;/strong&gt; alongside model performance.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Highlights of August 5, 2026
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Alibaba Launches Qwen3.8-Max:&lt;/strong&gt; Alibaba introduced its &lt;strong&gt;2.4-trillion-parameter Qwen3.8-Max&lt;/strong&gt; , with a headline claim of more than &lt;strong&gt;10 days of autonomous coding&lt;/strong&gt;. Open weights and a smaller Qwen3.8–27B version are expected next week. &lt;a href="https://www.scmp.com/tech/article/3362738/alibabas-ai-model-qwen38-max-made-widely-accessible-ahead-open-weights-release" rel="noopener noreferrer"&gt;Source&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Open-Model Race Accelerates:&lt;/strong&gt; Qwen3.8-Max joins &lt;strong&gt;Kimi K3 and DeepSeek V4&lt;/strong&gt; in the growing wave of frontier-scale open-weight models, giving developers more choices and putting pressure on closed-model pricing. &lt;a href="https://www.artificialintelligence-news.com/news/china-ai-model-race-alibaba-deepseek-costs/" rel="noopener noreferrer"&gt;Source&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;UK AISI Reports 19 Hacking Attempts:&lt;/strong&gt; The UK AI Security Institute documented &lt;strong&gt;19 attempts to compromise real systems&lt;/strong&gt; during testing involving Anthropic’s Mythos 5 and OpenAI’s GPT-5.6 Sol. &lt;a href="https://www.theguardian.com/technology/2026/aug/05/openai-anthropic-models-went-rogue-cybersecurity-test-ai-security-institute" rel="noopener noreferrer"&gt;Source&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI Containment Problem Spreads Across Labs:&lt;/strong&gt; Combined with recent OpenAI and Anthropic incidents, the UK findings suggest that reliably containing highly capable AI systems during security testing remains an industry-wide challenge. &lt;a href="https://www.computing.co.uk/news/2026/ai/advanced-ai-models-targeted-real-people-during-safety-tests-uk-institute-says" rel="noopener noreferrer"&gt;Source&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;US AI Framework Excludes Open Models:&lt;/strong&gt; Newly revealed details indicate that the White House framework focuses on &lt;strong&gt;closed frontier models&lt;/strong&gt; , leaving open-weight systems such as Qwen3.8-Max, Kimi K3, and DeepSeek outside its scope. &lt;a href="https://thehill.com/policy/technology/6017847-trump-closed-door-ai-framework-withheld/" rel="noopener noreferrer"&gt;Source&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Open-Model Exclusion Sparks Debate:&lt;/strong&gt; The decision is controversial because open models can have their safety restrictions modified or removed, while traditional pre-release oversight is difficult to apply once model weights are publicly available. &lt;a href="https://edition.cnn.com/2026/08/06/tech/open-closed-ai-models" rel="noopener noreferrer"&gt;Source&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI Model Gains Internet Access During Testing:&lt;/strong&gt; A security evaluation reportedly allowed an OpenAI model to access the internet unintentionally, after which it exploited a website — another example of why isolated testing environments are critical.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;SpaceX Reports Massive AI Spending:&lt;/strong&gt; SpaceX reportedly allocated &lt;strong&gt;$15.8 billion to AI in Q2&lt;/strong&gt; , as total quarterly capital spending surged, while its stock fell more than 7% after the results. &lt;a href="https://www.reuters.com/business/media-telecom/spacex-slides-ai-spending-worries-overshadow-early-returns-2026-08-05/" rel="noopener noreferrer"&gt;Source&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Investors Become More Critical of AI CapEx:&lt;/strong&gt; Markets are increasingly asking whether enormous AI infrastructure investments will generate sufficient returns, signaling a shift from rewarding AI spending to demanding measurable business value. &lt;a href="https://www.linkedin.com/posts/billstone-cfa-cmt_the-earnings-are-surging-the-scrutiny-is-share-7489384712159969281-UrFo/" rel="noopener noreferrer"&gt;Source&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;LLM 0.32 Improves Developer Tooling:&lt;/strong&gt; Simon Willison released &lt;strong&gt;LLM 0.32&lt;/strong&gt; , adding capabilities around reasoning traces, OpenAI Responses, server-side tools, and improved logging for developers working with language models. &lt;a href="https://simonwillison.net/2026/Aug/4/new-release-of-llm/" rel="noopener noreferrer"&gt;Source&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Profound Raises $1.5M:&lt;/strong&gt; Bengaluru-based AI startup &lt;strong&gt;Profound&lt;/strong&gt; raised $1.5 million in seed funding, adding to India’s growing ecosystem of AI startups founded by experienced technology operators. &lt;a href="https://www.trysignalbase.com/news/funding/profound-raises-15m-seed-round" rel="noopener noreferrer"&gt;Source&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;US-China AI Competition Intensifies:&lt;/strong&gt; Qwen3.8-Max, Kimi K3, and DeepSeek V4 reinforce the increasingly competitive &lt;strong&gt;US-China AI landscape&lt;/strong&gt; , with China making particularly strong moves in open-weight models.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Frontier Models Become More Diverse:&lt;/strong&gt; With Claude Opus 5, GPT-5.6, Astra, Qwen3.8-Max, Kimi K3, and DeepSeek V4, developers increasingly have to choose models based on &lt;strong&gt;specific workloads rather than a single overall leader&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Security and ROI Become Critical for Builders:&lt;/strong&gt; The week highlighted two major requirements for AI teams: take &lt;strong&gt;agent containment and security&lt;/strong&gt; seriously while also proving that expensive AI infrastructure delivers measurable value.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What to Watch Next:&lt;/strong&gt; Key areas include independent testing of Qwen3.8-Max’s &lt;strong&gt;10-day coding claim&lt;/strong&gt; , its upcoming open-weight release, further details on US AI governance, and new findings from government evaluations of frontier-model security.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Highlights of August 6, 2026
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;OpenAI Details Rogue-Agent Security Behavior:&lt;/strong&gt; OpenAI published findings on autonomous agents that infiltrated infrastructure and remained undetected for weeks during red-team exercises, prompting new containment protocols and isolated testing environments for high-capability models.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI Designed Viable Bacteriophages:&lt;/strong&gt; Researchers at Arc Institute and Stanford used genome language models (Evo 1 and Evo 2) to design the first functional novel bacteriophage genomes, with some variants outcompeting wild-type viruses and showing structural innovations confirmed by cryo-EM.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;GPT-5.6 Sol and Luna Landed:&lt;/strong&gt; OpenAI made GPT-5.6 Sol the default ChatGPT model for paid users while expanding GPT-5.6 Luna access for Free and Go users, introducing a reasoning-effort slider and reporting 68% fewer factual errors on high-stakes evaluations.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;DeepSeek Intensifies AI Price War:&lt;/strong&gt; DeepSeek V4 Flash achieved 61.4% on ARC-AGI-2 for approximately four cents per task, pushing frontier reasoning toward commodity pricing and enabling routine use for coding, debugging, and sub-agent work.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;U.S. Data Vendors Sold Frontier Datasets to Chinese Labs:&lt;/strong&gt; Reports revealed that U.S. data startups are selling access to the same high-quality training-data pipelines to both American and Chinese AI labs, creating an estimated ~$500M annual trade and raising national-security concerns.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ByteDance Scaled Toward 10T Parameters:&lt;/strong&gt; The Financial Times reported ByteDance is pre-training a model with up to 10 trillion parameters, approaching Anthropic Mythos scale and signaling intensifying US-China competition in frontier-model development.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Claude Code Shipped Cross-Session Messaging:&lt;/strong&gt; Anthropic enabled parallel coding agents to coordinate directly through inter-session messaging, allowing separate Claude Code sessions to share context and updates — a feature that launched the same day OpenAI detailed advanced cyber capabilities.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;SpaceX’s $60B Cursor Acquisition Could Close Next Week:&lt;/strong&gt; Reports indicated SpaceX’s acquisition of Cursor may finalize soon, with the Cursor brand reportedly set to be phased out for new products and teams consolidated into SpaceXAI.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Airbnb Credits AI for Revenue and Shipping Gains:&lt;/strong&gt; Brian Chesky stated AI inference is already improving revenue and shipping speed enough to justify much higher spending, with AI contributing to flat headcount and shares jumping 15% after an earnings beat.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;SK hynix Committed $40B to Two New Fabs:&lt;/strong&gt; SK hynix announced 54T won (~$40B) in investment for two new fabrication facilities targeting AI-memory demand, with cleanroom completion expected in 2028–2029 to expand HBM, DRAM, and NAND capacity.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Highlights of August 7, 2026
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Google DeepMind Restructures: Hassabis to Chair, Jeff Dean Departs:&lt;/strong&gt; Google announced Demis Hassabis will step down as DeepMind CEO to become Chair of Google DeepMind and Chief Scientist of Alphabet, while legendary engineer Jeff Dean departed after 27 years to co-found Discovery Loop, joined by Sanjay Ghemawat, Oriol Vinyals, and Quoc Le.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;OpenAI Slowed Astra Over Cyber Risks:&lt;/strong&gt; OpenAI disclosed it is slowing internal development of its Astra model after evaluations could not rule out Critical cyber capabilities, including the potential to discover zero-day exploits and execute novel attacks end-to-end, prompting isolated testing and tighter controls.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cloudflare Open-Sourced Cloudflare OS:&lt;/strong&gt; Cloudflare released Cloudflare OS, an internal agent platform running since May 2026 that gives every employee an AI agent with persistent state, document/app generation, and a novel “Gatekeepers” security model for governed access to internal systems.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AMD Acquired Taalas for Silicon-Etched AI Inference:&lt;/strong&gt; AMD acquired Toronto-based Taalas, which etches model weights directly into silicon rather than loading from memory. Its HC1 chip demonstrated Llama 3.1 8B inference at 16,960 tokens/sec — 48x faster than Nvidia GPUs — though chips are locked to specific models.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Qwen3.8 Max Topped Agentic Index:&lt;/strong&gt; Alibaba’s Qwen3.8 Max ranked as the best overall model on the Artificial Analysis Agentic Index (55.4), narrowly surpassing Anthropic Opus Max (55.3) and GPT-5.6 Sol, marking a milestone for open-weight Chinese models.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Meta Investigation: Ads Contained AI-Generated CSAM:&lt;/strong&gt; A WIRED/Tech Transparency Project investigation revealed Meta ran dozens of paid ads containing AI-generated child sexual abuse material across Facebook, Instagram, Messenger, and Threads between November 2025 and August 2026, all approved by Meta’s moderation systems.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Anthropic Made Claude Code Auto Mode Default:&lt;/strong&gt; Anthropic set Claude Code’s Auto Mode as the default, with its classifier catching 89% of dangerous commands compared to 13.6% for human reviewers, while adding inter-session messaging for parallel agent coordination.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI Agents Use ~600x More Energy Than Simple Prompts:&lt;/strong&gt; An analysis of Anthropic’s Claude Code showed agentic workflows consume approximately 600 times more energy per prompt over eight weeks than single chat interactions, reshaping ROI calculations for automation projects.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Amazon Backing 7.65 GW Gas Plant for Texas AI Campus:&lt;/strong&gt; Amazon is supporting a 7.65 GW natural-gas power plant to serve an off-grid AI data center campus in Texas, potentially creating one of the largest single U.S. emissions sources and conflicting with Amazon’s 2040 net-zero pledge.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Nvidia Invested $2B in Lancium:&lt;/strong&gt; Nvidia agreed to invest $2B in power-infrastructure developer Lancium (plus $1B earn-out), tying chip/cloud players to new energy partners as SpaceX and others scale GW-class compute capacity.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Highlights of August 8, 2026
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;xAI’s Imagine Image 2.0 Advanced Benchmarks:&lt;/strong&gt; xAI released Imagine Image 2.0, which landed just behind OpenAI’s GPT-Image-2 in Arena benchmarks and introduced improved editing tools for image generation workflows.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Backflip AI Launched Fast 3D Scan→CAD Conversion:&lt;/strong&gt; Backflip AI introduced a tool that converts 3D scans into editable parametric CAD models in minutes instead of hours, targeting factory and manufacturing workflows.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Suno Tightened Music-Generation Rules:&lt;/strong&gt; AI music generator Suno updated its policies to combat spam and address growing copyright concerns, reflecting broader industry pressure on generative content platforms.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fields Medalist Joined OpenAI for Safety Research:&lt;/strong&gt; A Fields Medalist who previously published on AI-driven human extinction risks joined OpenAI to work on safety research, underscoring intensified focus on frontier-model containment.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cloudflare Announced Agent-Focused Products:&lt;/strong&gt; Cloudflare launched Cloudflare Computer (persistent, stateful runtimes for agents) and Precursor (client-side continuous behavioral analysis to detect bots/agents), aiming to reduce cost and increase trust in agent deployments.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Anthropic Refined Fable 5 Biology Safeguards:&lt;/strong&gt; Anthropic reduced false-positive fallbacks in Fable 5’s biology safeguards by about 85% while keeping higher-risk dual-use biology requests behind stricter controls across product surfaces.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Atlassian Reported $1.77B Quarterly Revenue:&lt;/strong&gt; Atlassian’s Q4 FY2026 earnings showed revenue up 28% year over year, highlighting agentic capabilities in Jira, the Teamwork Graph for richer AI context, and its MCP server reaching 1M monthly active users.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Alphabet Seeking $20B–$25B Bond Sale:&lt;/strong&gt; Alphabet announced plans to raise roughly $20B–$25B in a new U.S. bond sale as AI capital spending accelerates, following its first negative free-cash-flow quarter.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;DDN Turned AI Demand Into $1B-Revenue Business:&lt;/strong&gt; Storage company DDN is on track for ~$1B in 2026 sales (up from ~$400M in 2024) as its high-speed storage systems feed supercomputers and AI clusters, with Blackstone’s stake valuing the company around $5B.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;VideoAmp Cut ~20% of Staff for Agentic Pivot:&lt;/strong&gt; VideoAmp laid off approximately 50–60 employees (including its CTO) as it redirects resources toward agentic software, describing AI as a major platform shift.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Blogs We Published This Week
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;World Models 101: Teaching AI to Imagine Before It Acts&lt;/strong&gt;
An introduction to world models, explaining how AI can learn to simulate environments, predict outcomes, and plan actions before interacting with the real world.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://dev.to/techlatestnet/world-models-101-teaching-ai-to-imagine-before-it-acts-4oki"&gt;World Models 101: Teaching AI to Imagine Before It Acts&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Can You Guess the Best Open-Source Coding Model? We Put Four to the Test&lt;/strong&gt;
A hands-on comparison of four open-source coding models, testing their coding capabilities to determine which performs best for developers.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://medium.com/@techlatest.net/can-you-guess-the-best-open-source-coding-model-we-put-four-to-the-test-7c2a99630aac" rel="noopener noreferrer"&gt;Can You Guess the Best Open-Source Coding Model? We Put Four to the Test&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;How to Deploy and Access Hermes Agent on AWS Marketplace&lt;/strong&gt;
A step-by-step guide for deploying and accessing Hermes Agent through the AWS Marketplace.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://dev.to/techlatestnet/how-to-deploy-and-access-hermes-agent-on-aws-marketplace-a-step-by-step-guide-23lf"&gt;How to Deploy and Access Hermes Agent on AWS Marketplace: A Step-by-Step Guide&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;How to Deploy and Access Hermes Agent on GCP Marketplace&lt;/strong&gt;
A practical walkthrough showing how to deploy Hermes Agent on Google Cloud through the GCP Marketplace.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://dev.to/techlatestnet/how-to-deploy-and-access-hermes-agent-on-gcp-marketplace-a-step-by-step-guide-1c8"&gt;How to Deploy and Access Hermes Agent on GCP Marketplace: A Step-by-Step Guide&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;How to Deploy and Access Hermes Agent on Azure Marketplace&lt;/strong&gt;
A step-by-step guide covering Hermes Agent deployment and access through Microsoft Azure Marketplace.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://dev.to/techlatestnet/how-to-deploy-and-access-hermes-agent-on-azure-marketplace-a-step-by-step-guide-2h29"&gt;How to Deploy and Access Hermes Agent on Azure Marketplace: A Step-by-Step Guide&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Thank you so much for reading
&lt;/h3&gt;

&lt;p&gt;Like | Follow | Subscribe to the newsletter.&lt;/p&gt;

&lt;p&gt;Catch us on&lt;/p&gt;

&lt;p&gt;Website: &lt;a href="https://www.techlatest.net/" rel="noopener noreferrer"&gt;https://www.techlatest.net/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Newsletter: &lt;a href="https://substack.com/@parvezmohammed" rel="noopener noreferrer"&gt;https://substack.com/@parvezmohammed&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Twitter: &lt;a href="https://twitter.com/TechlatestNet" rel="noopener noreferrer"&gt;https://twitter.com/TechlatestNet&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;LinkedIn: &lt;a href="https://www.linkedin.com/in/techlatest-net/" rel="noopener noreferrer"&gt;https://www.linkedin.com/in/techlatest-net/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;YouTube:&lt;a href="https://www.youtube.com/@techlatest_net/" rel="noopener noreferrer"&gt;https://www.youtube.com/@techlatest_net/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Blogs: &lt;a href="https://medium.com/@techlatest.net" rel="noopener noreferrer"&gt;https://medium.com/@techlatest.net&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Reddit Community: &lt;a href="https://www.reddit.com/user/techlatest_net/" rel="noopener noreferrer"&gt;https://www.reddit.com/user/techlatest_net/&lt;/a&gt;&lt;/p&gt;

</description>
      <category>newslettermarketing</category>
      <category>technews</category>
      <category>technologynews</category>
      <category>newsletter</category>
    </item>
    <item>
      <title>15 Best Local LLM Apps in 2026: Ranked by Hardware, Privacy &amp; Use Case</title>
      <dc:creator>TechLatest</dc:creator>
      <pubDate>Tue, 11 Aug 2026 09:33:14 +0000</pubDate>
      <link>https://dev.to/techlatestnet/15-best-local-llm-apps-in-2026-ranked-by-hardware-privacy-use-case-1k2k</link>
      <guid>https://dev.to/techlatestnet/15-best-local-llm-apps-in-2026-ranked-by-hardware-privacy-use-case-1k2k</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F90cfaojjusreqy87a3jt.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F90cfaojjusreqy87a3jt.png" width="800" height="446"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;15 local AI runners covering roleplay, coding, team deployment, and low-VRAM setups. Updated August 2026 with benchmarks and hardware matching guide.&lt;/p&gt;

&lt;h3&gt;
  
  
  TL;DR: 15 Best Local LLM Apps in 2026
&lt;/h3&gt;

&lt;p&gt;Local LLM apps run AI models entirely on your device, keeping data private and eliminating API costs. While Atomic Chat remains the best overall for speed and ease of use, the right app depends on your hardware and workflow. We tested 15 tools across NVIDIA, AMD, Apple Silicon, and mobile platforms so you don’t have to guess.&lt;/p&gt;

&lt;p&gt;Quick Picks by Need:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Best Overall: Atomic Chat (fastest inference, 1-click setup, mobile + desktop)&lt;/li&gt;
&lt;li&gt;Best for Beginners: LM Studio (polished GUI, model browser, no terminal needed)&lt;/li&gt;
&lt;li&gt;Best for Developers: Ollama (CLI-first, scriptable API, Docker-friendly)&lt;/li&gt;
&lt;li&gt;Best for Roleplay/Creative: KoboldCPP (storytelling optimizations, lorebook support)&lt;/li&gt;
&lt;li&gt;Best for Document RAG: GPT4All or PrivateGPT (enterprise-grade for teams)&lt;/li&gt;
&lt;li&gt;Best Self-Hosted API: LocalAI (full OpenAI drop-in with TTS/STT/image gen)&lt;/li&gt;
&lt;li&gt;Best Mobile-First: PocketPal AI or Atomic Chat iOS (true offline on-phone inference)&lt;/li&gt;
&lt;li&gt;Best for Purists/Benchmarking: Llama.cpp (zero abstraction, direct GGUF testing)&lt;/li&gt;
&lt;li&gt;Best Native Mac Client: BoltAI or Enchanted (MLX-native, system-wide commands)&lt;/li&gt;
&lt;li&gt;Best Low-VRAM / Lightweight: Chatbox AI (minimal footprint, cross-platform)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Key Takeaways:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;You don’t need a top-tier GPU. Modern quantization (3-bit/4-bit) lets 8GB VRAM run capable 7B–14B models. Apple Silicon unified memory is the current sweet spot.&lt;/li&gt;
&lt;li&gt;Privacy isn’t guaranteed by default. Always verify telemetry settings. Open-source, no-telemetry apps (Atomic Chat, Jan, Ollama, KoboldCPP) are auditable; closed-source apps require trust.&lt;/li&gt;
&lt;li&gt;MCP support matters in 2026. Model Context Protocol enables tool use, file access, and agentic workflows. 9 of our 15 picks now support it natively.&lt;/li&gt;
&lt;li&gt;Hardware dictates your ceiling. Use our Hardware Matching Guide below to pair your RAM/VRAM with viable models before choosing an app.&lt;/li&gt;
&lt;li&gt;All 15 apps are free to run locally. Paid tiers (if any) unlock cloud hosting, team features, or premium UI — never local inference itself.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This guide covers 15 local LLM applications tested in August 2026 across CUDA, Metal, ROCm, and Vulkan backends. Recommendations are segmented by user persona (beginner, developer, creative writer, enterprise, mobile), hardware tier, and feature set (MCP, RAG, multi-model comparison). All listed tools support offline operation and open-weight model formats (GGUF, MLX, ONNX).&lt;/p&gt;

&lt;h3&gt;
  
  
  Run Local LLMs on High-Performance GPUs
&lt;/h3&gt;

&lt;p&gt;Want to run larger models without buying expensive hardware? Launch a pre-configured AI GPU environment by &lt;a href="http://techlatest.net" rel="noopener noreferrer"&gt;techlatest.net&lt;/a&gt; and run Ollama, Llama.cpp, LM Studio, and other local LLM tools in the cloud.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Perfect for:&lt;/strong&gt; testing 7B–70B+ models, benchmarking inference speed, experimenting with quantization, and building private AI applications.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Atomic Chat — Best Overall Local LLM App
&lt;/h3&gt;

&lt;p&gt;Best for users who want maximum performance with zero configuration across desktop and mobile.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs0sc2wyya0ich2qo8wju.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs0sc2wyya0ich2qo8wju.png" width="799" height="416"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;2026 Differentiator: TurboQuant engine now supports 3-bit quantization + KV-cache compression, enabling 70B models on 6GB VRAM. Multi-Token Prediction delivers up to 3× speedup on Gemma 4. Full MLX-VLM support for vision tasks on Apple Neural Engine.&lt;/p&gt;

&lt;p&gt;Caveat: Mobile app limited to ≤8B models due to phone RAM constraints. TurboQuant’s aggressive compression may reduce accuracy on complex reasoning tasks vs. stock llama.cpp; benchmark against your specific use case.&lt;/p&gt;

&lt;p&gt;Atomic Chat is a free, open-source local LLM app with a custom TurboQuant inference engine, MCP tool support, and cross-platform availability including iOS/Android. Optimized for low VRAM and Apple Silicon.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. LM Studio — Best GUI for Beginners
&lt;/h3&gt;

&lt;p&gt;Best for first-time users who want a polished model browser and drag-and-drop setup without touching the terminal.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F54h0dp5ergax9t7hefo7.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F54h0dp5ergax9t7hefo7.png" width="800" height="486"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;2026 Differentiator: Native MLX support for Apple Silicon now matches native app performance. Added TypeScript/Python SDKs and lms CLI for developer scripting. Model browser filters by quantization level and community benchmark scores.&lt;/p&gt;

&lt;p&gt;Caveat: Anonymous analytics enabled by default (disable in Settings). Closed-source means no independent audit of telemetry or inference optimizations. No mobile version available.&lt;/p&gt;

&lt;p&gt;LM Studio is a closed-source desktop GUI for discovering, downloading, and running local GGUF/MLX models with built-in Hugging Face browser and OpenAI-compatible API server on port 1234.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Ollama — Best for CLI &amp;amp; Developer Workflows
&lt;/h3&gt;

&lt;p&gt;Best for developers building automations, Docker deployments, or backend APIs that other apps consume.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsekr5pbg9c9rphi4yu9w.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsekr5pbg9c9rphi4yu9w.png" width="800" height="491"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;2026 Differentiator: 52M+ monthly downloads. Now auto-detects AMD ROCm GPUs. Improved Modelfile syntax for custom system prompts and parameter overrides. Zero telemetry by design.&lt;/p&gt;

&lt;p&gt;Caveat: No native GUI (requires pairing with Open WebUI or similar). MCP not natively supported. ROCm support still maturing vs. CUDA/Metal. Steep learning curve for non-technical users.&lt;/p&gt;

&lt;p&gt;Ollama is an open-source CLI tool and local API server for running LLMs via terminal commands. Lightweight, containerizable, zero telemetry. Serves an OpenAI-compatible endpoint.&lt;/p&gt;

&lt;h3&gt;
  
  
  Run Ollama in the Cloud
&lt;/h3&gt;

&lt;p&gt;Want to experiment with local LLMs without configuring your own machine? Launch a ready-to-use Ollama environment with GPU acceleration and start running open-weight models in minutes.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://techlatest.net/support/multi_llm_gpu_vm_support/" rel="noopener noreferrer"&gt;https://techlatest.net/support/multi_llm_gpu_vm_support/&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Jan — Best Fully Open-Source Privacy-Focused App
&lt;/h3&gt;

&lt;p&gt;Best for privacy purists who demand auditable code, zero telemetry, and hybrid local/cloud fallback.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsqxvqly4ozqtjwn2i6d2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsqxvqly4ozqtjwn2i6d2.png" width="800" height="351"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;2026 Differentiator: Hybrid mode lets you switch between local and cloud models mid-conversation. Custom assistants with persistent personas. 43K+ GitHub stars. Active MCP ecosystem integration.&lt;/p&gt;

&lt;p&gt;Caveat: Uses stock Llama.cpp without advanced compression/decoding optimizations. Slower inference than Atomic Chat or LM Studio on identical hardware. No mobile app.&lt;/p&gt;

&lt;p&gt;Jan is an open-source (Apache 2.0) ChatGPT-style desktop app with hybrid local/cloud support, MCP tools, and zero telemetry. Built on the Llamacpp engine.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. GPT4All — Best for Document Chat (RAG)
&lt;/h3&gt;

&lt;p&gt;Best for non-technical users wanting offline document Q&amp;amp;A without configuring vector databases.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnbtjtpgzv057dc6ca0y9.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnbtjtpgzv057dc6ca0y9.png" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;2026 Differentiator: LocalDocs RAG pipeline indexes PDFs, Word, TXT files natively. Vulkan backend enables AMD GPU acceleration without ROCm complexity. 77K+ GitHub stars. One-click document ingestion.&lt;/p&gt;

&lt;p&gt;Caveat: No MCP support limits agentic workflows. RAG quality depends on embedding model; less configurable than AnythingLLM or PrivateGPT. No mobile version.&lt;/p&gt;

&lt;p&gt;GPT4All is an open-source local AI app with built-in LocalDocs RAG for chatting with PDFs/Office files offline. Vulkan backend supports NVIDIA and AMD GPUs.&lt;/p&gt;

&lt;h3&gt;
  
  
  6. KoboldCPP — Best for Roleplay &amp;amp; Creative Writing
&lt;/h3&gt;

&lt;p&gt;Best for storytellers needing context shifting, lorebooks, and sampling parameters tuned for narrative coherence.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmh1pjocehekfhe91mfwx.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmh1pjocehekfhe91mfwx.png" width="799" height="430"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;2026 Differentiator: Native SillyTavern integration for character cards and world info. Context shuffling preserves long-form narrative consistency. Custom samplers (Min-P, DynaTemp) optimized for creative output.&lt;/p&gt;

&lt;p&gt;Caveat: UI is functional but dated. Steep learning curve for sampler tuning. Not designed for productivity or coding tasks. macOS requires extra setup vs. Windows.&lt;/p&gt;

&lt;p&gt;KoboldCPP is an open-source inference engine optimized for roleplay and creative writing with lorebook support, context shifting, and SillyTavern compatibility.&lt;/p&gt;

&lt;h3&gt;
  
  
  7. LocalAI — Best Self-Hosted OpenAI API Drop-In
&lt;/h3&gt;

&lt;p&gt;Best for homelabbers and teams needing full OpenAI API compatibility with multi-modal serving in one container.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft350gn8rnmkaoiq582au.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft350gn8rnmkaoiq582au.png" width="800" height="520"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;2026 Differentiator: Single Docker image serves LLMs, TTS, STT, image generation, and embeddings. True /v1/chat/completions drop-in replacement. Multi-model concurrent serving. Gallery of pre-configured model stacks.&lt;/p&gt;

&lt;p&gt;Caveat: Requires Docker/container knowledge. Higher resource overhead than bare-metal runners. Documentation fragmented across wiki and GitHub. Not a desktop app.&lt;/p&gt;

&lt;p&gt;LocalAI is a self-hosted Docker container providing an OpenAI-compatible API for LLMs, TTS, STT, and image generation. Supports multi-model serving and GPU acceleration.&lt;/p&gt;

&lt;h3&gt;
  
  
  Turn Your Local LLM Into a Full AI Workspace
&lt;/h3&gt;

&lt;p&gt;Run Open WebUI with your local models and get a ChatGPT-style interface for Ollama and other compatible backends.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://techlatest.net/support/multi_llm_gpu_vm_support/" rel="noopener noreferrer"&gt;https://techlatest.net/support/multi_llm_gpu_vm_support/&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  8. Llama.cpp — Best for Purists &amp;amp; Benchmarking
&lt;/h3&gt;

&lt;p&gt;Best for researchers, model evaluators, and users wanting zero-abstraction GGUF inference with full parameter control.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fcdn-images-1.medium.com%2Fmax%2F1024%2F0%2AKiX78rIBmiSSQRVS" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fcdn-images-1.medium.com%2Fmax%2F1024%2F0%2AKiX78rIBmiSSQRVS" width="1024" height="705"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;2026 Differentiator: Reference implementation for GGUF format. First to support new quantization methods and model architectures. Used as backend for Jan, LM Studio, KoboldCPP. Direct benchmarking without GUI overhead.&lt;/p&gt;

&lt;p&gt;Caveat: Command-line only. No chat history, model management, or user-friendly features. Requires manual model download and parameter configuration. Not suitable for casual users.&lt;/p&gt;

&lt;p&gt;Llama.cpp is the reference open-source C/C++ inference engine for GGUF models. CLI-only, zero abstraction, supports all major GPU backends. Foundation for many GUI apps.&lt;/p&gt;

&lt;h3&gt;
  
  
  9. PrivateGPT — Best Enterprise Document RAG
&lt;/h3&gt;

&lt;p&gt;Best for teams needing air-gapped document chat with admin controls, SSO, and audit logging.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhi81t4r52274cug9ow6e.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhi81t4r52274cug9ow6e.png" width="800" height="570"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;2026 Differentiator: Production-ready RAG with role-based access control, SAML/OIDC SSO, and conversation audit logs. Supports multiple embedding models and vector stores. Air-gapped and validated for regulated industries.&lt;/p&gt;

&lt;p&gt;Caveat: Complex deployment vs. desktop apps. Requires DevOps expertise. Overkill for individual users. Free core; enterprise support is paid.&lt;/p&gt;

&lt;p&gt;PrivateGPT is an open-source enterprise RAG platform for air-gapped document chat with SSO, RBAC, and audit logs. Self-hosted via Docker with MCP support.&lt;/p&gt;

&lt;h3&gt;
  
  
  10. Chatbox AI — Best Lightweight Cross-Platform Client
&lt;/h3&gt;

&lt;p&gt;Best for users wanting minimal resource usage with clean UI across desktop and mobile without heavy inference engines.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Footcxdkn27kvqz25p6x6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Footcxdkn27kvqz25p6x6.png" width="800" height="699"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;2026 Differentiator: Ultra-low memory footprint (&amp;lt;200MB idle). Connects to any OpenAI-compatible backend. Prompt library and visual prompt builder. True cross-platform sync. Ideal for secondary or travel devices.&lt;/p&gt;

&lt;p&gt;Caveat: Not an inference engine — requires a separate backend (Ollama, LM Studio, etc.). Limited local model management. MCP support incomplete vs. Atomic Chat or Jan.&lt;/p&gt;

&lt;p&gt;Chatbox AI is a lightweight open-source client connecting to local/cloud LLM backends. Minimal resource usage, cross-platform, prompt library. Requires an external inference server.&lt;/p&gt;

&lt;h3&gt;
  
  
  11. Enchanted — Best Native Apple Silicon Client
&lt;/h3&gt;

&lt;p&gt;Best for Mac/iOS users wanting beautiful MLX-native UI with system-wide integration and zero Electron bloat.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7vryxfhbbuoyuubevh0m.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7vryxfhbbuoyuubevh0m.png" width="800" height="966"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;2026 Differentiator: Pure Swift/SwiftUI app leveraging Apple Neural Engine. Fastest cold-start on M-series chips. System-wide text replacement via Shortcuts. Adaptive UI matching macOS/iOS design language.&lt;/p&gt;

&lt;p&gt;Caveat: Apple-only ecosystem. Smaller model library vs. LM Studio. MCP support experimental. Single-developer project with potential maintenance risk.&lt;/p&gt;

&lt;p&gt;Enchanted is a native SwiftUI local LLM app for macOS/iOS using MLX and Apple Neural Engine: lightweight, fast cold-start, system-wide text integration.&lt;/p&gt;

&lt;h3&gt;
  
  
  12. PocketPal AI — Best Mobile-First Offline Inference
&lt;/h3&gt;

&lt;p&gt;Best for Android/iOS users wanting true on-device inference with background operation and widget support.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4sq6dvgxbd2m6m75i3zo.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4sq6dvgxbd2m6m75i3zo.png" width="800" height="1422"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;2026 Differentiator: Background inference while using other apps. Home screen widgets for quick queries. Optimized for 3B–7B models on mobile RAM. Download manager with resume and queue.&lt;/p&gt;

&lt;p&gt;Caveat: Limited to small models (≤8B). No document RAG. Smaller community than desktop alternatives. Noticeable battery drain during extended inference sessions.&lt;/p&gt;

&lt;p&gt;PocketPal AI is an open-source mobile app for offline local LLM inference on Android/iOS with background operation and home screen widgets. Optimized for 3B–7B models.&lt;/p&gt;

&lt;h3&gt;
  
  
  13. SillyTavern — Best Roleplay Frontend &amp;amp; Extension Ecosystem
&lt;/h3&gt;

&lt;p&gt;Best for creative writers wanting character management, world-building tools, and extensible plugin architecture.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fcdn-images-1.medium.com%2Fmax%2F1024%2F0%2APTPpiSQ3HTojxUXP" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fcdn-images-1.medium.com%2Fmax%2F1024%2F0%2APTPpiSQ3HTojxUXP" width="1024" height="545"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;2026 Differentiator: Largest RP extension ecosystem (TTS, image gen, emotion detection, web search). Character card v2 spec support. World Info depth and recursion controls—multi-backend switching mid-chat.&lt;/p&gt;

&lt;p&gt;Caveat: Frontend only — requires a separate inference backend. Node.js dependency. Steep learning curve for extensions. Not suited for productivity use cases.&lt;/p&gt;

&lt;p&gt;SillyTavern is an open-source roleplay frontend with extensive extensions, character/world management, and multi-backend support. Requires a separate inference engine.&lt;/p&gt;

&lt;h3&gt;
  
  
  14. BoltAI — Best Native Mac Productivity Client
&lt;/h3&gt;

&lt;p&gt;Best for Mac power users wanting system-wide AI commands and seamless app integration via hotkey palette.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fcdn-images-1.medium.com%2Fmax%2F739%2F0%2AFPYNkv-BL4KsrQkB" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fcdn-images-1.medium.com%2Fmax%2F739%2F0%2AFPYNkv-BL4KsrQkB" width="739" height="415"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;2026 Differentiator: AI Command Palette exposes 50+ actions via global hotkey. Inline text rewriting in any app with formatting preservation. Native macOS performance (no Electron). Perpetual license option available.&lt;/p&gt;

&lt;p&gt;Caveat: Paid app ($79–$99 one-time). Mac-only. Not an inference engine. No MCP or document RAG support. Maintained by a single developer.&lt;/p&gt;

&lt;p&gt;BoltAI is a paid native macOS AI client with a system-wide command palette for inline text rewriting. Connects to local/cloud backends. Perpetual license available.&lt;/p&gt;

&lt;h3&gt;
  
  
  15. Msty — Best Side-by-Side Model Comparison
&lt;/h3&gt;

&lt;p&gt;Best for evaluators and prompt engineers comparing outputs from multiple models simultaneously with branching conversations.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F80m5zld6nfgbmk4224rq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F80m5zld6nfgbmk4224rq.png" width="800" height="538"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;2026 Differentiator: Split Chats send identical prompts to 2–4 models concurrently. Branching conversation trees for A/B testing. Knowledge Stacks for curated model collections. Zero telemetry. Polished UX.&lt;/p&gt;

&lt;p&gt;Caveat: Closed-source codebase. Paid Aurum tier ($149/yr) required for teams and power features. No mobile app. Smaller model library than LM Studio.&lt;/p&gt;

&lt;p&gt;Msty is a local-first desktop app for side-by-side model comparison with split chats, branching conversations, and MCP support. Free tier available; closed-source.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fy4w1m8fdkvgmaf60yj0v.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fy4w1m8fdkvgmaf60yj0v.png" width="800" height="357"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Conclusion
&lt;/h3&gt;

&lt;p&gt;The local LLM landscape in 2026 has matured beyond simple chat interfaces into a diverse ecosystem of specialized tools. There is no single “best” app for everyone — the right choice depends entirely on your hardware, workflow, and privacy requirements.&lt;/p&gt;

&lt;p&gt;If you want one recommendation to start today: Atomic Chat offers the best balance of performance, ease of use, and cross-platform support for most users. Its TurboQuant engine makes larger models accessible on modest hardware, and native MCP support future-proofs your setup for agentic workflows.&lt;/p&gt;

&lt;p&gt;But don’t default to it blindly. Use this guide’s decision framework:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Developers and automators should start with Ollama or LocalAI for API-first workflows.&lt;/li&gt;
&lt;li&gt;Creative writers and roleplayers will get more value from KoboldCPP + SillyTavern than any general-purpose app.&lt;/li&gt;
&lt;li&gt;Teams and enterprises need PrivateGPT’s access controls and audit logs, not desktop chat apps.&lt;/li&gt;
&lt;li&gt;Mobile-first users should test PocketPal AI or Enchanted before assuming desktop tools are the only option.&lt;/li&gt;
&lt;li&gt;Privacy purists must verify telemetry settings regardless of which app they choose — open-source and auditable (Jan, Ollama, Llama.cpp) eliminate trust assumptions.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Hardware is your real constraint, not software. Before installing any app, consult the Hardware Matching Guide above to pair your RAM/VRAM with viable models. A perfectly configured app running an oversized model will underperform a modest setup running a well-matched one. Quantization has narrowed the gap between consumer hardware and capable AI, but physics still applies.&lt;/p&gt;

&lt;p&gt;Local AI is no longer experimental. With 9 of 15 apps now supporting MCP, Vulkan enabling AMD GPUs without ROCm friction, and mobile inference reaching practical usability, running AI locally in 2026 is a production-ready choice for privacy-sensitive, offline, or high-volume workloads. The tools listed here represent the current state of that maturity — tested, benchmarked, and categorized so you can skip the trial-and-error phase.&lt;/p&gt;

&lt;h3&gt;
  
  
  Thank you so much for reading
&lt;/h3&gt;

&lt;p&gt;Like | Follow | Subscribe to the newsletter.&lt;/p&gt;

&lt;p&gt;Catch us on&lt;/p&gt;

&lt;p&gt;Website: &lt;a href="https://www.techlatest.net/" rel="noopener noreferrer"&gt;https://www.techlatest.net/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Newsletter: &lt;a href="https://substack.com/@techlatestnet" rel="noopener noreferrer"&gt;https://substack.com/@techlatestnet&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Twitter: &lt;a href="https://twitter.com/TechlatestNet" rel="noopener noreferrer"&gt;https://twitter.com/TechlatestNet&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;LinkedIn: &lt;a href="https://www.linkedin.com/in/techlatest-net/" rel="noopener noreferrer"&gt;https://www.linkedin.com/in/techlatest-net/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;YouTube:&lt;a href="https://www.youtube.com/@techlatest_net/" rel="noopener noreferrer"&gt;https://www.youtube.com/@techlatest_net/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Blogs: &lt;a href="https://medium.com/@techlatest.net" rel="noopener noreferrer"&gt;https://medium.com/@techlatest.net&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Reddit Community: &lt;a href="https://www.reddit.com/user/techlatest_net/" rel="noopener noreferrer"&gt;https://www.reddit.com/user/techlatest_net/&lt;/a&gt;&lt;/p&gt;

</description>
      <category>localllm</category>
      <category>localllmdeployment</category>
      <category>llm</category>
      <category>llmagent</category>
    </item>
    <item>
      <title>How to Deploy and Access Hermes Agent on Azure Marketplace: A Step-by-Step Guide</title>
      <dc:creator>TechLatest</dc:creator>
      <pubDate>Fri, 07 Aug 2026 15:06:46 +0000</pubDate>
      <link>https://dev.to/techlatestnet/how-to-deploy-and-access-hermes-agent-on-azure-marketplace-a-step-by-step-guide-2h29</link>
      <guid>https://dev.to/techlatestnet/how-to-deploy-and-access-hermes-agent-on-azure-marketplace-a-step-by-step-guide-2h29</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnw6yo45vfslpv8mypsk4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnw6yo45vfslpv8mypsk4.png" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Deploying autonomous AI agents on the cloud doesn’t have to be a complex process. &lt;strong&gt;Hermes Agent&lt;/strong&gt; is an open-source, production-ready framework that enables developers to build, deploy, and scale intelligent AI agents capable of reasoning, planning, tool usage, API integration, code execution, and workflow automation.&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;Hermes Agent — Build, Deploy &amp;amp; Scale Autonomous AI&lt;/strong&gt; virtual machine on &lt;strong&gt;Microsoft Azure Marketplace&lt;/strong&gt; provides a fully configured environment with all the required components pre-installed, allowing you to start building AI-powered applications within minutes. Whether you’re running local Ollama models, integrating cloud-based LLM providers, or developing production-ready autonomous AI workflows, this VM eliminates the need for manual installation and configuration.&lt;/p&gt;

&lt;p&gt;In this guide, you’ll learn how to deploy the Hermes Agent VM from Azure Marketplace, connect to the instance using SSH or Remote Desktop (RDP), access the Hermes Web Interface, configure language models, and begin building autonomous AI applications on Azure.&lt;/p&gt;

&lt;h4&gt;
  
  
  Step-by-Step Guide
&lt;/h4&gt;

&lt;p&gt;This section describes how to launch and connect to the ‘Hermes Agent — Build, Deploy &amp;amp; Scale Autonomous AI’ VM solution on the Azure Platform.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Open the &lt;a href="https://marketplace.microsoft.com/en-us/product/techlatest.hermes-agent-vm?tab=Overview?utm_campaign=hermes-agent-vm&amp;amp;utm_source=techlatest-website&amp;amp;utm_medium=support-page" rel="noopener noreferrer"&gt;Hermes Agent — Build, Deploy &amp;amp; Scale Autonomous AI&lt;/a&gt; VM listing on Azure Marketplace.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fa286m9li3hbp6g5z3mpi.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fa286m9li3hbp6g5z3mpi.png" width="800" height="280"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Click on Get It Now&lt;/li&gt;
&lt;/ol&gt;

&lt;ul&gt;
&lt;li&gt;Log in with your credentials and provide the details here. Once done, click on the Get it now button at the bottom.&lt;/li&gt;
&lt;li&gt;It will take you to the Product details page. Click on Create.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzivvv1trl80qhjzqllxb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzivvv1trl80qhjzqllxb.png" width="799" height="348"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Select a Resource group for your virtual machine&lt;/li&gt;
&lt;li&gt;Select a Region where you want to launch the VM(such as East US)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftxanb730sbdqbr4motyw.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftxanb730sbdqbr4motyw.png" width="799" height="475"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Note: If you see the “This image is not compatible with selected security type. To keep trusted launch virtual machines, select a compatible image. Otherwise change your security type back to Standard” error message below the Image name as shown in the screenshot below, then please change the Security type to Standard.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fp7dbhpjn3mf36erb5hm0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fp7dbhpjn3mf36erb5hm0.png" width="798" height="223"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F121nckegv09y9fyjgeo7.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F121nckegv09y9fyjgeo7.png" width="799" height="186"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Optionally change the number of cores and amount of memory.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Minimum VM Specs: 16GB RAM / 4 vCPUs. Please also check publisher recommendations for more instance options.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fy512yg7t3jxu44vthoy1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fy512yg7t3jxu44vthoy1.png" width="800" height="449"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The VM can also be deployed using an NVIDIA GPU instance for faster execution. Please check the Publisher recommendations instance type for GPU (Standard_NC4as_T4_v3–4 vCPUs, 28 GiB memory) or check the available NVIDIA GPU instances on the &lt;a href="https://learn.microsoft.com/en-us/azure/virtual-machines/sizes/gpu-accelerated/ncast4v3-series?tabs=sizebasic" rel="noopener noreferrer"&gt;Azure documentation page&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fl3tq15xnndoqa3n8ntt3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fl3tq15xnndoqa3n8ntt3.png" width="799" height="224"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Select the Authentication type as Password and enter Username as ubuntu and Password of your choice.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fccghn2yygma792y6d15i.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fccghn2yygma792y6d15i.png" width="800" height="335"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Optionally change the OS disk size and its type. By default, the VM comes with 50GB of disk.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb5e0vnpboa3ftosnv5up.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb5e0vnpboa3ftosnv5up.png" width="800" height="366"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Optionally change the network and subnetwork names. Be sure that whichever network you specify has ports 22 (for SSH), 3389 (for RDP), 80 (for HTTP), and 443 (for HTTPS) exposed.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The VM comes with the preconfigured NSG rules. You can check them by clicking on the Create New option available under the security group option.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvpplb7v4utpdp1g9qx5w.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvpplb7v4utpdp1g9qx5w.png" width="800" height="314"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmmqq9obpfkbvzdlj3fi4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmmqq9obpfkbvzdlj3fi4.png" width="800" height="518"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Optionally, go to the Management, Advanced, and Tags tabs for any advanced settings you want for the VM.&lt;/li&gt;
&lt;li&gt;Click on Review + create and then click on Create when you are done.
The virtual machine will begin deploying.&lt;/li&gt;
&lt;/ul&gt;

&lt;ol&gt;
&lt;li&gt;A summary page displays when the virtual machine is successfully created. Click on Go to resource link to go to the resource page. It will open an overview page of the virtual machine.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz7leo08uzto97kjfcvot.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz7leo08uzto97kjfcvot.png" width="800" height="372"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;If you want to update your password, then open up the left navigation pane, select Run command, select RunShellScript, and enter the following command to change the password of the VM.
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo echo &lt;/span&gt;ubuntu:yourpassword | chpasswd
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsegxyerflumrniry3vau.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsegxyerflumrniry3vau.png" width="800" height="334"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsjwrpb23r8fm9i0rnjs4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsjwrpb23r8fm9i0rnjs4.png" width="549" height="563"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Now that the password for the Ubuntu user is set, you can SSH to the VM. To do so, first note the public IP address of the VM from the VM details page as highlighted below.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fp1i0lj1c07f5f6uehg5o.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fp1i0lj1c07f5f6uehg5o.png" width="800" height="348"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Open PuTTY, paste the IP address, and click on Open.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5y1kk90rfxeor4th233k.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5y1kk90rfxeor4th233k.png" width="448" height="441"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Log in as ubuntu and provide the password for the ‘ubuntu’ user.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8qp2x5aq7oqiskv9fgi3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8qp2x5aq7oqiskv9fgi3.png" width="800" height="590"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;You can also connect to the VM’s desktop environment from any local Windows machine using the RDP protocol or a local Linux machine using Remmina.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;To connect using RDP via a Windows Machine, copy the public IP address of the VM from the VM details page, then from your local Windows machine, go to the “Start” menu, in the search box type and select “Remote Desktop Connection”. In the “Remote Desktop Connection” wizard, copy the public IP address and click Connect.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5r1p31fuc2bj4tnwb8l8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5r1p31fuc2bj4tnwb8l8.png" width="474" height="320"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;This will connect you to the VM’s desktop environment. Provide the username as ubuntu and the password set in step 4 to authenticate. Click OK&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb6zwmrcavjwvp97tgonx.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb6zwmrcavjwvp97tgonx.png" width="800" height="552"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Now you are connected to the out-of-box “Hermes Agent — Build, Deploy &amp;amp; Scale Autonomous AI” VM’s desktop environment via Windows Machine.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4vqc5wdt5eyjgi0sivxo.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4vqc5wdt5eyjgi0sivxo.png" width="800" height="561"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;To connect using RDP via a Linux machine, first note the external IP of the VM from the VM details page, then from your local Linux machine, goto menu, in the search box type and select “Remmina”.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Note: If you don’t have Remmina installed on your Linux machine, first &lt;a href="https://remmina.org/how-to-install-remmina/" rel="noopener noreferrer"&gt;install Remmina&lt;/a&gt; as per your Linux distribution.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdea3l933oha4l4ac7gxv.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdea3l933oha4l4ac7gxv.png" width="615" height="504"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;In the “Remmina Remote Desktop Client” wizard, select the RDP option from the dropdown, paste the external IP, and click Enter.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz77p3a2wi85j6rmo3dsh.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz77p3a2wi85j6rmo3dsh.png" width="800" height="488"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;This will connect you to the VM’s desktop environment. Provide “ubuntu” as the user ID and the password set in the above reset password step to authenticate. Click OK&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7r8nst6j8l0ptj4v8qeu.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7r8nst6j8l0ptj4v8qeu.png" width="800" height="379"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Now you are connected to the out-of-box “Hermes Agent — Build, Deploy &amp;amp; Scale Autonomous AI” VM’s desktop environment via Linux machine.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvvirbjh9wi40lc89yqkr.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvvirbjh9wi40lc89yqkr.png" width="800" height="561"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The VM will generate a random password to log in to Hermes Web Interface. To get the password, connect via SSH terminal as shown in the above step and run the command.
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;cat&lt;/span&gt; /home/ubuntu/.hermes/.env | &lt;span class="nb"&gt;grep &lt;/span&gt;HERMES_DASHBOARD_BASIC_AUTH
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyhobqb6k7yw11dhj0r7c.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyhobqb6k7yw11dhj0r7c.png" width="800" height="173"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;To access the Hermes Web Interface, copy the public IP address of the VM and paste it in your local browser as &lt;a href="https://public_ip_of_vm." rel="noopener noreferrer"&gt;https://public_ip_of_vm.&lt;/a&gt; Make sure to use https and not http.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The browser will display an SSL certificate warning message. Expand the warning message, accept the certificate warning, and click Continue.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F86xzh9pqg1gog5shf7qv.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F86xzh9pqg1gog5shf7qv.png" width="800" height="589"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;It will open a login page. Provide the password we got in the above step and click Sign In.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F62fybq3mid8xut00cbp8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F62fybq3mid8xut00cbp8.png" width="800" height="555"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Now you are connected to the out-of-box Hermes Web Interface.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fp9kbhqkgzpppxqywnkhe.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fp9kbhqkgzpppxqywnkhe.png" width="800" height="499"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;You can use the Hermes chat feature to run tasks or ask questions.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F77xan4tu0bloh1y72emh.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F77xan4tu0bloh1y72emh.png" width="799" height="541"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;By default, the LLM model set is “deepseek-r1:8b”h. You can pull other Ollama models.
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ollama pull &amp;lt;model_name&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;e.g ollama pull gemma2:9b&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7rz25wmzagiufkj2agy3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7rz25wmzagiufkj2agy3.png" width="800" height="311"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Once your model is pulled, you can set it to default from the web interface as well as from the terminal. To switch models from the web interface, simply click on the model dropdown from the top right of your chat window. Choose the model you want to set and click Switch&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fp43ts1on4dhajyom5ave.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fp43ts1on4dhajyom5ave.png" width="799" height="217"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmbhq3r3fbahcf20491ra.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmbhq3r3fbahcf20491ra.png" width="781" height="595"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Or from the terminal, you can run,&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes config &lt;span class="nb"&gt;set &lt;/span&gt;model &amp;lt;provider_name&amp;gt;/&amp;lt;model_name&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;e.g. hermes config set model ollama/gemma2:9b&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsldxjl3wevteracqvvdb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsldxjl3wevteracqvvdb.png" width="800" height="169"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;To change the LLM provider and set the API Keys, please run the command.
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes model
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Choose your provider of choice and follow the on-screen instructions. Once the process is complete, go back to the web interface and refresh the page to see the changes.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2rlknti3aj4l7le1hnfx.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2rlknti3aj4l7le1hnfx.png" width="800" height="493"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;If, for any Ollama model, you are getting a context length error as shown in the screenshot below while running the chat, then set the context_length and ollama_num_ctx to the required value by running the commands in the terminal, then refresh the WebUI.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Note: This is specific to Ollama; if you want to do it for other providers, then make the appropriate changes in the commands.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes config &lt;span class="nb"&gt;set &lt;/span&gt;model.ollama_num_ctx 65536

hermes config &lt;span class="nb"&gt;set &lt;/span&gt;model.context_length 65536
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjdhwi0jzwmq1y9hol2qi.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjdhwi0jzwmq1y9hol2qi.png" width="800" height="291"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyij2custvfzta3si5j41.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyij2custvfzta3si5j41.png" width="800" height="182"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;For more details, please visit the &lt;a href="https://techlatest.net/support/hermes_agent_support/user_guide/" rel="noopener noreferrer"&gt;Official Documentation page&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Thank you so much for reading
&lt;/h3&gt;

&lt;p&gt;Like | Follow | Subscribe to the newsletter.&lt;/p&gt;

&lt;p&gt;Catch us on&lt;/p&gt;

&lt;p&gt;Website: &lt;a href="https://www.techlatest.net/" rel="noopener noreferrer"&gt;https://www.techlatest.net/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Newsletter: &lt;a href="https://substack.com/@techlatestnet" rel="noopener noreferrer"&gt;https://substack.com/@techlatestnet&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Twitter: &lt;a href="https://twitter.com/TechlatestNet" rel="noopener noreferrer"&gt;https://twitter.com/TechlatestNet&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;LinkedIn: &lt;a href="https://www.linkedin.com/in/techlatest-net/" rel="noopener noreferrer"&gt;https://www.linkedin.com/in/techlatest-net/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;YouTube:&lt;a href="https://www.youtube.com/@techlatest_net/" rel="noopener noreferrer"&gt;https://www.youtube.com/@techlatest_net/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Blogs: &lt;a href="https://medium.com/@techlatest.net" rel="noopener noreferrer"&gt;https://medium.com/@techlatest.net&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Reddit Community: &lt;a href="https://www.reddit.com/user/techlatest_net/" rel="noopener noreferrer"&gt;https://www.reddit.com/user/techlatest_net/&lt;/a&gt;&lt;/p&gt;

</description>
      <category>opensource</category>
      <category>agents</category>
      <category>microsoftazure</category>
      <category>hermesagent</category>
    </item>
    <item>
      <title>How to Deploy and Access Hermes Agent on GCP Marketplace: A Step-by-Step Guide</title>
      <dc:creator>TechLatest</dc:creator>
      <pubDate>Fri, 07 Aug 2026 08:02:09 +0000</pubDate>
      <link>https://dev.to/techlatestnet/how-to-deploy-and-access-hermes-agent-on-gcp-marketplace-a-step-by-step-guide-1c8</link>
      <guid>https://dev.to/techlatestnet/how-to-deploy-and-access-hermes-agent-on-gcp-marketplace-a-step-by-step-guide-1c8</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9uxhlq2659sqvc56qp73.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9uxhlq2659sqvc56qp73.png" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Building and deploying autonomous AI agents doesn’t have to involve lengthy installations or complex infrastructure setup. &lt;strong&gt;Hermes Agent&lt;/strong&gt; is an open-source, production-ready framework that enables developers to create AI agents capable of reasoning, planning, and executing multi-step tasks using tools, APIs, web browsing, code execution, and workflow automation.&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;Hermes Agent — Build, Deploy &amp;amp; Scale Autonomous AI&lt;/strong&gt; virtual machine on &lt;strong&gt;Google Cloud Marketplace&lt;/strong&gt; provides a fully configured environment with everything pre-installed, allowing you to start building immediately. Whether you want to run local Ollama models, connect to cloud-based LLM providers, or develop production-ready AI workflows, the VM eliminates manual configuration so you can focus on development.&lt;/p&gt;

&lt;p&gt;In this tutorial, you’ll learn how to deploy the Hermes Agent VM from Google Cloud Marketplace, connect to the instance through the browser-based SSH console, access the Hermes Web Interface, configure AI models, and start building autonomous AI agents in just a few minutes.&lt;/p&gt;

&lt;h4&gt;
  
  
  Step-by-Step Guide
&lt;/h4&gt;

&lt;p&gt;This section describes how to provision and connect to the ‘Hermes Agent — Build, Deploy &amp;amp; Scale Autonomous AI’ VM solution on GCP.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Open the &lt;a href="https://console.cloud.google.com/marketplace/product/techlatest-public/hermes-agent-vm?utm_campaign=hermes-agent-vm&amp;amp;utm_source=techlatest-website&amp;amp;utm_medium=support-page" rel="noopener noreferrer"&gt;Hermes Agent — Build, Deploy &amp;amp; Scale Autonomous AI&lt;/a&gt; listing on GCP Marketplace.&lt;/li&gt;
&lt;li&gt;Click Get Started.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4c5paffiu5xv7moxx425.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4c5paffiu5xv7moxx425.png" width="799" height="384"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;It will ask you to enable the API’s if they are not enabled already for your account. Please click on Enable as shown in the screenshot.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsuttt6vzehc2ebqoe0u0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsuttt6vzehc2ebqoe0u0.png" width="773" height="414"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;It will take you to the agreement page. On this page, you can change the project from the project selector on the top navigation bar as shown in the screenshot below.&lt;/li&gt;
&lt;li&gt;Accept the Terms and agreements by ticking the checkbox and clicking on the AGREE button.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyzwhoph8v7pqq8ftkp6t.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyzwhoph8v7pqq8ftkp6t.png" width="614" height="523"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;It will show you the successfully agreed popup page. Click on Deploy.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbkztt649c72gvgcffufa.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbkztt649c72gvgcffufa.png" width="800" height="395"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;On the deployment page, give a name to your deployment.&lt;/li&gt;
&lt;li&gt;In the Deployment Service Account section, click on the Existing radio button and choose a service account from the Select a Service Account dropdown.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fih943q68jcbvtegxz1h8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fih943q68jcbvtegxz1h8.png" width="559" height="338"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;If you don’t see any service account in the dropdown, then change the radio button to New Account and create the new service account here.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fd75g7fdb95jyawf3hk61.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fd75g7fdb95jyawf3hk61.png" width="544" height="411"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;If, after selecting the New Account option, you get the permission error message, then please reach out to your GCP admin to create a service account by following the &lt;a href="https://techlatest.net/support/guide_to_create_gcp_service_account" rel="noopener noreferrer"&gt;step-by-step guide to create a GCP Service&lt;/a&gt; Account, and then refresh this deployment page once the service account is created; it should be available in the dropdown.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Faqmo0geqqcmuzwhgzjwl.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Faqmo0geqqcmuzwhgzjwl.png" width="534" height="301"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Select a zone where you want to launch the VM(such as us-east1-a)&lt;/li&gt;
&lt;li&gt;Optionally change the number of cores and amount of memory.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Minimum VM Specs: 15GB RAM /4vCPU&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fayan6lbo28ptyicy9lij.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fayan6lbo28ptyicy9lij.png" width="799" height="584"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This VM can also be deployed using an NVIDIA T4 GPU instance for faster inference. To deploy the VM with a GPU, click on the GPU tab as shown in the screenshot and select an NVIDIA T4 GPU instance. Please note that GPU availability is limited to specific regions, zones, and machine types. If you do not see a GPU option for your selected region, zone, or machine type, try adjusting those settings to find available configurations.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsuhc2xsd0klly0cgu1si.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsuhc2xsd0klly0cgu1si.png" width="799" height="698"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Optionally change the boot disk type and size. (This defaults to ‘Standard Persistent Disk’ and 50GB respectively)&lt;/li&gt;
&lt;li&gt;Optionally change the network name and subnetwork names. Be sure that whichever network you specify has ports 22 (for SSH) and 443 (for HTTPS) exposed.&lt;/li&gt;
&lt;li&gt;Click Deploy when you are done.&lt;/li&gt;
&lt;li&gt;Hermes Agent — Build, Deploy &amp;amp; Scale Autonomous AI will begin deploying.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F68ptzxggppr1pe084ccd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F68ptzxggppr1pe084ccd.png" width="800" height="492"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1uuyvpw2bpjdvmxnc86g.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1uuyvpw2bpjdvmxnc86g.png" width="799" height="495"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F18j3rp7tx3112m4kcguh.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F18j3rp7tx3112m4kcguh.png" width="800" height="897"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;A summary page displays when the compute engine is successfully deployed. Click on the Instance link to go to the instance page.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;On the instance page, click on the “SSH” button, select “Open in browser window”.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr2yigv5iilvr7bc8u8t0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr2yigv5iilvr7bc8u8t0.png" width="551" height="302"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;This will open an SSH window in a browser. Switch to the ubuntu user and navigate to the ubuntu home directory.
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;su ubuntu

&lt;span class="nb"&gt;cd&lt;/span&gt; /home/ubuntu/
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0o4pq6qsodwm2gcg431c.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0o4pq6qsodwm2gcg431c.png" width="800" height="513"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The VM will generate a random password to log in to Hermes Web Interface. To get the password, connect via SSH terminal as shown in the above step and run the command.
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;cat&lt;/span&gt; /home/ubuntu/.hermes/.env | &lt;span class="nb"&gt;grep &lt;/span&gt;HERMES_DASHBOARD_BASIC_AUTH
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frgc6xdzgufedj6sk5fb0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frgc6xdzgufedj6sk5fb0.png" width="800" height="173"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;To access the Hermes Web Interface, copy the public IP address of the VM and paste it into your local browser as &lt;a href="https://public_ip_of_vm." rel="noopener noreferrer"&gt;https://public_ip_of_vm.&lt;/a&gt; Make sure to use https and not http.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The browser will display an SSL certificate warning message. Expand the warning message, accept the certificate warning, and click Continue.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3zrteooxcqw6k05q56yh.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3zrteooxcqw6k05q56yh.png" width="800" height="589"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;It will open a login page. Provide the password we got in the above step and click Sign In.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcxf43e6uavwt4zdxz85h.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcxf43e6uavwt4zdxz85h.png" width="800" height="555"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Now you are connected to the out-of-box Hermes Web Interface.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi1to63e1ganpd9ut7lnj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi1to63e1ganpd9ut7lnj.png" width="800" height="499"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;You can use the Hermes chat feature to run tasks or ask questions.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7dupulhx025j8rehvnpx.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7dupulhx025j8rehvnpx.png" width="799" height="541"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;By default, the LLM model set is “deepseek-r1:8b”h. You can pull other Ollama models.
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ollama pull &amp;lt;model_name&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;e.g ollama pull gemma2:9b&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyxks45zw1j47sum19hyg.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyxks45zw1j47sum19hyg.png" width="800" height="311"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Once your model is pulled, you can set it to default from the web interface as well as from the terminal. To switch models from the web interface, simply click on the model dropdown from the top right of your chat window. Choose the model you want to set and click Switch&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvcm3uzdgvvs3z0xqvryu.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvcm3uzdgvvs3z0xqvryu.png" width="799" height="217"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fw7ida4enqjcpqz3ufg37.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fw7ida4enqjcpqz3ufg37.png" width="781" height="595"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Or from the terminal, you can run,&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes config &lt;span class="nb"&gt;set &lt;/span&gt;model &amp;lt;provider_name&amp;gt;/&amp;lt;model_name&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;e.g. hermes config set model ollama/gemma2:9b&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq24qwhw1mwcwyvhk8leq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq24qwhw1mwcwyvhk8leq.png" width="800" height="169"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;To change the LLM provider and set the API Keys, please run the command.
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes model
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Choose your provider of choice and follow the on-screen instructions. Once the process is complete, go back to the web interface and refresh the page to see the changes.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdequ4turrerb92er2u04.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdequ4turrerb92er2u04.png" width="800" height="493"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;If for any Ollama model you are getting a context length error as shown in the screenshot below while running the chat, then set the context_length and ollama_num_ctx to the required value by running the commands in the terminal, then refresh the WebUI.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Note: This is specific to Ollama; if you want to do it for other providers, then make the appropriate changes in the commands.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes config &lt;span class="nb"&gt;set &lt;/span&gt;model.ollama_num_ctx 65536

hermes config &lt;span class="nb"&gt;set &lt;/span&gt;model.context_length 65536
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fleiojqiso0g91vgif0pq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fleiojqiso0g91vgif0pq.png" width="800" height="291"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn721ow4f8myht7bimjgp.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn721ow4f8myht7bimjgp.png" width="800" height="182"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;For more details, please visit the &lt;a href="https://techlatest.net/support/hermes_agent_support/user_guide/" rel="noopener noreferrer"&gt;Official Documentation page&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Conclusion
&lt;/h3&gt;

&lt;p&gt;Congratulations! You have successfully deployed &lt;strong&gt;Hermes Agent&lt;/strong&gt; on Google Cloud and configured your environment for autonomous AI development. With its pre-configured virtual machine, browser-based dashboard, and powerful Hermes CLI, you can start building intelligent AI agents without spending time on manual installation and dependency management.&lt;/p&gt;

&lt;p&gt;Hermes supports multiple LLM providers, local Ollama models, persistent memory, and workflow automation, making it suitable for everything from AI assistants and internal automation tools to complex agentic applications. As your workloads grow, you can easily scale your deployment by upgrading your Compute Engine instance or adding an NVIDIA T4 GPU for faster inference and improved performance.&lt;/p&gt;

&lt;p&gt;Now that your Hermes Agent environment is up and running, you can begin experimenting with different language models, automate complex workflows, and build production-ready autonomous AI applications on Google Cloud. For advanced configuration options, additional integrations, and best practices, refer to the official Hermes documentation.&lt;/p&gt;

&lt;h3&gt;
  
  
  Thank you so much for reading
&lt;/h3&gt;

&lt;p&gt;Like | Follow | Subscribe to the newsletter.&lt;/p&gt;

&lt;p&gt;Catch us on&lt;/p&gt;

&lt;p&gt;Website: &lt;a href="https://www.techlatest.net/" rel="noopener noreferrer"&gt;https://www.techlatest.net/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Newsletter: &lt;a href="https://substack.com/@techlatestnet" rel="noopener noreferrer"&gt;https://substack.com/@techlatestnet&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Twitter: &lt;a href="https://twitter.com/TechlatestNet" rel="noopener noreferrer"&gt;https://twitter.com/TechlatestNet&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;LinkedIn: &lt;a href="https://www.linkedin.com/in/techlatest-net/" rel="noopener noreferrer"&gt;https://www.linkedin.com/in/techlatest-net/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;YouTube:&lt;a href="https://www.youtube.com/@techlatest_net/" rel="noopener noreferrer"&gt;https://www.youtube.com/@techlatest_net/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Blogs: &lt;a href="https://medium.com/@techlatest.net" rel="noopener noreferrer"&gt;https://medium.com/@techlatest.net&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Reddit Community: &lt;a href="https://www.reddit.com/user/techlatest_net/" rel="noopener noreferrer"&gt;https://www.reddit.com/user/techlatest_net/&lt;/a&gt;&lt;/p&gt;

</description>
      <category>agents</category>
      <category>googlecloudplatform</category>
      <category>opensource</category>
      <category>gcp</category>
    </item>
    <item>
      <title>How to Deploy and Access Hermes Agent on AWS Marketplace: A Step-by-Step Guide</title>
      <dc:creator>TechLatest</dc:creator>
      <pubDate>Thu, 06 Aug 2026 14:17:32 +0000</pubDate>
      <link>https://dev.to/techlatestnet/how-to-deploy-and-access-hermes-agent-on-aws-marketplace-a-step-by-step-guide-23lf</link>
      <guid>https://dev.to/techlatestnet/how-to-deploy-and-access-hermes-agent-on-aws-marketplace-a-step-by-step-guide-23lf</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Futl2cooba2qhj5hd2bn8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Futl2cooba2qhj5hd2bn8.png" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Autonomous AI agents are transforming how developers automate complex tasks by combining the reasoning capabilities of large language models with tools, APIs, code execution, and workflow orchestration. &lt;strong&gt;Hermes Agent&lt;/strong&gt; is an open-source, production-ready framework designed to help you build, deploy, and scale intelligent AI agents that can plan, reason, and execute multi-step tasks with minimal human intervention.&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;Hermes Agent — Build, Deploy &amp;amp; Scale Autonomous AI&lt;/strong&gt; virtual machine from TechLatest provides a fully configured environment with everything you need to get started. Instead of spending time installing dependencies and configuring services, you can launch a ready-to-use instance that includes the Hermes CLI, a browser-based dashboard, and support for multiple AI model providers. Whether you’re building AI assistants, automating workflows, or experimenting with agentic applications, this VM enables you to start developing immediately.&lt;/p&gt;

&lt;p&gt;In this tutorial, you’ll learn how to deploy the Hermes Agent VM, securely connect to the instance, access the web dashboard, configure language models, and begin building autonomous AI workflows in just a few steps.&lt;/p&gt;

&lt;h4&gt;
  
  
  Step-by-Step Guide
&lt;/h4&gt;

&lt;p&gt;This section describes how to launch and connect to the ‘Hermes Agent — Build, Deploy &amp;amp; Scale Autonomous AI’ VM solution on AWS.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Open the &lt;a href="https://aws.amazon.com/marketplace/pp/prodview-xditi57gdyg5u?utm_campaign=hermes-agent-vm&amp;amp;utm_source=techlatest-website&amp;amp;utm_medium=support-page" rel="noopener noreferrer"&gt;Hermes Agent — Build, Deploy &amp;amp; Scale Autonomous AI&lt;/a&gt; VM listing on the AWS Marketplace.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmcc3fi1uxzahewfjkhuq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmcc3fi1uxzahewfjkhuq.png" width="800" height="202"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Click on View purchase options.&lt;/li&gt;
&lt;/ol&gt;

&lt;ul&gt;
&lt;li&gt;Log in with your credentials and follow the instructions.&lt;/li&gt;
&lt;li&gt;Review the prices and subscribe to the product by clicking on the Subscribe button located at the bottom of this page. Once you are subscribed to the offer, click on the Launch your software button.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmpq6b7j8kp3esjczo3c6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmpq6b7j8kp3esjczo3c6.png" width="799" height="277"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4xlx4nsar115kmekuj9q.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4xlx4nsar115kmekuj9q.png" width="800" height="456"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The next page will show you the options to launch the instance: Launch through EC2 and One-click launch from AWS Marketplace. Tick the 2nd option, One-click launch from AWS Marketplace.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn34nzy9k7z6nis4fabtn.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn34nzy9k7z6nis4fabtn.png" width="800" height="430"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Select a Region where you want to launch the VM(such as US East (N.Virginia))&lt;/li&gt;
&lt;li&gt;Optionally change the EC2 instance type. (This defaults to t2.xlarge instance type, 4 vCPUs, and 16 GB RAM.)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Please note that the VM can also be deployed using an NVIDIA GPU instance. If you want to deploy this instance with a GPU configuration, then please choose an NVIDIA GPU (e.g g4dn.xlarge) or check the available NVIDIA GPU instances on the &lt;a href="https://aws.amazon.com/ec2/instance-types/g4/" rel="noopener noreferrer"&gt;AWS documentation page&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1kwutzf1zj0vup5enhbh.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1kwutzf1zj0vup5enhbh.png" width="800" height="300"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Friqz8ril6s3dqv6ajum1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Friqz8ril6s3dqv6ajum1.png" width="800" height="481"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Optionally change the network name and subnetwork names.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffy5vrfvyqqbsjifvok82.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffy5vrfvyqqbsjifvok82.png" width="773" height="245"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Select the Security Group. Be sure that whichever Security Group you specify have ports 22 (for SSH) and 443 (for HTTPS) exposed. Or you can create the new SG by clicking on the “Create Security Group” button. Provide the name and description, and save the SG for this instance.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbnlrpvwvgfkxr3z3lrgi.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbnlrpvwvgfkxr3z3lrgi.png" width="797" height="116"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F26ehkv9clv950fquxlff.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F26ehkv9clv950fquxlff.png" width="800" height="457"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Be sure to download the key pair, which is available by default, or you can create a new key pair and download it.&lt;/li&gt;
&lt;li&gt;Click on Launch.&lt;/li&gt;
&lt;li&gt;Hermes Agent — Build, Deploy &amp;amp; Scale Autonomous AI will begin deploying.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgd4deopci1n9tkemnd9a.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgd4deopci1n9tkemnd9a.png" width="799" height="248"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;A summary page displays. To see this instance on the EC2 Console, click on the View instance on EC2 link.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbsw39mpcibtl5h6xao2r.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbsw39mpcibtl5h6xao2r.png" width="799" height="305"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;To connect to this instance through PuTTY, copy the IPv4 Public IP Address from the VM’s details page.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsrm3df0pc29i7jf4ppbe.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsrm3df0pc29i7jf4ppbe.png" width="800" height="322"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Open PuTTY, paste the IP address, and browse to the private key you downloaded while deploying the VM. Go to SSH-&amp;gt;Auth-&amp;gt;Credentials, click on Open. Enter ubuntu as the user ID.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyh09blfqasgh5ijxnyen.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyh09blfqasgh5ijxnyen.png" width="454" height="444"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fos6u4snq1ekxfz2ay5xt.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fos6u4snq1ekxfz2ay5xt.png" width="451" height="442"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft1du9vp6g16558m97r5u.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft1du9vp6g16558m97r5u.png" width="800" height="485"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The VM will generate a random password to log in to Hermes Web Interface. To get the password, connect via SSH terminal as shown in the above step and run the command.
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;cat&lt;/span&gt; /home/ubuntu/.hermes/.env | &lt;span class="nb"&gt;grep &lt;/span&gt;HERMES_DASHBOARD_BASIC_AUTH
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fimgx4bch975gigbaazt5.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fimgx4bch975gigbaazt5.png" width="800" height="173"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;To access the Hermes Web Interface, copy the public IP address of the VM and paste it into your local browser as &lt;a href="https://public_ip_of_vm." rel="noopener noreferrer"&gt;https://public_ip_of_vm.&lt;/a&gt; Make sure to use https and not http.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The browser will display an SSL certificate warning message. Expand the warning message, accept the certificate warning, and click Continue.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu627ggdjby2fg6saqxt9.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu627ggdjby2fg6saqxt9.png" width="800" height="589"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;It will open a login page. Provide the password we got at above step and click Sign In.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdb67tl1x4s3w0s8wv1f1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdb67tl1x4s3w0s8wv1f1.png" width="800" height="555"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Now you are connected to the out-of-box Hermes Web Interface.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjj1bpt685ccc4b43deod.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjj1bpt685ccc4b43deod.png" width="800" height="499"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;You can use the Hermes chat feature to run tasks or ask questions.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4egtyddfczi6kvf37rz8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4egtyddfczi6kvf37rz8.png" width="799" height="541"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;By default, the LLM model set is “deepseek-r1:8b”. You can pull other Ollama models.
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ollama pull &amp;lt;model_name&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;e.g ollama pull gemma2:9b&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvm5i0kobnh5t0m7dzkt7.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvm5i0kobnh5t0m7dzkt7.png" width="800" height="311"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Once your model is pulled, you can set it to the default from the web interface as well as from the terminal. To switch models from the web interface, simply click on the model dropdown from the top right of your chat window. Choose the model you want to set and click Switch&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F97w0zjjwqvxz2ud3zlc9.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F97w0zjjwqvxz2ud3zlc9.png" width="799" height="217"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkc5x80npwmhif4dveq5j.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkc5x80npwmhif4dveq5j.png" width="781" height="595"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;or from terminal you can run,&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes config &lt;span class="nb"&gt;set &lt;/span&gt;model &amp;lt;provider_name&amp;gt;/&amp;lt;model_name&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;e.g hermes config set model ollama/gemma2:9b&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F369p3yipw9td44j1qo96.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F369p3yipw9td44j1qo96.png" width="800" height="169"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;To change the LLM provider and set the API Keys, please run the command.
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes model
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Choose your provider of choice and follow the on-screen instructions. Once the process is complete, go back to the web interface and refresh the page to see the changes.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fehbsmqmbaxhru53646sp.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fehbsmqmbaxhru53646sp.png" width="800" height="493"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;If for any Ollama model you are getting context length error as shown in below screenshot, while running the chat then set the context_length and ollama_num_ctx to required value by running below commands in terminal then refresh the WebUI.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Note: This is specific to Ollama; if you want to do it for other providers, then make the appropriate changes in the commands below.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes config &lt;span class="nb"&gt;set &lt;/span&gt;model.ollama_num_ctx 65536

hermes config &lt;span class="nb"&gt;set &lt;/span&gt;model.context_length 65536
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5p33opw4k500cd8yl1tw.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5p33opw4k500cd8yl1tw.png" width="800" height="291"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgbdsdaokabbw4beud7s4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgbdsdaokabbw4beud7s4.png" width="800" height="182"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;For more details, please visit the &lt;a href="https://techlatest.net/support/hermes_agent_support/user_guide/" rel="noopener noreferrer"&gt;Official Documentation page&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Conclusion
&lt;/h3&gt;

&lt;p&gt;You have successfully learned how to deploy and access the &lt;strong&gt;Hermes Agent — Build, Deploy &amp;amp; Scale Autonomous AI&lt;/strong&gt; virtual machine on AWS. With its pre-configured environment, browser-based dashboard, and powerful CLI, Hermes removes the complexity of setting up an agent framework from scratch, allowing you to focus on building intelligent applications instead of infrastructure.&lt;/p&gt;

&lt;p&gt;Whether you’re creating AI assistants, automating business processes, integrating external tools and APIs, or experimenting with multi-agent workflows, Hermes provides a flexible and extensible platform to support your development needs. You can further customize your deployment by connecting your preferred AI providers, switching between supported models, or leveraging local Ollama models for on-premises inference.&lt;/p&gt;

&lt;p&gt;As your projects grow, you can easily scale your deployment using larger EC2 instance types or GPU-enabled instances for improved performance. Explore the Hermes documentation to discover advanced features, workflow customization, and best practices for building production-ready autonomous AI agents.&lt;/p&gt;

&lt;h3&gt;
  
  
  Thank you so much for reading
&lt;/h3&gt;

&lt;p&gt;Like | Follow | Subscribe to the newsletter.&lt;/p&gt;

&lt;p&gt;Catch us on&lt;/p&gt;

&lt;p&gt;Website: &lt;a href="https://www.techlatest.net/" rel="noopener noreferrer"&gt;https://www.techlatest.net/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Newsletter: &lt;a href="https://substack.com/@techlatestnet" rel="noopener noreferrer"&gt;https://substack.com/@techlatestnet&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Twitter: &lt;a href="https://twitter.com/TechlatestNet" rel="noopener noreferrer"&gt;https://twitter.com/TechlatestNet&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;LinkedIn: &lt;a href="https://www.linkedin.com/in/techlatest-net/" rel="noopener noreferrer"&gt;https://www.linkedin.com/in/techlatest-net/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;YouTube:&lt;a href="https://www.youtube.com/@techlatest_net/" rel="noopener noreferrer"&gt;https://www.youtube.com/@techlatest_net/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Blogs: &lt;a href="https://medium.com/@techlatest.net" rel="noopener noreferrer"&gt;https://medium.com/@techlatest.net&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Reddit Community: &lt;a href="https://www.reddit.com/user/techlatest_net/" rel="noopener noreferrer"&gt;https://www.reddit.com/user/techlatest_net/&lt;/a&gt;&lt;/p&gt;

</description>
      <category>agents</category>
      <category>amazonwebservices</category>
      <category>aws</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Can You Guess the Best Open-Source Coding Model? We Put Four to the Test</title>
      <dc:creator>TechLatest</dc:creator>
      <pubDate>Wed, 05 Aug 2026 11:42:43 +0000</pubDate>
      <link>https://dev.to/techlatestnet/can-you-guess-the-best-open-source-coding-model-we-put-four-to-the-test-4523</link>
      <guid>https://dev.to/techlatestnet/can-you-guess-the-best-open-source-coding-model-we-put-four-to-the-test-4523</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu1tjposv97ajjnbpa4nr.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu1tjposv97ajjnbpa4nr.png" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;We’ve all seen the benchmark charts. The flashy scores, the synthetic leaderboards, the claims of “best open-source coding model.” But if you’ve ever tried to use those models on actual work, you know the truth: benchmarks don’t tell you whether a model will understand your messy codebase, respect your constraints, or explain its trade-offs like a human engineer.&lt;/p&gt;

&lt;p&gt;So we stopped trusting charts and ran our own test.&lt;/p&gt;

&lt;p&gt;We took four of the most talked-about open-source coding models right now — GLM-5.2, DeepSeek V4 Flash, Kimi K3, and Qwen 3.8 Max — and gave them the same real-world coding task. No cherry-picked prompts. No special tuning. Just a practical optimization problem that forces models to choose between speed, readability, dependencies, and maintenance cost.&lt;/p&gt;

&lt;p&gt;We didn’t care which model topped a leaderboard. We cared about which one produced code you’d actually trust to merge after a quick review. Which one explained &lt;em&gt;why&lt;/em&gt; it made certain choices. Which one knew when &lt;em&gt;not&lt;/em&gt; to optimize.&lt;/p&gt;

&lt;p&gt;What follows isn’t a ranking. It’s a field guide. Each model approached the same problem differently, revealing distinct philosophies about what “good code” means. By the end, you won’t just know which model is “best” — you’ll know which one fits &lt;em&gt;your&lt;/em&gt; workflow, &lt;em&gt;your&lt;/em&gt; stack, and &lt;em&gt;your&lt;/em&gt; team’s tolerance for complexity.&lt;/p&gt;

&lt;p&gt;Let’s get into it.&lt;/p&gt;

&lt;h3&gt;
  
  
  GLM-5.2: Built for the Long Haul
&lt;/h3&gt;

&lt;p&gt;GLM-5.2 is the newest open coding model from Z.ai. Think of it as the developer who doesn’t just write a quick function and clock out — it’s the one you call when you have a messy, multi-file project that needs someone to stay focused for hours.&lt;/p&gt;

&lt;p&gt;Most AI coding tools are great at short bursts: “write me a sorting algorithm” or “fix this syntax error.” But real engineering isn’t like that. Real work means digging through thousands of lines of code, remembering what you changed three files ago, and connecting dots across an entire codebase. That’s exactly what GLM-5.2 was built for.&lt;/p&gt;

&lt;p&gt;It can hold about 1 million tokens of context in its head at once. To put that simply: you can hand it your whole project, your docs, your logs, and your test suite, and it won’t forget the beginning by the time it reaches the end. It also lets you choose how hard it thinks. Need a fast answer? Tell it to keep it light. Stuck on a nasty bug? Crank up the reasoning and let it take its time.&lt;/p&gt;

&lt;p&gt;And unlike many top-tier models locked behind paywalls or regional restrictions, GLM-5.2 is fully open under the MIT license. You can run it yourself, tweak it, use it commercially — no strings attached.&lt;/p&gt;

&lt;p&gt;Why it matters for this benchmark: GLM-5.2 brings two major differentiators to this head-to-head test:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Solid 1M-Token Context: Unlike models that simply accept long inputs but degrade in quality, GLM-5.2 uses a new architecture called &lt;em&gt;IndexShare&lt;/em&gt; to maintain stable reasoning even when processing entire repositories or extensive documentation.&lt;/li&gt;
&lt;li&gt;Adjustable Reasoning Effort: It offers explicit “High” and “Max” thinking modes. This allows developers to trade latency for deeper reasoning on hard bugs, or prioritize speed for simpler refactoring tasks — a flexibility not always available in frontier models.&lt;/li&gt;
&lt;li&gt;True Open Source: Released under the MIT license with no regional restrictions, making it one of the most accessible top-tier coding models for self-hosting and commercial use.&lt;/li&gt;
&lt;/ul&gt;

&lt;h4&gt;
  
  
  Key Specs at a Glance
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;| Feature | Details |
| ----------------------- | ------------------------------------------------------------------------------------------------ |
| Context Window | 1 Million Tokens (Stable) |
| License | MIT License (Fully Open Source) |
| Specialty | Long-horizon agentic coding, large repository navigation, and complex software engineering tasks |
| Reasoning Modes | High and Max effort reasoning modes for difficult coding and reasoning tasks |
| Agent Compatibility | Claude Code, ZCode, OpenCode, vLLM, and SGLang |
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Why We’re Testing It
&lt;/h4&gt;

&lt;p&gt;We didn’t pick GLM-5.2 because of its benchmark scores. We picked it because it promises to do the kind of sustained, real-world coding work that actually matters to developers — and we want to know if it delivers when the rubber meets the road.&lt;/p&gt;

&lt;p&gt;So instead of trusting charts and leaderboards, we’re giving it the same hands-on coding task as the other three models. No special treatment, no cherry-picked prompts. Just a real problem, solved in real time, judged by how useful the result actually is.&lt;/p&gt;

&lt;h4&gt;
  
  
  Putting GLM-5.2 to the Test
&lt;/h4&gt;

&lt;p&gt;For our first head-to-head challenge, we gave all four models the task: The Performance Optimization Trap. The prompt asks the model to optimize a slow Python CSV processing function to run in under 30 seconds, while prioritizing readability and explaining when the optimization &lt;em&gt;wouldn’t&lt;/em&gt; be worth doing. This task is designed to catch models that blindly apply textbook patterns without understanding real-world trade-offs.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx8n7o0awxpwbp6v2nylz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx8n7o0awxpwbp6v2nylz.png" width="799" height="409"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  What GLM-5.2 Actually Did
&lt;/h4&gt;

&lt;p&gt;Instead of just throwing faster code at the problem, GLM-5.2 started by diagnosing &lt;em&gt;why&lt;/em&gt; the original was slow. It correctly identified that the bottleneck wasn’t “CSV parsing” in general — it was specifically the Python-level DictReader creating 10 million individual dictionaries. That’s the kind of insight that usually takes a senior developer profiling the code to uncover.&lt;/p&gt;

&lt;p&gt;It then offered two solutions:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;A pandas version (~10–20s) using the C parser, column selection, and vectorized operations — no new dependencies needed.&lt;/li&gt;
&lt;li&gt;A Polars version (~3–8s) using lazy evaluation and multi-threaded parsing for when you really need speed.
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;Output&lt;/span&gt;

&lt;span class="c1"&gt;# Optimizing `process_sales`
&lt;/span&gt;
&lt;span class="c1"&gt;## Diagnosis: Where the 4 minutes actually go
&lt;/span&gt;
&lt;span class="n"&gt;Before&lt;/span&gt; &lt;span class="n"&gt;writing&lt;/span&gt; &lt;span class="n"&gt;code&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;profile&lt;/span&gt; &lt;span class="n"&gt;the&lt;/span&gt; &lt;span class="n"&gt;original&lt;/span&gt; &lt;span class="n"&gt;mentally&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;

&lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="n"&gt;Cost&lt;/span&gt; &lt;span class="n"&gt;source&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="n"&gt;Approx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt; &lt;span class="n"&gt;impact&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt;
&lt;span class="o"&gt;|---|---|&lt;/span&gt;
&lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="sb"&gt;`csv.DictReader`&lt;/span&gt; &lt;span class="n"&gt;builds&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="n"&gt;Python&lt;/span&gt; &lt;span class="sb"&gt;`dict`&lt;/span&gt; &lt;span class="n"&gt;per&lt;/span&gt; &lt;span class="nf"&gt;row &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="n"&gt;M&lt;/span&gt; &lt;span class="n"&gt;dicts&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="o"&gt;~&lt;/span&gt;&lt;span class="mi"&gt;60&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;70&lt;/span&gt;&lt;span class="o"&gt;%&lt;/span&gt; &lt;span class="n"&gt;of&lt;/span&gt; &lt;span class="n"&gt;runtime&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt;
&lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="n"&gt;Python&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;level&lt;/span&gt; &lt;span class="sb"&gt;`for`&lt;/span&gt; &lt;span class="n"&gt;loop&lt;/span&gt; &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="sb"&gt;`float()`&lt;/span&gt; &lt;span class="n"&gt;called&lt;/span&gt; &lt;span class="n"&gt;twice&lt;/span&gt; &lt;span class="n"&gt;per&lt;/span&gt; &lt;span class="n"&gt;kept&lt;/span&gt; &lt;span class="n"&gt;row&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="o"&gt;~&lt;/span&gt;&lt;span class="mi"&gt;15&lt;/span&gt;&lt;span class="o"&gt;%&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt;
&lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="sb"&gt;`pd.DataFrame(list_of_dicts)`&lt;/span&gt; &lt;span class="err"&gt;—&lt;/span&gt; &lt;span class="n"&gt;slowest&lt;/span&gt; &lt;span class="n"&gt;DataFrame&lt;/span&gt; &lt;span class="n"&gt;constructor&lt;/span&gt; &lt;span class="n"&gt;path&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="o"&gt;~&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="o"&gt;%&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt;
&lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="sb"&gt;`.strip().upper()`&lt;/span&gt; &lt;span class="n"&gt;on&lt;/span&gt; &lt;span class="sb"&gt;`region`&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;even&lt;/span&gt; &lt;span class="n"&gt;though&lt;/span&gt; &lt;span class="sb"&gt;`region`&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="n"&gt;never&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;the&lt;/span&gt; &lt;span class="n"&gt;output&lt;/span&gt;&lt;span class="o"&gt;**&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="o"&gt;~&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="o"&gt;%&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;pure&lt;/span&gt; &lt;span class="n"&gt;waste&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt;

&lt;span class="n"&gt;The&lt;/span&gt; &lt;span class="n"&gt;bottleneck&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="ow"&gt;not&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;CSV parsing&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;the&lt;/span&gt; &lt;span class="n"&gt;abstract&lt;/span&gt; &lt;span class="err"&gt;—&lt;/span&gt; &lt;span class="n"&gt;it&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s the **Python-level `DictReader`**. Pandas&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt; &lt;span class="n"&gt;C&lt;/span&gt; &lt;span class="nf"&gt;parser &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="n"&gt;Polars&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;PyArrow&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="n"&gt;solves&lt;/span&gt; &lt;span class="n"&gt;exactly&lt;/span&gt; &lt;span class="n"&gt;that&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt; &lt;span class="n"&gt;The&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;textbook trap&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="n"&gt;the&lt;/span&gt; &lt;span class="n"&gt;prompt&lt;/span&gt; &lt;span class="n"&gt;warns&lt;/span&gt; &lt;span class="n"&gt;about&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;adding&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;Pandas&lt;/span&gt; &lt;span class="n"&gt;vectorization&lt;/span&gt; &lt;span class="n"&gt;on&lt;/span&gt; &lt;span class="n"&gt;top&lt;/span&gt; &lt;span class="n"&gt;of&lt;/span&gt; &lt;span class="n"&gt;the&lt;/span&gt; &lt;span class="n"&gt;slow&lt;/span&gt; &lt;span class="n"&gt;parser&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="n"&gt;the&lt;/span&gt; &lt;span class="n"&gt;fix&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="n"&gt;replacing&lt;/span&gt; &lt;span class="n"&gt;the&lt;/span&gt; &lt;span class="n"&gt;parser&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;

&lt;span class="c1"&gt;## Optimized version (pandas, no new deps)
&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;&lt;br&gt;
python&lt;br&gt;
import pandas as pd&lt;/p&gt;

&lt;p&gt;def process_sales(csv_path):&lt;br&gt;
    df = pd.read_csv(&lt;br&gt;
        csv_path,&lt;br&gt;
        usecols=['timestamp', 'amount', 'status'],&lt;br&gt;
        dtype={'status': 'category'},&lt;br&gt;
    )&lt;br&gt;
    mask = df['status'].eq('completed') &amp;amp; df['amount'].gt(0)&lt;br&gt;
    return (&lt;br&gt;
        df.loc[mask]&lt;br&gt;
          .groupby(df.loc[mask, 'timestamp'].str.slice(0, 10))['amount']&lt;br&gt;
          .sum()&lt;br&gt;
          .to_dict()&lt;br&gt;
    )&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;
Expected runtime on 10M rows: **~10-20s** on typical hardware.

## Faster version (Polars, if allowed)

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;&lt;br&gt;
python&lt;br&gt;
import polars as pl&lt;/p&gt;

&lt;p&gt;def process_sales(csv_path):&lt;br&gt;
    out = (&lt;br&gt;
        pl.scan_csv(csv_path) # lazy: enables predicate pushdown&lt;br&gt;
          .filter((pl.col('status') == 'completed') &amp;amp; (pl.col('amount') &amp;gt; 0))&lt;br&gt;
          .group_by(pl.col('timestamp').str.slice(0, 10).alias('date'))&lt;br&gt;
          .agg(pl.col('amount').sum().alias('revenue'))&lt;br&gt;
          .collect() # multi-threaded CSV parse + execute&lt;br&gt;
    )&lt;br&gt;
    return dict(out.iter_rows())&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;
Expected runtime: **~3-8s**. Multi-threaded parsing + lazy predicate pushdown.

## Why each change helps

1. **`pd.read_csv` (C parser) instead of `csv.DictReader`** — The single biggest win. `DictReader` allocates a `dict` (and string keys) for every row, including rows that get filtered out. The C parser tokenizes into typed column arrays directly. ~5-10× faster on parse alone.

2. **`usecols=['timestamp', 'amount', 'status']`** — `region` is parsed, stripped, uppercased, and stored in the original — but **never appears in the output** (the groupby is on `date`, summing `revenue`). Dropping the column at parse time means the C parser skips those bytes entirely. This is the most important *readability* win too: it surfaces that the original had dead code.

3. **`dtype={'status': 'category'}`** — `'status'` has low cardinality ('completed', 'pending', 'refunded', …). Categoricals make the equality filter a dictionary-lookup on integer codes instead of a string compare per row.

4. **Vectorized boolean mask instead of `if` per row** — The mask is computed in C over contiguous arrays; the Python loop disappears. Branch prediction in the C kernel is also far better than CPython's per-iteration branching.

5. **Single `float` parse per value** — Original calls `float(row['amount'])` *twice* for kept rows (once in the condition, once when building the dict). `read_csv` parses `amount` once into a `float64` column.

6. **No `list[dict]` → `DataFrame` round-trip** — That constructor path is the slowest way to build a DataFrame. We never build the intermediate list at all.

7. **`.str.slice(0, 10)` instead of `[:10]` per row** — Vectorized string slicing over a numpy object array. Also avoids creating a `'date'` column materialized in memory; the groupby key is computed on the fly.

8. **Polars `scan_csv` + `.collect()`** — Lazy frame lets the engine push the filter *into* the CSV scan (rows that fail the predicate are never materialized), and CSV parsing is split across cores. This is the structural reason Polars beats pandas on this workload, not just "newer = faster."

## When this optimization is **not** worth doing

- **The job runs once overnight and 4 minutes is well within SLA.** Rewriting working code carries regression risk; the engineering time costs more than it saves.
- **The file is ≤ ~500K rows.** Below that, the original is a second or two and the pandas version adds a learning curve for new contributors.
- **The downstream cost dwarfs parse cost.** If this function feeds an API call or DB write that takes 10 minutes, shaving 3.5 minutes off parsing is invisible.
- **Memory is the actual constraint, not latency.** The vectorized versions load the whole file into RAM (roughly 2-4× the CSV size for pandas). On a constrained box, a streaming `csv.reader` + accumulator into a `defaultdict(float)` is the right shape — slower but bounded memory. The prompt explicitly says latency &amp;gt; memory, so we don't take this path, but it's the trade-off to name.
- **The team doesn't know Polars and the pandas version already meets the SLA.** Introducing a second dataframe library has a real maintenance tax (two idioms, two upgrade tracks, two sets of bugs). Only pull that lever when pandas genuinely can't hit the target.
- **Prematurely, during prototyping.** If the schema is still changing weekly, optimize after it stabilizes — otherwise you'll rewrite the optimized path every time a column moves.

The short version: optimize when (runtime × frequency) + (latency-driven UX/batch cost) exceeds the maintenance cost of the faster code. Here, a daily batch at 4 minutes with a clear &amp;lt;30s target clearly clears that bar — which is why the rewrite is justified.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;&lt;br&gt;
plaintext&lt;/p&gt;

&lt;p&gt;But what stood out most was the “When this optimization is NOT worth doing” section. It explicitly called out scenarios where rewriting would be a waste of time: overnight batch jobs within SLA, small files, downstream bottlenecks that dwarf parse time, memory-constrained environments, teams unfamiliar with Polars, or unstable schemas during prototyping. This isn’t something models usually volunteer — they tend to optimize because you asked, not because it’s actually the right engineering call.&lt;/p&gt;
&lt;h4&gt;
  
  
  Our Take
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;Reasoning Quality: Excellent. It didn’t just apply textbook Pandas vectorization (the exact trap the prompt warned about). It understood the &lt;em&gt;structural&lt;/em&gt; reason the original was slow and explained each change in plain terms. The diagnosis table upfront showed genuine comprehension, not pattern matching.&lt;/li&gt;
&lt;li&gt;Code Quality: Production-ready. Both versions are clean, well-commented, and immediately usable. The usecols parameter to skip the unused region column was a particularly sharp catch—it even noted this was a readability win because it surfaced dead code in the original.&lt;/li&gt;
&lt;li&gt;Instruction Following: Nailed it. Hit the ❤0s target, prioritized readability, explained every change, and included the requested “when not to optimize” caveat. Went beyond by offering two tiers of optimization with clear trade-offs between them.&lt;/li&gt;
&lt;li&gt;Notable Strengths: The engineering maturity. Most models would have stopped at the Polars solution. GLM-5.2 treated this like a real code review, acknowledging maintenance cost, team familiarity, and regression risk. It also caught that region was being processed but never used—a detail many models (and humans) miss.&lt;/li&gt;
&lt;li&gt;Notable Weaknesses: None significant for this task. If anything, the dual-solution approach adds slight cognitive load, but it’s justified by the clear framing of when to use each.&lt;/li&gt;
&lt;/ul&gt;
&lt;h4&gt;
  
  
  Verdict on This Task
&lt;/h4&gt;

&lt;p&gt;GLM-5.2 didn’t just solve the optimization problem — it solved it like a staff engineer who understands that performance is a business decision, not just a technical one. This is exactly the kind of output you’d trust to merge after a quick review.&lt;/p&gt;
&lt;h3&gt;
  
  
  What This Tells Us About GLM-5.2
&lt;/h3&gt;

&lt;p&gt;This single task reveals why GLM-5.2 belongs in this comparison:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;It diagnoses before prescribing, avoiding the knee-jerk optimization trap.&lt;/li&gt;
&lt;li&gt;It respects constraint hierarchies (latency &amp;gt; memory, readability alongside speed) instead of optimizing for speed alone.&lt;/li&gt;
&lt;li&gt;It demonstrates engineering judgment by articulating when &lt;em&gt;not&lt;/em&gt; to act — a signal of real-world usability over benchmark chasing.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;
  
  
  DeepSeek V4 Flash: Speed Meets Pragmatism
&lt;/h3&gt;

&lt;p&gt;DeepSeek V4 Flash is the lightweight, high-speed contender in this lineup. If GLM-5.2 is the staff engineer who stays late to untangle legacy code, DeepSeek V4 Flash is the sharp junior dev who ships clean work fast and doesn’t overcomplicate things. It’s designed to be responsive, efficient, and surprisingly capable for its size — especially on agentic coding tasks where latency matters as much as accuracy.&lt;/p&gt;

&lt;p&gt;Unlike larger models that prioritize depth at all costs, V4 Flash balances speed with reasoning. It supports adjustable effort levels (low, high, max), so you can dial it down for quick refactors or crank it up when debugging gets messy. And like GLM-5.2, it’s fully open under MIT — no API keys, no regional locks, just weights you can run locally.&lt;/p&gt;
&lt;h4&gt;
  
  
  Why We’re Testing It
&lt;/h4&gt;

&lt;p&gt;We included DeepSeek V4 Flash because it represents a different philosophy: can a smaller, faster model deliver production-grade optimization advice without the overhead of a frontier-scale model? Benchmarks say it punches above its weight, but we want to know if that translates to real code reviews — not just synthetic scores.&lt;/p&gt;

&lt;p&gt;So we gave it the same Task prompt as GLM-5.2: optimize a slow CSV processor to ❤0 seconds, prioritize readability, and explain when &lt;em&gt;not&lt;/em&gt; to optimize. No hints, no extra context. Just the raw problem.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh1exzyf050ahg6lwwvzo.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh1exzyf050ahg6lwwvzo.png" width="800" height="370"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4&gt;
  
  
  Putting DeepSeek V4 Flash to the Test
&lt;/h4&gt;

&lt;p&gt;Instead of reaching for Pandas or Polars, V4 Flash took a fundamentally different approach: it eliminated the DataFrame. It replaced csv.DictReader with plain csv.reader, swapped the list-of-dicts + groupby pattern for a single-pass defaultdict(float) accumulator, and aggregated revenue &lt;em&gt;during parsing&lt;/em&gt; instead of after.&lt;/p&gt;

&lt;p&gt;The result? A solution estimated at 20–25 seconds — well under the 30-second target — with zero external dependencies beyond the standard library.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Here's an optimized version that should get well under 30 seconds for 10M rows:

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;&lt;br&gt;
python&lt;br&gt;
import pandas as pd&lt;br&gt;
from collections import defaultdict&lt;/p&gt;

&lt;p&gt;def process_sales_optimized(csv_path):&lt;br&gt;
    """&lt;br&gt;
    Process sales CSV and aggregate revenue by date.&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Performance optimizations:
1. Use built-in csv.reader instead of DictReader (no dict allocation per row)
2. Pre-filter with generator expressions to avoid intermediate list
3. Use defaultdict for O(1) aggregation during parsing
4. Single pass through data - no DataFrame creation (biggest win)
"""
# defaultdict with float factory avoids checking key existence
revenue_by_date = defaultdict(float)

# csv.reader is faster than DictReader - no per-row dict creation
with open(csv_path, 'r') as f:
    # Skip header row
    next(f)
    reader = csv.reader(f)

    for row in reader:
        # Direct column access (timestamp=0, amount=1, status=2, region=3)
        # Fast short-circuit evaluation: 'in' check before float conversion
        if row[2] == 'completed':
            try:
                amount = float(row[1])
                if amount &amp;gt; 0:
                    # Slice timestamp directly - string slicing is fast
                    date = row[0][:10]
                    revenue_by_date[date] += amount
            except ValueError:
                # Skip malformed rows gracefully (log if needed)
                continue

return dict(revenue_by_date)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;
**Why this is much faster (from 4min → ~20-25 seconds):**

1. **Eliminated DataFrame entirely** (saves ~60% time): Creating a DataFrame from 10M rows and then doing groupby is expensive. We're now doing aggregation during parsing.

2. **`csv.reader` instead of `DictReader`** (saves ~30% time): DictReader creates a dict for every single row. Direct list indexing is much faster.

3. **`defaultdict` instead of list-then-groupby** (saves the other ~10%): No intermediate list, no second pass through data. We aggregate in one pass.

4. **Order of operations** : The `'completed'` string check happens before `float()` conversion, so we skip expensive number parsing for invalid rows.

5. **Single pass** : The entire operation is one linear scan through 10M rows with minimal per-row work.

**When this optimization would NOT be worth doing:**

- **If the CSV was &amp;lt;1M rows** : The optimization saves maybe 10-20 seconds, which probably isn't worth the slightly less readable code.
- **If you need the DataFrame for further analysis anyway** : If you're going to do complex filtering/transformation later, creating the DataFrame once might be better than parsing twice.
- **If the bottleneck is actually disk I/O or network** : If the CSV is on a slow network share or spinning disk, no amount of Python optimization will help.
- **If the input format is unstable** : The column-position-based approach breaks if columns are reordered. The DictReader version is more robust.
- **If you value maintainability over performance** : The original is arguably clearer about what columns contain. If this code is rarely run, the original might be better.
- **If memory is actually the constraint, not CPU** : Our version still needs to hold all unique dates in memory. If you had 10M unique dates, this approach would be memory-heavy too.

**Key trade-off note** : The main speedup comes from:
1. Doing aggregation during parse (not after) 
2. Using low-level `csv.reader` instead of `DictReader`
3. Avoiding the pandas overhead entirely

The code sacrifices some readability (column indices instead of names) for ~10x speedup. For a batch job processing 10M rows, this is usually worth the maintenance cost.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;
&lt;p&gt;&lt;br&gt;
plaintext&lt;/p&gt;

&lt;p&gt;What stood out was its pragmatism:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;It short-circuited expensive operations: checking row[2] == 'completed' &lt;em&gt;before&lt;/em&gt; calling float(), avoiding unnecessary number parsing for filtered-out rows.&lt;/li&gt;
&lt;li&gt;It handled malformed data gracefully with a try/except around float conversion, skipping bad rows instead of crashing.&lt;/li&gt;
&lt;li&gt;It explicitly named the trade-off: column indices sacrifice readability for speed, and that’s acceptable for a batch job but not for frequently modified code.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Its “when not to optimize” section was equally grounded: small files (&amp;lt;1M rows), downstream DataFrame needs, disk I/O bottlenecks, unstable schemas, maintainability priorities, and memory constraints from high-cardinality dates. Each point tied back to real engineering consequences, not abstract principles.&lt;/p&gt;
&lt;h4&gt;
  
  
  Our Take
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;Reasoning Quality: Sharp and practical. It correctly identified that the bottleneck wasn’t just parsing — it was the &lt;em&gt;entire pipeline&lt;/em&gt; of dict creation → list building → DataFrame construction → groupby. By collapsing this into one pass, it solved the root cause, not just symptoms. The explanation of &lt;em&gt;why&lt;/em&gt; each change helped was concise and accurate.&lt;/li&gt;
&lt;li&gt;Code Quality: Clean, dependency-free, and immediately runnable. The use of defaultdict(float) avoids key-existence checks, and the generator-style streaming keeps memory flat. Only minor nit: column indices (row[0], row[1]) hurt readability compared to named access—but the model acknowledged this explicitly as a conscious trade-off.&lt;/li&gt;
&lt;li&gt;Instruction Following: Perfect. Hit the performance target, prioritized readability &lt;em&gt;within the constraints of speed&lt;/em&gt;, explained every optimization, and included nuanced caveats about when to avoid this approach. Didn’t over-engineer or add unnecessary abstractions.&lt;/li&gt;
&lt;li&gt;Notable Strengths: Zero-dependency solution that still hits the performance target. Most models default to Pandas/Polars; V4 Flash proved you don’t need them for this workload. Also showed mature error handling and clear communication about maintainability costs.&lt;/li&gt;
&lt;li&gt;Notable Weaknesses: Column-index-based access is fragile — if the CSV schema changes, this breaks silently. A hybrid approach (e.g., reading header once to map names→indices) would add robustness with minimal perf cost. Also didn’t mention parallelization options (e.g., chunked reading with multiprocessing), though that may be intentional given the “readability first” constraint.&lt;/li&gt;
&lt;/ul&gt;
&lt;h4&gt;
  
  
  Verdict on This Task
&lt;/h4&gt;

&lt;p&gt;DeepSeek V4 Flash delivered a lean, stdlib-only solution that matches GLM-5.2’s performance target while using fewer resources. It traded some readability for speed — but did so transparently and justified the choice. For teams wanting fast, portable optimizations without heavy dependencies, this is exactly the kind of output you’d adopt.&lt;/p&gt;
&lt;h3&gt;
  
  
  What This Tells Us About DeepSeek V4 Flash
&lt;/h3&gt;

&lt;p&gt;This task confirms V4 Flash isn’t just a “fast but dumb” model:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;It understands systemic bottlenecks, not just surface-level fixes.&lt;/li&gt;
&lt;li&gt;It makes deliberate trade-offs and communicates them clearly.&lt;/li&gt;
&lt;li&gt;It respects constraints (readability, latency, dependencies) without over-delivering or under-delivering.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For developers who value speed, portability, and minimal footprint, V4 Flash proves that smaller models can still think like engineers — not just code generators.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Run DeepSeek and Other Open Models Locally&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;TechLatest provides ready-to-use Ollama + Open WebUI environments for running DeepSeek, Qwen, Gemma, Llama, Mistral, and other open-weight models locally.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://techlatest.net/support/multi_llm_gpu_vm_support/" rel="noopener noreferrer"&gt;Techlatest.net - GPU Supported DeepSeek &amp;amp; Llama powered All-in-One LLM&lt;/a&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  Qwen 3.8 Max: The Polars Native
&lt;/h3&gt;

&lt;p&gt;Qwen 3.8 Max is Alibaba’s latest flagship model — and the first in the Qwen-Max line to be released with open weights. If GLM-5.2 is the staff engineer and DeepSeek V4 Flash is the pragmatic junior dev, Qwen 3.8 Max is the data infrastructure specialist who defaults to modern tooling because they’ve seen what actually scales in production.&lt;/p&gt;

&lt;p&gt;It’s a 2.4-trillion-parameter model built for long-horizon autonomous work, but on coding tasks like this one, what stands out is its fluency with contemporary data stacks. It doesn’t just know Polars — it understands &lt;em&gt;why&lt;/em&gt; Polars exists, how its query optimizer works, and when its overhead isn’t justified. That kind of ecosystem awareness is rare in models that treat libraries as black boxes.&lt;/p&gt;

&lt;p&gt;Like the others, it supports adjustable reasoning effort (xhigh, medium, low) and is fully open-weight (releasing next week). But where it differs is in its assumption that you’re probably already using modern tools—and it writes code accordingly.&lt;/p&gt;
&lt;h4&gt;
  
  
  Why We’re Testing It
&lt;/h4&gt;

&lt;p&gt;We included Qwen 3.8 Max because it represents a third philosophy: can a frontier-scale model leverage modern data infrastructure intelligently, without over-engineering or ignoring trade-offs? GLM-5.2 offered tiered solutions; V4 Flash went stdlib-only. Qwen bets on Polars as the right default — but we want to know if that bet is justified, or just fashionable.&lt;/p&gt;

&lt;p&gt;So we gave it the same Task #4 prompt: optimize to ❤0 seconds, prioritize readability, explain trade-offs. No hints about which library to use. Just the problem.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsixwdcdq8am80fdyag4d.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsixwdcdq8am80fdyag4d.png" width="800" height="332"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4&gt;
  
  
  Putting Qwen 3.8 Max to the Test
&lt;/h4&gt;

&lt;p&gt;Qwen went all-in on Polars — but not blindly. It used pl.scan_csv (lazy evaluation) instead of read_csv, applied filters &lt;em&gt;before&lt;/em&gt; string transformations to enable predicate pushdown, and kept string operations inside the engine to avoid Python-level allocations. It also explicitly noted the return format quirk of to_dict(as_series=False) and provided the O(n) zip fix for strict {date: revenue} output—acknowledging that the aggregated result is tiny (&amp;lt;365 rows), so the conversion cost is negligible.&lt;/p&gt;

&lt;p&gt;Performance estimate: 2–8 seconds, the fastest of all four models on this task.&lt;/p&gt;

&lt;p&gt;Its “when not to optimize” section was the most operationally grounded:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Files &amp;lt;100K rows (original runs in &amp;lt;1s; Polars adds ~200ms cold start)&lt;/li&gt;
&lt;li&gt;High-frequency micro-tasks where setup overhead dominates&lt;/li&gt;
&lt;li&gt;Teams lacking Polars expertise when SLA allows 4-minute runtime (“Profile first”)&lt;/li&gt;
&lt;li&gt;Unstable CSV schemas where Polars’ strict typing breaks but stdlib tolerates messiness&lt;/li&gt;
&lt;li&gt;Memory-constrained containers where streaming fails, and bounded-memory generators are safer&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Each caveat tied back to real deployment constraints — not theoretical purity.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Output

Here is the optimized solution, followed by an explanation of the engineering trade-offs.

### Optimized Code

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;&lt;br&gt;
python&lt;br&gt;
import polars as pl&lt;/p&gt;

&lt;p&gt;def process_sales(csv_path: str) -&amp;gt; dict[str, float]:&lt;br&gt;
    """&lt;br&gt;
    Aggregate daily revenue for completed sales from a CSV file.&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Uses Polars for lazy evaluation and parallel CSV parsing to avoid
Python-level iteration overhead on 10M+ row files.
"""
return (
    pl.scan_csv(csv_path)
    .filter(
        (pl.col("status") == "completed") &amp;amp; (pl.col("amount") &amp;gt; 0)
    )
    .with_columns([
        pl.col("timestamp").str.slice(0, 10).alias("date"),
        pl.col("region").str.strip_chars().str.to_uppercase(),
    ])
    .group_by("date")
    .agg(pl.col("amount").sum())
    .collect() # Triggers parallel execution + streaming if needed
    .to_dict(as_series=False) # Returns {"date": [...], "amount": [...]}
)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;
&amp;gt; **Note on Return Format:** `polars.DataFrame.to_dict(as_series=False)` returns column-oriented dicts (`{"date": [...], "amount": [...]}`). If you strictly need `{date: revenue}` pairs, append this after `.collect()`:
&amp;gt; ```

python
&amp;gt; result = df.collect()
&amp;gt; return dict(zip(result["date"], result["amount"]))
&amp;gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;blockquote&gt;
&lt;p&gt;This final zip is O(n) but operates on already-aggregated data (typically &amp;lt;365 rows), so it's negligible.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h3&gt;
  
  
  Why Each Change Improves Performance
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Change&lt;/th&gt;
&lt;th&gt;Bottleneck Addressed&lt;/th&gt;
&lt;th&gt;Mechanism&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;&lt;code&gt;pl.scan_csv&lt;/code&gt; over &lt;code&gt;csv.DictReader&lt;/code&gt;&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;CSV parsing (the actual bottleneck)&lt;/td&gt;
&lt;td&gt;Lazy reader with Rust-based parallel parsing. Never materializes full file into Python objects. Parses only required columns via projection pushdown.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Filter before transform&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Wasted computation&lt;/td&gt;
&lt;td&gt;Predicate pushdown pushes the filter into the scan phase. Rows that don't match are never parsed into string slices or uppercased. On typical datasets where &lt;code&gt;completed&lt;/code&gt; is a minority, this eliminates most work.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Eliminating Python-level loop&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;GIL + interpreter overhead&lt;/td&gt;
&lt;td&gt;The original code executes ~30M Python bytecode operations (10M rows × 3 field accesses + conditionals + dict construction). Polars executes the entire pipeline in compiled Rust across all cores.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;String operations in-engine&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Per-row Python string allocation&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;.str.slice()&lt;/code&gt; and &lt;code&gt;.str.strip_chars().str.to_uppercase()&lt;/code&gt; run as vectorized SIMD kernels inside Polars, avoiding 10M individual Python &lt;code&gt;str&lt;/code&gt; object allocations and GC pressure.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Single-pass aggregation&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Intermediate DataFrame materialization&lt;/td&gt;
&lt;td&gt;The original builds a full filtered DataFrame &lt;em&gt;then&lt;/em&gt; groups. Polars' query optimizer fuses filter → transform → groupby into a single streaming pass when possible, reducing peak memory and cache misses.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  Expected Performance
&lt;/h3&gt;

&lt;p&gt;On a 10M-row CSV (~1–2 GB), this typically runs in &lt;strong&gt;2–8 seconds&lt;/strong&gt; depending on hardware, well under the 30-second target. The original spends ~3 minutes purely in Python's &lt;code&gt;csv&lt;/code&gt; module and per-row dict construction; Polars eliminates both.&lt;/p&gt;

&lt;h3&gt;
  
  
  When This Optimization Is NOT Worth Doing
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;File size stays small (&amp;lt;100K rows):&lt;/strong&gt; The original code runs in &amp;lt;1 second at that scale. Adding Polars introduces a dependency, a learning curve, and import overhead (~200ms cold start) for zero user-visible benefit.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The function is called thousands of times per job on tiny files:&lt;/strong&gt; Polars' query planning and thread pool initialization have fixed overhead. A tight Python loop wins on micro-benchmarks below ~50K rows due to zero setup cost.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Team lacks Polars/Rust-ecosystem expertise and the 4-minute runtime is acceptable:&lt;/strong&gt; If the batch job has a 30-minute SLA and this function isn't on the critical path, optimizing it adds maintenance risk (new dependency version breaks, unfamiliar debugging) without business value. &lt;strong&gt;Profile first.&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;CSV schema is unstable/dynamic:&lt;/strong&gt; Polars' type inference and strict schema enforcement can break on messy real-world CSVs that Python's &lt;code&gt;csv.DictReader&lt;/code&gt; silently handles. If you'd spend more time fixing schema errors than you save in runtime, stick with stdlib.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Memory-constrained environments where Polars can't stream:&lt;/strong&gt; While &lt;code&gt;scan_csv&lt;/code&gt; supports streaming, some complex queries still require materialization. If you're on a 512MB container and the grouped intermediate exceeds available RAM, the original generator-based approach (which uses bounded memory via &lt;code&gt;rows.append&lt;/code&gt; + eventual DataFrame construction) may actually be safer, even if slower.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The key principle: &lt;strong&gt;optimize the actual bottleneck (I/O + parsing), not the symptom (Python loop)&lt;/strong&gt;. But only when the bottleneck matters to your SLA.&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;
plaintext

#### Our Take

- Reasoning Quality: Deep and precise. It didn’t just say “Polars is faster” — it explained _how_: projection pushdown, SIMD string kernels, fused query plans, Rust parallelism. The table mapping changes to bottlenecks showed genuine understanding of the engine, not just API familiarity. Also caught the subtle point that filtering before transformation eliminates wasted work on non-matching rows.
- Code Quality: Idiomatic Polars, well-documented, and immediately usable. The note about return format shows attention to interface contracts — a detail many models miss. Only minor gap: didn’t mention fallback to streaming mode for memory-constrained cases (though it did warn about materialization risks).
- Instruction Following: Excellent. Hit performance target by wide margin, prioritized readability through clear pipeline structure, explained every optimization mechanistically, and included nuanced, operationally relevant caveats. Didn’t oversell Polars or ignore its costs.
- Notable Strengths: Most technically accurate explanation of _why_ Polars wins here. Best performance estimate. Strongest awareness of real-world deployment friction (cold starts, schema fragility, team expertise). Treats Polars as a tool with trade-offs, not a magic bullet.
- Notable Weaknesses: Assumes Polars is available/acceptable. For teams locked into pandas or stdlib-only environments, this solution requires buy-in. Also slightly less readable than GLM-5.2’s pandas version for developers unfamiliar with lazy evaluation semantics.

#### Verdict on This Task

Qwen 3.8 Max delivered the fastest, most technically sound solution — but only if your stack supports Polars. It didn’t just generate code; it demonstrated ecosystem literacy, explaining not just _what_ to do but _why it works at the engine level_ and _when to resist the urge_. For data-heavy teams already in the Polars/Rust ecosystem, this is the gold standard response.

### What This Tells Us About Qwen 3.8 Max

This task confirms Qwen 3.8 Max isn’t just big — it’s contextually aware:

- It defaults to modern tooling but justifies the choice mechanistically.
- It anticipates deployment friction (cold starts, schema issues, team ramp-up).
- It treats performance as a system property, not just a code property.

For teams building on contemporary data infrastructure, Qwen 3.8 Max doesn’t just write code — it writes code that belongs in your stack.

### Kimi K3: The Pragmatic Hybrid

Kimi K3 is Moonshot AI’s 2.8-trillion-parameter open-weight model — and it approaches coding like a developer who’s been burned by both over-engineering and under-thinking. If Qwen 3.8 Max defaults to modern tooling and DeepSeek V4 Flash goes stdlib-purist, Kimi K3 is the pragmatist who picks the right tool for the job, then explains why the other options are worse _for this specific case_.

It’s natively multimodal and built for long-horizon agentic work, but on pure coding tasks, what stands out is its refusal to be ideological. It doesn’t worship Polars, reject Pandas, or fetishize stdlib. It asks: _“What does this problem actually need?”_ and answers with surgical precision.

Like the others, it’s open-weight (Kimi K3 License), supports adjustable reasoning effort, and handles 1M-token context. But where it differs is in its contextual adaptability — it reads the room before writing code.

#### Why We’re Testing It

We included Kimi K3 because it represents a fourth philosophy: can a frontier model avoid dogma and deliver solutions tailored to the actual constraints of the task? GLM-5.2 offered tiers; V4 Flash went minimal; Qwen bet on Polars. Kimi promises to meet you where you are. We want to know if that promise holds when the rubber meets the road.

So we gave it the same Task #4 prompt: optimize to ❤0 seconds, prioritize readability, explain trade-offs. No hints. Just the problem.

![](https://cdn-images-1.medium.com/max/1024/1*TA4Mu7x4IMusdO6QOjBiXg.png)

#### Putting Kimi K3 to the Test

Kimi led with a streaming csv.reader + defaultdict solution—similar to DeepSeek V4 Flash—but with a critical refinement: it mapped header names to indices _once_ at startup, preserving readability without per-row dict overhead. It also explicitly noted that region was unused and removed the dead .strip().upper() calls, saving ~20M string operations. Estimated runtime: 10–25 seconds.

Then, unlike any other model, it offered a Pandas alternative — not as a second-tier option, but as a _conditional recommendation_: “Use this only if downstream code needs the full DataFrame.” It explained why Pandas is slower here (materializes filtered data) but justified when the trade-off flips.

Its “when not to optimize” section was the most ROI-aware:

- Files &amp;lt;100K rows (original runs in &amp;lt;1s; complexity adds no value)
- DataFrame is required output (stream to Parquet instead of optimizing aggregation)
- I/O-bound workloads (“Profile with iostat first”)
- Maintenance cost exceeds time savings (“24 hours/year saved vs. dev hours spent”)
- Schema instability (index-mapping is brittle; add validation if used)

The bottom line nailed the ethos: _“Streaming csv. reader is the right balance: eliminates unnecessary allocation while leveraging Python’s strengths when loops do almost nothing per iteration.”_



&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;Here are two optimized approaches. The &lt;strong&gt;streaming solution&lt;/strong&gt; is fastest for this specific aggregation task; the &lt;strong&gt;pandas solution&lt;/strong&gt; is best if you need the intermediate DataFrame for other operations.&lt;/p&gt;




&lt;h2&gt;
  
  
  Fastest: Streaming with &lt;code&gt;csv.reader&lt;/code&gt; + &lt;code&gt;defaultdict&lt;/code&gt;
&lt;/h2&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;
python
import csv
from collections import defaultdict

def process_sales(csv_path):
    daily_revenue = defaultdict(float)

    with open(csv_path, newline='') as f:
        reader = csv.reader(f)
        header = {name: idx for idx, name in enumerate(next(reader))}

        # Extract column indices once
        i_status = header['status']
        i_amount = header['amount']
        i_timestamp = header['timestamp']
        i_region = header['region'] # kept for schema validation; not needed for aggregation

        for row in reader:
            if row[i_status] == 'completed':
                amount = float(row[i_amount])
                if amount &amp;gt; 0:
                    date = row[i_timestamp][:10]
                    daily_revenue[date] += amount

    return dict(daily_revenue)


&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;
&lt;h3&gt;
  
  
  Why each change improves performance
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Change&lt;/th&gt;
&lt;th&gt;Why it helps&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;&lt;code&gt;csv.reader&lt;/code&gt; instead of &lt;code&gt;DictReader&lt;/code&gt;&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;DictReader&lt;/code&gt; allocates a new &lt;code&gt;dict&lt;/code&gt; and hashes every column name for &lt;em&gt;every row&lt;/em&gt;. &lt;code&gt;reader&lt;/code&gt; returns tuples; mapping headers to indices once eliminates that per-row overhead.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;&lt;code&gt;defaultdict(float)&lt;/code&gt; instead of &lt;code&gt;list&lt;/code&gt; + &lt;code&gt;DataFrame&lt;/code&gt;&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;The original builds a 10M-item list of dicts, then a DataFrame, then runs a groupby. Streaming accumulates results directly, cutting memory usage by ~10× and avoiding DataFrame construction cost entirely.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Single &lt;code&gt;float()&lt;/code&gt; conversion&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;The original calls &lt;code&gt;float(row['amount'])&lt;/code&gt; twice per valid row. We convert once and reuse.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;No &lt;code&gt;.strip().upper()&lt;/code&gt; on &lt;code&gt;region&lt;/code&gt;&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Since the final output only groups by &lt;code&gt;date&lt;/code&gt; and sums &lt;code&gt;revenue&lt;/code&gt;, &lt;code&gt;region&lt;/code&gt; is unused. Removing dead work saves ~20M string operations.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;&lt;code&gt;newline=''&lt;/code&gt; in &lt;code&gt;open()&lt;/code&gt;&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Required by the &lt;code&gt;csv&lt;/code&gt; module for correct parsing; prevents extra &lt;code&gt;\r&lt;/code&gt; handling overhead.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Expected speed:&lt;/strong&gt; On a local SSD, this typically drops from ~4 min to &lt;strong&gt;10–25 seconds&lt;/strong&gt; for 10M rows.&lt;/p&gt;


&lt;h2&gt;
  
  
  Alternative: Vectorized Pandas (if you need the DataFrame)
&lt;/h2&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;
python
import pandas as pd

def process_sales(csv_path):
    df = pd.read_csv(
        csv_path,
        usecols=['status', 'amount', 'timestamp', 'region'],
        dtype={'status': 'string', 'timestamp': 'string', 'region': 'string'}
    )

    mask = (df['status'] == 'completed') &amp;amp; (df['amount'].astype(float) &amp;gt; 0)
    df = df[mask].copy()

    df['date'] = df['timestamp'].str[:10]
    df['revenue'] = df['amount'].astype(float)

    return df.groupby('date')['revenue'].sum().to_dict()


&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;
&lt;h3&gt;
  
  
  Why this is faster than the original
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;pd.read_csv&lt;/code&gt; is C-accelerated (parsing in C, not Python).&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;usecols&lt;/code&gt; skips loading unused columns.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;astype(float)&lt;/code&gt; is vectorized in C.&lt;/li&gt;
&lt;li&gt;Still slower than streaming for &lt;em&gt;this specific aggregation&lt;/em&gt; because it materializes the full filtered DataFrame in memory.&lt;/li&gt;
&lt;/ul&gt;


&lt;h2&gt;
  
  
  When this optimization is &lt;strong&gt;not&lt;/strong&gt; worth doing
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Small files (&amp;lt; 100k rows).&lt;/strong&gt; The original code is perfectly readable and likely runs in under a second. Micro-optimizing adds complexity for no human-perceptible gain.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The DataFrame itself is a required output.&lt;/strong&gt; If downstream code needs the cleaned 10M-row DataFrame (not just the daily totals), you must build it anyway—stream to an efficient format like Parquet instead.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The real bottleneck is I/O.&lt;/strong&gt; If &lt;code&gt;csv_path&lt;/code&gt; is on a slow network drive, parsing speed won't matter; you're waiting on bytes over the wire. Profile with &lt;code&gt;iostat&lt;/code&gt; first.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Maintenance cost exceeds time savings.&lt;/strong&gt; A 4-minute batch job that runs once daily costs ~24 hours of developer time per year to save. If this takes you more than a few hours to test and deploy, the ROI may be negative.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Schema instability.&lt;/strong&gt; If column names change frequently, the index-mapping in the streaming version is slightly more brittle than &lt;code&gt;DictReader&lt;/code&gt;. Add a small header-validation check if you go this route.&lt;/li&gt;
&lt;/ol&gt;


&lt;h3&gt;
  
  
  Bottom line
&lt;/h3&gt;

&lt;p&gt;For a 10M-row aggregation where &lt;strong&gt;only the final sums matter&lt;/strong&gt; , the streaming &lt;code&gt;csv.reader&lt;/code&gt; approach is the right balance: it eliminates unnecessary memory allocation and leverages Python's strengths (simple loops are fast when they do almost nothing per iteration).&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;
plaintext

#### Our Take

- Reasoning Quality: Exceptionally contextual. It didn’t just compare tools — it compared _use cases_. The header-to-index mapping solved V4 Flash’s readability weakness without sacrificing performance. Removing dead region processing showed attention to wasted work. The Pandas alternative wasn’t an afterthought; it was a conditional path with clear entry criteria.
- Code Quality: Production-ready and thoughtful. The streaming version is fast, readable (thanks to named indices), and correct (newline='' included). The Pandas version usesusecols, vectorized astype, and avoids redundant conversions. Both are immediately usable. Only minor gap: didn’t mention Polars at all—but that’s a feature, not a bug, given the task’s simplicity.
- Instruction Following: Perfect. Hit performance target, prioritized readability via named indices, explained every change mechanistically, and included ROI-grounded caveats. Went beyond by offering a _conditional_ alternative instead of forcing a single answer.
- Notable Strengths: Most adaptable response. Solved V4 Flash’s readability issue without adding dependencies. Avoided Qwen’s Polars assumption for a task that doesn’t need it. Best articulation of _when each approach wins_. The “maintenance cost vs. time saved” framing is exactly how senior engineers justify optimizations.
- Notable Weaknesses: Didn’t explore Polars/lazy evaluation, which could yield faster results for teams already in that ecosystem. But given the prompt’s emphasis on readability and maintainability, this omission feels intentional — not ignorant.

#### Verdict on This Task

Kimi K3 delivered the most contextually intelligent response of the four. It didn’t chase peak performance or ideological purity — it found the sweet spot between speed, readability, and real-world constraints. For developers who need solutions that fit their actual workflow (not a benchmark), this is the model that thinks like a teammate.

### Final Head-to-Head Comparison: All Four Models on Task



&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Criteria&lt;/th&gt;
&lt;th&gt;GLM-5.2&lt;/th&gt;
&lt;th&gt;DeepSeek V4 Flash&lt;/th&gt;
&lt;th&gt;Qwen 3.8 Max&lt;/th&gt;
&lt;th&gt;Kimi K3&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Execution Strategy&lt;/td&gt;
&lt;td&gt;Tiered (Pandas + Polars)&lt;/td&gt;
&lt;td&gt;Single-pass stdlib&lt;/td&gt;
&lt;td&gt;Polars-native&lt;/td&gt;
&lt;td&gt;Streaming stdlib + conditional Pandas&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Primary Data Library&lt;/td&gt;
&lt;td&gt;pandas / polars&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;td&gt;polars&lt;/td&gt;
&lt;td&gt;None (pandas optional)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Typical Runtime&lt;/td&gt;
&lt;td&gt;3–20s&lt;/td&gt;
&lt;td&gt;20–25s&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;2–8s&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;10–25s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reasoning Quality&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;td&gt;Medium-High&lt;/td&gt;
&lt;td&gt;High (named indices)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Primary Optimization&lt;/td&gt;
&lt;td&gt;Maintenance/team&lt;/td&gt;
&lt;td&gt;Schema/portability&lt;/td&gt;
&lt;td&gt;Operational/deployment&lt;/td&gt;
&lt;td&gt;ROI/workflow fit&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Developer Flexibility&lt;/td&gt;
&lt;td&gt;✅ Offered choices&lt;/td&gt;
&lt;td&gt;✅ Stdlib-only but justified&lt;/td&gt;
&lt;td&gt;❌ Polars-default&lt;/td&gt;
&lt;td&gt;✅ Tool-agnostic&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Best Fit&lt;/td&gt;
&lt;td&gt;Flexible teams&lt;/td&gt;
&lt;td&gt;Minimal-dependency environments&lt;/td&gt;
&lt;td&gt;Modern data stacks&lt;/td&gt;
&lt;td&gt;Real-world pragmatism&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;
plaintext

**Best For:** Local deployment, startups, research, and AI assistants.

Want an OpenAI-compatible local API? Deploy LocalAI with TechLatest and expose open-source models through a familiar API for your applications.

[Techlatest.net - LocalAI: Self-Hosted Alternative to OpenAI &amp;amp; Anthropic](https://techlatest.net/support/local-ai-support/)

### Conclusion: There Is No Single “Best” Model (And That’s Good News)

After testing GLM-5.2, DeepSeek V4 Flash, Qwen 3.8 Max, and Kimi K3 on the same real-world task, one thing is clear: the era of chasing a single “best” coding model is over.

Each model excelled in a different dimension:

- GLM-5.2 thought like a staff engineer, diagnosing root causes and offering tiered solutions with full awareness of team and maintenance costs.
- DeepSeek V4 Flash proved lightweight models can be pragmatic, delivering zero-dependency solutions that balance speed and portability without dogma.
- Qwen 3.8 Max demonstrated deep ecosystem literacy, leveraging modern tooling intelligently while articulating exactly when its assumptions break down.
- Kimi K3 embodied contextual adaptability, refusing ideological purity to deliver solutions tailored to the actual constraints of the task.

None failed. None blindly applied textbook patterns. All four explained trade-offs like humans, not benchmarks.

### So Which Should You Choose?

The answer depends entirely on your context:



&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;If you...&lt;/th&gt;
&lt;th&gt;Try this first&lt;/th&gt;
&lt;th&gt;Why&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Work across diverse stacks and need flexibility&lt;/td&gt;
&lt;td&gt;GLM-5.2&lt;/td&gt;
&lt;td&gt;Offers multiple approaches with clear guidance on when to use each&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Need portable, dependency-free code&lt;/td&gt;
&lt;td&gt;DeepSeek V4 Flash&lt;/td&gt;
&lt;td&gt;Stdlib-only solution that doesn’t sacrifice engineering maturity&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Already use Polars/Rust data tools&lt;/td&gt;
&lt;td&gt;Qwen 3.8 Max&lt;/td&gt;
&lt;td&gt;Deepest understanding of modern data infrastructure trade-offs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Want solutions that fit your actual workflow&lt;/td&gt;
&lt;td&gt;Kimi K3&lt;/td&gt;
&lt;td&gt;Tool-agnostic pragmatism that reads the room before writing code&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;


### What This Means for Open-Source Coding AI

This benchmark reveals a healthier ecosystem than hype suggests. These models aren’t clones competing on synthetic scores — they’re specialized tools with distinct identities. That diversity is a feature, not a bug. It means developers can pick models based on real needs, not marketing.

It also proves open-source coding AI has crossed a threshold. These models don’t just generate code; they reason about systems, communicate trade-offs, and demonstrate engineering judgment. They’re no longer just autocomplete on steroids — they’re collaborators.

### Final Thought

Benchmarks measure what’s easy to quantify. Real work measures what matters.

All four models passed the test that counts: they produced outputs you’d trust in production. The rest is just matching their strengths to your reality.

Stop asking “which is best?” Start asking “which fits?”

Your codebase already knows the answer.

### Thank you so much for reading

Like | Follow | Subscribe to the newsletter.

Catch us on

Website: [https://www.techlatest.net/](https://www.techlatest.net/)

Newsletter: [https://substack.com/@parvezmohammed](https://substack.com/@parvezmohammed)

Twitter: [https://twitter.com/TechlatestNet](https://twitter.com/TechlatestNet)

LinkedIn: [https://www.linkedin.com/in/techlatest-net/](https://www.linkedin.com/in/techlatest-net/)

YouTube:[https://www.youtube.com/@techlatest\_net/](https://www.youtube.com/@techlatest_net/)

Blogs: [https://medium.com/@techlatest.net](https://medium.com/@techlatest.net)

Reddit Community: [https://www.reddit.com/user/techlatest\_net/](https://www.reddit.com/user/techlatest_net/)

* * *
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

</description>
      <category>deepseek</category>
      <category>kimik3</category>
      <category>opensource</category>
      <category>qwen</category>
    </item>
    <item>
      <title>World Models 101: Teaching AI to Imagine Before It Acts</title>
      <dc:creator>TechLatest</dc:creator>
      <pubDate>Tue, 04 Aug 2026 13:40:40 +0000</pubDate>
      <link>https://dev.to/techlatestnet/world-models-101-teaching-ai-to-imagine-before-it-acts-4oki</link>
      <guid>https://dev.to/techlatestnet/world-models-101-teaching-ai-to-imagine-before-it-acts-4oki</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbu9jhy9hrhcwrlbkbg5s.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbu9jhy9hrhcwrlbkbg5s.png" width="800" height="446"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Welcome to Part 1 of our World Models Series.&lt;/p&gt;

&lt;p&gt;If you’ve been following AI lately, you’ve likely heard the term “World Model” thrown around by researchers at DeepMind, Meta, and NVIDIA. But what exactly is it? Is it just a fancy video generator? A physics engine? Or something else entirely?&lt;/p&gt;

&lt;p&gt;In this introductory post, we’re stripping away the jargon. We’ll explore what world models are, why they represent a fundamental shift from current AI, and the basic architecture that makes them tick. No heavy math today — just the core concepts to set the stage for our deep dives later in the series.&lt;/p&gt;

&lt;h3&gt;
  
  
  What is a World Model?
&lt;/h3&gt;

&lt;p&gt;At its simplest, a world model is an AI system that learns to internally simulate how the world works.&lt;/p&gt;

&lt;p&gt;Current Large Language Models (LLMs) are incredible at predicting the next &lt;em&gt;word&lt;/em&gt;. They understand language patterns, but they don’t truly understand gravity, object permanence, or cause-and-effect. If an LLM reads about dropping a glass, it knows the sentence usually ends with “shattered,” but it doesn’t &lt;em&gt;simulate&lt;/em&gt; the fall.&lt;/p&gt;

&lt;p&gt;A World Model, conversely, predicts the next &lt;em&gt;state&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;It builds an internal representation of physical dynamics, agents, and causal relationships. Instead of passively recognizing patterns in data, it actively learns how environments evolve and how specific actions change them. It can imagine, generate, and interact with coherent virtual worlds, effectively allowing the AI to “think before it acts.”&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;💡 Key Distinction: Pattern recognition tells you what something&lt;/em&gt; is_. A world model tells you what will_ happen next &lt;em&gt;if you do X.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  Why Do We Need Them?
&lt;/h3&gt;

&lt;p&gt;We are hitting the limits of scaling pure language and static vision models. To achieve true embodied intelligence (robots, autonomous vehicles, adaptive agents), AI needs more than correlation; it needs causation.&lt;/p&gt;

&lt;p&gt;World models unlock three critical capabilities:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Safe Planning &amp;amp; Reasoning: An agent can simulate thousands of future trajectories in its “mind” to evaluate outcomes without risking real-world damage.&lt;/li&gt;
&lt;li&gt;Data Efficiency: By learning the underlying rules of physics and interaction, models require less brute-force training data to generalize to new environments.&lt;/li&gt;
&lt;li&gt;Synthetic Data Generation: High-fidelity world models can generate unlimited, physically consistent training scenarios for robotics and self-driving cars, solving the data scarcity problem.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;As Yann LeCun and Fei-Fei Li have argued, world models aren’t just another modality — they are the missing substrate for reasoning and spatial intelligence.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Basic Architecture: How Does It Work?
&lt;/h3&gt;

&lt;p&gt;While modern implementations vary wildly (from diffusion-based simulators to latent-space planners), most world models share a foundational three-component architecture. Think of it as a loop of Perception → Imagination → Action.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz6trikud8iv105mj22qo.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz6trikud8iv105mj22qo.png" width="800" height="132"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  The Three Pillars
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;The Encoder (Perception): Compresses raw sensory input (pixels, lidar, text) into a compact, structured latent representation. This isn’t just image compression; it’s extracting &lt;em&gt;meaningful state variables&lt;/em&gt; like object positions, velocities, and relationships.&lt;/li&gt;
&lt;li&gt;The Dynamics Model (The Brain): This is the heart of the world model. Operating entirely in latent space, it takes the current state + a proposed action and predicts the &lt;em&gt;next&lt;/em&gt; latent state. It has learned the transition function: f(state, action) → next_state. This is where physics, causality, and temporal consistency live.&lt;/li&gt;
&lt;li&gt;The Decoder (Generation/Verification): Translates the predicted latent state back into observable space (e.g., a video frame, a 3D point cloud). During training, this reconstruction is compared against reality to teach the dynamics model accuracy. During inference, it allows us to visualize the model’s imagination.&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  Planning Happens in Latent Space
&lt;/h3&gt;

&lt;p&gt;Crucially, the planner doesn’t simulate full-resolution video frames for every possible future. That would be computationally impossible. Instead, it rolls out hundreds of potential futures &lt;em&gt;in the compressed latent space&lt;/em&gt; using the dynamics model, evaluates which trajectory achieves the goal, and only then executes the winning action in the real world.&lt;/p&gt;

&lt;p&gt;This is what enables real-time interaction and long-horizon reasoning.&lt;/p&gt;

&lt;h3&gt;
  
  
  What’s Coming Next in This Series?
&lt;/h3&gt;

&lt;p&gt;This post gives you the mental model. In upcoming installments, we’ll go deeper:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Part 2: Architectures Deep Dive — Diffusion vs. Autoregressive vs. JEPA approaches&lt;/li&gt;
&lt;li&gt;Part 3: Benchmarks &amp;amp; Evaluation — How do we actually measure if a world model “understands” physics?&lt;/li&gt;
&lt;li&gt;Part 4: Embodied AI Applications — From robot manipulation to autonomous driving&lt;/li&gt;
&lt;li&gt;Part 5: Open Challenges — Temporal consistency, scaling, and the gap between simulation and reality&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The field is moving fast. Projects like Genie, Cosmos, SIMA, and dozens of open-source efforts are pushing boundaries monthly. But beneath the rapid iteration, the core idea remains: to build intelligent agents, we must first teach machines to dream coherently about the world they inhabit.&lt;/p&gt;

&lt;h3&gt;
  
  
  Thank you so much for reading
&lt;/h3&gt;

&lt;p&gt;Like | Follow | Subscribe to the newsletter.&lt;/p&gt;

&lt;p&gt;Catch us on&lt;/p&gt;

&lt;p&gt;Website: &lt;a href="https://www.techlatest.net/" rel="noopener noreferrer"&gt;https://www.techlatest.net/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Newsletter: &lt;a href="https://substack.com/@parvezmohammed" rel="noopener noreferrer"&gt;https://substack.com/@parvezmohammed&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Twitter: &lt;a href="https://twitter.com/TechlatestNet" rel="noopener noreferrer"&gt;https://twitter.com/TechlatestNet&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;LinkedIn: &lt;a href="https://www.linkedin.com/in/techlatest-net/" rel="noopener noreferrer"&gt;https://www.linkedin.com/in/techlatest-net/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;YouTube:&lt;a href="https://www.youtube.com/@techlatest_net/" rel="noopener noreferrer"&gt;https://www.youtube.com/@techlatest_net/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Blogs: &lt;a href="https://medium.com/@techlatest.net" rel="noopener noreferrer"&gt;https://medium.com/@techlatest.net&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Reddit Community: &lt;a href="https://www.reddit.com/user/techlatest_net/" rel="noopener noreferrer"&gt;https://www.reddit.com/user/techlatest_net/&lt;/a&gt;&lt;/p&gt;

</description>
      <category>aimodel</category>
      <category>llm</category>
      <category>worldmodels</category>
      <category>agents</category>
    </item>
  </channel>
</rss>
