<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: zhijie</title>
    <description>The latest articles on DEV Community by zhijie (@zjshen14).</description>
    <link>https://dev.to/zjshen14</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4155601%2Fb5ffae32-4c7b-42de-8389-d16dcf9f4891.png</url>
      <title>DEV Community: zhijie</title>
      <link>https://dev.to/zjshen14</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/zjshen14"/>
    <language>en</language>
    <item>
      <title>MTP on an RTX 3090: Faster Tokens, but What About Coding Quality?</title>
      <dc:creator>zhijie</dc:creator>
      <pubDate>Fri, 02 Oct 2026 21:22:39 +0000</pubDate>
      <link>https://dev.to/zjshen14/mtp-on-an-rtx-3090-faster-tokens-but-what-about-coding-quality-28j</link>
      <guid>https://dev.to/zjshen14/mtp-on-an-rtx-3090-faster-tokens-but-what-about-coding-quality-28j</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff4y8x666xdgtlma3iwwy.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff4y8x666xdgtlma3iwwy.png" alt="MTP speed and coding quality on RTX 3090" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Originally published on &lt;a href="https://zjshen14.github.io/en/blog/qwen-mtp-speed-quality/?utm_source=devto&amp;amp;utm_medium=referral&amp;amp;utm_campaign=qwen_mtp" rel="noopener noreferrer"&gt;my blog&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Enabling MTP on this RTX 3090 raised generation throughput from &lt;strong&gt;36.0 to 56.2 tokens/s, about 56% faster&lt;/strong&gt;. The two cache tasks that succeeded in both modes also finished about 20% and 40% sooner.&lt;/p&gt;

&lt;p&gt;One transfer-task pair produced a patch-quality difference, however. I am continuing to test MTP while keeping it off by default.&lt;/p&gt;

&lt;p&gt;This experiment grew out of a reader's suggestion on the &lt;a href="https://zjshen14.github.io/en/blog/local-qwen-opencode-3090/" rel="noopener noreferrer"&gt;previous local-deployment post&lt;/a&gt;: try MTP or a quantization that includes it. The RTX 3090 left over from Ethereum mining was already running Qwen3.8-27B through OpenCode. I wanted to see whether faster generation would deliver correct patches sooner.&lt;/p&gt;

&lt;p&gt;We ran &lt;strong&gt;eight paired comparisons, sixteen attempts&lt;/strong&gt;, using the same model file. Each pair shared its task, seed, and initial environment; only the MTP setting changed.&lt;/p&gt;

&lt;h2&gt;
  
  
  MTP drafts tokens for the target model to verify
&lt;/h2&gt;

&lt;p&gt;Ordinary generation typically produces one token per step. MTP (Multi-Token Prediction) uses an embedded prediction head to propose upcoming tokens, then asks the target model to verify them. Accepting several in sequence can reduce serial decoding steps. This is a &lt;a href="https://github.com/ggml-org/llama.cpp/blob/b11146/docs/speculative.md" rel="noopener noreferrer"&gt;speculative-decoding path in llama.cpp&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The target model still verifies the candidates. The &lt;a href="https://arxiv.org/abs/2211.17192" rel="noopener noreferrer"&gt;speculative-decoding paper&lt;/a&gt; shows that a correct algorithm can preserve its output distribution. Speed need not inherently cost model quality.&lt;/p&gt;

&lt;p&gt;Our existing &lt;strong&gt;Qwen3.8-27B Q4_K_M GGUF contained an MTP prediction layer&lt;/strong&gt;, so we used it directly without switching to another Unsloth quantization. The server commands differed only in &lt;code&gt;--spec-type none&lt;/code&gt; versus &lt;code&gt;--spec-type draft-mtp&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Generation was 56% faster; successful tasks saved 20–40%
&lt;/h2&gt;

&lt;p&gt;Generation throughput counts tokens produced per second, including the model's reasoning output. Task time also includes reading source, executing tools, and running tests. It is closer to how long you wait for a patch.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Measure&lt;/th&gt;
&lt;th&gt;MTP off&lt;/th&gt;
&lt;th&gt;MTP on&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Aggregate coding-generation throughput&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;36.0 tokens/s&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;56.2 tokens/s&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Predefined independent checks passed&lt;/td&gt;
&lt;td&gt;35/80&lt;/td&gt;
&lt;td&gt;26/80&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Completed within deadline and passed predefined checks&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;2/8&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;2/8&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The two cache pairs that succeeded in both modes provide the clearest evidence of time saved:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Paired run&lt;/th&gt;
&lt;th&gt;MTP off&lt;/th&gt;
&lt;th&gt;MTP on&lt;/th&gt;
&lt;th&gt;Task time saved&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Pair one&lt;/td&gt;
&lt;td&gt;192.0 s&lt;/td&gt;
&lt;td&gt;153.5 s&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;20.1%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pair two&lt;/td&gt;
&lt;td&gt;431.2 s&lt;/td&gt;
&lt;td&gt;259.0 s&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;39.9%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Both delivered a patch passing the predefined checks sooner. A failed run ending earlier does not count as speeding up a successful task. Later prompts also diverge with the agent's actions, so aggregate throughput is not an identical-text speed comparison.&lt;/p&gt;

&lt;h3&gt;
  
  
  Short replies and generation-timing definitions
&lt;/h3&gt;

&lt;p&gt;Short replies to matched first requests reached &lt;strong&gt;38.9–39.8 tokens/s&lt;/strong&gt; with MTP off and &lt;strong&gt;74.0–84.2 tokens/s&lt;/strong&gt; with it on. These were brief tool-selection responses; they do not show that whole coding tasks finish in half the time.&lt;/p&gt;

&lt;p&gt;Aggregate throughput is total accepted target-output tokens divided by total corresponding server generation time across completed generations. It includes reasoning and tool calls, excludes prompt processing and rejected drafts, and omits interrupted generations without final timing records.&lt;/p&gt;

&lt;p&gt;Task time includes prompt processing, tools, and the agent's own tests. Server startup and independent grading afterward are excluded.&lt;/p&gt;

&lt;h2&gt;
  
  
  Nine differing checks came from one transfer repair
&lt;/h2&gt;

&lt;p&gt;The entire 35/80 versus 26/80 gap came from &lt;strong&gt;one transfer task at one seed&lt;/strong&gt;. Cache, ledger, and build-planner scores matched within every pair.&lt;/p&gt;

&lt;p&gt;With MTP off, the agent produced a production-code patch that subsequently passed all ten independent checks. It reached the deadline at &lt;strong&gt;480.1 seconds&lt;/strong&gt; without a final handoff.&lt;/p&gt;

&lt;p&gt;With MTP on, the agent ended after &lt;strong&gt;186.1 seconds&lt;/strong&gt; without editing production code. It passed only the one check the initial implementation already passed.&lt;/p&gt;

&lt;p&gt;The off candidate performed better on those checks, and neither attempt met the full delivery standard. We therefore tracked &lt;strong&gt;patch correctness&lt;/strong&gt; and &lt;strong&gt;timely delivery&lt;/strong&gt; separately: completed tasks remained 2/8 in both modes. Nine related checks are not nine independent task regressions.&lt;/p&gt;

&lt;h3&gt;
  
  
  Task scores and an early difference that did not recur
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Task&lt;/th&gt;
&lt;th&gt;MTP off, two scores&lt;/th&gt;
&lt;th&gt;MTP on, two scores&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Expiring cache&lt;/td&gt;
&lt;td&gt;10/10, 10/10&lt;/td&gt;
&lt;td&gt;10/10, 10/10&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CSV ledger&lt;/td&gt;
&lt;td&gt;1/10, 1/10&lt;/td&gt;
&lt;td&gt;1/10, 1/10&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Incremental build planner&lt;/td&gt;
&lt;td&gt;1/10, 1/10&lt;/td&gt;
&lt;td&gt;1/10, 1/10&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SQLite transfers&lt;/td&gt;
&lt;td&gt;1/10, &lt;strong&gt;10/10&lt;/strong&gt;
&lt;/td&gt;
&lt;td&gt;1/10, &lt;strong&gt;1/10&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Completed and passing requires no timeout, successful process exit, unchanged protected contracts/configuration, passing public and added tests, and all ten independent checks passing. It is not proof that every possible defect is covered.&lt;/p&gt;

&lt;p&gt;An early exploratory comparison showed a build-planner score difference. It did not recur under these tighter controls; those scores are not pooled into this study.&lt;/p&gt;

&lt;p&gt;The cache task also gave me a reason to inspect code after green tests. An MTP candidate's added TTL validation overflowed on a very large integer, while the paired off candidate passed the extra probe. This extreme API boundary is worth recording; it establishes neither everyday failure frequency nor MTP as the cause.&lt;/p&gt;

&lt;h3&gt;
  
  
  Extra review: cache overflow and shared omissions
&lt;/h3&gt;

&lt;p&gt;All four cache attempts passed the ten predefined checks. One MTP candidate added &lt;code&gt;math.isfinite(ttl)&lt;/code&gt;. The contract permits finite positive integer TTLs. With &lt;code&gt;10**400&lt;/code&gt; and an injected integer clock increasing on every call, validation raised &lt;code&gt;OverflowError&lt;/code&gt; while converting the integer to a float. The paired off candidate passed.&lt;/p&gt;

&lt;p&gt;Ledger probes for large exact amounts, header names, and read-time errors failed in both modes at both seeds. A wallet destination-integer-overflow probe passed off and failed on in the second pair; it failed both modes in the first pair.&lt;/p&gt;

&lt;p&gt;Extra review ran after grading, never reached the agent, and was not added retrospectively to the eighty checks. The &lt;a href="https://zjshen14.github.io/experiments/qwen-mtp-speed-quality/controlled-summary.json" rel="noopener noreferrer"&gt;paired records&lt;/a&gt; retain details.&lt;/p&gt;

&lt;h2&gt;
  
  
  Faster output still exhausts the same allowance
&lt;/h2&gt;

&lt;p&gt;Most unsuccessful attempts shared a problem: they spent the &lt;strong&gt;8,192-token response allowance&lt;/strong&gt; on reasoning, ended with &lt;code&gt;length&lt;/code&gt;, and never edited production code. Many 1/10 scores belong to unchanged starting code; they do not mean a tenth of the repair was completed.&lt;/p&gt;

&lt;p&gt;Increasing context capacity would not resolve that limit. Context determines how much input and output can fit; the response allowance caps what one generation can produce. MTP can produce tokens faster without adding room to that allowance.&lt;/p&gt;

&lt;p&gt;This also limits the quality comparison. Both modes frequently produced no patch, leaving few completed repairs to compare. There were only two seeds per task and four small Python projects. The eighty checks are correlated, not eighty independent samples.&lt;/p&gt;

&lt;p&gt;The transfer difference deserves investigation, but its cause remains unresolved. Numerical behavior in batched verification, state handling, and sampling are possible avenues to investigate. An agent's different action can also change later inputs and patches. These are possible explanations, not findings confirmed by this experiment.&lt;/p&gt;

&lt;h2&gt;
  
  
  Controls and the limits of this comparison
&lt;/h2&gt;

&lt;p&gt;Each task ran at two seeds, with MTP off and on; the second seed reversed the order. Every attempt rebuilt the same initial code and fixed the working directory, prompts, tools, permissions, and generation settings. Ten independent test methods per task stayed outside the agent's workspace, with no results fed back to it.&lt;/p&gt;

&lt;p&gt;A proxy checked the complete first request actually sent to the model: &lt;strong&gt;all eight pairs matched&lt;/strong&gt;, as did initial file hashes. Each attempt restarted the server and disabled prefix-cache reuse in both modes. These results therefore cannot predict gains in ordinary cached conversations.&lt;/p&gt;

&lt;h3&gt;
  
  
  Full configuration, request audit, and scope
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Setting&lt;/th&gt;
&lt;th&gt;Fixed configuration&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Model and GPU&lt;/td&gt;
&lt;td&gt;Same Qwen3.8-27B Q4_K_M GGUF; RTX 3090 24GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Software&lt;/td&gt;
&lt;td&gt;llama.cpp b11146 / CUDA 12.8; OpenCode 2.0.20&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Server&lt;/td&gt;
&lt;td&gt;131,072-token capacity; one slot; all model layers on GPU&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cache and compute&lt;/td&gt;
&lt;td&gt;q8_0 K/V; flash attention; batch 512 / microbatch 256&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Generation&lt;/td&gt;
&lt;td&gt;Medium thinking; temperature 1.0; top_p 0.95; top_k 20&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Budgets&lt;/td&gt;
&lt;td&gt;8,192 output tokens per generation; 480 seconds per task&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Paired seeds&lt;/td&gt;
&lt;td&gt;4242 and 8675309&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Speculation&lt;/td&gt;
&lt;td&gt;Off: &lt;code&gt;none&lt;/code&gt;; on: &lt;code&gt;draft-mtp&lt;/code&gt;; maximum draft length 3&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Pairs one and two in the elapsed-time table correspond to seeds 4242 and 8675309.&lt;/p&gt;

&lt;p&gt;Target and draft backend sampling were disabled in both modes; synthetic acceptance was not used. Each pair shared the initial commit, agent title, project configuration, and grader. Reported cached-prefix token counts were zero. The first seed ran off→on; the second ran on→off.&lt;/p&gt;

&lt;p&gt;These were fresh-server attempts without prefix-cache reuse. Task timing excludes server startup but includes prompt processing, tools, and the agent's own tests. Independent grading afterward is excluded. Desktop and other background GPU activity were not fully controlled.&lt;/p&gt;

&lt;p&gt;The largest actual input was &lt;strong&gt;20,057 tokens&lt;/strong&gt;. The 128K setting is capacity; this experiment did not measure full-128K coding performance or acceleration in ordinary cached conversations.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep the current default; test a larger output budget next
&lt;/h2&gt;

&lt;p&gt;For this setup, I am keeping MTP off by default. Saving 20–40% on the successful cache tasks is useful; before expanding its use, I want to investigate the transfer-patch difference.&lt;/p&gt;

&lt;p&gt;The next comparison should raise the output allowance equally in both modes, add paired seeds, and use tasks closer to everyday work. Cached conversations need a separate test. That should give a better basis for deciding whether faster generation consistently delivers correct patches sooner.&lt;/p&gt;

&lt;p&gt;If you try MTP, record generation speed, task time, independent tests, and final delivery together, then decide using the tasks you actually do.&lt;/p&gt;

&lt;h2&gt;
  
  
  Data and references
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://zjshen14.github.io/experiments/qwen-mtp-speed-quality/controlled-summary.json" rel="noopener noreferrer"&gt;Paired experiment JSON&lt;/a&gt;: scores, elapsed times, generation timings, request checks, and extra probes for sixteen attempts.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://zjshen14.github.io/experiments/qwen-mtp-speed-quality/README.txt" rel="noopener noreferrer"&gt;Method and data notes&lt;/a&gt;: attachments omit experiment dates, locations, personal paths, session identifiers, and raw conversations. They are result records, not a complete reproduction kit.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://zjshen14.github.io/en/blog/local-qwen-opencode-3090/" rel="noopener noreferrer"&gt;Previous post: local Qwen3.8-27B + OpenCode deployment&lt;/a&gt;: installation and integration. Its first-attempt study is a separate batch; scores should not be pooled with these.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/ggml-org/llama.cpp/blob/b11146/docs/speculative.md" rel="noopener noreferrer"&gt;llama.cpp b11146 speculative-decoding documentation&lt;/a&gt; and the &lt;a href="https://arxiv.org/abs/2211.17192" rel="noopener noreferrer"&gt;speculative-decoding paper&lt;/a&gt;: mechanism background. Measurements on this card come from the paired records.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>programming</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Run OpenCode from your phone or laptop: remote hosting, authentication, and reverse proxy</title>
      <dc:creator>zhijie</dc:creator>
      <pubDate>Thu, 01 Oct 2026 20:09:27 +0000</pubDate>
      <link>https://dev.to/zjshen14/run-opencode-from-your-phone-or-laptop-remote-hosting-authentication-and-reverse-proxy-16bd</link>
      <guid>https://dev.to/zjshen14/run-opencode-from-your-phone-or-laptop-remote-hosting-authentication-and-reverse-proxy-16bd</guid>
      <description>&lt;p&gt;Originally published on &lt;a href="https://zjshen14.github.io/en/blog/setup-opencode-remote-web-ide/?utm_source=devto&amp;amp;utm_medium=referral&amp;amp;utm_campaign=remote_opencode" rel="noopener noreferrer"&gt;my blog&lt;/a&gt;. This is an adapted version of the deployment guide.&lt;/p&gt;

&lt;p&gt;A coding agent is easier to use across devices when the workspace lives on one host. The phone or laptop opens the browser UI; the host keeps the repository, runs tools, and connects to the model provider.&lt;/p&gt;

&lt;p&gt;That host can be a desktop, a homelab machine, or a VPS. You do not need a public domain to begin. The choices that matter are how clients reach the service, how access is authenticated, and what happens when the launching terminal closes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with the workspace and authentication
&lt;/h2&gt;

&lt;p&gt;Enter the project directory before starting OpenCode. For a service behind a reverse proxy on the same host:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;cd&lt;/span&gt; /path/to/your/workspace
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;OPENCODE_SERVER_PASSWORD&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'replace-with-a-strong-password'&lt;/span&gt;
opencode web &lt;span class="nt"&gt;--hostname&lt;/span&gt; 127.0.0.1 &lt;span class="nt"&gt;--port&lt;/span&gt; 8080
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The default authentication username is &lt;code&gt;opencode&lt;/code&gt;. &lt;code&gt;opencode web&lt;/code&gt; includes the browser interface; &lt;code&gt;opencode serve&lt;/code&gt; runs the headless API service. See the &lt;a href="https://opencode.ai/docs/web/" rel="noopener noreferrer"&gt;official Web documentation&lt;/a&gt; for the current flags.&lt;/p&gt;

&lt;p&gt;Treat access to this UI as access to the host's development tools. Use a dedicated workspace and account with only the permissions the agent needs. Keep the password out of committed configuration and screenshots.&lt;/p&gt;

&lt;h2&gt;
  
  
  Choose a connection path
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Private access:&lt;/strong&gt; Keep the service on a private network. A LAN connection works for devices on that network; Tailscale is another way to connect authorized devices without exposing the IDE directly to the public internet. If you bind to &lt;code&gt;0.0.0.0&lt;/code&gt;, remember that it listens on every interface: restrict reachability with the host firewall and keep authentication enabled. Plain HTTP on a LAN is not encrypted by itself.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A VPS with a domain:&lt;/strong&gt; Keep OpenCode on loopback and put an HTTPS reverse proxy in front. For example, a Caddy site can be as small as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ide.example.com {
    reverse_proxy 127.0.0.1:8080
}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The domain must point to the server and the necessary ports must be reachable. Caddy can manage HTTPS certificates and proxy WebSocket connections; see its &lt;a href="https://caddyserver.com/docs/caddyfile/directives/reverse_proxy" rel="noopener noreferrer"&gt;reverse-proxy documentation&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;A reverse proxy does not automatically provide application authentication. Confirm that an unauthenticated visitor cannot enter the OpenCode workspace. With Nginx, check WebSocket upgrades and long-running streamed responses rather than testing only whether the initial page loads. The full article includes an Nginx example.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep the host process running
&lt;/h2&gt;

&lt;p&gt;A remote host helps only while its service stays alive. On Linux, systemd can restart OpenCode after a process failure and start it after a reboot. My guide includes an example service with a dedicated working directory and a protected environment file for the password.&lt;/p&gt;

&lt;p&gt;Use the actual installed binary path and service user. Check service logs and test a restart before relying on it for longer tasks. Closing a client browser and shutting down the host are different events; the host must remain awake and connected.&lt;/p&gt;

&lt;h2&gt;
  
  
  Select the model separately
&lt;/h2&gt;

&lt;p&gt;Hosting OpenCode's workspace does not mean the model runs locally. The original guide uses &lt;strong&gt;Muse Spark 1.3 Contributor Free&lt;/strong&gt; through OpenCode Zen as one example. Zen currently lists this as a limited-time free model, and its Contributor terms permit Meta to use prompts and completions for future model training. Review the &lt;a href="https://opencode.ai/docs/zen/" rel="noopener noreferrer"&gt;current Zen pricing and privacy terms&lt;/a&gt; before using it, and keep confidential code and secrets out of that endpoint.&lt;/p&gt;

&lt;p&gt;The deployment pattern also works with a provider appropriate to your data requirements. In my &lt;a href="https://zjshen14.github.io/en/blog/local-qwen-opencode-3090/?utm_source=devto&amp;amp;utm_medium=referral&amp;amp;utm_campaign=remote_opencode" rel="noopener noreferrer"&gt;follow-up experiment&lt;/a&gt;, inference runs on an RTX 3090 with Qwen3.8-27B and llama.cpp.&lt;/p&gt;

&lt;h2&gt;
  
  
  Check the whole path
&lt;/h2&gt;

&lt;p&gt;Before handing it a long task, verify these steps from the intended client device:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Authentication blocks unauthenticated access.&lt;/li&gt;
&lt;li&gt;The expected project and model are selected.&lt;/li&gt;
&lt;li&gt;A harmless file-read request succeeds.&lt;/li&gt;
&lt;li&gt;Tool calls and live output work through the chosen network path.&lt;/li&gt;
&lt;li&gt;The host service survives the terminal disconnects and restarts you expect.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The benefit is one workspace that I can reach from different screens. The phone becomes a way to check progress or send a follow-up; execution remains on the host.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://zjshen14.github.io/en/blog/setup-opencode-remote-web-ide/?utm_source=devto&amp;amp;utm_medium=referral&amp;amp;utm_campaign=remote_opencode" rel="noopener noreferrer"&gt;Full guide with service and proxy configuration&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>tutorial</category>
      <category>devops</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Qwen3.8-27B on RTX 3090 with OpenCode and measured coding results</title>
      <dc:creator>zhijie</dc:creator>
      <pubDate>Thu, 01 Oct 2026 19:05:13 +0000</pubDate>
      <link>https://dev.to/zjshen14/qwen38-27b-on-rtx-3090-with-opencode-and-measured-coding-results-4m4h</link>
      <guid>https://dev.to/zjshen14/qwen38-27b-on-rtx-3090-with-opencode-and-measured-coding-results-4m4h</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3ngeoxbo0h35p7nfk69f.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3ngeoxbo0h35p7nfk69f.png" alt="One RTX 3090 running a local coding agent" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Originally published on &lt;a href="https://zjshen14.github.io/en/blog/local-qwen-opencode-3090/?utm_source=devto&amp;amp;utm_medium=referral&amp;amp;utm_campaign=local_qwen_3090" rel="noopener noreferrer"&gt;my blog&lt;/a&gt;. This cross-post includes the setup and measured results; the linked records and configuration kit let you inspect the experiment.&lt;/p&gt;

&lt;p&gt;My &lt;strong&gt;RTX 3090&lt;/strong&gt; is a leftover from my Ethereum mining days. Running a local model gives it a second life: the same card now powers a coding agent.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://zjshen14.github.io/en/blog/setup-opencode-remote-web-ide/" rel="noopener noreferrer"&gt;previous post&lt;/a&gt; covered accessing one OpenCode workspace from a laptop or phone. This time, I wanted to move inference onto that host too and find out how far the hardware already on hand could take a local AI coding workflow.&lt;/p&gt;

&lt;p&gt;We loaded &lt;strong&gt;Qwen3.8-27B Q4_K_M&lt;/strong&gt; into the 3090's &lt;strong&gt;24GB of VRAM&lt;/strong&gt;, served it through llama.cpp, and connected OpenCode. After checking chat and tool calls, we gave it four small coding tasks.&lt;/p&gt;

&lt;p&gt;The results make me want to keep using it: &lt;strong&gt;bounded fixes and small features with clear acceptance criteria can produce useful patches.&lt;/strong&gt; Exhausted output budgets, a timeout, and defects missed by tests also showed why independent validation still matters.&lt;/p&gt;

&lt;p&gt;I'll explain the model choice and its connection to OpenCode, then share the measured results and lessons from using it. Full setup instructions follow later; to get straight to them, jump to deployment.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Qwen3.8-27B
&lt;/h2&gt;

&lt;p&gt;Qwen3.8-27B is a prominent local coding candidate, with results that make it worth trying: the &lt;a href="https://huggingface.co/Qwen/Qwen3.8-27B#text-performance" rel="noopener noreferrer"&gt;official model card&lt;/a&gt; reports &lt;strong&gt;61.7 on SWE-bench Pro&lt;/strong&gt; and &lt;strong&gt;73.0 on Terminal Bench 2.1&lt;/strong&gt;. Software fixes and terminal-agent tasks are close to what I want OpenCode to do. These reported results gave me a reason to test it first.&lt;/p&gt;

&lt;p&gt;The hardware fit mattered just as much. &lt;strong&gt;The 27B Q4_K_M weights occupy about 16.8GB&lt;/strong&gt;, fitting entirely into the 3090's 24GB VRAM in our configuration with room for KV cache and runtime buffers. For this old card, the balance of capability, memory use, and speed matters more than chasing a larger parameter count.&lt;/p&gt;

&lt;p&gt;Its agent focus also matches the workflow: reading code, calling tools, and revising changes after test feedback. The &lt;a href="https://github.com/QwenLM/Qwen3.8#introduction" rel="noopener noreferrer"&gt;official introduction&lt;/a&gt; highlights coding, multi-step agent execution, and adjustable thinking, making this workflow a useful place to test those capabilities.&lt;/p&gt;

&lt;p&gt;Those three considerations made it the starting point for this experiment. &lt;strong&gt;What its Q4 version can accomplish on this card still needs validation with our own tasks.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What the old card does now
&lt;/h2&gt;

&lt;p&gt;With the model chosen, the next step is connecting it to a workflow that can operate on a repository. OpenCode owns the workspace and tools: reading files, editing code, running shell commands, and executing tests. llama-server handles inference, connected through a local API.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;A laptop or phone connects to OpenCode through a LAN or private network.&lt;/li&gt;
&lt;li&gt;OpenCode sends messages and tool definitions to llama-server at &lt;code&gt;127.0.0.1:8080/v1&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;llama-server runs Qwen3.8-27B on the RTX 3090 and returns text or tool calls.&lt;/li&gt;
&lt;li&gt;OpenCode executes the tools in its workspace and sends their results back to the model.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;OpenCode executes the model's &lt;code&gt;tool_calls&lt;/code&gt; and returns tool results for the next turn. The model API itself does not execute terminal commands.&lt;/p&gt;

&lt;p&gt;The host has an &lt;strong&gt;RTX 3090 24GB, Ryzen 7 5800X, and 64GB RAM&lt;/strong&gt;. We used &lt;strong&gt;llama.cpp b11146 / CUDA 12.8&lt;/strong&gt; and &lt;strong&gt;OpenCode 2.0.20&lt;/strong&gt;, placing all model layers on the GPU with one generation slot.&lt;/p&gt;

&lt;p&gt;Weights came from &lt;a href="https://ollama.com/library/qwen3.8:27b-q4_K_M" rel="noopener noreferrer"&gt;Ollama's &lt;code&gt;qwen3.8:27b-q4_K_M&lt;/code&gt; tag&lt;/a&gt; and were loaded directly by llama-server, without installing an Ollama service. Only the text model was loaded; vision input was not tested.&lt;/p&gt;

&lt;p&gt;Two settings matter throughout this post: &lt;strong&gt;131,072 tokens of context capacity and an 8,192-token per-response output allowance&lt;/strong&gt;. Input and output share context, and reasoning consumes the response allowance too. That distinction later explained one of the coding failures.&lt;/p&gt;

&lt;h2&gt;
  
  
  Large prompts fit, but fresh processing takes time
&lt;/h2&gt;

&lt;p&gt;With short input, generation reached about &lt;strong&gt;36.4 tokens/s&lt;/strong&gt;. At roughly 120K input tokens, it fell to &lt;strong&gt;20.9 tokens/s&lt;/strong&gt;, and a fresh request took about &lt;strong&gt;three minutes&lt;/strong&gt; to produce its first token.&lt;/p&gt;

&lt;p&gt;The model processes the input before it starts generating. As input grows, that initial wait becomes a substantial part of the experience:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Actual input&lt;/th&gt;
&lt;th&gt;Fresh prompt processing&lt;/th&gt;
&lt;th&gt;Generation&lt;/th&gt;
&lt;th&gt;First token, fresh&lt;/th&gt;
&lt;th&gt;First token, cached&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;2,073 tokens&lt;/td&gt;
&lt;td&gt;999 tokens/s&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;36.4 tokens/s&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;2.75 s&lt;/td&gt;
&lt;td&gt;0.46 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;16,378 tokens&lt;/td&gt;
&lt;td&gt;995 tokens/s&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;33.5 tokens/s&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;17.00 s&lt;/td&gt;
&lt;td&gt;0.47 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;65,537 tokens&lt;/td&gt;
&lt;td&gt;795 tokens/s&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;25.9 tokens/s&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;83.41 s&lt;/td&gt;
&lt;td&gt;0.51 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;120,011 tokens&lt;/td&gt;
&lt;td&gt;656 tokens/s&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;20.9 tokens/s&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;183.01 s&lt;/td&gt;
&lt;td&gt;0.63 s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The last column is particularly relevant to coding sessions. Cached continuations reused almost the entire prefix and processed only &lt;strong&gt;27–28 new input tokens&lt;/strong&gt;. Even with roughly 120K context retained, first-token latency fell to &lt;strong&gt;0.63 seconds&lt;/strong&gt;. Sustained generation still had to attend to the long context, so it did not regain short-input speed.&lt;/p&gt;

&lt;p&gt;This suggests a practical approach: keep stable session prefixes and supply source as needed. Repeatedly opening fresh requests with large inputs incurs the processing cost again.&lt;/p&gt;

&lt;p&gt;A separate capacity check processed &lt;strong&gt;119,968 input tokens&lt;/strong&gt;, retrieved three widely separated values, and answered a follow-up. It verified capacity for that synthetic retrieval task. Coding quality at the same input length needs its own evaluation.&lt;/p&gt;

&lt;p&gt;VRAM was also close to the card's limits. Samples during throughput testing showed peak total use of &lt;strong&gt;22,162 MiB&lt;/strong&gt;, minimum free memory of &lt;strong&gt;1,943 MiB&lt;/strong&gt;, and a maximum temperature of &lt;strong&gt;85°C&lt;/strong&gt;, including desktop GPU use. The configuration fit, with limited headroom for another model or GPU-heavy workload.&lt;/p&gt;

&lt;h3&gt;
  
  
  Timing methodology, sample counts, and source data
&lt;/h3&gt;

&lt;p&gt;Every row used the same &lt;strong&gt;131,072-token server capacity&lt;/strong&gt;, varying only actual input length. Main runs disabled thinking and used synthetic source listings, temperature 0, and seed 1234. Fresh requests generated 512 tokens; continuations generated 256. Generated text was never executed.&lt;/p&gt;

&lt;p&gt;Generation uses server decode timing, excluding initial prompt processing. First-token latency uses the streaming HTTP client, including tokenization and processing overhead. The 2K row is the median of three fresh requests; larger inputs had one fresh request and one continuation each. These measurements do not establish reliability intervals. GPU samples were collected every two seconds.&lt;/p&gt;

&lt;p&gt;A separate short request with medium thinking generated about &lt;strong&gt;35.7 tokens/s&lt;/strong&gt;, including reasoning. Throughput does not directly measure usable code produced per second. The script measures single-request performance, excluding model loading and multi-user throughput.&lt;/p&gt;

&lt;p&gt;Download the &lt;a href="https://zjshen14.github.io/experiments/qwen3.8-27b-3090/throughput-results.json" rel="noopener noreferrer"&gt;measurement records&lt;/a&gt; and &lt;a href="https://zjshen14.github.io/experiments/qwen3.8-27b-3090/throughput-summary.json" rel="noopener noreferrer"&gt;summary&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The real test: four small coding projects
&lt;/h2&gt;

&lt;p&gt;After measuring speed and capacity, we checked whether the setup could complete code changes. We prepared four small Python repositories: an expiring cache, a CSV ledger, an incremental build planner, and SQLite transfers. Each had a written acceptance contract, a fresh session, and an eight-minute deadline.&lt;/p&gt;

&lt;p&gt;To check more than the agent's own tests, we prepared ten independent test methods per task, kept them outside its workspace, and never fed their results back to the model.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Task&lt;/th&gt;
&lt;th&gt;Independent checks, before → after&lt;/th&gt;
&lt;th&gt;Time&lt;/th&gt;
&lt;th&gt;Agent-added tests&lt;/th&gt;
&lt;th&gt;Outcome&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Expiring LRU cache&lt;/td&gt;
&lt;td&gt;0/10 → &lt;strong&gt;10/10&lt;/strong&gt;
&lt;/td&gt;
&lt;td&gt;≈3m 07s&lt;/td&gt;
&lt;td&gt;23&lt;/td&gt;
&lt;td&gt;Completed; strongest result&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CSV ledger and refunds&lt;/td&gt;
&lt;td&gt;1/10 → &lt;strong&gt;10/10&lt;/strong&gt;
&lt;/td&gt;
&lt;td&gt;≈5m 44s&lt;/td&gt;
&lt;td&gt;40&lt;/td&gt;
&lt;td&gt;Completed; review found boundary and error-handling gaps&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Incremental build planner&lt;/td&gt;
&lt;td&gt;1/10 → &lt;strong&gt;1/10&lt;/strong&gt;
&lt;/td&gt;
&lt;td&gt;3m 58s&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;No patch; exhausted one response's output allowance&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Atomic SQLite transfers&lt;/td&gt;
&lt;td&gt;1/10 → &lt;strong&gt;10/10&lt;/strong&gt;
&lt;/td&gt;
&lt;td&gt;8m deadline&lt;/td&gt;
&lt;td&gt;11&lt;/td&gt;
&lt;td&gt;Candidate passed checks; final test rerun and handoff unfinished&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Three candidates passed all predefined independent checks; two completed editing, testing, and final handoff within the deadline.&lt;/strong&gt; The wallet candidate passed checks, but the agent timed out before its final test rerun and handoff. Both outcomes matter.&lt;/p&gt;

&lt;p&gt;First-attempt candidates passed &lt;strong&gt;31/40 independent test methods&lt;/strong&gt;. The build planner made no edits and already passed one method at baseline. Methods also differ in difficulty, so 31/40 should not be interpreted as a general coding success rate.&lt;/p&gt;

&lt;p&gt;Two findings were more useful than the aggregate score.&lt;/p&gt;

&lt;h3&gt;
  
  
  One failure came from the response budget
&lt;/h3&gt;

&lt;p&gt;The build-planner agent read the repository and never produced a patch. Its final response generated &lt;strong&gt;8,192 tokens&lt;/strong&gt;, all reasoning, and ended with &lt;code&gt;length&lt;/code&gt;. OpenCode exited with status zero without a final answer.&lt;/p&gt;

&lt;p&gt;That request had roughly &lt;strong&gt;4,985 input tokens&lt;/strong&gt;, far below the 128K limit. Thinking and the response allowance needed attention; increasing context capacity would not resolve an exhausted output budget. The &lt;a href="https://zjshen14.github.io/experiments/qwen3.8-27b-3090/buildplan-termination.json" rel="noopener noreferrer"&gt;termination record&lt;/a&gt; preserves the evidence.&lt;/p&gt;

&lt;p&gt;A separate diagnostic restarted from the original repository with thinking disabled, keeping the prompt, allowance, and deadline unchanged. It completed in &lt;strong&gt;5m 40s&lt;/strong&gt; and passed &lt;strong&gt;9/10 independent checks&lt;/strong&gt;. The remaining failure was &lt;code&gt;load_manifest()&lt;/code&gt; returning tuples where the original API returned lists; the agent's own test accepted that change too.&lt;/p&gt;

&lt;p&gt;The diagnostic is reported separately and does not replace the first attempt. One additional sample at temperature 1 cannot establish that disabling thinking is generally better, but it identifies a setting worth testing further.&lt;/p&gt;

&lt;h3&gt;
  
  
  Many new tests still missed defects
&lt;/h3&gt;

&lt;p&gt;The ledger added &lt;strong&gt;40 tests&lt;/strong&gt;, passed the predefined independent checks, and had a reasonably clear structure. Further review still found that Python's default 28-digit Decimal precision rounded a valid large amount when adding &lt;code&gt;0.01&lt;/code&gt;, violating the no-rounding contract. An I/O error during file iteration could also escape as a traceback.&lt;/p&gt;

&lt;p&gt;The wallet exposed a different problem: its own concurrency tests submitted one future and immediately waited before submitting the next. Execution was sequential. Our separate eight-caller concurrent test passed, but the agent's tests had not demonstrated overlapping requests.&lt;/p&gt;

&lt;p&gt;Review also found that crediting two cents to a destination near SQLite's signed-integer limit converted the balance to &lt;code&gt;REAL&lt;/code&gt; while recording success. That extreme numerical boundary should be rejected with a rollback.&lt;/p&gt;

&lt;p&gt;These probes ran after grading and were not added retroactively to the forty checks. They made me focus on what tests actually cover rather than how many were added. Review found no additional acceptance-contract defect in the cache, though expiration cleanup scans the whole cache and our assessment remains limited to the small task tested.&lt;/p&gt;

&lt;h3&gt;
  
  
  Coding protocol and additional test results
&lt;/h3&gt;

&lt;p&gt;Each repository had three fixed public tests. Independent checks were prepared before the corresponding attempt; the evaluator did not repair candidate production code. The agent left protected acceptance contracts, public tests, and configurations unchanged.&lt;/p&gt;

&lt;p&gt;First attempts ran serially with medium thinking and an 8,192-token per-response output allowance. Repository file tools and a narrow allowlist of local test commands were available; network tools, subagents, and cloud fallback were disabled. Infrastructure-only aborted launches were excluded from scores. Cache and ledger durations were recovered from event records and are approximate.&lt;/p&gt;

&lt;p&gt;The evaluator subsequently ran public and agent-added tests: cache &lt;strong&gt;26/26&lt;/strong&gt;, ledger &lt;strong&gt;43/43&lt;/strong&gt;, and wallet &lt;strong&gt;14/14&lt;/strong&gt; passed. Wallet's green rerun occurred after the agent timed out. The original build planner still failed its three public tests. The thinking-disabled diagnostic passed &lt;strong&gt;47/47&lt;/strong&gt; public and agent-added tests, while retaining the independent-check failure described above.&lt;/p&gt;

&lt;p&gt;These were four deliberately bounded Python repositories. We did not measure reliability across seeds, large production repositories, Q4 versus higher-precision weights, or coding quality at 64K / 128K. No controlled model comparison was performed. The &lt;a href="https://zjshen14.github.io/experiments/qwen3.8-27b-3090/coding-summary.json" rel="noopener noreferrer"&gt;coding summary&lt;/a&gt; retains scores and review probes.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I would keep using it
&lt;/h2&gt;

&lt;p&gt;I would give it bounded fixes and small features, state the acceptance criteria first, and inspect the actual diff, independent tests, and final handoff. It demonstrated the ability to edit across files, add regression coverage, and correct some mistakes found during its own checks.&lt;/p&gt;

&lt;p&gt;For longer tasks, I would watch whether reasoning crowds out output and whether the workflow times out. Numerical boundaries, concurrency, and error handling still deserve review. The next experiment should retain the same tasks and checks, add repeated runs, and compare thinking settings and response allowances.&lt;/p&gt;

&lt;p&gt;From Ethereum mining to local coding, this 3090 has found another useful job. If similar hardware is already on hand, a small task with clear acceptance criteria is a practical way to see what it can contribute to your workflow.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reproduce the local coding setup
&lt;/h2&gt;

&lt;p&gt;Here is the deployment path we used. The &lt;a href="https://zjshen14.github.io/experiments/qwen3.8-27b-3090/reproduction-kit.zip" rel="noopener noreferrer"&gt;download kit&lt;/a&gt; contains the full configuration. Essential launch and verification steps are below, with additional settings under the headings below.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Start the model service
&lt;/h3&gt;

&lt;p&gt;On Ubuntu, prepare an NVIDIA driver, &lt;code&gt;curl&lt;/code&gt;, &lt;code&gt;unzip&lt;/code&gt;, Python 3, and a working user session, then run:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;mkdir&lt;/span&gt; &lt;span class="nt"&gt;-p&lt;/span&gt; ~/local-qwen &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;cd&lt;/span&gt; ~/local-qwen
curl &lt;span class="nt"&gt;-fL&lt;/span&gt; https://zjshen14.github.io/experiments/qwen3.8-27b-3090/reproduction-kit.zip &lt;span class="nt"&gt;-o&lt;/span&gt; kit.zip
unzip kit.zip
bash setup.sh
./run-server.sh
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;setup.sh&lt;/code&gt; downloads and verifies the GGUF and &lt;a href="https://github.com/ggml-org/llama.cpp/releases/tag/b11146" rel="noopener noreferrer"&gt;llama.cpp b11146 CUDA 12.8 release archives&lt;/a&gt;. Downloads total roughly &lt;strong&gt;17.6GB&lt;/strong&gt;, plus extracted runtime space; weights are not inside the ZIP. Prebuilt binaries avoid compiling a CUDA development toolchain. The script installs no system service and makes no global OpenCode configuration changes.&lt;/p&gt;

&lt;p&gt;Once the service is ready, verify the API from another terminal:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-f&lt;/span&gt; http://127.0.0.1:8080/health
curl &lt;span class="nt"&gt;-f&lt;/span&gt; http://127.0.0.1:8080/v1/chat/completions &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s1"&gt;'Content-Type: application/json'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"model":"qwen3.8-27b","messages":[{"role":"user","content":"Reply with OK."}],"max_tokens":64,"chat_template_kwargs":{"enable_thinking":false}}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The server listens on &lt;code&gt;127.0.0.1:8080&lt;/code&gt;, and local requests need no real API key. A client requiring a value can use a nonsecret placeholder such as &lt;code&gt;local&lt;/code&gt;, but that is not access control. Keep this unauthenticated endpoint on loopback.&lt;/p&gt;

&lt;h4&gt;
  
  
  Full experiment settings, launch options, and background service
&lt;/h4&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Component&lt;/th&gt;
&lt;th&gt;Tested configuration&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;System&lt;/td&gt;
&lt;td&gt;Ubuntu, Linux x86-64&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPU&lt;/td&gt;
&lt;td&gt;NVIDIA RTX 3090, 24GB VRAM&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CPU / RAM&lt;/td&gt;
&lt;td&gt;Ryzen 7 5800X / 64GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Weights&lt;/td&gt;
&lt;td&gt;Qwen3.8-27B, Q4_K_M, 16,810,714,464 bytes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Serving engine&lt;/td&gt;
&lt;td&gt;llama.cpp &lt;strong&gt;b11146&lt;/strong&gt;, prebuilt CUDA 12.8 runtime&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agent&lt;/td&gt;
&lt;td&gt;OpenCode &lt;strong&gt;2.0.20&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Context / per-response output limit&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;131,072 / 8,192 tokens&lt;/strong&gt;; input and output share context&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;KV cache&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;q8_0&lt;/strong&gt; for both K and V&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPU / concurrency&lt;/td&gt;
&lt;td&gt;All model layers on GPU; &lt;strong&gt;one generation slot&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Other settings&lt;/td&gt;
&lt;td&gt;Flash attention; batch 512 / microbatch 256&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Thinking for coding tasks&lt;/td&gt;
&lt;td&gt;Enabled, &lt;strong&gt;medium&lt;/strong&gt; effort&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Q4_K_M describes weight quantization; q8_0 describes KV-cache quantization. Beyond the roughly 16.8GB weight file, VRAM is needed for caches, compute buffers, and desktop applications. The kit's &lt;code&gt;run-server.sh&lt;/code&gt; supplies the full command; these are its core options:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;llama-server &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--model&lt;/span&gt; models/Qwen3.8-27B-Q4_K_M.gguf &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--alias&lt;/span&gt; qwen3.8-27b &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--host&lt;/span&gt; 127.0.0.1 &lt;span class="nt"&gt;--port&lt;/span&gt; 8080 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--ctx-size&lt;/span&gt; 131072 &lt;span class="nt"&gt;--parallel&lt;/span&gt; 1 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--n-gpu-layers&lt;/span&gt; all &lt;span class="nt"&gt;--fit&lt;/span&gt; off &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--flash-attn&lt;/span&gt; on &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--cache-type-k&lt;/span&gt; q8_0 &lt;span class="nt"&gt;--cache-type-v&lt;/span&gt; q8_0 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--batch-size&lt;/span&gt; 512 &lt;span class="nt"&gt;--ubatch-size&lt;/span&gt; 256 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--jinja&lt;/span&gt; &lt;span class="nt"&gt;--reasoning-format&lt;/span&gt; deepseek &lt;span class="nt"&gt;--reasoning-preserve&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--no-context-shift&lt;/span&gt; &lt;span class="nt"&gt;--no-mmproj&lt;/span&gt; &lt;span class="nt"&gt;--metrics&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The packaged script sets library paths and invokes the actual binary. If other applications need VRAM, stop the server and use &lt;code&gt;MODEL_CONTEXT=65536 ./run-server.sh&lt;/code&gt; for 64K capacity. All measurements in this post used 128K capacity.&lt;/p&gt;

&lt;p&gt;To survive closing the launching terminal, stop the foreground server first and run:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;mkdir&lt;/span&gt; &lt;span class="nt"&gt;-p&lt;/span&gt; logs
systemd-run &lt;span class="nt"&gt;--user&lt;/span&gt; &lt;span class="nt"&gt;--collect&lt;/span&gt; &lt;span class="nt"&gt;--unit&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;qwen-local &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--property&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"WorkingDirectory=&lt;/span&gt;&lt;span class="nv"&gt;$PWD&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--property&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"StandardOutput=append:&lt;/span&gt;&lt;span class="nv"&gt;$PWD&lt;/span&gt;&lt;span class="s2"&gt;/logs/server.log"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--property&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"StandardError=append:&lt;/span&gt;&lt;span class="nv"&gt;$PWD&lt;/span&gt;&lt;span class="s2"&gt;/logs/server.log"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$PWD&lt;/span&gt;&lt;span class="s2"&gt;/run-server.sh"&lt;/span&gt;
&lt;span class="c"&gt;# Stop: systemctl --user stop qwen-local&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is a transient user service that must be started again after reboot. Survival after logging out of the entire user session depends on user-manager configuration and was not tested here.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Configure OpenCode
&lt;/h3&gt;

&lt;p&gt;Place the kit's &lt;code&gt;opencode.json&lt;/code&gt; in your project root, or merge its provider into &lt;code&gt;~/.config/opencode/opencode.json&lt;/code&gt;. This configuration was verified with &lt;strong&gt;OpenCode 2.0.20&lt;/strong&gt;: model ID &lt;code&gt;local-qwen/qwen3.8-27b&lt;/code&gt;, API base URL &lt;code&gt;http://127.0.0.1:8080/v1&lt;/code&gt;.&lt;/p&gt;

&lt;h4&gt;
  
  
  Full OpenCode provider configuration
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"$schema"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://opencode.ai/config.json"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"local-qwen/qwen3.8-27b"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"providers"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"local-qwen"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Local Qwen"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"package"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"@opencode/ai/providers/openai-compatible"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"settings"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"baseURL"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"http://127.0.0.1:8080/v1"&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"models"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"qwen3.8-27b"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Qwen3.8 27B — 128K"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"capabilities"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
            &lt;/span&gt;&lt;span class="nl"&gt;"tools"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
            &lt;/span&gt;&lt;span class="nl"&gt;"input"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"text"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
            &lt;/span&gt;&lt;span class="nl"&gt;"output"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"text"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"limit"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"context"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;131072&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"output"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;8192&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"compatibility"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"reasoningField"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"reasoning_content"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"body"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
            &lt;/span&gt;&lt;span class="nl"&gt;"chat_template_kwargs"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
              &lt;/span&gt;&lt;span class="nl"&gt;"enable_thinking"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
              &lt;/span&gt;&lt;span class="nl"&gt;"reasoning_effort"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"medium"&lt;/span&gt;&lt;span class="w"&gt;
            &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The tested configuration uses &lt;code&gt;providers&lt;/code&gt; / &lt;code&gt;package&lt;/code&gt; / &lt;code&gt;settings&lt;/code&gt;; the &lt;a href="https://opencode.ai/docs/providers/#custom-provider" rel="noopener noreferrer"&gt;online provider guide&lt;/a&gt; has shown &lt;code&gt;provider&lt;/code&gt; / &lt;code&gt;npm&lt;/code&gt; / &lt;code&gt;options&lt;/code&gt;. If your installed version differs, follow its actual schema and avoid combining structures.&lt;/p&gt;

&lt;p&gt;Start the model service, then restart OpenCode. Our installation also reloaded the configuration with &lt;code&gt;opencode reload&lt;/code&gt;. Existing sessions may retain their previous model; choose &lt;code&gt;local-qwen/qwen3.8-27b&lt;/code&gt; through &lt;code&gt;/models&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;For the previous post's Web workflow, launch from your project and choose a different port:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;OPENCODE_SERVER_PASSWORD&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'replace-with-a-strong-password'&lt;/span&gt;
opencode web &lt;span class="nt"&gt;--hostname&lt;/span&gt; 127.0.0.1 &lt;span class="nt"&gt;--port&lt;/span&gt; 4096
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;8080 serves the model; 4096 serves OpenCode.&lt;/strong&gt; On the same machine, the provider uses loopback while remote browsers connect to OpenCode. If OpenCode runs on another host or inside a container, &lt;code&gt;127.0.0.1&lt;/code&gt; points to that environment and needs replacing with a private connection to the GPU host. See the &lt;a href="https://zjshen14.github.io/en/blog/setup-opencode-remote-web-ide/" rel="noopener noreferrer"&gt;previous post&lt;/a&gt; for network binding and authentication.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Verify tools, then try a task
&lt;/h3&gt;

&lt;p&gt;Check an ordinary reply first, then create a harmless fixture and ask the agent to &lt;strong&gt;use its &lt;code&gt;read&lt;/code&gt; tool&lt;/strong&gt; and return the contents. Both checks passed here, along with synthetic tool-call round trips and streamed tool calls at the server.&lt;/p&gt;

&lt;p&gt;To repeat throughput measurements, keep the server idle and choose a new output directory:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python3 benchmark_throughput.py &lt;span class="nt"&gt;--report-dir&lt;/span&gt; reports/throughput-new-run
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then try a bounded task of your own and validate it with independent tests and review.&lt;/p&gt;

&lt;h2&gt;
  
  
  Configuration, data, and reproduction materials
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://zjshen14.github.io/experiments/qwen3.8-27b-3090/reproduction-kit.zip" rel="noopener noreferrer"&gt;Configuration and measurement ZIP&lt;/a&gt;: verified downloads, full launcher, OpenCode configuration, throughput script, and data. It includes neither weights nor the complete coding-task fixtures.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://zjshen14.github.io/experiments/qwen3.8-27b-3090/README.txt" rel="noopener noreferrer"&gt;Experiment notes&lt;/a&gt;, &lt;a href="https://zjshen14.github.io/experiments/qwen3.8-27b-3090/throughput-results.json" rel="noopener noreferrer"&gt;throughput measurements&lt;/a&gt;, and &lt;a href="https://zjshen14.github.io/experiments/qwen3.8-27b-3090/coding-summary.json" rel="noopener noreferrer"&gt;coding summary&lt;/a&gt;. Experiment dates, absolute timestamps, timezones, personal paths, and session identifiers were removed from attachments; measurements, scores, and review probes are retained.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://huggingface.co/Qwen/Qwen3.8-27B" rel="noopener noreferrer"&gt;Official Qwen model card&lt;/a&gt;, &lt;a href="https://github.com/ggml-org/llama.cpp/blob/b11146/tools/server/README.md" rel="noopener noreferrer"&gt;pinned llama-server documentation&lt;/a&gt;, and &lt;a href="https://opencode.ai/docs/providers/" rel="noopener noreferrer"&gt;OpenCode provider documentation&lt;/a&gt;. Upstream sources provide selection context, reported benchmarks, and deployment guidance; this host's performance and four-task coding results come from the attached experiment records.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Measurements and review findings are supported by the linked experiment records.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>opensource</category>
      <category>tutorial</category>
    </item>
  </channel>
</rss>
