<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Ajay Vishwakarma</title>
    <description>The latest articles on DEV Community by Ajay Vishwakarma (@ajayvish01).</description>
    <link>https://dev.to/ajayvish01</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4159066%2F5b9aa6de-8161-4fef-988d-752790f0e876.jpeg</url>
      <title>DEV Community: Ajay Vishwakarma</title>
      <link>https://dev.to/ajayvish01</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/ajayvish01"/>
    <language>en</language>
    <item>
      <title>"Stop swapping SSH keys for two GitLab accounts. Use host aliases"</title>
      <dc:creator>Ajay Vishwakarma</dc:creator>
      <pubDate>Wed, 07 Oct 2026 09:30:11 +0000</pubDate>
      <link>https://dev.to/ajayvish01/stop-swapping-ssh-keys-for-two-gitlab-accounts-use-host-aliases-5f1h</link>
      <guid>https://dev.to/ajayvish01/stop-swapping-ssh-keys-for-two-gitlab-accounts-use-host-aliases-5f1h</guid>
      <description>&lt;p&gt;Two GitLab accounts on one laptop, and &lt;code&gt;git push&lt;/code&gt; to the work repo fails with &lt;code&gt;Permission denied (publickey)&lt;/code&gt;, or authenticates as the wrong account.&lt;/p&gt;

&lt;p&gt;The usual fix is swapping keys by hand: rename &lt;code&gt;~/.ssh/id_rsa&lt;/code&gt;, or re-add the other key before every push. It works until you forget.&lt;/p&gt;

&lt;p&gt;The rule: &lt;strong&gt;the repository's remote URL should select the SSH identity. One host alias per account, one key per alias, nothing to switch by hand.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Scope: gitlab.com, one work and one personal account, Windows and macOS. Not CI keys or key rotation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why does Git use the wrong GitLab account?
&lt;/h2&gt;

&lt;p&gt;Both accounts live on the same server, so both remotes look the same to SSH:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;git@gitlab.com:company/project.git
git@gitlab.com:username/project.git
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Nothing in either URL says which account you mean. SSH connects to &lt;code&gt;gitlab.com&lt;/code&gt; as user &lt;code&gt;git&lt;/code&gt; both times. If your agent holds both keys, SSH can offer more than one, and GitLab authenticates whichever it accepts first.&lt;/p&gt;

&lt;p&gt;The model behind key swapping is &lt;strong&gt;"SSH has one identity, and I change it before I push."&lt;/strong&gt; That turns the active account into hidden global state. Nothing in the repository records which account it belongs to, so every push depends on what you did last.&lt;/p&gt;

&lt;h2&gt;
  
  
  What should choose the key instead?
&lt;/h2&gt;

&lt;p&gt;Treat it as routing, not switching. The repository names an alias, the alias names a key, and the key belongs to one account.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu6tr065kpyfs5oqt7qv2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu6tr065kpyfs5oqt7qv2.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;git pull
  ↓
origin = git@gitlab-work:company/project.git
  ↓
SSH config: Host gitlab-work
  ↓
IdentityFile ~/.ssh/id_gitlab_work
  ↓
local proof made with the private key
  ↓
GitLab verifies it with the public key on the work account
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;gitlab-work&lt;/code&gt; is not a DNS hostname. It is a name in your SSH config that maps to &lt;code&gt;gitlab.com&lt;/code&gt; plus one specific key.&lt;/p&gt;

&lt;p&gt;Two facts keep this model honest:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The private key never leaves your machine.&lt;/strong&gt; SSH uses it locally to produce a proof (a signature). GitLab checks that proof with the public key you registered.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;ssh-agent&lt;/code&gt; is optional.&lt;/strong&gt; It keeps an unlocked key available locally so you don't retype the passphrase. It does not send the key anywhere.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq9o33ghmzgglhl3r9bet.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq9o33ghmzgglhl3r9bet.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  How do you set up one key and one alias per account?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Generate two keys and load them.&lt;/strong&gt; Give each a passphrase. The &lt;code&gt;-C&lt;/code&gt; comment is only a label; it does not decide which account gets the key.&lt;/p&gt;

&lt;p&gt;Windows (PowerShell; the two service lines need an elevated session):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight powershell"&gt;&lt;code&gt;&lt;span class="n"&gt;New-Item&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-ItemType&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;Directory&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-Force&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-Path&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$&lt;/span&gt;&lt;span class="nn"&gt;env&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="nv"&gt;USERPROFILE&lt;/span&gt;&lt;span class="s2"&gt;\.ssh"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="n"&gt;ssh-keygen&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-t&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;ed25519&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-C&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"work@example.com"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-f&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$&lt;/span&gt;&lt;span class="nn"&gt;env&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="nv"&gt;USERPROFILE&lt;/span&gt;&lt;span class="s2"&gt;\.ssh\id_gitlab_work"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="n"&gt;ssh-keygen&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-t&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;ed25519&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-C&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"personal@example.com"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-f&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$&lt;/span&gt;&lt;span class="nn"&gt;env&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="nv"&gt;USERPROFILE&lt;/span&gt;&lt;span class="s2"&gt;\.ssh\id_gitlab_personal"&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;span class="n"&gt;Get-Service&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;ssh-agent&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;Set-Service&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-StartupType&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;Automatic&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="n"&gt;Start-Service&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;ssh-agent&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;span class="n"&gt;ssh-add&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$&lt;/span&gt;&lt;span class="nn"&gt;env&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="nv"&gt;USERPROFILE&lt;/span&gt;&lt;span class="s2"&gt;\.ssh\id_gitlab_work"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="n"&gt;ssh-add&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$&lt;/span&gt;&lt;span class="nn"&gt;env&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="nv"&gt;USERPROFILE&lt;/span&gt;&lt;span class="s2"&gt;\.ssh\id_gitlab_personal"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="n"&gt;ssh-add&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-l&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;macOS:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ssh-keygen &lt;span class="nt"&gt;-t&lt;/span&gt; ed25519 &lt;span class="nt"&gt;-C&lt;/span&gt; &lt;span class="s2"&gt;"work@example.com"&lt;/span&gt; &lt;span class="nt"&gt;-f&lt;/span&gt; ~/.ssh/id_gitlab_work
ssh-keygen &lt;span class="nt"&gt;-t&lt;/span&gt; ed25519 &lt;span class="nt"&gt;-C&lt;/span&gt; &lt;span class="s2"&gt;"personal@example.com"&lt;/span&gt; &lt;span class="nt"&gt;-f&lt;/span&gt; ~/.ssh/id_gitlab_personal

ssh-add &lt;span class="nt"&gt;--apple-use-keychain&lt;/span&gt; ~/.ssh/id_gitlab_work
ssh-add &lt;span class="nt"&gt;--apple-use-keychain&lt;/span&gt; ~/.ssh/id_gitlab_personal
ssh-add &lt;span class="nt"&gt;-l&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A running agent and a loaded key are different states. Trust &lt;code&gt;ssh-add -l&lt;/code&gt;, not the service status.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Register each public key with the matching account.&lt;/strong&gt; &lt;code&gt;id_gitlab_work.pub&lt;/code&gt; goes into the work account's SSH key settings, &lt;code&gt;id_gitlab_personal.pub&lt;/code&gt; into the personal one. Never the file without &lt;code&gt;.pub&lt;/code&gt;. (&lt;a href="https://docs.gitlab.com/user/ssh/" rel="noopener noreferrer"&gt;GitLab SSH docs&lt;/a&gt;)&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Write &lt;code&gt;~/.ssh/config&lt;/code&gt;.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Host gitlab-work
    HostName gitlab.com
    User git
    IdentityFile ~/.ssh/id_gitlab_work
    IdentitiesOnly yes

Host gitlab-personal
    HostName gitlab.com
    User git
    IdentityFile ~/.ssh/id_gitlab_personal
    IdentitiesOnly yes
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;Host&lt;/code&gt; is the alias you will put in the remote URL.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;HostName&lt;/code&gt; is the real server.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;User git&lt;/code&gt; is the SSH user GitLab expects.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;IdentityFile&lt;/code&gt; is the private key for this alias.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;IdentitiesOnly yes&lt;/code&gt; tells SSH to use the identity configured here instead of offering every key the agent holds. It does not turn the agent off.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;On macOS, add &lt;code&gt;AddKeysToAgent yes&lt;/code&gt; and &lt;code&gt;UseKeychain yes&lt;/code&gt; under each &lt;code&gt;Host&lt;/code&gt; for Keychain integration. Those two lines are macOS-only.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Point each repository at its alias.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git remote set-url origin git@gitlab-work:company/project.git
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Personal repositories use &lt;code&gt;git@gitlab-personal:username/project.git&lt;/code&gt;. From here on, the remote is part of the authentication design, not just metadata.&lt;/p&gt;

&lt;h2&gt;
  
  
  How do you prove each layer before the first push?
&lt;/h2&gt;

&lt;p&gt;Test SSH directly before you debug Git. If &lt;code&gt;ssh -T&lt;/code&gt; fails, Git is not the problem yet.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Command&lt;/th&gt;
&lt;th&gt;Expect&lt;/th&gt;
&lt;th&gt;What it proves&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;ssh -G gitlab-work&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;hostname gitlab.com&lt;/code&gt;, &lt;code&gt;user git&lt;/code&gt;, your work &lt;code&gt;identityfile&lt;/code&gt;, &lt;code&gt;identitiesonly yes&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;SSH reads the alias the way you meant&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;ssh -T git@gitlab-work&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;GitLab's welcome message for your work username&lt;/td&gt;
&lt;td&gt;The work key authenticates as the work account&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;ssh -T git@gitlab-personal&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;The welcome message for your personal username&lt;/td&gt;
&lt;td&gt;The personal key authenticates as the personal account&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;git remote -v&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;git@gitlab-work:...&lt;/code&gt; for fetch and push&lt;/td&gt;
&lt;td&gt;This repository routes through the right alias&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;git fetch&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;No error, no prompt for the other key&lt;/td&gt;
&lt;td&gt;The whole chain works through Git&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;On the first connection, check the host key fingerprint against &lt;a href="https://docs.gitlab.com/user/gitlab_com/" rel="noopener noreferrer"&gt;GitLab's published fingerprints&lt;/a&gt; before you type &lt;code&gt;yes&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why can &lt;code&gt;ssh -T&lt;/code&gt; pass while Git still prompts or fails?
&lt;/h2&gt;

&lt;p&gt;This one cost me the most time, on Windows.&lt;/p&gt;

&lt;p&gt;My machine had MSYS2's OpenSSH installed next to Windows OpenSSH. Both keys were loaded in the Windows agent, and I had already cleaned up PATH. Git still kept asking for the key passphrase.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr809avv60vzq31500aze.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr809avv60vzq31500aze.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The lesson: &lt;strong&gt;the &lt;code&gt;ssh&lt;/code&gt; your terminal resolves is not proof of which SSH command Git runs.&lt;/strong&gt; Check both:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight powershell"&gt;&lt;code&gt;&lt;span class="n"&gt;Get-Command&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;ssh&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="n"&gt;where.exe&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;ssh&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="n"&gt;git&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;config&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;--show-origin&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;--get&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;core.sshCommand&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="nv"&gt;$&lt;/span&gt;&lt;span class="nn"&gt;env&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="nv"&gt;GIT_SSH_COMMAND&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On macOS the first two are &lt;code&gt;command -v ssh&lt;/code&gt; and &lt;code&gt;which -a ssh&lt;/code&gt;. If &lt;code&gt;GIT_SSH_COMMAND&lt;/code&gt; is set, it overrides &lt;code&gt;core.sshCommand&lt;/code&gt; (&lt;a href="https://git-scm.com/docs/git-config" rel="noopener noreferrer"&gt;git-config&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;What fixed it on that machine was telling Git exactly which client to launch:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight powershell"&gt;&lt;code&gt;&lt;span class="n"&gt;git&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;config&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;--global&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;core.sshCommand&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"C:/Windows/System32/OpenSSH/ssh.exe"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After that, Git used the same OpenSSH as the Windows agent, and the prompts stopped. This setting does not change the protocol, move keys, or replace your SSH config. It only makes Git's choice of SSH executable explicit.&lt;/p&gt;

&lt;p&gt;That is one machine, not a rule. If your default setup already works, leave it alone.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where else does this break?
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;An old remote.&lt;/strong&gt; A repository still on &lt;code&gt;git@gitlab.com:...&lt;/code&gt; bypasses both aliases. Run &lt;code&gt;git remote -v&lt;/code&gt; first.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The IDE terminal.&lt;/strong&gt; An IDE can inherit a different PATH and different agent access than your system terminal. Run the same checks inside it before blaming the config.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Commit author.&lt;/strong&gt; The SSH identity decides which account authenticates. &lt;code&gt;user.name&lt;/code&gt; and &lt;code&gt;user.email&lt;/code&gt; decide what gets written into commits. Fixing one does nothing for the other.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;macOS.&lt;/strong&gt; I ran the Windows path on my own machine. The macOS steps are checked against current documentation, not run by me on a Mac. If &lt;code&gt;--apple-use-keychain&lt;/code&gt; behaves differently on your version, read &lt;code&gt;ssh-add&lt;/code&gt;'s own help.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Shortcuts that make it worse.&lt;/strong&gt; &lt;code&gt;StrictHostKeyChecking=no&lt;/code&gt; and a passphrase stored in &lt;code&gt;.env&lt;/code&gt; both make the prompt go away by removing the protection.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What should you check in the next 20 minutes?
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv5s47se0wvbenr0l8ooo.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv5s47se0wvbenr0l8ooo.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Run &lt;code&gt;git remote -v&lt;/code&gt; in one work repo and one personal repo. Is anything still on &lt;code&gt;git@gitlab.com:&lt;/code&gt;?&lt;/li&gt;
&lt;li&gt;Run &lt;code&gt;ssh -G gitlab-work&lt;/code&gt;. Do &lt;code&gt;identityfile&lt;/code&gt; and &lt;code&gt;identitiesonly&lt;/code&gt; say what you intended?&lt;/li&gt;
&lt;li&gt;Run &lt;code&gt;ssh -T&lt;/code&gt; against each alias. Does each one greet the right username?&lt;/li&gt;
&lt;li&gt;Run &lt;code&gt;git config --show-origin --get core.sshCommand&lt;/code&gt; and check &lt;code&gt;GIT_SSH_COMMAND&lt;/code&gt;. Do you know which &lt;code&gt;ssh&lt;/code&gt; Git launches?&lt;/li&gt;
&lt;li&gt;Run &lt;code&gt;git config user.email&lt;/code&gt; in each repo. Does the commit identity match the account?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;em&gt;Next: debugging &lt;code&gt;Permission denied (publickey)&lt;/code&gt; one layer at a time.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Has Git ever run a different &lt;code&gt;ssh&lt;/code&gt; than your terminal did? What was installed, and how did you catch it?&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>git</category>
      <category>gitlab</category>
      <category>security</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>"LLM streaming works in the demo. These 4 hops break it in prod"</title>
      <dc:creator>Ajay Vishwakarma</dc:creator>
      <pubDate>Wed, 07 Oct 2026 06:58:28 +0000</pubDate>
      <link>https://dev.to/ajayvish01/llm-streaming-works-in-the-demo-these-4-hops-break-it-in-prod-50pp</link>
      <guid>https://dev.to/ajayvish01/llm-streaming-works-in-the-demo-these-4-hops-break-it-in-prod-50pp</guid>
      <description>&lt;p&gt;Words appear one by one on localhost, and LLM streaming looks done. In production the answer lands in one lump, the model keeps generating after the user leaves, and half a sentence shows up as the full answer.&lt;/p&gt;

&lt;p&gt;None of these are LLM problems. They live in the plumbing between the model and the browser.&lt;/p&gt;

&lt;p&gt;The rule: &lt;strong&gt;your server is not a pipe. It sits between two connections, and either one can misbehave.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Scope: text chat over &lt;code&gt;fetch&lt;/code&gt;, examples in Node/Express. The ideas carry over to other stacks.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where does a streamed answer actually break?
&lt;/h2&gt;

&lt;p&gt;Here's the flow:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Browser  ──POST /chat──►  Your API  ──stream:true──►  LLM provider
   ▲                          │                            │
   └──── tokens as events ◄───┴──────── tokens ◄───────────┘
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The 30-line demo teaches the pipe model: tokens go in one side and come out the other. That took me a while to unlearn.&lt;/p&gt;

&lt;p&gt;Your server sits in the middle of two connections. Something in front of it can buffer. The client behind it can leave. The provider underneath it can fail. And the browser can misread what arrives.&lt;/p&gt;

&lt;p&gt;![Four hops where a streamed LLM answer can break: browser, proxy, API, provider]&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fiorowlymaqenlw6730yu.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fiorowlymaqenlw6730yu.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Most streaming bugs are one of these four hops going wrong. The sections below take them in order.&lt;/p&gt;

&lt;p&gt;I use &lt;code&gt;fetch&lt;/code&gt; with a streamed response for all of this. A chat request needs a message history and an auth token, and &lt;code&gt;fetch&lt;/code&gt; gives me a POST body, headers, and abort support. The alternatives are near the end.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why does the answer arrive all at once in production?
&lt;/h2&gt;

&lt;p&gt;Reverse proxies like Nginx &lt;strong&gt;collect the upstream response and send it in bigger chunks&lt;/strong&gt; by default. That's great for normal pages and terrible for a stream. Your tokens sit in a buffer, and the user sees nothing until it fills or the response ends.&lt;/p&gt;

&lt;p&gt;I've seen this called out as the most frequent production surprise with streaming, and it matches what the infrastructure docs say.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;In your app&lt;/strong&gt;, send headers that tell proxies to back off:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;writeHead&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Content-Type&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;text/event-stream&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Cache-Control&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;no-cache, no-transform&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Connection&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;keep-alive&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;X-Accel-Buffering&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;no&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;// Nginx-specific: don't buffer this response&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;In Nginx&lt;/strong&gt;, for the streaming route:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight nginx"&gt;&lt;code&gt;&lt;span class="k"&gt;location&lt;/span&gt; &lt;span class="n"&gt;/api/chat&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kn"&gt;proxy_pass&lt;/span&gt; &lt;span class="s"&gt;http://app&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;proxy_http_version&lt;/span&gt; &lt;span class="mf"&gt;1.1&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;proxy_set_header&lt;/span&gt; &lt;span class="s"&gt;Connection&lt;/span&gt; &lt;span class="s"&gt;""&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;proxy_buffering&lt;/span&gt; &lt;span class="no"&gt;off&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;      &lt;span class="c1"&gt;# the important one&lt;/span&gt;
    &lt;span class="kn"&gt;gzip&lt;/span&gt; &lt;span class="no"&gt;off&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;                 &lt;span class="c1"&gt;# compression can re-buffer the stream&lt;/span&gt;
    &lt;span class="kn"&gt;proxy_read_timeout&lt;/span&gt; &lt;span class="s"&gt;300s&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;  &lt;span class="c1"&gt;# longer than your slowest answer&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A CDN, load balancer, or API gateway in front of you may have its own buffering and idle timeout settings. I can't tell you the right config for each one, so check your provider's docs. What works everywhere is a test:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-N&lt;/span&gt; &lt;span class="nt"&gt;-X&lt;/span&gt; POST https://your-domain.com/api/chat &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"messages":[{"role":"user","content":"count slowly from 1 to 20"}]}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;-N&lt;/code&gt; turns off curl's own buffering. If lines appear gradually, the whole path streams. If they arrive in one lump, something is buffering. Run it against the &lt;strong&gt;real production URL&lt;/strong&gt;, and again after any infrastructure change.&lt;/p&gt;

&lt;p&gt;The same hop kills quiet streams. When the model is thinking or a tool call takes several seconds, some proxies and load balancers see silence and decide the connection is dead. Raise timeouts above your slowest realistic answer on every layer, and send a heartbeat. In SSE, a line starting with &lt;code&gt;:&lt;/code&gt; is a comment that clients ignore:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;heartbeat&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;setInterval&lt;/span&gt;&lt;span class="p"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;write&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;: ping&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="mi"&gt;15000&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;on&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;close&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;clearInterval&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;heartbeat&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If your answer involves tools, also stream status events like &lt;code&gt;{"status": "searching documents"}&lt;/code&gt;. Ten seconds of silence feels broken. Ten seconds with "Looking that up…" feels normal.&lt;/p&gt;

&lt;p&gt;Deploys count too. A rolling restart can kill streams mid-answer, so give your servers a graceful shutdown period at least as long as your longest answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  What happens when the user closes the tab?
&lt;/h2&gt;

&lt;p&gt;This is the hop I think is most underrated.&lt;/p&gt;

&lt;p&gt;A user asks a question, sees it start answering, and closes the tab. In a lot of apps the server does nothing. The loop reading from the LLM keeps going, the provider keeps generating, and you pay for output nobody sees. A Stop button that only hides text in the UI is the same problem with a friendlier face.&lt;/p&gt;

&lt;p&gt;A real cancel has to travel the whole chain:&lt;/p&gt;

&lt;p&gt;![Without cancellation the model keeps generating after the user leaves; with cancellation one abort signal stops it end to end]&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2zpbm6rw0zdji5ulmhyr.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2zpbm6rw0zdji5ulmhyr.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Browser:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;controller&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;AbortController&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;/api/chat&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;method&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;POST&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Content-Type&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;application/json&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="na"&gt;body&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;messages&lt;/span&gt; &lt;span class="p"&gt;}),&lt;/span&gt;
  &lt;span class="na"&gt;signal&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;controller&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;signal&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="c1"&gt;// Stop button:&lt;/span&gt;
&lt;span class="nx"&gt;stopButton&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;onclick&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;controller&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;abort&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Server:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;/api/chat&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;upstream&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;AbortController&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

  &lt;span class="c1"&gt;// If the client leaves before we finish, cancel the LLM call.&lt;/span&gt;
  &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;on&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;close&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;writableEnded&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nx"&gt;upstream&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;abort&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;

  &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;writeHead&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Content-Type&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;text/event-stream&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="cm"&gt;/* + headers above */&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;

  &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;stream&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;llm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
      &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;stream&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
      &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;signal&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;upstream&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;signal&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;);&lt;/span&gt;

    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="k"&gt;await &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;chunk&lt;/span&gt; &lt;span class="k"&gt;of&lt;/span&gt; &lt;span class="nx"&gt;stream&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;text&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]?.&lt;/span&gt;&lt;span class="nx"&gt;delta&lt;/span&gt;&lt;span class="p"&gt;?.&lt;/span&gt;&lt;span class="nx"&gt;content&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
      &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;write&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`data: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;text&lt;/span&gt; &lt;span class="p"&gt;})}&lt;/span&gt;&lt;span class="s2"&gt;\n\n`&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;write&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`data: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;done&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt; &lt;span class="p"&gt;})}&lt;/span&gt;&lt;span class="s2"&gt;\n\n`&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;err&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;upstream&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;signal&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;aborted&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;write&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`data: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;error&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;generation_failed&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;})}&lt;/span&gt;&lt;span class="s2"&gt;\n\n`&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;finally&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;end&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Choices I'd make again:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Listen on the response's &lt;code&gt;close&lt;/code&gt; event, and check &lt;code&gt;writableEnded&lt;/code&gt;.&lt;/strong&gt; That separates "the client left early" from "we finished normally." In recent Node versions, I wouldn't rely on the request's &lt;code&gt;close&lt;/code&gt; event for this.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pass the same signal to everything the request started.&lt;/strong&gt; If the question also triggered a vector search, a reranker, or tool calls, they should stop too. Otherwise you cancel the visible part and keep paying for the invisible part.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Don't assume the provider stops billing immediately.&lt;/strong&gt; Aborting the connection usually stops generation, but behavior varies by provider. I'd test it with a short script and check the usage numbers, instead of trusting the docs alone.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Decide what to do with partial answers.&lt;/strong&gt; If a user stops mid-answer, do you save what was generated? Choose on purpose.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  How do you report an error after you've already sent 200 OK?
&lt;/h2&gt;

&lt;p&gt;Normally errors are status codes: 400, 429, 500. In a stream, &lt;strong&gt;you've already sent &lt;code&gt;200 OK&lt;/code&gt; and some tokens before the problem happens.&lt;/strong&gt; The provider times out, you hit a rate limit, or a tool call throws. The status code is already gone.&lt;/p&gt;

&lt;p&gt;So errors have to travel &lt;strong&gt;inside the stream&lt;/strong&gt; as events. Two rules I follow, both in the server code above:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Send an explicit error event&lt;/strong&gt;, so the UI can say something useful.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Always send a final &lt;code&gt;done&lt;/code&gt; event on success.&lt;/strong&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The second one is easy to skip and it matters. Without &lt;code&gt;done&lt;/code&gt;, the client can't tell "complete" from "the connection dropped after the third sentence." Users will read the truncated answer as the real answer.&lt;/p&gt;

&lt;p&gt;In the UI I treat these as three separate end states:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Complete:&lt;/strong&gt; got &lt;code&gt;done&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Failed:&lt;/strong&gt; got &lt;code&gt;error&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Interrupted:&lt;/strong&gt; the stream ended without either (show "response was cut off, try again")&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I'd also avoid silently auto-retrying a half-finished generation. A retry costs a second generation, and it might give a different answer from the one the user already half-read.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why does the browser show broken or half-parsed text?
&lt;/h2&gt;

&lt;p&gt;The client code looks simple. Two bugs got me.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Bug 1: network chunks don't line up with your events.&lt;/strong&gt; One read can contain half an event, or three and a half. If you parse each chunk directly, you'll eventually hit a &lt;code&gt;JSON.parse&lt;/code&gt; error on a partial line. Keep a buffer, split on the blank line that ends each event, and carry the leftover:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;reader&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getReader&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;decoder&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;TextDecoder&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;buffer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="dl"&gt;""&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;while &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;value&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;done&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;reader&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;done&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;break&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

  &lt;span class="c1"&gt;// stream: true keeps multi-byte characters intact across chunks&lt;/span&gt;
  &lt;span class="nx"&gt;buffer&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="nx"&gt;decoder&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;decode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;value&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;stream&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;

  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;events&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;buffer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="nx"&gt;buffer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;events&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;pop&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt; &lt;span class="c1"&gt;// the last piece may be incomplete, keep it&lt;/span&gt;

  &lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;evt&lt;/span&gt; &lt;span class="k"&gt;of&lt;/span&gt; &lt;span class="nx"&gt;events&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;evt&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;startsWith&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;data: &lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="k"&gt;continue&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="c1"&gt;// skips heartbeat comments&lt;/span&gt;
    &lt;span class="nf"&gt;handle&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;parse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;evt&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;slice&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;6&lt;/span&gt;&lt;span class="p"&gt;)));&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Bug 2: broken characters.&lt;/strong&gt; An emoji or a non-English character can be several bytes, and a chunk can end in the middle of one. Decoding each chunk separately gives you &lt;code&gt;�&lt;/code&gt;. The &lt;code&gt;{ stream: true }&lt;/code&gt; option on &lt;code&gt;TextDecoder&lt;/code&gt; fixes it. This one is nasty because it passes every English-only test.&lt;/p&gt;

&lt;p&gt;Rendering has the same trap. If you re-render the whole message on every token, especially if you re-parse Markdown each time, the page starts to stutter. What helps:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Batch updates.&lt;/strong&gt; Update the screen once per animation frame, not once per token.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Treat the streaming message as append-only.&lt;/strong&gt; Avoid rebuilding earlier content.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Plan for half-finished Markdown.&lt;/strong&gt; A code fence that hasn't closed yet can make a renderer flicker.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Respect the user's scroll.&lt;/strong&gt; If they scrolled up to read, don't pull them back down on every token.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And measure the right thing. Users will wait through a long answer if &lt;strong&gt;the first token arrives quickly&lt;/strong&gt;. Time to first token is worth logging more than total duration.&lt;/p&gt;

&lt;h2&gt;
  
  
  When is this the wrong approach?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;When the transport doesn't fit.&lt;/strong&gt; You have three realistic options:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Option&lt;/th&gt;
&lt;th&gt;Good for&lt;/th&gt;
&lt;th&gt;Catch&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;SSE via &lt;code&gt;EventSource&lt;/code&gt;&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Simple server-to-browser streams&lt;/td&gt;
&lt;td&gt;GET only, no custom headers, auto-reconnects by itself&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Streamed response via &lt;code&gt;fetch&lt;/code&gt;&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Chat: POST body, auth headers, abort support&lt;/td&gt;
&lt;td&gt;You parse the stream yourself&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;WebSocket&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;True two-way, low latency (voice, live collaboration)&lt;/td&gt;
&lt;td&gt;More infrastructure, harder to scale and debug&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;code&gt;EventSource&lt;/code&gt; reconnecting on its own sounds helpful. For LLMs it can mean &lt;strong&gt;a second paid generation for the same question&lt;/strong&gt; that nobody planned for. If you use it, make retries something you decide.&lt;/p&gt;

&lt;p&gt;I'd reach for WebSockets when audio or real two-way traffic is involved. I've used one for a realtime voice assistant, and it's the right tool there. "The user interrupted" is just cancellation with a microphone.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;When you shouldn't stream at all.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Structured output&lt;/strong&gt; (JSON, function arguments) is useless half-finished. Wait and validate the whole thing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Server-to-server calls&lt;/strong&gt; gain nothing from streaming.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Answers you need to check before showing&lt;/strong&gt; (safety filters, "does this match the tool results?") can't be checked until they're done. Streaming and validation pull in opposite directions, and you have to pick which matters more.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Where my own experience runs out.&lt;/strong&gt; I haven't run streaming at massive scale. Three things I'm still figuring out:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Resumable streams.&lt;/strong&gt; If a user refreshes mid-answer, can they pick up where they left off? I've read about approaches that write chunks to something like Redis Streams so a reconnect can resume, but I haven't built this myself yet.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Scaling beyond one server.&lt;/strong&gt; Long-lived connections behave differently under load balancing and autoscaling, and I'm still learning where the limits are.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Backpressure.&lt;/strong&gt; If the client reads slower than you write, buffers grow on your side. I know the concept, but I haven't pushed it hard enough in practice to have strong opinions.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What should you check before you ship?
&lt;/h2&gt;

&lt;p&gt;One check per hop, plus one metric:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Proxy:&lt;/strong&gt; run the &lt;code&gt;curl -N&lt;/code&gt; command against your real production URL. Do lines arrive gradually, or in one lump?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Your API:&lt;/strong&gt; start a long answer, close the tab, and watch your server logs. Does the LLM call stop? Do the tools and searches stop with it?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Provider:&lt;/strong&gt; force an error mid-answer. Does the UI say "failed" or "cut off", or does it show half an answer as if it were complete?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Browser:&lt;/strong&gt; ask for an answer full of emoji and non-English text. Any &lt;code&gt;�&lt;/code&gt;?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Metric:&lt;/strong&gt; log time to first token, not just total duration.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;em&gt;Next post: keeping LLM output trustworthy when you can't validate it before it streams, covering structured outputs, retries, and what to do with almost-valid JSON.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Which of these four hops broke first for you in production, and what was sitting in front of your server when it did?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Four things to do before you publish:&lt;/p&gt;

&lt;p&gt;Images: upload where-streams-break.png and cancellation-chain.png in the DEV editor and replace the two REPLACE-WITH-DEV-URL placeholders. Upload the cover separately at 1000×420; your current cover carries the old title.&lt;br&gt;
Front matter: the block between the --- lines sets title, description and tags in DEV's Markdown editor. If your editor shows separate Title and Tags boxes instead, delete that block and type those values into the boxes.&lt;br&gt;
Draft flag: it is set to published: false so you can preview first. Flip it to true when you publish.&lt;br&gt;
AI disclosure: set the field to AI-Assisted.&lt;/p&gt;

&lt;p&gt;The same text is in dev-llm-streaming-full.md. This is the single-post route; the shorter cancellation-only post from earlier is still the alternative if you decide to run it as a series.&lt;/p&gt;

&lt;p&gt;Dev llm streaming full&lt;br&gt;
Document·MD&amp;nbsp;&lt;/p&gt;

&lt;p&gt;Lastly,&lt;br&gt;
Convert this whole text into a single copy-paste format so that I can just copy and paste it onto Dev Community and follow the Dev Community copy-paste algorithm. The formatting should be structured, according to Dev Community.&lt;/p&gt;

&lt;p&gt;Here is the whole post as one block in DEV's Markdown format: front matter on top (title, description, four tags), then the body starting at H2 with language-tagged code blocks. I removed the two image lines so nothing in it needs editing before you paste.&lt;/p&gt;
&lt;h2&gt;
  
  
  markdown
&lt;/h2&gt;

&lt;p&gt;title: "LLM streaming works in the demo. These 4 hops break it in prod"&lt;br&gt;
published: false&lt;br&gt;
description: "Your server is not a pipe. It sits between two connections, and four hops can fail: proxy buffering, cancellation, mid-stream errors, client parsing."&lt;/p&gt;
&lt;h2&gt;
  
  
  tags: ai, node, backend, architecture
&lt;/h2&gt;

&lt;p&gt;Words appear one by one on localhost, and LLM streaming looks done. In production the answer lands in one lump, the model keeps generating after the user leaves, and half a sentence shows up as the full answer.&lt;/p&gt;

&lt;p&gt;None of these are LLM problems. They live in the plumbing between the model and the browser.&lt;/p&gt;

&lt;p&gt;The rule: &lt;strong&gt;your server is not a pipe. It sits between two connections, and either one can misbehave.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Scope: text chat over &lt;code&gt;fetch&lt;/code&gt;, examples in Node/Express. The ideas carry over to other stacks.&lt;/p&gt;
&lt;h2&gt;
  
  
  Where does a streamed answer actually break?
&lt;/h2&gt;

&lt;p&gt;Here's the flow:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Browser  ──POST /chat──►  Your API  ──stream:true──►  LLM provider
   ▲                          │                            │
   └──── tokens as events ◄───┴──────── tokens ◄───────────┘
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The 30-line demo teaches the pipe model: tokens go in one side and come out the other. That took me a while to unlearn.&lt;/p&gt;

&lt;p&gt;Your server sits in the middle of two connections. Something in front of it can buffer. The client behind it can leave. The provider underneath it can fail. And the browser can misread what arrives.&lt;/p&gt;

&lt;p&gt;Most streaming bugs are one of these four hops going wrong. The sections below take them in order.&lt;/p&gt;

&lt;p&gt;I use &lt;code&gt;fetch&lt;/code&gt; with a streamed response for all of this. A chat request needs a message history and an auth token, and &lt;code&gt;fetch&lt;/code&gt; gives me a POST body, headers, and abort support. The alternatives are near the end.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why does the answer arrive all at once in production?
&lt;/h2&gt;

&lt;p&gt;Reverse proxies like Nginx &lt;strong&gt;collect the upstream response and send it in bigger chunks&lt;/strong&gt; by default. That's great for normal pages and terrible for a stream. Your tokens sit in a buffer, and the user sees nothing until it fills or the response ends.&lt;/p&gt;

&lt;p&gt;I've seen this called out as the most frequent production surprise with streaming, and it matches what the infrastructure docs say.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;In your app&lt;/strong&gt;, send headers that tell proxies to back off:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;writeHead&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Content-Type&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;text/event-stream&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Cache-Control&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;no-cache, no-transform&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Connection&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;keep-alive&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;X-Accel-Buffering&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;no&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;// Nginx-specific: don't buffer this response&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;In Nginx&lt;/strong&gt;, for the streaming route:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight nginx"&gt;&lt;code&gt;&lt;span class="k"&gt;location&lt;/span&gt; &lt;span class="n"&gt;/api/chat&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kn"&gt;proxy_pass&lt;/span&gt; &lt;span class="s"&gt;http://app&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;proxy_http_version&lt;/span&gt; &lt;span class="mf"&gt;1.1&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;proxy_set_header&lt;/span&gt; &lt;span class="s"&gt;Connection&lt;/span&gt; &lt;span class="s"&gt;""&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;proxy_buffering&lt;/span&gt; &lt;span class="no"&gt;off&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;      &lt;span class="c1"&gt;# the important one&lt;/span&gt;
    &lt;span class="kn"&gt;gzip&lt;/span&gt; &lt;span class="no"&gt;off&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;                 &lt;span class="c1"&gt;# compression can re-buffer the stream&lt;/span&gt;
    &lt;span class="kn"&gt;proxy_read_timeout&lt;/span&gt; &lt;span class="s"&gt;300s&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;  &lt;span class="c1"&gt;# longer than your slowest answer&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A CDN, load balancer, or API gateway in front of you may have its own buffering and idle timeout settings. I can't tell you the right config for each one, so check your provider's docs. What works everywhere is a test:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-N&lt;/span&gt; &lt;span class="nt"&gt;-X&lt;/span&gt; POST https://your-domain.com/api/chat &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"messages":[{"role":"user","content":"count slowly from 1 to 20"}]}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;-N&lt;/code&gt; turns off curl's own buffering. If lines appear gradually, the whole path streams. If they arrive in one lump, something is buffering. Run it against the &lt;strong&gt;real production URL&lt;/strong&gt;, and again after any infrastructure change.&lt;/p&gt;

&lt;p&gt;The same hop kills quiet streams. When the model is thinking or a tool call takes several seconds, some proxies and load balancers see silence and decide the connection is dead. Raise timeouts above your slowest realistic answer on every layer, and send a heartbeat. In SSE, a line starting with &lt;code&gt;:&lt;/code&gt; is a comment that clients ignore:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;heartbeat&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;setInterval&lt;/span&gt;&lt;span class="p"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;write&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;: ping&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="mi"&gt;15000&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;on&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;close&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;clearInterval&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;heartbeat&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If your answer involves tools, also stream status events like &lt;code&gt;{"status": "searching documents"}&lt;/code&gt;. Ten seconds of silence feels broken. Ten seconds with "Looking that up…" feels normal.&lt;/p&gt;

&lt;p&gt;Deploys count too. A rolling restart can kill streams mid-answer, so give your servers a graceful shutdown period at least as long as your longest answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  What happens when the user closes the tab?
&lt;/h2&gt;

&lt;p&gt;This is the hop I think is most underrated.&lt;/p&gt;

&lt;p&gt;A user asks a question, sees it start answering, and closes the tab. In a lot of apps the server does nothing. The loop reading from the LLM keeps going, the provider keeps generating, and you pay for output nobody sees. A Stop button that only hides text in the UI is the same problem with a friendlier face.&lt;/p&gt;

&lt;p&gt;A real cancel has to travel the whole chain:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Browser:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;controller&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;AbortController&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;/api/chat&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;method&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;POST&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Content-Type&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;application/json&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="na"&gt;body&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;messages&lt;/span&gt; &lt;span class="p"&gt;}),&lt;/span&gt;
  &lt;span class="na"&gt;signal&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;controller&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;signal&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="c1"&gt;// Stop button:&lt;/span&gt;
&lt;span class="nx"&gt;stopButton&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;onclick&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;controller&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;abort&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Server:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;/api/chat&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;upstream&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;AbortController&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

  &lt;span class="c1"&gt;// If the client leaves before we finish, cancel the LLM call.&lt;/span&gt;
  &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;on&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;close&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;writableEnded&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nx"&gt;upstream&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;abort&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;

  &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;writeHead&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Content-Type&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;text/event-stream&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="cm"&gt;/* + headers above */&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;

  &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;stream&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;llm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
      &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;stream&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
      &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;signal&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;upstream&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;signal&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;);&lt;/span&gt;

    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="k"&gt;await &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;chunk&lt;/span&gt; &lt;span class="k"&gt;of&lt;/span&gt; &lt;span class="nx"&gt;stream&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;text&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]?.&lt;/span&gt;&lt;span class="nx"&gt;delta&lt;/span&gt;&lt;span class="p"&gt;?.&lt;/span&gt;&lt;span class="nx"&gt;content&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
      &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;write&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`data: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;text&lt;/span&gt; &lt;span class="p"&gt;})}&lt;/span&gt;&lt;span class="s2"&gt;\n\n`&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;write&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`data: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;done&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt; &lt;span class="p"&gt;})}&lt;/span&gt;&lt;span class="s2"&gt;\n\n`&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;err&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;upstream&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;signal&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;aborted&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;write&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`data: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;error&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;generation_failed&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;})}&lt;/span&gt;&lt;span class="s2"&gt;\n\n`&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;finally&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;end&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Choices I'd make again:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Listen on the response's &lt;code&gt;close&lt;/code&gt; event, and check &lt;code&gt;writableEnded&lt;/code&gt;.&lt;/strong&gt; That separates "the client left early" from "we finished normally." In recent Node versions, I wouldn't rely on the request's &lt;code&gt;close&lt;/code&gt; event for this.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pass the same signal to everything the request started.&lt;/strong&gt; If the question also triggered a vector search, a reranker, or tool calls, they should stop too. Otherwise you cancel the visible part and keep paying for the invisible part.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Don't assume the provider stops billing immediately.&lt;/strong&gt; Aborting the connection usually stops generation, but behavior varies by provider. I'd test it with a short script and check the usage numbers, instead of trusting the docs alone.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Decide what to do with partial answers.&lt;/strong&gt; If a user stops mid-answer, do you save what was generated? Choose on purpose.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  How do you report an error after you've already sent 200 OK?
&lt;/h2&gt;

&lt;p&gt;Normally errors are status codes: 400, 429, 500. In a stream, &lt;strong&gt;you've already sent &lt;code&gt;200 OK&lt;/code&gt; and some tokens before the problem happens.&lt;/strong&gt; The provider times out, you hit a rate limit, or a tool call throws. The status code is already gone.&lt;/p&gt;

&lt;p&gt;So errors have to travel &lt;strong&gt;inside the stream&lt;/strong&gt; as events. Two rules I follow, both in the server code above:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Send an explicit error event&lt;/strong&gt;, so the UI can say something useful.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Always send a final &lt;code&gt;done&lt;/code&gt; event on success.&lt;/strong&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The second one is easy to skip and it matters. Without &lt;code&gt;done&lt;/code&gt;, the client can't tell "complete" from "the connection dropped after the third sentence." Users will read the truncated answer as the real answer.&lt;/p&gt;

&lt;p&gt;In the UI I treat these as three separate end states:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Complete:&lt;/strong&gt; got &lt;code&gt;done&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Failed:&lt;/strong&gt; got &lt;code&gt;error&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Interrupted:&lt;/strong&gt; the stream ended without either (show "response was cut off, try again")&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I'd also avoid silently auto-retrying a half-finished generation. A retry costs a second generation, and it might give a different answer from the one the user already half-read.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why does the browser show broken or half-parsed text?
&lt;/h2&gt;

&lt;p&gt;The client code looks simple. Two bugs got me.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Bug 1: network chunks don't line up with your events.&lt;/strong&gt; One read can contain half an event, or three and a half. If you parse each chunk directly, you'll eventually hit a &lt;code&gt;JSON.parse&lt;/code&gt; error on a partial line. Keep a buffer, split on the blank line that ends each event, and carry the leftover:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;reader&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getReader&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;decoder&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;TextDecoder&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;buffer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="dl"&gt;""&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;while &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;value&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;done&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;reader&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;done&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;break&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

  &lt;span class="c1"&gt;// stream: true keeps multi-byte characters intact across chunks&lt;/span&gt;
  &lt;span class="nx"&gt;buffer&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="nx"&gt;decoder&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;decode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;value&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;stream&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;

  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;events&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;buffer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="nx"&gt;buffer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;events&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;pop&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt; &lt;span class="c1"&gt;// the last piece may be incomplete, keep it&lt;/span&gt;

  &lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;evt&lt;/span&gt; &lt;span class="k"&gt;of&lt;/span&gt; &lt;span class="nx"&gt;events&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;evt&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;startsWith&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;data: &lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="k"&gt;continue&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="c1"&gt;// skips heartbeat comments&lt;/span&gt;
    &lt;span class="nf"&gt;handle&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;parse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;evt&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;slice&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;6&lt;/span&gt;&lt;span class="p"&gt;)));&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Bug 2: broken characters.&lt;/strong&gt; An emoji or a non-English character can be several bytes, and a chunk can end in the middle of one. Decoding each chunk separately gives you &lt;code&gt;�&lt;/code&gt;. The &lt;code&gt;{ stream: true }&lt;/code&gt; option on &lt;code&gt;TextDecoder&lt;/code&gt; fixes it. This one is nasty because it passes every English-only test.&lt;/p&gt;

&lt;p&gt;Rendering has the same trap. If you re-render the whole message on every token, especially if you re-parse Markdown each time, the page starts to stutter. What helps:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Batch updates.&lt;/strong&gt; Update the screen once per animation frame, not once per token.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Treat the streaming message as append-only.&lt;/strong&gt; Avoid rebuilding earlier content.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Plan for half-finished Markdown.&lt;/strong&gt; A code fence that hasn't closed yet can make a renderer flicker.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Respect the user's scroll.&lt;/strong&gt; If they scrolled up to read, don't pull them back down on every token.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And measure the right thing. Users will wait through a long answer if &lt;strong&gt;the first token arrives quickly&lt;/strong&gt;. Time to first token is worth logging more than total duration.&lt;/p&gt;

&lt;h2&gt;
  
  
  When is this the wrong approach?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;When the transport doesn't fit.&lt;/strong&gt; You have three realistic options:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Option&lt;/th&gt;
&lt;th&gt;Good for&lt;/th&gt;
&lt;th&gt;Catch&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;SSE via &lt;code&gt;EventSource&lt;/code&gt;&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Simple server-to-browser streams&lt;/td&gt;
&lt;td&gt;GET only, no custom headers, auto-reconnects by itself&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Streamed response via &lt;code&gt;fetch&lt;/code&gt;&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Chat: POST body, auth headers, abort support&lt;/td&gt;
&lt;td&gt;You parse the stream yourself&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;WebSocket&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;True two-way, low latency (voice, live collaboration)&lt;/td&gt;
&lt;td&gt;More infrastructure, harder to scale and debug&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;code&gt;EventSource&lt;/code&gt; reconnecting on its own sounds helpful. For LLMs it can mean &lt;strong&gt;a second paid generation for the same question&lt;/strong&gt; that nobody planned for. If you use it, make retries something you decide.&lt;/p&gt;

&lt;p&gt;I'd reach for WebSockets when audio or real two-way traffic is involved. I've used one for a realtime voice assistant, and it's the right tool there. "The user interrupted" is just cancellation with a microphone.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;When you shouldn't stream at all.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Structured output&lt;/strong&gt; (JSON, function arguments) is useless half-finished. Wait and validate the whole thing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Server-to-server calls&lt;/strong&gt; gain nothing from streaming.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Answers you need to check before showing&lt;/strong&gt; (safety filters, "does this match the tool results?") can't be checked until they're done. Streaming and validation pull in opposite directions, and you have to pick which matters more.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Where my own experience runs out.&lt;/strong&gt; I haven't run streaming at massive scale. Three things I'm still figuring out:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Resumable streams.&lt;/strong&gt; If a user refreshes mid-answer, can they pick up where they left off? I've read about approaches that write chunks to something like Redis Streams so a reconnect can resume, but I haven't built this myself yet.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Scaling beyond one server.&lt;/strong&gt; Long-lived connections behave differently under load balancing and autoscaling, and I'm still learning where the limits are.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Backpressure.&lt;/strong&gt; If the client reads slower than you write, buffers grow on your side. I know the concept, but I haven't pushed it hard enough in practice to have strong opinions.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What should you check before you ship?
&lt;/h2&gt;

&lt;p&gt;One check per hop, plus one metric:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Proxy:&lt;/strong&gt; run the &lt;code&gt;curl -N&lt;/code&gt; command against your real production URL. Do lines arrive gradually, or in one lump?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Your API:&lt;/strong&gt; start a long answer, close the tab, and watch your server logs. Does the LLM call stop? Do the tools and searches stop with it?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Provider:&lt;/strong&gt; force an error mid-answer. Does the UI say "failed" or "cut off", or does it show half an answer as if it were complete?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Browser:&lt;/strong&gt; ask for an answer full of emoji and non-English text. Any &lt;code&gt;�&lt;/code&gt;?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Metric:&lt;/strong&gt; log time to first token, not just total duration.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;em&gt;Next post: keeping LLM output trustworthy when you can't validate it before it streams, covering structured outputs, retries, and what to do with almost-valid JSON.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Which of these four hops broke first for you in production, and what was sitting in front of your server when it did?&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>api</category>
      <category>backend</category>
      <category>javascript</category>
      <category>llm</category>
    </item>
  </channel>
</rss>
