<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Amaresh Pelleti</title>
    <description>The latest articles on DEV Community by Amaresh Pelleti (@amareswer).</description>
    <link>https://dev.to/amareswer</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3978481%2Fbef1aa2c-c07a-414a-bb88-ea788ca39ba2.jpg</url>
      <title>DEV Community: Amaresh Pelleti</title>
      <link>https://dev.to/amareswer</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/amareswer"/>
    <language>en</language>
    <item>
      <title>Ollama Model Library, By the Numbers</title>
      <dc:creator>Amaresh Pelleti</dc:creator>
      <pubDate>Sun, 06 Sep 2026 14:21:50 +0000</pubDate>
      <link>https://dev.to/amareswer/ollama-model-library-by-the-numbers-43n1</link>
      <guid>https://dev.to/amareswer/ollama-model-library-by-the-numbers-43n1</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;Originally published on &lt;a href="https://devtoolhub.com/ollama-model-library-by-the-numbers/" rel="noopener noreferrer"&gt;DevToolHub&lt;/a&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Someone asks "what's the best Ollama model" in a Discord or a Reddit thread almost every day. The answer usually comes from memory, or a single favorite. So &lt;a href="https://ollama.com/library" rel="noopener noreferrer"&gt;we pulled the actual data&lt;/a&gt; instead. Every model in the Ollama model library, its pull count, its capability tags, and when it was last updated. Current as of September 5, 2026 — the library changes weekly, so treat these counts as a snapshot, not a permanent ranking.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's in the Ollama model library right now
&lt;/h2&gt;

&lt;p&gt;The Ollama model library holds 239 models as of this pull. That spans general chat models, embeddings, vision models, and a growing "cloud" category that runs on Ollama's hosted infrastructure instead of your GPU. Combined, every model in the library has been pulled just over 1.04 billion times.&lt;/p&gt;

&lt;p&gt;That total tells you less than it looks like it should. Pulls aren't spread evenly. A handful of general-purpose chat models account for most of it. The long tail of specialized or older models barely registers.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Ollama model library pulls actually concentrate
&lt;/h2&gt;

&lt;p&gt;The top 10 models by pull count account for 56.4% of all pulls in the entire library. Meta's Llama family and DeepSeek's reasoning models lead by a wide margin:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Rank&lt;/th&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Pulls&lt;/th&gt;
&lt;th&gt;Capabilities&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;&lt;code&gt;llama3.1&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;119.2M&lt;/td&gt;
&lt;td&gt;tools&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;&lt;code&gt;deepseek-r1&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;92.4M&lt;/td&gt;
&lt;td&gt;tools, thinking&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;&lt;code&gt;nomic-embed-text&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;84.6M&lt;/td&gt;
&lt;td&gt;embedding&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;&lt;code&gt;llama3.2&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;82.4M&lt;/td&gt;
&lt;td&gt;tools&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;&lt;code&gt;gemma3&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;40.1M&lt;/td&gt;
&lt;td&gt;vision&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;&lt;code&gt;qwen2.5&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;39.3M&lt;/td&gt;
&lt;td&gt;tools&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;&lt;code&gt;qwen3&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;36.3M&lt;/td&gt;
&lt;td&gt;tools, thinking&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;&lt;code&gt;mistral&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;33.3M&lt;/td&gt;
&lt;td&gt;tools&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;td&gt;&lt;code&gt;gemma2&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;31.8M&lt;/td&gt;
&lt;td&gt;text only&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;&lt;code&gt;llama3&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;25.2M&lt;/td&gt;
&lt;td&gt;text only&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Here's the thing about that list: it's mostly older releases. &lt;code&gt;llama3.1&lt;/code&gt; and &lt;code&gt;llama3.2&lt;/code&gt; alone outpull every newer Llama version combined. &lt;code&gt;llama3&lt;/code&gt; — several generations behind — still sits at 25.2M pulls. That's because pull counts accumulate over a model's lifetime. An older model with a long track record will always look bigger than a newer one that's objectively better. Don't read this table as "what to use today." Read it as "what people have been defaulting to."&lt;/p&gt;

&lt;p&gt;&lt;code&gt;nomic-embed-text&lt;/code&gt; at #3 breaks the chat-model pattern. It's a pure embedding model. Its position this high says a lot about how much RAG and semantic-search tooling runs on Ollama underneath the chat interfaces people actually see.&lt;/p&gt;

&lt;h2&gt;
  
  
  How many Ollama models actually support tool calling
&lt;/h2&gt;

&lt;p&gt;If you're building an agent, this number matters more than pull count. Only 93 of 239 models (38.9%) carry the &lt;code&gt;tools&lt;/code&gt; capability tag. The rest either don't support structured tool calling, or Ollama hasn't tagged them for it.&lt;/p&gt;

&lt;p&gt;The full capability breakdown across the library:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Tools (function calling):&lt;/strong&gt; 93 models — 38.9%&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Thinking (extended reasoning):&lt;/strong&gt; 42 models — 17.6%&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Vision:&lt;/strong&gt; 38 models — 15.9%&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cloud (hosted, not local):&lt;/strong&gt; 18 models — 7.5%&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Embedding:&lt;/strong&gt; 12 models — 5.0%&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No special tag (plain text completion):&lt;/strong&gt; 118 models — 49.4%&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These overlap — a model like &lt;code&gt;gemma4&lt;/code&gt; carries vision, tools, thinking, and cloud tags at once. But the practical read is simple. Half the library is plain text completion. If your use case needs tool calling specifically, you're choosing from under 40% of what's listed. Check a model's tags on its &lt;a href="https://ollama.com/library" rel="noopener noreferrer"&gt;library page&lt;/a&gt; before you build around it, not after.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the Ollama model library carries 7,359 tag variants
&lt;/h2&gt;

&lt;p&gt;Across all 239 models, there are 7,359 individual tags. Each one is a different quantization, parameter size, or context configuration of the same base model. That averages out to roughly 30.8 variants per model, though it's lopsided — &lt;code&gt;llama3.1&lt;/code&gt; alone carries 93 tags, while a narrow embedding model might have 3.&lt;/p&gt;

&lt;p&gt;This ties into the same issue covered in the &lt;a href="https://devtoolhub.com/ollama-vs-lm-studio/" rel="noopener noreferrer"&gt;Ollama vs LM Studio comparison&lt;/a&gt;. Ollama's model names default to a specific quantization, usually &lt;code&gt;Q4_K_M&lt;/code&gt;, a 4-bit quant, without making that choice obvious on the model card. With 30+ tagged variants per model on average, "the model" isn't one thing — it's a family. The tag you don't specify is a decision Ollama makes for you. If you're benchmarking against another tool, pin the exact tag (&lt;code&gt;ollama pull llama3.1:8b-instruct-q8_0&lt;/code&gt;, for example) instead of the bare model name.&lt;/p&gt;

&lt;h2&gt;
  
  
  How much of the Ollama model library is still maintained
&lt;/h2&gt;

&lt;p&gt;Freshness in the library skews old. Of the 239 models:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;174 (72.8%)&lt;/strong&gt; were last updated over a year ago&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;51 (21.3%)&lt;/strong&gt; were updated sometime this year&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;9 (3.8%)&lt;/strong&gt; were updated this month&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;5 (2.1%)&lt;/strong&gt; were updated this week&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Nearly three-quarters of the library hasn't been touched in over a year. That's not automatically a red flag. A model's weights don't need updates the way a CLI tool does, and a good 2024 model is often still a good model. But it does mean the library is mostly an archive with a small, actively-tended front section. If a model card shows no update in over a year, and a same-family successor has recent activity, the successor is usually the safer default.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this means if you're picking a model to run
&lt;/h2&gt;

&lt;p&gt;Two practical takeaways come out of this data. First, pull count is a popularity signal, not a quality signal. It rewards models that have been available the longest. That's why year-old Llama releases still outrank newer, often better models. Cross-check pull count against a model's actual release date before treating it as a recommendation.&lt;/p&gt;

&lt;p&gt;Second, the capability tags are the fastest filter for narrowing 239 models down to the ones that fit your use case. Building an agent? Filter to the 93 &lt;code&gt;tools&lt;/code&gt;-tagged models first. Need a local embedding model for RAG? &lt;code&gt;nomic-embed-text&lt;/code&gt; and &lt;code&gt;mxbai-embed-large&lt;/code&gt; are the two with real usage behind them. For everything else — running the model once you've picked it — the &lt;a href="https://devtoolhub.com/ollama-hardware-requirements/" rel="noopener noreferrer"&gt;hardware requirements guide&lt;/a&gt; and the &lt;a href="https://devtoolhub.com/ollama-api-guide/" rel="noopener noreferrer"&gt;Ollama API guide&lt;/a&gt; cover the setup.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Q: How many models are in the Ollama library?&lt;/strong&gt;&lt;br&gt;
A: 239 models as of September 5, 2026. The Ollama model library is updated regularly, so this number changes — check &lt;a href="https://ollama.com/library" rel="noopener noreferrer"&gt;ollama.com/library&lt;/a&gt; directly for the current count.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: What is the most popular model in the Ollama library?&lt;/strong&gt;&lt;br&gt;
A: &lt;code&gt;llama3.1&lt;/code&gt;, with 119.2 million pulls, ahead of &lt;code&gt;deepseek-r1&lt;/code&gt; at 92.4 million and &lt;code&gt;nomic-embed-text&lt;/code&gt; at 84.6 million. All three are well over a year old, which is typical — pull counts accumulate over time and favor established models.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: How many Ollama models support function calling?&lt;/strong&gt;&lt;br&gt;
A: 93 of 239 models (38.9%) carry the &lt;code&gt;tools&lt;/code&gt; capability tag, which indicates support for structured function calling. Check the individual model's page on ollama.com to confirm before building an agent around it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Why does Ollama have so many tags for one model?&lt;/strong&gt;&lt;br&gt;
A: Each tag is a different quantization, parameter size, or configuration of the same base model. The library averages about 30.8 tags per model, and the untagged default usually points to a 4-bit quantization rather than the highest-quality version available.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quick Summary:&lt;/strong&gt;&lt;br&gt;
– 239 models in the Ollama model library as of September 5, 2026, with 1.04 billion cumulative pulls.&lt;br&gt;
– The top 10 models account for 56.4% of all pulls — usage concentrates hard in a handful of established releases.&lt;br&gt;
– Only 38.9% of models carry the &lt;code&gt;tools&lt;/code&gt; (function-calling) tag; 49.4% have no special capability tag at all.&lt;br&gt;
– The library averages 30.8 tag variants per model, so the untagged default is a quantization choice you're making without realizing it.&lt;br&gt;
– 72.8% of models haven't been updated in over a year — the library is mostly an archive with a small active front section.&lt;/p&gt;

&lt;p&gt;If you've picked a model from this list and need to know whether your hardware can run it, the &lt;a href="https://devtoolhub.com/ollama-hardware-requirements/" rel="noopener noreferrer"&gt;Ollama hardware requirements guide&lt;/a&gt; breaks down RAM and VRAM by model size. And if local hardware runs out before your model does, the &lt;a href="https://devtoolhub.com/ollama-cloud-free-vs-pro-limits-pricing-2026/" rel="noopener noreferrer"&gt;Ollama Cloud pricing and limits guide&lt;/a&gt; covers the hosted tier.&lt;/p&gt;

</description>
      <category>ollama</category>
      <category>localllm</category>
      <category>aitools</category>
      <category>opensourcemodels</category>
    </item>
    <item>
      <title>Ollama API: A Practical Guide with Examples</title>
      <dc:creator>Amaresh Pelleti</dc:creator>
      <pubDate>Fri, 04 Sep 2026 11:35:17 +0000</pubDate>
      <link>https://dev.to/amareswer/ollama-api-a-practical-guide-with-examples-4di9</link>
      <guid>https://dev.to/amareswer/ollama-api-a-practical-guide-with-examples-4di9</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;Originally published on &lt;a href="https://devtoolhub.com/ollama-api-guide/" rel="noopener noreferrer"&gt;DevToolHub&lt;/a&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Every Ollama install runs a local HTTP server on port &lt;code&gt;11434&lt;/code&gt;, and that server is the real interface to the models. The &lt;code&gt;ollama run&lt;/code&gt; command is a thin client on top of it. Once you know the two main endpoints, the streaming format, and the options object, you can wire a local model into any application.&lt;/p&gt;

&lt;p&gt;There is also an OpenAI-compatible route, so existing code that talks to OpenAI can point at Ollama with a base-URL change.&lt;/p&gt;

&lt;h2&gt;
  
  
  How the Ollama API works
&lt;/h2&gt;

&lt;p&gt;The Ollama API is a plain REST API served at &lt;code&gt;http://localhost:11434&lt;/code&gt;. You send JSON with &lt;code&gt;POST&lt;/code&gt;, and by default you get a stream of newline-delimited JSON objects back. No API key is required for local access.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Endpoint&lt;/th&gt;
&lt;th&gt;Method&lt;/th&gt;
&lt;th&gt;Purpose&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;/api/generate&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;POST&lt;/td&gt;
&lt;td&gt;Single-prompt text completion&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;/api/chat&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;POST&lt;/td&gt;
&lt;td&gt;Multi-turn chat with message history and tools&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;/api/embed&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;POST&lt;/td&gt;
&lt;td&gt;Generate embeddings&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;/api/tags&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;GET&lt;/td&gt;
&lt;td&gt;List installed models&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;/api/ps&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;GET&lt;/td&gt;
&lt;td&gt;List models currently loaded in memory&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;/api/pull&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;POST&lt;/td&gt;
&lt;td&gt;Download a model&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Check the server with &lt;code&gt;curl http://localhost:11434/api/version&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Generating text: /api/generate and /api/chat
&lt;/h2&gt;

&lt;p&gt;Use &lt;code&gt;/api/generate&lt;/code&gt; for a single prompt with no history:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl http://localhost:11434/api/generate &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
  "model": "llama3.1",
  "prompt": "Summarize in one sentence: Ollama serves a local HTTP API on port 11434.",
  "stream": false
}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Use &lt;code&gt;/api/chat&lt;/code&gt; for turn-by-turn context or tool calling. Pass a &lt;code&gt;messages&lt;/code&gt; array with &lt;code&gt;role&lt;/code&gt; values of &lt;code&gt;system&lt;/code&gt;, &lt;code&gt;user&lt;/code&gt;, &lt;code&gt;assistant&lt;/code&gt;, or &lt;code&gt;tool&lt;/code&gt;, and send the whole history on each request:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl http://localhost:11434/api/chat &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
  "model": "llama3.1",
  "messages": [
    {"role": "system", "content": "You answer in one short sentence."},
    {"role": "user", "content": "What is the KV cache?"}
  ],
  "stream": false
}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For application code, &lt;code&gt;/api/chat&lt;/code&gt; is the better default even for single questions.&lt;/p&gt;

&lt;h2&gt;
  
  
  Streaming and response metrics
&lt;/h2&gt;

&lt;p&gt;By default &lt;code&gt;stream&lt;/code&gt; is &lt;code&gt;true&lt;/code&gt;, and Ollama returns one JSON object per chunk. The final chunk has &lt;code&gt;"done": true&lt;/code&gt; plus timing data:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"llama3.1"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"response"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"done"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
 &lt;/span&gt;&lt;span class="nl"&gt;"total_duration"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;4883583458&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"prompt_eval_count"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;26&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
 &lt;/span&gt;&lt;span class="nl"&gt;"eval_count"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;298&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"eval_duration"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;3789981000&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;All durations are in nanoseconds. Tokens per second is &lt;code&gt;eval_count / eval_duration * 1e9&lt;/code&gt;. Set &lt;code&gt;"stream": false&lt;/code&gt; for a single response object.&lt;/p&gt;

&lt;h2&gt;
  
  
  Setting model options
&lt;/h2&gt;

&lt;p&gt;The &lt;code&gt;options&lt;/code&gt; object tunes sampling and context per request:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl http://localhost:11434/api/chat &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
  "model": "llama3.1",
  "messages": [{"role": "user", "content": "Name three container runtimes."}],
  "stream": false,
  "options": {"temperature": 0.2, "num_ctx": 8192, "num_predict": 200, "seed": 42, "stop": ["\n\n"]}
}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;temperature&lt;/code&gt; — lower is more deterministic&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;num_ctx&lt;/code&gt; — context window in tokens for this request&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;num_predict&lt;/code&gt; — cap on generated tokens&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;seed&lt;/code&gt; — with &lt;code&gt;temperature: 0&lt;/code&gt;, gives repeatable output&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;stop&lt;/code&gt; — strings that end generation&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Setting &lt;code&gt;num_ctx&lt;/code&gt; above what your hardware holds forces a partial CPU offload. Confirm with &lt;code&gt;ollama ps&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Structured JSON output from the Ollama API
&lt;/h2&gt;

&lt;p&gt;Set &lt;code&gt;format&lt;/code&gt; to &lt;code&gt;"json"&lt;/code&gt; for any valid JSON, or pass a JSON schema object to force a shape:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl http://localhost:11434/api/chat &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
  "model": "llama3.1",
  "messages": [{"role": "user", "content": "List two Linux distros with release years. Respond in JSON."}],
  "stream": false,
  "format": {
    "type": "object",
    "properties": {
      "distros": {"type": "array", "items": {
        "type": "object",
        "properties": {"name": {"type": "string"}, "year": {"type": "integer"}},
        "required": ["name", "year"]
      }}
    },
    "required": ["distros"]
  }
}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Keep the word "JSON" in the prompt and use a low temperature.&lt;/p&gt;

&lt;h2&gt;
  
  
  Tool calling with the Ollama API
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;/api/chat&lt;/code&gt; supports function calling through a &lt;code&gt;tools&lt;/code&gt; array. The model replies with a &lt;code&gt;tool_calls&lt;/code&gt; entry instead of text when it decides to use one:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl http://localhost:11434/api/chat &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
  "model": "llama3.1",
  "messages": [{"role": "user", "content": "What is the weather in Toronto?"}],
  "stream": false,
  "tools": [{
    "type": "function",
    "function": {
      "name": "get_weather",
      "description": "Get the current weather for a city",
      "parameters": {"type": "object", "properties": {"city": {"type": "string"}}, "required": ["city"]}
    }
  }]
}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run the function, then send the result back as a message with &lt;code&gt;"role": "tool"&lt;/code&gt;. Tool support depends on the model.&lt;/p&gt;

&lt;h2&gt;
  
  
  Calling the Ollama API from Python
&lt;/h2&gt;

&lt;p&gt;Install with &lt;code&gt;pip install ollama&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;ollama&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;chat&lt;/span&gt;

&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;llama3.1&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Why is the sky blue?&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;])&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Streaming:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;ollama&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;chat&lt;/span&gt;

&lt;span class="n"&gt;stream&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;llama3.1&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Explain the KV cache in two sentences.&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
    &lt;span class="n"&gt;stream&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;chunk&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;stream&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;message&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;end&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;''&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;flush&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Remote host:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;ollama&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Client&lt;/span&gt;
&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Client&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;host&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;http://192.168.1.50:11434&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There is an &lt;code&gt;AsyncClient&lt;/code&gt; with the same methods, plus &lt;code&gt;embed()&lt;/code&gt;, &lt;code&gt;list()&lt;/code&gt;, &lt;code&gt;ps()&lt;/code&gt;, and &lt;code&gt;pull()&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The OpenAI-compatible endpoint
&lt;/h2&gt;

&lt;p&gt;Ollama serves an OpenAI-style API at &lt;code&gt;http://localhost:11434/v1&lt;/code&gt; with &lt;code&gt;/v1/chat/completions&lt;/code&gt;, &lt;code&gt;/v1/completions&lt;/code&gt;, &lt;code&gt;/v1/embeddings&lt;/code&gt;, and &lt;code&gt;/v1/models&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;http://localhost:11434/v1&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;ollama&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;llama3.1&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Hello&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;api_key&lt;/code&gt; is required by the SDK but ignored by Ollama. Use &lt;code&gt;/v1&lt;/code&gt; for compatibility and &lt;code&gt;/api&lt;/code&gt; for full features.&lt;/p&gt;

&lt;h2&gt;
  
  
  Securing the Ollama API for remote access
&lt;/h2&gt;

&lt;p&gt;The Ollama API has no built-in authentication. Anyone who can reach port &lt;code&gt;11434&lt;/code&gt; can use, pull, or delete your models. Keep it bound to &lt;code&gt;127.0.0.1&lt;/code&gt; and put a reverse proxy in front with auth and TLS:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight nginx"&gt;&lt;code&gt;&lt;span class="k"&gt;server&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kn"&gt;listen&lt;/span&gt; &lt;span class="mi"&gt;443&lt;/span&gt; &lt;span class="s"&gt;ssl&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;server_name&lt;/span&gt; &lt;span class="s"&gt;ollama.example.com&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;location&lt;/span&gt; &lt;span class="n"&gt;/&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="kn"&gt;proxy_pass&lt;/span&gt; &lt;span class="s"&gt;http://127.0.0.1:11434&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
        &lt;span class="kn"&gt;proxy_set_header&lt;/span&gt; &lt;span class="s"&gt;Host&lt;/span&gt; &lt;span class="nf"&gt;localhost&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;11434&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
        &lt;span class="kn"&gt;auth_basic&lt;/span&gt; &lt;span class="s"&gt;"Ollama"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
        &lt;span class="kn"&gt;auth_basic_user_file&lt;/span&gt; &lt;span class="n"&gt;/etc/nginx/.htpasswd&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Setting &lt;code&gt;OLLAMA_HOST=0.0.0.0&lt;/code&gt; without a proxy puts an unauthenticated model server on the open network. Only do that inside a private network or behind an IP-restricted firewall.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Q: What port does the Ollama API use?&lt;/strong&gt;&lt;br&gt;
A: Port &lt;code&gt;11434&lt;/code&gt; on &lt;code&gt;127.0.0.1&lt;/code&gt; by default. Change it with &lt;code&gt;OLLAMA_HOST&lt;/code&gt;, for example &lt;code&gt;OLLAMA_HOST=0.0.0.0:11434&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Does the Ollama API need an API key?&lt;/strong&gt;&lt;br&gt;
A: No, not for local use. Ollama's hosted cloud models use a key; self-hosted remote access should sit behind a reverse proxy that adds authentication.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: What is the difference between /api/generate and /api/chat?&lt;/strong&gt;&lt;br&gt;
A: &lt;code&gt;/api/generate&lt;/code&gt; takes a single &lt;code&gt;prompt&lt;/code&gt; string. &lt;code&gt;/api/chat&lt;/code&gt; takes a &lt;code&gt;messages&lt;/code&gt; array with roles and supports tool calling. Use &lt;code&gt;/api/chat&lt;/code&gt; for application code.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: How do I get JSON output from the Ollama API?&lt;/strong&gt;&lt;br&gt;
A: Set &lt;code&gt;format&lt;/code&gt; to &lt;code&gt;"json"&lt;/code&gt; or to a JSON schema object. Keep the word "JSON" in your prompt and use a low temperature.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Can I use the OpenAI Python SDK with Ollama?&lt;/strong&gt;&lt;br&gt;
A: Yes. Point &lt;code&gt;base_url&lt;/code&gt; at &lt;code&gt;http://localhost:11434/v1&lt;/code&gt; and pass any non-empty &lt;code&gt;api_key&lt;/code&gt;.&lt;/p&gt;

</description>
      <category>ollama</category>
      <category>api</category>
      <category>llm</category>
      <category>python</category>
    </item>
    <item>
      <title>Ollama Hardware Requirements: RAM, VRAM and GPU</title>
      <dc:creator>Amaresh Pelleti</dc:creator>
      <pubDate>Fri, 04 Sep 2026 11:33:50 +0000</pubDate>
      <link>https://dev.to/amareswer/ollama-hardware-requirements-ram-vram-and-gpu-268m</link>
      <guid>https://dev.to/amareswer/ollama-hardware-requirements-ram-vram-and-gpu-268m</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;Originally published on &lt;a href="https://devtoolhub.com/ollama-hardware-requirements/" rel="noopener noreferrer"&gt;DevToolHub&lt;/a&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Ollama runs a model by loading its weights into memory. So the size of the model file is the number that decides what hardware you need. For example, a 4-bit 8B model is about 5 GB on disk. It needs roughly that much free memory plus overhead to run. But push to a 70B model and you are looking at 40 GB or more. This guide gives the real numbers by model size, explains GPU versus CPU, and covers the setting most people miss: context length.&lt;/p&gt;

&lt;p&gt;The Ollama hardware requirements come down to one formula. Memory needed is roughly the model file size, plus 1 to 2 GB of overhead, plus the KV cache for your context window. Everything below is that formula applied.&lt;/p&gt;

&lt;h2&gt;
  
  
  What are the Ollama hardware requirements?
&lt;/h2&gt;

&lt;p&gt;The baseline is simple. First, you need enough free RAM or VRAM to hold the model file plus about 20% headroom. Ollama's model tags like &lt;code&gt;llama3.1&lt;/code&gt; or &lt;code&gt;gemma3&lt;/code&gt; default to &lt;code&gt;Q4_K_M&lt;/code&gt;, a 4-bit quantization, so the &lt;a href="https://ollama.com/library" rel="noopener noreferrer"&gt;download sizes on ollama.com&lt;/a&gt; are already compressed.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://github.com/ollama/ollama/blob/main/README.md" rel="noopener noreferrer"&gt;official README&lt;/a&gt; sets the baseline: &lt;strong&gt;at least 8 GB of RAM for 7B models, 16 GB for 13B models, and 32 GB for 33B models&lt;/strong&gt;. That holds up in practice for 4-bit quants at a modest context length. However, larger context windows and running multiple models at once push those numbers higher.&lt;/p&gt;

&lt;p&gt;Storage matters too. Models live in &lt;code&gt;~/.ollama/models&lt;/code&gt;, or &lt;code&gt;/usr/share/ollama/.ollama/models&lt;/code&gt; on a Linux service install. A single 70B model is about 40 GB on disk. So it is easy to fill a drive if you pull several large models to compare them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Ollama hardware requirements by model size
&lt;/h2&gt;

&lt;p&gt;These figures assume the default 4-bit quantization and a small-to-moderate context window. Add memory for larger context (see below).&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model size&lt;/th&gt;
&lt;th&gt;Approx. file size (Q4)&lt;/th&gt;
&lt;th&gt;Minimum RAM (CPU)&lt;/th&gt;
&lt;th&gt;Comfortable GPU VRAM&lt;/th&gt;
&lt;th&gt;Example GPU&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1B–3B&lt;/td&gt;
&lt;td&gt;0.7–2 GB&lt;/td&gt;
&lt;td&gt;8 GB&lt;/td&gt;
&lt;td&gt;4 GB&lt;/td&gt;
&lt;td&gt;Most modern GPUs, integrated graphics&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;7B–8B&lt;/td&gt;
&lt;td&gt;4.5–5 GB&lt;/td&gt;
&lt;td&gt;8 GB&lt;/td&gt;
&lt;td&gt;6–8 GB&lt;/td&gt;
&lt;td&gt;RTX 3060, RTX 4060&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;12B–14B&lt;/td&gt;
&lt;td&gt;7–9 GB&lt;/td&gt;
&lt;td&gt;16 GB&lt;/td&gt;
&lt;td&gt;12 GB&lt;/td&gt;
&lt;td&gt;RTX 3060 12GB, RTX 4070&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;27B–32B&lt;/td&gt;
&lt;td&gt;16–20 GB&lt;/td&gt;
&lt;td&gt;32 GB&lt;/td&gt;
&lt;td&gt;24 GB&lt;/td&gt;
&lt;td&gt;RTX 3090, RTX 4090&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;70B&lt;/td&gt;
&lt;td&gt;40–43 GB&lt;/td&gt;
&lt;td&gt;64 GB&lt;/td&gt;
&lt;td&gt;48 GB&lt;/td&gt;
&lt;td&gt;RTX 6000 Ada, 2× RTX 3090&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;120B+ / MoE&lt;/td&gt;
&lt;td&gt;65 GB and up&lt;/td&gt;
&lt;td&gt;128 GB+&lt;/td&gt;
&lt;td&gt;80 GB+&lt;/td&gt;
&lt;td&gt;A100, H100, multi-GPU&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Planning to serve a model to more than one request at a time? Then multiply the context memory by the number of parallel slots. Ollama's &lt;code&gt;OLLAMA_NUM_PARALLEL&lt;/code&gt; controls that count, and each slot needs its own slice of KV cache. For a broader look at which models are worth running, see the &lt;a href="https://devtoolhub.com/best-open-source-llms-2025/" rel="noopener noreferrer"&gt;best open-source LLMs roundup&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Do you need a GPU to run Ollama?
&lt;/h2&gt;

&lt;p&gt;No. Ollama runs on CPU alone, and a 7B model on a modern multi-core CPU with fast RAM produces a few tokens per second. That is usable for scripts and batch jobs but slow for interactive chat.&lt;/p&gt;

&lt;p&gt;A GPU changes the experience. When a compatible GPU fits the whole model in VRAM, Ollama loads every layer there. So inference is often 10 to 30 times faster than CPU. But when the model does not fit, Ollama does a &lt;strong&gt;partial offload&lt;/strong&gt;. It puts as many layers as possible on the GPU and runs the rest on the CPU. That works, however speed drops sharply because every token now waits on the slow path.&lt;/p&gt;

&lt;p&gt;⚠️ Note: a partial offload can be slower than pure CPU in some cases because of the constant copying between GPU and system RAM. If &lt;code&gt;ollama ps&lt;/code&gt; shows a CPU/GPU split, either use a smaller model or a smaller quantization so the whole thing fits in VRAM.&lt;/p&gt;

&lt;p&gt;For CPU-only inference, RAM bandwidth is the bottleneck, not core count. DDR5 and dual-channel memory help more than extra cores. The &lt;a href="https://devtoolhub.com/run-ollama-on-digitalocean-droplet/" rel="noopener noreferrer"&gt;DigitalOcean droplet setup guide&lt;/a&gt; shows CPU-only sizing for a cloud server.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which GPUs does Ollama support?
&lt;/h2&gt;

&lt;p&gt;Ollama supports three GPU paths, and the requirements differ:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;NVIDIA&lt;/strong&gt;: compute capability 5.0 and higher, with driver 550 or newer (570+ for the oldest supported cards). This covers everything from the GTX 900 series through the RTX 50 series, and you can confirm a card on the &lt;a href="https://developer.nvidia.com/cuda-gpus" rel="noopener noreferrer"&gt;NVIDIA CUDA GPUs list&lt;/a&gt;. NVIDIA is the most tested path.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AMD&lt;/strong&gt;: on Linux, ROCm v7 drivers are required. Supported cards include the Radeon RX 9000 and RX 7000 series, parts of the RX 6000 series, Radeon PRO W-series, and Instinct accelerators. Windows support is narrower.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Apple Silicon&lt;/strong&gt;: Ollama uses the Metal API on M-series Macs with no extra setup. Unified memory means the GPU can address most of system RAM (more on this below).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;To limit Ollama to specific GPUs, set &lt;code&gt;CUDA_VISIBLE_DEVICES&lt;/code&gt; (NVIDIA) or &lt;code&gt;ROCR_VISIBLE_DEVICES&lt;/code&gt; (AMD) to the device UUIDs. Set &lt;code&gt;CUDA_VISIBLE_DEVICES=-1&lt;/code&gt; to force CPU-only mode. The full compatibility table is in the &lt;a href="https://github.com/ollama/ollama/blob/main/docs/gpu.mdx" rel="noopener noreferrer"&gt;official GPU documentation&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  How context length changes Ollama hardware requirements
&lt;/h2&gt;

&lt;p&gt;Context length is the setting that quietly breaks memory planning. The KV cache holds the attention state for every token in the context window, and it grows with both the window size and the model size. A large context can add several gigabytes on top of the model weights.&lt;/p&gt;

&lt;p&gt;Recent Ollama versions set the default context length based on available VRAM rather than a fixed value:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Under 24 GB VRAM: 4K tokens&lt;/li&gt;
&lt;li&gt;24–48 GB VRAM: 32K tokens&lt;/li&gt;
&lt;li&gt;48 GB or more: 256K tokens&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You override this with &lt;code&gt;OLLAMA_CONTEXT_LENGTH&lt;/code&gt; on the server or &lt;code&gt;num_ctx&lt;/code&gt; per request:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;OLLAMA_CONTEXT_LENGTH&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;16384 ollama serve
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two settings reduce the cost. First, enabling flash attention with &lt;code&gt;OLLAMA_FLASH_ATTENTION=1&lt;/code&gt; cuts KV cache memory as context grows. Also, you can cap context per request instead of raising the server default for every model. So if you need long context for agents or coding tools, budget for it in your VRAM math from the start. Running a model across a cluster is a different problem, covered in the &lt;a href="https://devtoolhub.com/deploy-llm-kubernetes-guide/" rel="noopener noreferrer"&gt;deploy an LLM on Kubernetes guide&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to check if Ollama is using your GPU
&lt;/h2&gt;

&lt;p&gt;Run &lt;code&gt;ollama ps&lt;/code&gt; while a model is loaded. The &lt;code&gt;PROCESSOR&lt;/code&gt; column tells you exactly where the model is running:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;ollama ps
NAME               ID              SIZE     PROCESSOR    UNTIL
llama3.1:8b        365c0bd3c000    6.7 GB   100% GPU     4 minutes from now
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;100% GPU&lt;/code&gt; means the whole model is in VRAM. &lt;code&gt;100% CPU&lt;/code&gt; means no GPU acceleration. A split like &lt;code&gt;35%/65% CPU/GPU&lt;/code&gt; means partial offload, and that is your signal to drop to a smaller model or quant.&lt;/p&gt;

&lt;p&gt;If you expect GPU use and see CPU, work through three checks. First, confirm the driver meets the minimum version. Then check that &lt;code&gt;nvidia-smi&lt;/code&gt; or &lt;code&gt;rocminfo&lt;/code&gt; sees the card. Finally, make sure another process is not already holding the VRAM.&lt;/p&gt;

&lt;h2&gt;
  
  
  Apple Silicon: unified memory changes the math
&lt;/h2&gt;

&lt;p&gt;On an M-series Mac, the CPU and GPU share one pool of memory, so there is no separate VRAM number. A MacBook with 64 GB of unified memory can run models that would need a 48 GB workstation GPU on a PC. This makes high-memory Macs one of the most cost-effective ways to run 70B models locally.&lt;/p&gt;

&lt;p&gt;The catch is that macOS reserves part of that memory for the system. By default the GPU can use roughly 65 to 75 percent of total RAM. That is usually fine. But if a large model fails to load on a Mac that seems to have enough memory, the wired memory limit is the likely cause.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Q: What are the minimum Ollama hardware requirements?&lt;/strong&gt;&lt;br&gt;
A: A machine with 8 GB of RAM runs small models (1B–3B) and quantized 7B models on CPU. For a smooth experience with 7B models, use 16 GB of RAM or a GPU with 6–8 GB of VRAM. No GPU is required, but one makes inference 10 to 30 times faster.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: How much VRAM do I need for a 70B model in Ollama?&lt;/strong&gt;&lt;br&gt;
A: About 40 to 43 GB for the 4-bit weights, plus a few GB for context. A single 48 GB GPU or two 24 GB GPUs handles it. On CPU, plan for 64 GB of system RAM and expect slow output.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Can Ollama run on a server with no GPU?&lt;/strong&gt;&lt;br&gt;
A: Yes. CPU-only inference works for any model that fits in RAM. Expect a few tokens per second for 7B models. Fast, dual-channel DDR5 memory matters more than the number of CPU cores.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Why is my model running on CPU when I have a GPU?&lt;/strong&gt;&lt;br&gt;
A: Common causes are a driver below the minimum version, too little free VRAM, another process holding VRAM, or an unsupported card. Run &lt;code&gt;ollama ps&lt;/code&gt; to confirm, then &lt;code&gt;nvidia-smi&lt;/code&gt; to check the card and driver.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Does a bigger context window need more memory?&lt;/strong&gt;&lt;br&gt;
A: Yes. The KV cache scales with context length and model size and can add several GB. Enable &lt;code&gt;OLLAMA_FLASH_ATTENTION=1&lt;/code&gt; to reduce it, and only raise &lt;code&gt;OLLAMA_CONTEXT_LENGTH&lt;/code&gt; as far as your hardware allows.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quick Summary:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Memory needed ≈ model file size + 1–2 GB overhead + KV cache for your context window.&lt;/li&gt;
&lt;li&gt;Official minimums: 8 GB RAM for 7B, 16 GB for 13B, 32 GB for 33B, all at 4-bit quantization.&lt;/li&gt;
&lt;li&gt;A 70B model needs about 40–43 GB of VRAM or 64 GB of system RAM.&lt;/li&gt;
&lt;li&gt;No GPU is required, but a GPU that fits the whole model is 10–30× faster than CPU; a partial CPU/GPU split is much slower.&lt;/li&gt;
&lt;li&gt;NVIDIA needs compute capability 5.0+ and driver 550+; AMD needs ROCm v7 on Linux; Apple Silicon uses unified memory and punches above its price for large models.&lt;/li&gt;
&lt;li&gt;Run &lt;code&gt;ollama ps&lt;/code&gt; and read the &lt;code&gt;PROCESSOR&lt;/code&gt; column to confirm where a model is actually running.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So once you have matched the Ollama hardware requirements to your model, the rest is setup. But if local hardware caps out below the model you need, compare the hosted tier in the &lt;a href="https://devtoolhub.com/ollama-cloud-free-vs-pro-limits-pricing-2026/" rel="noopener noreferrer"&gt;Ollama Cloud pricing and limits guide&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ollama</category>
      <category>localllm</category>
      <category>gpu</category>
      <category>hardware</category>
    </item>
    <item>
      <title>Ollama vs LM Studio: Which One Should You Use</title>
      <dc:creator>Amaresh Pelleti</dc:creator>
      <pubDate>Fri, 04 Sep 2026 11:33:04 +0000</pubDate>
      <link>https://dev.to/amareswer/ollama-vs-lm-studio-which-one-should-you-use-1oii</link>
      <guid>https://dev.to/amareswer/ollama-vs-lm-studio-which-one-should-you-use-1oii</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;Originally published on &lt;a href="https://devtoolhub.com/ollama-vs-lm-studio/" rel="noopener noreferrer"&gt;DevToolHub&lt;/a&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;Ollama vs LM Studio&lt;/strong&gt; is a choice most people run into within a week of trying local models. Both tools run open-weight language models on your own machine, pull from the same pool of GGUF model files, and expose an OpenAI-compatible API. The real split is how you work: from a terminal and scripts, or from a window with a chat box. This comparison covers the differences in interface, licensing, platform support, and performance. Then it gives a verdict for the two most common situations.&lt;/p&gt;

&lt;p&gt;The short version: pick Ollama if you are wiring a model into a server, a container, or your own code. Pick LM Studio if you want to browse, download, and chat with models from a desktop app. They also run side by side without conflict, which is a valid answer too.&lt;/p&gt;

&lt;h2&gt;
  
  
  Ollama vs LM Studio at a glance
&lt;/h2&gt;

&lt;p&gt;Here is how the two tools line up on the decisions that actually matter.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Factor&lt;/th&gt;
&lt;th&gt;Ollama&lt;/th&gt;
&lt;th&gt;LM Studio&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Primary interface&lt;/td&gt;
&lt;td&gt;Command line, plus a basic desktop app&lt;/td&gt;
&lt;td&gt;Full desktop GUI with chat&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;License&lt;/td&gt;
&lt;td&gt;Open source (MIT)&lt;/td&gt;
&lt;td&gt;Proprietary, free for personal &lt;strong&gt;and&lt;/strong&gt; commercial use&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Platforms&lt;/td&gt;
&lt;td&gt;macOS, Windows, Linux&lt;/td&gt;
&lt;td&gt;macOS (Apple Silicon + Intel), Windows (x64/ARM64), Linux (x64)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model source&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;ollama.com&lt;/code&gt; library, plus GGUF import&lt;/td&gt;
&lt;td&gt;Hugging Face browser built into the app&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;API&lt;/td&gt;
&lt;td&gt;Native REST on &lt;code&gt;:11434&lt;/code&gt;, plus OpenAI-compatible &lt;code&gt;/v1&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;OpenAI- and Anthropic-compatible on &lt;code&gt;:1234&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Inference engine&lt;/td&gt;
&lt;td&gt;llama.cpp-based, plus a newer native engine for some models&lt;/td&gt;
&lt;td&gt;llama.cpp and Apple MLX&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Headless / server mode&lt;/td&gt;
&lt;td&gt;Built in (&lt;code&gt;ollama serve&lt;/code&gt;, systemd unit)&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;lms&lt;/code&gt; CLI and a headless daemon&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Document RAG&lt;/td&gt;
&lt;td&gt;Not built in&lt;/td&gt;
&lt;td&gt;Built into the app&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Best for&lt;/td&gt;
&lt;td&gt;Servers, automation, embedding in apps&lt;/td&gt;
&lt;td&gt;Desktop experimentation, non-CLI users&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Both tools use llama.cpp under the hood for GGUF models, so raw token throughput on the same model and quantization is close. The differences are in packaging, not the core math.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Ollama does best
&lt;/h2&gt;

&lt;p&gt;Ollama is built to run as a service. After install, it starts a background server on port &lt;code&gt;11434&lt;/code&gt; and stays out of the way. So you pull a model with &lt;code&gt;ollama pull llama3.1&lt;/code&gt;, then call it from any language over HTTP. That design makes it the default choice for anything programmatic.&lt;/p&gt;

&lt;p&gt;On Linux, the install script sets up a &lt;code&gt;systemd&lt;/code&gt; unit so the server survives reboots:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-fsSL&lt;/span&gt; https://ollama.com/install.sh | sh
&lt;span class="nb"&gt;sudo &lt;/span&gt;systemctl &lt;span class="nb"&gt;enable&lt;/span&gt; &lt;span class="nt"&gt;--now&lt;/span&gt; ollama
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You configure it with environment variables through &lt;code&gt;systemctl edit ollama&lt;/code&gt; — bind address, model directory, how long a model stays in memory, how many run at once. There is an official Docker image, so dropping Ollama into a container stack is a few lines of Compose. The &lt;a href="https://devtoolhub.com/run-ollama-on-digitalocean-droplet/" rel="noopener noreferrer"&gt;Ollama on a DigitalOcean droplet walkthrough&lt;/a&gt; shows the full server setup.&lt;/p&gt;

&lt;p&gt;Ollama is also open source under the MIT license. That matters if you need to vendor it, audit it, or ship it inside a product without checking terms.&lt;/p&gt;

&lt;h2&gt;
  
  
  What LM Studio does best
&lt;/h2&gt;

&lt;p&gt;LM Studio is the better tool for exploring models. The app has a Hugging Face search built in. You can find a model, read its card, check quantization options with size estimates, and download it without leaving the window. The chat interface shows tokens per second, lets you edit the system prompt live, and keeps conversation history.&lt;/p&gt;

&lt;p&gt;It ships two inference engines: &lt;a href="https://github.com/ggml-org/llama.cpp" rel="noopener noreferrer"&gt;llama.cpp&lt;/a&gt; on every platform, and &lt;a href="https://github.com/ml-explore/mlx" rel="noopener noreferrer"&gt;Apple MLX&lt;/a&gt; on Apple Silicon. For MLX-format models on an M-series Mac, that path is often faster than the llama.cpp build. Ollama does not use MLX.&lt;/p&gt;

&lt;p&gt;LM Studio also has built-in document chat. You attach PDFs or text files and the app handles the retrieval step, which Ollama leaves to you. Since July 2025, LM Studio is &lt;a href="https://lmstudio.ai/blog/free-for-work" rel="noopener noreferrer"&gt;free for use at work&lt;/a&gt; with no form to fill out, though it stays closed source.&lt;/p&gt;

&lt;p&gt;For teams that want a GUI, the &lt;a href="https://devtoolhub.com/install-lm-studio-ollama/" rel="noopener noreferrer"&gt;LM Studio and Ollama setup guide&lt;/a&gt; covers installing both on the same machine.&lt;/p&gt;

&lt;h2&gt;
  
  
  Ollama vs LM Studio: performance and model support
&lt;/h2&gt;

&lt;p&gt;On the same GGUF file and quantization level, expect similar speed from both tools, because both call llama.cpp. The practical performance gaps come from three things:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Apple MLX&lt;/strong&gt;: LM Studio can run MLX-optimized models on Apple Silicon, which often beats the GGUF path on the same Mac. This is LM Studio's clearest speed advantage.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Model loading defaults&lt;/strong&gt;: Ollama keeps a model in memory for 5 minutes after the last request by default, then unloads it. A cold request after that pays the load cost again. You change this with &lt;code&gt;OLLAMA_KEEP_ALIVE&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Context length defaults&lt;/strong&gt;: recent Ollama versions set the default context window based on available VRAM rather than a fixed 4096 tokens. Check &lt;code&gt;ollama ps&lt;/code&gt; to confirm a model is on the GPU and not partly on CPU.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Model availability is close to even. Ollama's &lt;code&gt;ollama.com/library&lt;/code&gt; is curated and versioned. LM Studio pulls from all of Hugging Face, so obscure or brand-new quants show up there first. Both let you load your own GGUF files.&lt;/p&gt;

&lt;p&gt;⚠️ Note: Ollama's model names like &lt;code&gt;llama3.1&lt;/code&gt; default to a 4-bit quantization (&lt;code&gt;Q4_K_M&lt;/code&gt;). If you compare it against a Q8 model in LM Studio, you are measuring quantization, not the tools.&lt;/p&gt;

&lt;h2&gt;
  
  
  Ollama vs LM Studio: which one should you use?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Use Ollama if&lt;/strong&gt; you are building something. That covers a backend service, a CLI tool, a RAG pipeline, a coding assistant in your editor, or anything running in Docker or on a Linux box. Ollama's always-on server and clean REST API are the right fit. The &lt;a href="https://devtoolhub.com/ollama-cloud-free-vs-pro-limits-pricing-2026/" rel="noopener noreferrer"&gt;Ollama Cloud free vs Pro limits guide&lt;/a&gt; covers the hosted option when local hardware runs out.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use LM Studio if&lt;/strong&gt; you want to work with models directly. Comparing outputs across models, testing prompts, chatting with a local model, or running quick document Q&amp;amp;A all go faster in the GUI. It also helps on a Mac, where the MLX engine speeds up supported models. You never touch a terminal.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;For most developers&lt;/strong&gt;, the honest answer is both. In practice, that means LM Studio on your workstation for testing and model discovery, but Ollama on your server for the actual application. They do not conflict — different ports, different jobs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Can you run both together?
&lt;/h2&gt;

&lt;p&gt;Yes. Ollama listens on &lt;code&gt;11434&lt;/code&gt; and LM Studio on &lt;code&gt;1234&lt;/code&gt;, so there is no port clash. A common setup is LM Studio on a laptop for prototyping and Ollama on a home server or VPS for anything that needs to stay up. Both expose an OpenAI-compatible endpoint, so pointing your code from one to the other is a base-URL change.&lt;/p&gt;

&lt;p&gt;The one resource they share is your GPU. Running a large model in both at once will exhaust VRAM. Load models in one tool at a time unless you have headroom to spare.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Q: Is Ollama or LM Studio faster?&lt;/strong&gt;&lt;br&gt;
A: On the same model file and quantization, they perform about the same because both use llama.cpp. LM Studio is faster for MLX-format models on Apple Silicon Macs, since it supports Apple's MLX engine and Ollama does not.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Is LM Studio free for commercial use?&lt;/strong&gt;&lt;br&gt;
A: Yes. Since July 2025, LM Studio is free for both personal and work use with no license form required. It remains closed source, and some enterprise features are paid.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Can I use LM Studio models in Ollama?&lt;/strong&gt;&lt;br&gt;
A: Both tools run GGUF files. A model downloaded by LM Studio imports into Ollama with a Modelfile that points at the &lt;code&gt;.gguf&lt;/code&gt; path. Any GGUF from Hugging Face works in both.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Does Ollama have a GUI?&lt;/strong&gt;&lt;br&gt;
A: Ollama ships a basic desktop app for chatting with models, added in 2025. It is far simpler than LM Studio's interface and has no model browser or document RAG.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Which one is better for a server?&lt;/strong&gt;&lt;br&gt;
A: Ollama. It runs as a background service with a &lt;code&gt;systemd&lt;/code&gt; unit, has an official Docker image, and is configured through environment variables. LM Studio can run headless with its daemon, but Ollama is built for that job.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quick Summary:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Ollama is open source, CLI-first, and built to run as a background API server — the pick for servers, containers, and code.&lt;/li&gt;
&lt;li&gt;LM Studio is a closed-source but free desktop GUI with a Hugging Face model browser, live chat, and built-in document RAG — the pick for hands-on experimentation.&lt;/li&gt;
&lt;li&gt;Both use llama.cpp, so token speed is similar; LM Studio adds Apple MLX for faster inference on Apple Silicon.&lt;/li&gt;
&lt;li&gt;Ollama serves on port &lt;code&gt;11434&lt;/code&gt;, LM Studio on &lt;code&gt;1234&lt;/code&gt; — they run side by side without conflict.&lt;/li&gt;
&lt;li&gt;Ollama defaults to 4-bit (&lt;code&gt;Q4_K_M&lt;/code&gt;) quantization; match quant levels before comparing output quality.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The Ollama vs LM Studio decision really is that simple: Ollama for anything a program calls, LM Studio for anything you click. If you are deciding where to run models once local hardware is maxed out, read the &lt;a href="https://devtoolhub.com/ollama-cloud-free-vs-pro-limits-pricing-2026/" rel="noopener noreferrer"&gt;Ollama Cloud pricing and limits breakdown&lt;/a&gt; next.&lt;/p&gt;

</description>
      <category>ollama</category>
      <category>lmstudio</category>
      <category>localllm</category>
      <category>aitools</category>
    </item>
    <item>
      <title>Docker Desktop Licensing Cost: Pay or Switch?</title>
      <dc:creator>Amaresh Pelleti</dc:creator>
      <pubDate>Thu, 03 Sep 2026 11:56:20 +0000</pubDate>
      <link>https://dev.to/amareswer/docker-desktop-licensing-cost-pay-or-switch-16b9</link>
      <guid>https://dev.to/amareswer/docker-desktop-licensing-cost-pay-or-switch-16b9</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;Originally published on &lt;a href="https://devtoolhub.com/docker-desktop-licensing-cost/" rel="noopener noreferrer"&gt;DevToolHub&lt;/a&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Docker Desktop is free until your company crosses &lt;strong&gt;250 employees or $10 million in annual revenue&lt;/strong&gt; — then every developer running it owes a paid subscription, and at the Business tier that is &lt;strong&gt;$288 per seat per year&lt;/strong&gt; with no annual discount. For a 200-developer engineering org that is &lt;strong&gt;$57,600 a year&lt;/strong&gt; for a tool most people forget is even installed.&lt;/p&gt;

&lt;p&gt;This is the cost math nobody runs before the renewal quote lands — the kind of line item &lt;a href="https://devtoolhub.com/finops-apptio-cloudability-vs-vantage/" rel="noopener noreferrer"&gt;FinOps reviews&lt;/a&gt; tend to catch a year too late. Here is the real Docker Desktop licensing cost at different team sizes, what the free alternatives cost once you count migration time, and where the line is that makes switching worth it.&lt;/p&gt;

&lt;h2&gt;
  
  
  When does Docker Desktop require a paid subscription?
&lt;/h2&gt;

&lt;p&gt;The rule is in Section 3.2 of the &lt;a href="https://www.docker.com/legal/docker-subscription-service-agreement/" rel="noopener noreferrer"&gt;Docker Subscription Service Agreement&lt;/a&gt;. Docker Desktop use without a paid subscription is restricted to:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"(i) use for a non-commercial open source project and/or (ii) use in a commercial undertaking with fewer than 250 employees and less than US $10,000,000 (or equivalent local currency) in annual revenue."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Read the operators carefully:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The two thresholds inside clause (ii) are joined by &lt;strong&gt;and&lt;/strong&gt;. You stay free only if you are under 250 employees &lt;strong&gt;and&lt;/strong&gt; under $10M revenue. Cross &lt;strong&gt;either&lt;/strong&gt; one and clause (ii) no longer applies to you.&lt;/li&gt;
&lt;li&gt;A 40-person startup that raised a round and is now doing $12M ARR needs to pay. A 300-person company doing $4M does too.&lt;/li&gt;
&lt;li&gt;This covers &lt;strong&gt;Docker Desktop specifically&lt;/strong&gt; — the macOS/Windows application and its bundled Linux VM. Docker Engine (&lt;code&gt;dockerd&lt;/code&gt;) on a Linux host is Apache 2.0 and always free. So is the &lt;code&gt;docker&lt;/code&gt; CLI itself.&lt;/li&gt;
&lt;li&gt;Docker Hub pull limits are a separate policy with separate thresholds. This article is only about the Desktop subscription.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Enforcement is contractual, not technical. Docker Desktop will not stop working if you are out of compliance. But it phones home, Docker has sent true-up invoices before, and an unlicensed deployment is exactly the kind of thing that surfaces during acquisition due diligence or a SOC 2 audit.&lt;/p&gt;

&lt;h2&gt;
  
  
  Docker Desktop licensing cost: the 2026 price tiers
&lt;/h2&gt;

&lt;p&gt;Docker raised prices on &lt;strong&gt;December 10, 2024&lt;/strong&gt; — Pro went up about 80%, Team about 67%. Current per-seat pricing:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Plan&lt;/th&gt;
&lt;th&gt;Monthly&lt;/th&gt;
&lt;th&gt;Annual (per month)&lt;/th&gt;
&lt;th&gt;Annual (per year)&lt;/th&gt;
&lt;th&gt;Seat cap&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Personal&lt;/td&gt;
&lt;td&gt;$0&lt;/td&gt;
&lt;td&gt;$0&lt;/td&gt;
&lt;td&gt;$0&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pro&lt;/td&gt;
&lt;td&gt;$11&lt;/td&gt;
&lt;td&gt;$9&lt;/td&gt;
&lt;td&gt;$108&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Team&lt;/td&gt;
&lt;td&gt;$16&lt;/td&gt;
&lt;td&gt;$15&lt;/td&gt;
&lt;td&gt;$180&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Business&lt;/td&gt;
&lt;td&gt;$24&lt;/td&gt;
&lt;td&gt;$24&lt;/td&gt;
&lt;td&gt;$288&lt;/td&gt;
&lt;td&gt;unlimited&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two things to note. &lt;strong&gt;Team caps at 100 seats.&lt;/strong&gt; Any org that needs more than 100 licenses is on Business, whether it wants the SSO and hardened-image features or not. And &lt;strong&gt;Business has no annual discount&lt;/strong&gt; — $24/month is the only price, $288/year.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Prices verified against docker.com/pricing in September 2026. Docker's last change was December 2024; check the current figures before you budget.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The real Docker Desktop licensing cost by team size
&lt;/h2&gt;

&lt;p&gt;"Seats" here means people who actually launch Docker Desktop, not total headcount. In most orgs that is the backend, platform, and QA engineers plus some data folks — call it 30-60% of an R&amp;amp;D team.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Docker Desktop seats&lt;/th&gt;
&lt;th&gt;Docker Team (annual)&lt;/th&gt;
&lt;th&gt;Docker Business&lt;/th&gt;
&lt;th&gt;OrbStack Pro&lt;/th&gt;
&lt;th&gt;Podman / Rancher / Colima&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;$1,800&lt;/td&gt;
&lt;td&gt;$2,880&lt;/td&gt;
&lt;td&gt;$960&lt;/td&gt;
&lt;td&gt;$0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;50&lt;/td&gt;
&lt;td&gt;$9,000&lt;/td&gt;
&lt;td&gt;$14,400&lt;/td&gt;
&lt;td&gt;$4,800&lt;/td&gt;
&lt;td&gt;$0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;$18,000&lt;/td&gt;
&lt;td&gt;$28,800&lt;/td&gt;
&lt;td&gt;$9,600&lt;/td&gt;
&lt;td&gt;$0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;200&lt;/td&gt;
&lt;td&gt;— (over cap)&lt;/td&gt;
&lt;td&gt;$57,600&lt;/td&gt;
&lt;td&gt;$19,200&lt;/td&gt;
&lt;td&gt;$0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;500&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;$144,000&lt;/td&gt;
&lt;td&gt;$48,000&lt;/td&gt;
&lt;td&gt;$0&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;a href="https://orbstack.dev/pricing" rel="noopener noreferrer"&gt;OrbStack&lt;/a&gt; Pro is $8/user/month or $96/year for commercial use, but it is &lt;strong&gt;macOS only&lt;/strong&gt; — a mixed Mac/Windows/Linux fleet cannot standardize on it. Podman Desktop, Rancher Desktop, and Colima are open source with no company-size clause: $0 at any scale, on any OS.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the free alternatives actually cost
&lt;/h2&gt;

&lt;p&gt;The sticker price of Podman Desktop is zero. The switch is not.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Migration engineering time.&lt;/strong&gt; Budget per developer:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;1-2 hours&lt;/strong&gt; if your workflow is &lt;code&gt;docker build&lt;/code&gt; / &lt;code&gt;docker run&lt;/code&gt; / &lt;code&gt;docker compose up&lt;/code&gt; on standard images. Podman and nerdctl are near drop-in; &lt;code&gt;alias docker=podman&lt;/code&gt; covers most of it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;4-8 hours&lt;/strong&gt; if you have &lt;a href="https://devtoolhub.com/docker-compose-recreates-containers/" rel="noopener noreferrer"&gt;Compose files&lt;/a&gt; using build secrets, custom networks, &lt;code&gt;depends_on&lt;/code&gt; health conditions, or host networking; if devs script against the Docker socket; or if you rely on Docker Desktop's file-sharing performance tuning on macOS.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Days, not hours&lt;/strong&gt; if you have Testcontainers-based test suites, dev containers wired to the Docker API, or bespoke Docker Desktop extensions.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Platform team time.&lt;/strong&gt; Updating onboarding docs, the &lt;code&gt;README&lt;/code&gt;, CI base images, and the internal setup script — figure 1-2 weeks of one engineer for a mid-size org, more if you support multiple OSes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A worked example — 50 seats:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Switching cost: 50 devs x 4 hours x $75/hr loaded ≈ &lt;strong&gt;$15,000&lt;/strong&gt;, plus ~60 hours of platform time ≈ &lt;strong&gt;$4,500&lt;/strong&gt;. Call it &lt;strong&gt;~$20,000 one-time&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Docker Business for 50 seats: &lt;strong&gt;$14,400/year&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Payback: &lt;strong&gt;~1.4 years.&lt;/strong&gt; After that it is $14,400/year saved, every year.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The bigger the team, the faster it pays back. The switching cost scales with team size, but so does the license you are avoiding. Past 100 seats you are forced onto Business anyway. At 200 seats the switch pays back in well under a year.&lt;/p&gt;

&lt;p&gt;For a &lt;strong&gt;10-person team&lt;/strong&gt;, the math flips: ~$5,000 to switch versus $1,800-$2,880/year. Only worth it if you were going to grow past 50 soon, or if you want off the licensing-risk treadmill entirely.&lt;/p&gt;

&lt;h2&gt;
  
  
  The honest trade-offs
&lt;/h2&gt;

&lt;p&gt;Switching is not free of downside:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Podman on macOS/Windows&lt;/strong&gt; runs a Linux VM, just like Docker Desktop did. Its file-sync performance for large bind mounts has historically been a notch behind. Test your actual repo before committing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Docker Desktop's GUI&lt;/strong&gt; for inspecting containers, images, and volumes is genuinely good. Podman Desktop has closed most of the gap; Rancher Desktop's is more basic.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Testcontainers&lt;/strong&gt; officially supports Podman now, but expect to set &lt;code&gt;DOCKER_HOST&lt;/code&gt; and &lt;code&gt;TESTCONTAINERS_RYUK_DISABLED&lt;/code&gt; and debug a few CI runs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Docker Business features you would lose&lt;/strong&gt; — Hardened Docker Desktop, registry access management, SSO enforcement on the desktop app. If you bought Business for the compliance controls rather than just the right to run it, an alternative does not replace that.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;OrbStack&lt;/strong&gt; is the least-effort paid switch: it is Docker-socket-compatible, noticeably faster and lighter than Docker Desktop, and $96/seat/year undercuts Business by two thirds. But macOS only.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For a deeper feature-by-feature look at the leading alternative, see our &lt;a href="https://devtoolhub.com/docker-vs-podman-which-container-tool-should-you-use/" rel="noopener noreferrer"&gt;Docker vs Podman comparison&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The verdict
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Under 250 employees and under $10M revenue:&lt;/strong&gt; Docker Desktop is free. Do nothing. Revisit if you raise a large round or an acquisition is on the table.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Over a threshold, fewer than ~30 seats:&lt;/strong&gt; pay for Docker Team ($180/seat/year). Not worth the disruption yet.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Over a threshold, 30-100 seats:&lt;/strong&gt; run the payback math. A switch to Podman or Rancher Desktop typically pays back in 1-2 years and removes the licensing-risk line item permanently. All-Mac shops should price OrbStack Pro first — it is the cheapest low-effort option.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Over 100 seats:&lt;/strong&gt; you are on Docker Business at $288/seat/year with no discount. The switch is close to a no-brainer here. At 200 seats you are choosing between $57,600 a year forever and a one-time project that costs less than a single year of it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Whatever you decide, decide it deliberately. The worst outcome on Docker Desktop licensing cost is the one you never chose. That is finding out during an audit that 180 engineers have run Docker Desktop unlicensed for two years, then paying the back bill at list price. If cost control across the stack is the real goal, the same logic applies to &lt;a href="https://devtoolhub.com/ai-finops-kubernetes-cost-optimization/" rel="noopener noreferrer"&gt;Kubernetes spend&lt;/a&gt; and every other per-seat developer tool.&lt;/p&gt;

</description>
      <category>docker</category>
      <category>podman</category>
      <category>containers</category>
      <category>finops</category>
    </item>
    <item>
      <title>Docker Compose Recreates Containers After 5.5 Upgrade</title>
      <dc:creator>Amaresh Pelleti</dc:creator>
      <pubDate>Wed, 02 Sep 2026 11:51:56 +0000</pubDate>
      <link>https://dev.to/amareswer/docker-compose-recreates-containers-after-55-upgrade-429j</link>
      <guid>https://dev.to/amareswer/docker-compose-recreates-containers-after-55-upgrade-429j</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;Originally published on &lt;a href="https://devtoolhub.com/docker-compose-recreates-containers/" rel="noopener noreferrer"&gt;DevToolHub&lt;/a&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Docker Compose recreates containers you didn't touch the first time you run &lt;code&gt;docker compose up&lt;/code&gt; after upgrading to 5.5. That's expected, not a bug. Compose 5.5 &lt;a href="https://github.com/docker/compose/releases/tag/v5.5.0" rel="noopener noreferrer"&gt;overhauls image digest reconciliation&lt;/a&gt; — the logic that decides whether an existing container matches what your &lt;code&gt;docker-compose.yml&lt;/code&gt; currently describes. The release notes say it plainly: "existing containers may be recreated the first time you run &lt;code&gt;compose up&lt;/code&gt; after upgrading, as image digests are re-evaluated using the new logic."&lt;/p&gt;

&lt;p&gt;It's a one-time event, not a recurring problem. But if you weren't expecting it, a mass container restart looks like an incident. Here's what actually changed, why it happened once and won't again, and what else shipped in the same release.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Docker Compose Recreates Containers on the First Upgrade Run
&lt;/h2&gt;

&lt;p&gt;Compose decides whether to recreate a container by comparing a hash of its current configuration against what's already running, and image digests are part of that hash. Version 5.5 changes how those digests get evaluated — more accurately, according to the release notes, which is the whole point of the change. But "more accurate" means containers created under the old evaluation logic look different from what the new logic expects, even if nothing in your compose file actually changed.&lt;/p&gt;

&lt;p&gt;So on your first &lt;code&gt;compose up&lt;/code&gt; after upgrading, Compose reconciles every service against the new digest logic. Anything that reads as "different" gets recreated once. After that first pass, your containers are baselined against the new logic and stay put on subsequent runs, the same way Compose has always behaved for unchanged services.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Recurring Complaint Behind the Overhaul
&lt;/h2&gt;

&lt;p&gt;This overhaul didn't come out of nowhere. Complaints about containers getting recreated on &lt;code&gt;compose up&lt;/code&gt; when nothing changed have followed the v2 rewrite from the start. &lt;a href="https://github.com/docker/compose/issues/10307" rel="noopener noreferrer"&gt;A 2023 GitHub issue&lt;/a&gt; documents the pattern exactly: a user changed one service's image tag, ran &lt;code&gt;compose up -d&lt;/code&gt; again with zero further changes, and watched Compose recreate every dependent container anyway — "I did NOT make any changes here, I simply ran compose up again, and again it recreated the containers!"&lt;/p&gt;

&lt;p&gt;That specific bug was a v2.16 regression, and Docker fixed it within weeks in v2.17.0. It also involved config-hash comparison generally, not digest reconciliation specifically. But it's the same category of problem, and it kept resurfacing in new forms: Compose's decision about whether a container needs recreating didn't match what a human expected, and each occurrence got a one-off fix. The 5.5 overhaul reads like Docker rebuilding that logic instead of patching around it again.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Else Shipped in Docker Compose 5.5
&lt;/h2&gt;

&lt;p&gt;The digest change is the headline, but a few other fixes matter if you run Compose in CI or with build-heavy setups.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;compose pull&lt;/code&gt; now honors &lt;code&gt;pull_policy&lt;/code&gt; refresh windows&lt;/strong&gt; — &lt;code&gt;daily&lt;/code&gt;, &lt;code&gt;weekly&lt;/code&gt;, and &lt;code&gt;every_N&lt;/code&gt; settings actually control when Compose re-pulls an image, instead of pulling on every invocation regardless of policy.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Build-only services skip pulling a default image reference&lt;/strong&gt; they were never going to use, which was previously wasted network time on every &lt;code&gt;up&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Compose stopped pruning every dangling image in a project&lt;/strong&gt; on cleanup, a behavior that was too aggressive for anyone sharing base images across services.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The file watcher now skips unreadable directories&lt;/strong&gt; instead of failing the entire watch session, and Compose tolerates containers whose image record has gone missing instead of erroring out.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of these force a workflow change on their own. They're the kind of fixes you only notice once they stop happening. If you're still deciding between Compose-managed Docker and a daemonless alternative, the &lt;a href="https://devtoolhub.com/docker-vs-podman-which-container-tool-should-you-use/" rel="noopener noreferrer"&gt;Docker vs Podman comparison&lt;/a&gt; is worth reading before you standardize either way.&lt;/p&gt;

&lt;p&gt;[IMAGE: articles/images/2026-08-25-docker-compose-recreates-containers-diagram.png | alt: "image digest reconciliation flow in Compose 5.5, from compose up to a stable container state"]&lt;/p&gt;

&lt;h2&gt;
  
  
  What to Check Before You Upgrade
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Expect one restart cycle, not zero.&lt;/strong&gt; If your services hold in-memory state or long-lived connections, plan the first post-upgrade &lt;code&gt;compose up&lt;/code&gt; like a deploy, not a routine command.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Check &lt;code&gt;pull_policy&lt;/code&gt; settings if you rely on scheduled pulls.&lt;/strong&gt; The refresh-window fix means &lt;code&gt;daily&lt;/code&gt;/&lt;code&gt;weekly&lt;/code&gt;/&lt;code&gt;every_N&lt;/code&gt; now actually gate pulls — if you were working around the old always-pull behavior with a wrapper script, that workaround may now be redundant or conflicting.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Run &lt;code&gt;docker compose config --hash "*"&lt;/code&gt; before and after upgrading&lt;/strong&gt; to see exactly which services Compose considers changed, so the recreation isn't a surprise mid-deploy.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Don't run the upgrade for the first time against production.&lt;/strong&gt; Because the recreation is real, not cosmetic, test it against staging first if your containers have any startup cost or connection-draining behavior worth protecting.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;⚠️ &lt;strong&gt;Important:&lt;/strong&gt; if any service uses &lt;code&gt;restart: always&lt;/code&gt; alongside a health check with a short grace period, a mass recreation event can trigger cascading restart attempts before dependent services are back up. Stagger the upgrade across environments rather than rolling every host to 5.5 at once. The &lt;a href="https://devtoolhub.com/top-10-docker-errors-and-fixes-devops-guide-2024/" rel="noopener noreferrer"&gt;Docker errors and fixes guide&lt;/a&gt; covers what a cascading restart loop looks like if you want to recognize it fast.&lt;/p&gt;

&lt;h2&gt;
  
  
  Verdict: Should You Upgrade?
&lt;/h2&gt;

&lt;p&gt;Yes, for most setups. The one-time recreation is a known, explained side effect, not a regression, and the underlying fix addresses a category of complaint that's followed Compose v2 for years. Teams running build-only services or relying on &lt;code&gt;pull_policy&lt;/code&gt; refresh windows get a direct, measurable improvement, not just a bug fix.&lt;/p&gt;

&lt;p&gt;The exception is anything with strict uptime requirements and no rolling-deploy story for Compose-managed services — for those, schedule the upgrade like you would a dependency bump that touches container lifecycle, with a maintenance window and a rollback plan, rather than pulling 5.5 straight into a production host. If you're layering Compose on top of a base you're still learning, the &lt;a href="https://devtoolhub.com/docker-beginners-guide-2025/" rel="noopener noreferrer"&gt;Docker beginner's guide&lt;/a&gt; and the &lt;a href="https://devtoolhub.com/docker-best-practices-a-beginners-guide/" rel="noopener noreferrer"&gt;Docker best practices guide&lt;/a&gt; are worth a read before you change upgrade habits on production systems.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Q: Will Docker Compose keep recreating my containers every time I run &lt;code&gt;compose up&lt;/code&gt;?&lt;/strong&gt;&lt;br&gt;
A: No. The recreation happens once, on your first &lt;code&gt;compose up&lt;/code&gt; after upgrading to 5.5, while Compose re-evaluates image digests under the new logic. After that first pass, unchanged services stay running exactly as before.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Do I need to change my &lt;code&gt;docker-compose.yml&lt;/code&gt; to avoid the recreation?&lt;/strong&gt;&lt;br&gt;
A: No changes are required. The recreation is triggered by the new digest evaluation logic itself, not by anything in your compose file. There's no configuration flag to skip it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Does the &lt;code&gt;pull_policy&lt;/code&gt; refresh window change affect existing &lt;code&gt;daily&lt;/code&gt;/&lt;code&gt;weekly&lt;/code&gt; settings?&lt;/strong&gt;&lt;br&gt;
A: It makes them work as documented. Before 5.5, &lt;code&gt;compose pull&lt;/code&gt; didn't reliably respect those refresh windows; now it does, so a pull that was firing more often than intended should slow down to match your configured policy.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Is this related to the old "Compose v2 recreates containers for no reason" complaints?&lt;/strong&gt;&lt;br&gt;
A: It's the same category of problem — Compose deciding a container needs recreating when a human wouldn't expect it — but a different specific mechanism. This overhaul targets digest evaluation, not general config-hash comparison.&lt;/p&gt;

&lt;h2&gt;
  
  
  Quick Summary
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Docker Compose 5.5 overhauls image digest reconciliation, and existing containers get recreated once on the first &lt;code&gt;compose up&lt;/code&gt; after upgrading&lt;/li&gt;
&lt;li&gt;After that first pass, Compose behaves normally — unchanged services stay running on subsequent runs&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;compose pull&lt;/code&gt; now actually honors &lt;code&gt;pull_policy&lt;/code&gt; refresh windows (&lt;code&gt;daily&lt;/code&gt;, &lt;code&gt;weekly&lt;/code&gt;, &lt;code&gt;every_N&lt;/code&gt;), which previously didn't reliably apply&lt;/li&gt;
&lt;li&gt;Build-only services no longer pull an unused default image reference, and Compose no longer prunes every dangling image in a project on cleanup&lt;/li&gt;
&lt;li&gt;Test the upgrade in staging first if any service has real startup cost, in-memory state, or strict health-check timing&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Run &lt;code&gt;docker compose config --hash "*"&lt;/code&gt; before you upgrade so you know in advance which services Compose considers changed. Docker Compose recreates containers exactly once during this transition, and treating that first post-upgrade &lt;code&gt;compose up&lt;/code&gt; like a deploy rather than a routine command is the difference between a known side effect and an incident.&lt;/p&gt;

</description>
      <category>docker</category>
      <category>dockercompose</category>
      <category>containers</category>
      <category>devops</category>
    </item>
    <item>
      <title>CKA vs CKAD vs CKS: Which Cert to Take First</title>
      <dc:creator>Amaresh Pelleti</dc:creator>
      <pubDate>Tue, 01 Sep 2026 18:04:27 +0000</pubDate>
      <link>https://dev.to/amareswer/cka-vs-ckad-vs-cks-which-cert-to-take-first-31jo</link>
      <guid>https://dev.to/amareswer/cka-vs-ckad-vs-cks-which-cert-to-take-first-31jo</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;Originally published on &lt;a href="https://devtoolhub.com/cka-vs-ckad-vs-cks/" rel="noopener noreferrer"&gt;DevToolHub&lt;/a&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Three Kubernetes certifications exist, and each one tests a different job. The CKA is for people who run clusters. The CKAD is for people who deploy apps onto them. Securing those clusters is what the CKS covers. When you are choosing, CKA vs CKAD vs CKS comes down to your actual role, not which exam sounds hardest.&lt;/p&gt;

&lt;p&gt;This guide covers what each exam tests, the format, the cost, the big 2025 CKA changes, and how long prep really takes.&lt;/p&gt;

&lt;p&gt;[IMAGE: articles/images/2026-08-28-cka-vs-ckad-vs-cks-featured.png | alt: "CKA vs CKAD vs CKS Kubernetes certification comparison"]&lt;/p&gt;

&lt;h2&gt;
  
  
  CKA vs CKAD vs CKS: what's the difference?
&lt;/h2&gt;

&lt;p&gt;Cluster operations are the CKA's focus: installation, upgrades, networking, storage, and fixing broken nodes. The CKAD is about application delivery, so it covers pod design, config, health probes, and multi-container patterns. Security hardening drives the CKS, which means RBAC, network policies, supply-chain checks, and runtime detection. All three are hands-on, run from a terminal, and use the same Kubernetes version. The CKS also requires a passed CKA first.&lt;/p&gt;

&lt;p&gt;Here is the split by curriculum weight, taken from the current &lt;a href="https://www.cncf.io/training/certification/cka/" rel="noopener noreferrer"&gt;CNCF certification pages&lt;/a&gt;.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Exam&lt;/th&gt;
&lt;th&gt;Who it's for&lt;/th&gt;
&lt;th&gt;Heaviest domains&lt;/th&gt;
&lt;th&gt;Prerequisite&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;CKA&lt;/td&gt;
&lt;td&gt;Cluster admins, SREs, platform engineers&lt;/td&gt;
&lt;td&gt;Troubleshooting 30%, Cluster architecture 25%, Services &amp;amp; networking 20%&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CKAD&lt;/td&gt;
&lt;td&gt;Backend and app developers shipping to Kubernetes&lt;/td&gt;
&lt;td&gt;App environment, config &amp;amp; security 25%, Design &amp;amp; build 20%, Deployment 20%&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CKS&lt;/td&gt;
&lt;td&gt;Security engineers, platform teams with a security remit&lt;/td&gt;
&lt;td&gt;Microservice vulnerabilities 20%, Supply chain 20%, Monitoring &amp;amp; runtime 20%&lt;/td&gt;
&lt;td&gt;Passed CKA&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The CKA and CKAD overlap on services and networking. Past that, they pull apart. A developer rarely drains a node, and an admin rarely writes a readiness probe from scratch.&lt;/p&gt;

&lt;h2&gt;
  
  
  CKA vs CKAD vs CKS: which one should you take first?
&lt;/h2&gt;

&lt;p&gt;Take the CKA first if you run infrastructure, and take the CKAD first if you write application code. Most people should start with the CKA. It is the broadest of the three, it has no prerequisite, and it is the one hiring managers recognize. The CKS is not a starting point, because you must pass the CKA before you can attempt it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Platform or DevOps engineer:&lt;/strong&gt; take the CKA. It maps to what you already do, and it unlocks the CKS later.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Backend developer:&lt;/strong&gt; take the CKAD. Your time goes into manifests and rollouts, not &lt;code&gt;etcd&lt;/code&gt; backups.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Security is your job title:&lt;/strong&gt; CKA, then CKS right after. Budget for both from the start.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Verdict:&lt;/strong&gt; the CKA is the default answer for anyone unsure. It carries the most weight on a resume and opens the most doors, the CKS included. Our &lt;a href="https://devtoolhub.com/the-ultimate-guide-to-popular-devops-certifications/" rel="noopener noreferrer"&gt;guide to DevOps certifications&lt;/a&gt; shows how these sit next to the AWS and Azure tracks.&lt;/p&gt;

&lt;h2&gt;
  
  
  CKA vs CKAD vs CKS exam format, scoring, and rules
&lt;/h2&gt;

&lt;p&gt;Every exam is a two-hour, online-proctored, performance-based test taken from a command line. There is no multiple choice. You solve real tasks on a live cluster: create resources, fix broken configs, and verify your work. The Linux Foundation delivers all three through &lt;a href="https://docs.linuxfoundation.org/tc-docs/certification/faq-cka-ckad-cks" rel="noopener noreferrer"&gt;PSI's "Bridge" proctoring platform&lt;/a&gt; with the PSI Secure Browser, so you need a quiet room, a clear desk, and a webcam ID check.&lt;/p&gt;

&lt;p&gt;Passing scores are fixed:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;CKA:&lt;/strong&gt; 66%&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;CKAD:&lt;/strong&gt; 66%&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;CKS:&lt;/strong&gt; 67%&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The Linux Foundation does not publish an official task count for any of the three exams. The Killer.sh simulator bundled with each exam runs 17 questions, and test-takers generally report a similar number on the real exam. Time management matters more than raw knowledge here, because each task is weighted and partial credit is common.&lt;/p&gt;

&lt;p&gt;Every exam is open-book, but only for a fixed list of sites. During the CKA and CKAD you may open &lt;code&gt;kubernetes.io/docs&lt;/code&gt;, &lt;code&gt;kubernetes.io/blog&lt;/code&gt;, and &lt;code&gt;helm.sh/docs&lt;/code&gt;, and the CKA also allows &lt;code&gt;gateway-api.sigs.k8s.io&lt;/code&gt;. The CKS adds &lt;code&gt;falco.org/docs&lt;/code&gt;, &lt;code&gt;etcd.io/docs&lt;/code&gt;, &lt;code&gt;docs.cilium.io&lt;/code&gt;, &lt;code&gt;istio.io/latest/docs/&lt;/code&gt;, the NGINX Ingress controller docs, and the Kubernetes SIGs &lt;code&gt;bom&lt;/code&gt; docs. You can use the search on &lt;code&gt;kubernetes.io/docs&lt;/code&gt;, but you cannot open an external search result. The full list lives on the &lt;a href="https://docs.linuxfoundation.org/tc-docs/certification/certification-resources-allowed" rel="noopener noreferrer"&gt;allowed-resources page&lt;/a&gt;. Practice navigating the docs fast, because tab-hunting mid-exam burns minutes you do not have.&lt;/p&gt;

&lt;p&gt;⚠️ Note: certifications earned on or after April 1, 2024 are valid for &lt;strong&gt;2 years&lt;/strong&gt; (older ones keep their 3-year term). Renew by retaking and passing the current exam before yours expires.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the CKA, CKAD, and CKS certifications cost
&lt;/h2&gt;

&lt;p&gt;Each exam costs &lt;strong&gt;$445&lt;/strong&gt; on its own. The Linux Foundation also sells two bundles: exam plus a THRIVE-ONE annual subscription for &lt;strong&gt;$625&lt;/strong&gt;, or exam plus the matching instructor course (LFS258 for CKA, LFD259 for CKAD, LFS260 for CKS) for &lt;strong&gt;$645&lt;/strong&gt;. Every purchase includes one free retake and a 12-month window to schedule and sit the exam.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Purchase&lt;/th&gt;
&lt;th&gt;Price (USD)&lt;/th&gt;
&lt;th&gt;Includes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Exam only&lt;/td&gt;
&lt;td&gt;$445&lt;/td&gt;
&lt;td&gt;Two attempts (one retake), two Killer.sh simulator sessions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Exam + THRIVE-ONE&lt;/td&gt;
&lt;td&gt;$625&lt;/td&gt;
&lt;td&gt;The above, plus a year of Linux Foundation training access&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Exam + course&lt;/td&gt;
&lt;td&gt;$645&lt;/td&gt;
&lt;td&gt;The above, plus the official prep course&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Wait for a sale. The Linux Foundation runs discount codes on certification bundles regularly, especially around Black Friday, CyberMonday, and KubeCon. Check for an active coupon before you buy — paying full price when one exists is the most common money mistake with these exams. Your enrollment does not expire for a year, so buy on sale and schedule later.&lt;/p&gt;

&lt;h2&gt;
  
  
  What changed in the CKA exam in 2025
&lt;/h2&gt;

&lt;p&gt;The CKA changed on February 18, 2025. The five domains and their weights stayed the same. The competencies inside them were reworked. The &lt;a href="https://training.linuxfoundation.org/certified-kubernetes-administrator-cka-program-changes/" rel="noopener noreferrer"&gt;official program-changes page&lt;/a&gt; confirms what got added: Gateway API (&lt;code&gt;GatewayClass&lt;/code&gt;, &lt;code&gt;Gateway&lt;/code&gt;, &lt;code&gt;HTTPRoute&lt;/code&gt;), Helm and Kustomize for installing cluster components, and CRDs and operators. Troubleshooting is the single heaviest domain at 30%.&lt;/p&gt;

&lt;p&gt;For prep, that means three things. First, you have to know &lt;a href="https://devtoolhub.com/helm-commands-templates-best-practices/" rel="noopener noreferrer"&gt;Helm basics&lt;/a&gt; and Kustomize overlays, not just raw manifests. Second, Gateway API is now testable alongside Ingress, so learn both. Third, the troubleshooting weight rewards speed at reading logs, describing resources, and checking &lt;code&gt;kubectl get events&lt;/code&gt; before you change anything. CNCF versions the curriculum per Kubernetes release (currently v1.35), so download the current curriculum PDF from the &lt;a href="https://github.com/cncf/curriculum" rel="noopener noreferrer"&gt;CNCF curriculum repo&lt;/a&gt; before you book.&lt;/p&gt;

&lt;p&gt;If you studied from a pre-2025 course, your material has no Gateway API, Helm, Kustomize, or CRD coverage at all.&lt;/p&gt;

&lt;p&gt;[IMAGE: articles/images/2026-08-28-cka-vs-ckad-vs-cks-diagram.png | alt: "Study path from fundamentals through the Killer.sh simulator to booking the Kubernetes exam"]&lt;/p&gt;

&lt;h2&gt;
  
  
  How long CKA, CKAD, and CKS preparation takes
&lt;/h2&gt;

&lt;p&gt;There is no official prep-time guidance. Real numbers swing widely with how much Kubernetes you already run day to day. Community reports cluster around 40 to 60 hours for the CKA from a working knowledge of Kubernetes, and 20 to 30 hours for the CKAD once you hold the CKA. One engineer who passed the CKAD in January 2026 logged about 20 hours, having taken the CKA the same month. A candidate preparing for the CKA around a full-time job reported 38 to 44 hours over roughly six weeks. Treat these as rough anchors, not targets.&lt;/p&gt;

&lt;p&gt;The CKS usually needs less raw prep time — community estimates land around 25 to 35 hours — but only because it assumes everything the CKA already taught you. Do not book it until CKA reflexes are automatic.&lt;/p&gt;

&lt;p&gt;Two things decide pass or fail beyond knowledge:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;kubectl speed.&lt;/strong&gt; Use imperative commands such as &lt;code&gt;kubectl run --dry-run=client -o yaml&lt;/code&gt;, and set up aliases and shell autocompletion in the first two minutes of the exam.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Remote-desktop fluency.&lt;/strong&gt; The exam console has its own copy-paste behavior. Practice it in Killer.sh until it is muscle memory, because fighting the clipboard mid-task has failed people who knew the material cold.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  A study path for your first Kubernetes exam
&lt;/h2&gt;

&lt;p&gt;This path works for the CKA or the CKAD. Swap the course for CKS once your CKA is done.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Learn the fundamentals first.&lt;/strong&gt; If pods, deployments, and services still feel fuzzy, start with our &lt;a href="https://devtoolhub.com/kubernetes-complete-guide/" rel="noopener noreferrer"&gt;Kubernetes guide&lt;/a&gt; and the &lt;a href="https://devtoolhub.com/kubernetes-commands/" rel="noopener noreferrer"&gt;kubectl commands&lt;/a&gt; reference before any exam prep.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Take one structured course.&lt;/strong&gt; KodeKloud and the official Linux Foundation courses both track the current curriculum. Pick one, then finish it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Do every lab twice.&lt;/strong&gt; Hands-on repetition beats re-watching videos. Rebuild each scenario from an empty cluster.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Practice &lt;a href="https://devtoolhub.com/kubernetes-rbac-tutorial/" rel="noopener noreferrer"&gt;RBAC&lt;/a&gt; and network policies by hand.&lt;/strong&gt; Both show up across all three exams and trip people who only read about them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sit both Killer.sh sessions.&lt;/strong&gt; The simulator is harder than the real exam by design. Score above the pass line on the second attempt before you book.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Book the real exam within a week of your second simulator.&lt;/strong&gt; Skills fade fast once you stop drilling.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;For the CKS specifically, add hands-on time with Falco for runtime security, Trivy image scanning, and pod security standards.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Q: Is CKA or CKAD harder?&lt;/strong&gt;&lt;br&gt;
A: Most people find the CKA harder. It covers more ground, its troubleshooting section is unpredictable, and it now includes Helm, Kustomize, and Gateway API. The CKAD is narrower and more forgiving on time. Someone who already holds the CKA can usually pass the CKAD with about 20 hours of focused prep.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Do I need the CKA before the CKS?&lt;/strong&gt;&lt;br&gt;
A: Yes. The official rule is that you must have taken and passed the CKA before attempting the CKS. There is no such prerequisite for the CKA or the CKAD, which you can take in any order.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: How long are the certifications valid?&lt;/strong&gt;&lt;br&gt;
A: Certifications earned on or after April 1, 2024 are valid for two years. Older certifications keep their original three-year term. You renew by passing the current version of the exam before your certification expires.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Can I use the Kubernetes documentation during the exam?&lt;/strong&gt;&lt;br&gt;
A: Yes, but only specific sites. The CKA and CKAD allow &lt;code&gt;kubernetes.io/docs&lt;/code&gt;, &lt;code&gt;kubernetes.io/blog&lt;/code&gt;, and &lt;code&gt;helm.sh/docs&lt;/code&gt;. The CKS allows several more, including the Falco, Cilium, and etcd docs. You can search within those domains, but you cannot open an external search result.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Which certification helps most with getting hired?&lt;/strong&gt;&lt;br&gt;
A: The CKA. It is the most widely recognized of the three and it maps to the largest number of job descriptions. The CKAD matters for developer roles specifically, and the CKS is a specialist add-on that means little without the CKA behind it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quick Summary:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The CKA is for cluster operators, the CKAD for app developers, and the CKS for security engineers; start with the CKA unless you only write application code.&lt;/li&gt;
&lt;li&gt;All three are two-hour, performance-based, command-line exams on the same Kubernetes version, delivered through PSI Bridge.&lt;/li&gt;
&lt;li&gt;Passing scores are 66% for the CKA and CKAD and 67% for the CKS; each exam is $445 with one free retake, and bundles run $625 to $645.&lt;/li&gt;
&lt;li&gt;The CKA competencies were revised on February 18, 2025 to add Gateway API, Helm, Kustomize, and CRDs; troubleshooting is the heaviest domain at 30%.&lt;/li&gt;
&lt;li&gt;Budget 40 to 60 hours for the CKA and 20 to 30 for the CKAD after it, and buy during a Linux Foundation sale rather than at full price.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Still deciding between CKA vs CKAD vs CKS? Default to the CKA, set up &lt;code&gt;kubectl&lt;/code&gt; autocompletion and imperative-command aliases today, and use them for every task this week so they are automatic on exam day.&lt;/p&gt;

</description>
      <category>kubernetes</category>
      <category>certification</category>
      <category>cka</category>
      <category>devops</category>
    </item>
    <item>
      <title>Database Provisioning Is Now an AI Agent's Job</title>
      <dc:creator>Amaresh Pelleti</dc:creator>
      <pubDate>Tue, 01 Sep 2026 10:58:46 +0000</pubDate>
      <link>https://dev.to/amareswer/database-provisioning-is-now-an-ai-agents-job-56j3</link>
      <guid>https://dev.to/amareswer/database-provisioning-is-now-an-ai-agents-job-56j3</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;Originally published on &lt;a href="https://devtoolhub.com/database-provisioning-ai-agents/" rel="noopener noreferrer"&gt;DevToolHub&lt;/a&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;When Neon's serverless Postgres platform hit general availability in 2024, roughly 30% of new databases on it were created by AI agents instead of humans. &lt;a href="https://www.databricks.com/blog/databricks-neon" rel="noopener noreferrer"&gt;By the time Databricks acquired Neon&lt;/a&gt; in May 2025 — reportedly for about $1 billion — that number had passed 80%, and Databricks' January 2026 State of AI Agents report says it's still above that line. Database provisioning used to mean a ticket, a Terraform apply, and a wait. Now it's something an agent does mid-task, without asking anyone.&lt;/p&gt;

&lt;p&gt;This isn't a vector-search story. It's a plumbing story. The part of the stack that decides how fast a database can appear, fork, and disappear again is being rebuilt around agents instead of people. That changes a few things about how you should design access to your own databases.&lt;/p&gt;

&lt;h2&gt;
  
  
  From 30% to 80%: Database Provisioning Is Now Automated
&lt;/h2&gt;

&lt;p&gt;Databricks cited this shift as a direct reason for acquiring Neon. OLTP databases are a market built on tools designed decades before anything called an "agent" wrote a line of SQL. Per Databricks' &lt;a href="https://www.databricks.com/resources/ebook/state-of-ai-agents" rel="noopener noreferrer"&gt;2026 State of AI Agents report&lt;/a&gt;, which draws on telemetry from over 20,000 organizations, 97% of database branches on the platform now get created through natural-language agent requests instead of a human running a CLI command. Multi-agent workflow usage on Databricks grew 327% between June and October 2025 alone.&lt;/p&gt;

&lt;p&gt;That growth is why provisioning speed suddenly matters. A developer opening a ticket can tolerate a 10-minute wait for a new database. An agent mid-task cannot. It's spinning up an isolated environment to test a schema change, and it either gets a database in under a second or the whole workflow stalls.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Database Provisioning Had to Get Faster for Agents
&lt;/h2&gt;

&lt;p&gt;Neon's pitch has always been instant, isolated Postgres instances instead of a shared staging database everyone fights over. Databricks CEO Ali Ghodsi put the agent version of that problem bluntly to TechCrunch: "Because these agents are super fast. They just spin up lots of databases, much faster than humans can, but you don't want to go bankrupt doing that." Copy-on-write branching lets them try a change against real production data and throw it away if it's wrong, without touching the primary environment.&lt;/p&gt;

&lt;p&gt;That's a different requirement than the one traditional provisioning tools were built for. Terraform, CloudFormation, and RDS snapshots all assume a human is planning ahead. An agent making a database decision inside a single tool call needs infrastructure fast enough that provisioning stops being a step it has to plan around at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  How Lakebase Handles Database Provisioning in Under 500ms
&lt;/h2&gt;

&lt;p&gt;Databricks rebuilt Neon's internals under the name Lakebase. The architecture is why the speed claim holds up. It decouples compute from storage and moves storage onto object storage. Then it adds two components to make that fast enough for a database.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Safekeepers&lt;/strong&gt;, built on the Paxos consensus algorithm, handle low-latency writes. &lt;strong&gt;Page servers&lt;/strong&gt; compensate for object storage's latency by serving reads quickly, without losing transactional consistency.&lt;/p&gt;

&lt;p&gt;The result, per Databricks' Data + AI Summit 2026 keynote: branch creation lands under 500 milliseconds, new-instance creation under a second. The platform handles 12 million database launches a day in production. A newer addition, &lt;a href="https://www.databricks.com/blog/announcing-lakebase-search-agent-native-retrieval-built-lakebase-postgres" rel="noopener noreferrer"&gt;Lakebase Search&lt;/a&gt;, adds hybrid vector and full-text retrieval with 32x index compression, and can index over a billion vectors — it now ships as &lt;code&gt;lakebase_vector&lt;/code&gt; and &lt;code&gt;lakebase_text&lt;/code&gt; extensions you install with a plain &lt;code&gt;CREATE EXTENSION&lt;/code&gt;. That folds search into the same provisioning layer instead of bolting it on as a separate service.&lt;/p&gt;

&lt;p&gt;[IMAGE: articles/images/2026-08-25-database-provisioning-ai-agents-diagram.png | alt: "database provisioning pipeline showing an AI agent request flowing through safekeepers and page servers to a new branch"]&lt;/p&gt;

&lt;h2&gt;
  
  
  Tiger Data's Agentic Postgres Takes a Different Approach
&lt;/h2&gt;

&lt;p&gt;Neon isn't the only vendor rebuilding around this. Tiger Data — the company behind TimescaleDB — &lt;a href="https://infoq.com/news/2025/12/agentic-postgres-fast-forking/" rel="noopener noreferrer"&gt;shipped Agentic Postgres&lt;/a&gt; on a storage layer it calls Fluid Storage. It's built for zero-copy forks of production data in seconds. An agent can spin up a full copy of production, benchmark a new index against it, and discard the fork without ever touching the live database.&lt;/p&gt;

&lt;p&gt;Search is where Tiger Data diverges from Neon. Instead of leaning on &lt;code&gt;pgvector&lt;/code&gt; alone, Agentic Postgres pairs &lt;code&gt;pgvectorscale&lt;/code&gt; for higher-throughput vector indexing with a new &lt;code&gt;pg_textsearch&lt;/code&gt; extension. That extension implements BM25 ranked keyword search, aimed squarely at hybrid retrieval for agent workflows. It also ships an MCP server, so an agent can provision and configure a database from a plain-language prompt instead of a schema migration script. Tiger Data has since opened a free tier for it, which makes it cheap to test the fork-per-agent pattern before committing.&lt;/p&gt;

&lt;p&gt;Two well-funded companies independently reached the same conclusion. The bottleneck for agent-driven work isn't query speed. It's how fast you can hand an agent its own disposable copy of the database. That's a specific, testable claim, not a marketing line.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Database Provisioning Breaks: pgvector and Branching Limits
&lt;/h2&gt;

&lt;p&gt;None of this is friction-free. &lt;code&gt;pgvector&lt;/code&gt; runs entirely on CPU — there's no GPU offload. And the memory math gets unforgiving at scale: 50 million vectors at 768 dimensions is roughly 150GB of raw vector data alone, before the HNSW index built on top of it, which in practice can more than double the total footprint. Where the ceiling sits depends on your hardware and recall target, but if your agents run concurrent vector search over embedding sets in the tens of millions, memory and CPU — not provisioning speed — become the constraint.&lt;/p&gt;

&lt;p&gt;Branching has a narrower limit too. Neon's branch-and-fork model was designed for development and CI workflows — a handful of long-lived branches per project. Fleet-scale agent workloads want the opposite pattern. They need hundreds of lightweight, ephemeral branches created and torn down per minute, each isolated to a single agent run. That's closer to per-request sandboxing than to a Git-style branch. It's not what Neon's original design, or most managed Postgres services, were built to sustain at volume.&lt;/p&gt;

&lt;p&gt;⚠️ &lt;strong&gt;Important:&lt;/strong&gt; if you're letting agents provision their own databases or branches, put a hard quota on branch count and lifetime per agent first. An agent stuck in a retry loop can create branches far faster than a human ever would. And "scale to zero" only helps your bill if something eventually tears the branch down.&lt;/p&gt;

&lt;h2&gt;
  
  
  What This Means for Your Database Access Patterns
&lt;/h2&gt;

&lt;p&gt;There's a broader trend backing all of this. Native vector support is moving back into general-purpose relational databases instead of staying in dedicated vector stores. SQL Server 2025 ships &lt;a href="https://learn.microsoft.com/en-us/sql/t-sql/data-types/vector-data-type?view=sql-server-ver17" rel="noopener noreferrer"&gt;a native &lt;code&gt;VECTOR&lt;/code&gt; data type&lt;/a&gt; built into the core engine. It supports up to 1,998 dimensions, stores data in optimized binary, and pairs with &lt;a href="https://learn.microsoft.com/en-us/sql/relational-databases/vectors/vectors-sql-server" rel="noopener noreferrer"&gt;DiskANN-based approximate nearest-neighbor indexes&lt;/a&gt;. Oracle's AI Database 26ai ships AI Vector Search as a built-in feature, not a bolt-on. RDS and Cloud SQL already lean on &lt;code&gt;pgvector&lt;/code&gt; for the same job. Database provisioning and the vector search agents need are converging into the same box.&lt;/p&gt;

&lt;p&gt;That has practical consequences. Stop treating "provision a database" as a privileged, human-gated action if agents are already doing it in your stack. Design the access boundary around the agent's identity and scope, not around a shared service account. Budget for cleanup, too — an agent that can create a branch in 500ms can also forget to delete ten of them. And if you're choosing where embeddings live, native vector support in Postgres now has real production backing at scale. That changes the calculus against standing up a separate vector database by default.&lt;/p&gt;

&lt;p&gt;None of this requires a migration if you're already running &lt;a href="https://devtoolhub.com/postgresql-18-new-features/" rel="noopener noreferrer"&gt;PostgreSQL 18&lt;/a&gt;. Lakebase and Agentic Postgres are both Postgres underneath, so the SQL you write doesn't change. What changes is who issues the &lt;code&gt;CREATE DATABASE&lt;/code&gt; call, and how fast they expect it back. The &lt;a href="https://devtoolhub.com/mcp-model-context-protocol-guide/" rel="noopener noreferrer"&gt;Model Context Protocol guide&lt;/a&gt; covers the layer actually issuing these provisioning requests. And if you're weighing where a given workload belongs at all, the breakdown of &lt;a href="https://devtoolhub.com/database-types-guide/" rel="noopener noreferrer"&gt;different database types&lt;/a&gt; is still the right starting point before assuming agent-driven Postgres fits everything.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Q: Do I need to switch to Neon or Tiger Data to let AI agents provision databases?&lt;/strong&gt;&lt;br&gt;
A: No. Both are specific implementations of a broader shift toward fast branching and forking on top of Postgres. If your managed Postgres already supports quick snapshots or read replicas, you can build a scoped-down version of the same pattern. You just won't get sub-second provisioning without a storage layer built for it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Is &lt;code&gt;pgvector&lt;/code&gt; no longer good enough for AI workloads?&lt;/strong&gt;&lt;br&gt;
A: It's fine for moderate scale. The CPU bottleneck shows up specifically at large vector counts with high concurrent query volume — most teams won't hit it. Test with your actual embedding count and query rate before assuming you need a different engine.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: What's the real risk of letting agents provision their own databases?&lt;/strong&gt;&lt;br&gt;
A: Runaway resource creation, not data loss. An agent in a retry loop can create branches or instances far faster than a human, so the fix is a hard quota on branch count and lifetime per agent identity, not a ban on the capability.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Does this trend replace dedicated vector databases entirely?&lt;/strong&gt;&lt;br&gt;
A: Not entirely, but it removes the default reason to reach for one. When Postgres, SQL Server, and Oracle all ship native vector search, a standalone vector database has to justify itself on something beyond "it does vector search" — usually scale or a specific indexing algorithm you can't get natively yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  Quick Summary
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Over 80% of databases provisioned on Neon are created by AI agents, up from roughly 30% at its 2024 GA, per Databricks&lt;/li&gt;
&lt;li&gt;Databricks' Lakebase architecture provisions new Postgres branches in under 500ms (new instances in under a second) using Paxos-based safekeepers and separate page servers, handling 12 million launches a day&lt;/li&gt;
&lt;li&gt;Tiger Data's Agentic Postgres takes a different route to the same problem: zero-copy forks on "Fluid Storage" plus &lt;code&gt;pgvectorscale&lt;/code&gt; and a new BM25-based &lt;code&gt;pg_textsearch&lt;/code&gt; extension&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;pgvector&lt;/code&gt; is CPU-bound, and memory becomes the real ceiling in the tens of millions of vectors; Neon's original branching model wasn't built for hundreds of ephemeral agent branches per minute&lt;/li&gt;
&lt;li&gt;SQL Server 2025, Oracle AI Database 26ai, and RDS/Cloud SQL all now ship native vector search, pulling that workload back into general-purpose relational databases&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If agents already touch your database layer, put a branch quota and a teardown policy in place before you give them database provisioning access — the infrastructure is fast enough now that the bottleneck is entirely on your side.&lt;/p&gt;

</description>
      <category>database</category>
      <category>aiagents</category>
      <category>postgres</category>
      <category>neon</category>
    </item>
    <item>
      <title>GitHub Actions and Git: A Complete Workflow Guide</title>
      <dc:creator>Amaresh Pelleti</dc:creator>
      <pubDate>Fri, 28 Aug 2026 17:50:20 +0000</pubDate>
      <link>https://dev.to/amareswer/github-actions-and-git-a-complete-workflow-guide-23jl</link>
      <guid>https://dev.to/amareswer/github-actions-and-git-a-complete-workflow-guide-23jl</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;Originally published on &lt;a href="https://devtoolhub.com/git-github-actions-guide/" rel="noopener noreferrer"&gt;DevToolHub&lt;/a&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Git and GitHub Actions are two tools most teams adopt without ever deciding how they fit together. You pick a branching habit, wire up a workflow file, and six months later nobody remembers why deploys are slow or why one pull request skipped its tests. So this guide covers the decisions that actually matter — branching model, rebase timing, workflow structure, secret handling, and the failure modes that show up once real traffic hits your pipelines. Each section leads with the short answer, then the detail behind it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What does a Git and GitHub Actions workflow actually look like?
&lt;/h2&gt;

&lt;p&gt;A working setup has three layers. Git tracks your code history on short-lived branches. GitHub hosts the shared repository and enforces review rules. GitHub Actions runs automated jobs — tests, builds, deploys — triggered by events like a push or a pull request. The branch you push decides which workflows run. The review rules decide whether that branch can merge. Align those three layers and most CI pain disappears.&lt;/p&gt;

&lt;p&gt;Here is the chain every run follows. An event fires: a &lt;code&gt;push&lt;/code&gt;, a &lt;code&gt;pull_request&lt;/code&gt;, a &lt;code&gt;schedule&lt;/code&gt;, or a manual &lt;code&gt;workflow_dispatch&lt;/code&gt;. GitHub matches that event against the &lt;code&gt;on:&lt;/code&gt; block of every workflow file in &lt;code&gt;.github/workflows/&lt;/code&gt;. Each matching workflow starts its jobs. By default, jobs run in parallel on separate runners, and &lt;code&gt;needs:&lt;/code&gt; forces an order. Each job runs its steps in sequence on one machine. A step is either a shell command (&lt;code&gt;run:&lt;/code&gt;) or a packaged action (&lt;code&gt;uses:&lt;/code&gt;).&lt;/p&gt;

&lt;p&gt;[IMAGE: articles/images/2026-08-28-git-github-actions-guide-diagram.png | alt: "Flow from a git event to workflow match to parallel jobs on runners to sequential steps"]&lt;/p&gt;

&lt;p&gt;New to the split between the version control tool and the hosting platform? The &lt;a href="https://devtoolhub.com/git-vs-github-for-beginners/" rel="noopener noreferrer"&gt;difference between Git and GitHub for beginners&lt;/a&gt; is worth ten minutes first. Everything below assumes you know which is which.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Verdict:&lt;/strong&gt; if you cannot draw this chain from memory, every debugging session starts from zero. Learn it once and the error messages start making sense.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which Git branching model should you use?
&lt;/h2&gt;

&lt;p&gt;For most teams, use GitHub Flow: one long-lived &lt;code&gt;main&lt;/code&gt; branch, short-lived feature branches, and a merge through a pull request once checks pass. Trunk-based development goes further, with everyone committing to &lt;code&gt;main&lt;/code&gt; several times a day behind feature flags, and it suits teams that ship continuously. Git Flow, with its &lt;code&gt;develop&lt;/code&gt; and &lt;code&gt;release&lt;/code&gt; branches, only earns its complexity if you ship versioned software to customers who run several versions at once.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Branch structure&lt;/th&gt;
&lt;th&gt;Best for&lt;/th&gt;
&lt;th&gt;Main cost&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GitHub Flow&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;main&lt;/code&gt; + short feature branches&lt;/td&gt;
&lt;td&gt;Most web teams, continuous deployment&lt;/td&gt;
&lt;td&gt;Needs solid PR checks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Trunk-based&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;main&lt;/code&gt; only, feature flags&lt;/td&gt;
&lt;td&gt;Fast-moving teams, many deploys a day&lt;/td&gt;
&lt;td&gt;Requires a fast, trusted test suite&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Git Flow&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;main&lt;/code&gt;, &lt;code&gt;develop&lt;/code&gt;, &lt;code&gt;release/*&lt;/code&gt;, &lt;code&gt;hotfix/*&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Versioned/on-prem software, parallel releases&lt;/td&gt;
&lt;td&gt;Heavy branch bookkeeping&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;GitHub Flow works because there is only one integration point. A branch is either merged or it is not. The risk is that a weak test suite lets broken code into &lt;code&gt;main&lt;/code&gt;, so the model leans on branch protection to hold the line.&lt;/p&gt;

&lt;p&gt;Trunk-based development removes even the feature branch. Work goes straight to &lt;code&gt;main&lt;/code&gt; behind a flag that is off in production until the feature is ready. In practice, this keeps merge conflicts tiny because nothing diverges for long. But the price is discipline. Your pipeline has to catch regressions on every commit, because there is no staging branch to catch them later.&lt;/p&gt;

&lt;p&gt;Git Flow is the odd one out in 2026. It was designed for scheduled, versioned releases, and it still fits that case. But for a team deploying a web app several times a week, the &lt;code&gt;develop&lt;/code&gt;-to-&lt;code&gt;release&lt;/code&gt;-to-&lt;code&gt;main&lt;/code&gt; promotion path is bureaucracy with no payoff. For example commands per model, see the &lt;a href="https://devtoolhub.com/git-workflows-gitflow-githubflow-trunk-based/" rel="noopener noreferrer"&gt;full breakdown of Git Flow, GitHub Flow, and trunk-based development&lt;/a&gt;. Pair it with a review of your &lt;a href="https://devtoolhub.com/git-branching-merging-strategies/" rel="noopener noreferrer"&gt;branching and merging strategies&lt;/a&gt;. The &lt;a href="https://devtoolhub.com/git-best-practices-branching-approvals/" rel="noopener noreferrer"&gt;branching and approval best practices&lt;/a&gt; guide covers how to enforce whichever one you pick.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Verdict:&lt;/strong&gt; default to GitHub Flow. Move to trunk-based only when your test suite is fast and trustworthy enough to gate &lt;code&gt;main&lt;/code&gt; directly. Reach for Git Flow only if you genuinely support multiple released versions.&lt;/p&gt;

&lt;h2&gt;
  
  
  When should you rebase instead of merge?
&lt;/h2&gt;

&lt;p&gt;Rebase to clean up your own local commits before you share them. Merge to combine branches that other people have already pulled. The Pro Git book states the rule plainly: "Do not rebase commits that exist outside your repository and that people may have based work on." Break it and teammates get duplicate commits, a broken history, and merge conflicts that should not exist.&lt;/p&gt;

&lt;p&gt;Rebase and merge do different things. A merge performs a three-way merge between the two branch tips and their common ancestor, then records a new merge commit. A rebase takes the diffs your branch introduced, resets your branch to the target, and replays each diff on top. As the book puts it: "Rebasing replays changes from one line of work onto another in the order they were introduced, whereas merging takes the endpoints and merges them together."&lt;/p&gt;

&lt;p&gt;The everyday use is keeping a feature branch current:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git checkout feature/login
git rebase main
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That replays your login work on top of the latest &lt;code&gt;main&lt;/code&gt;, so the eventual pull request is a clean diff with no merge commits in the middle. You can make &lt;code&gt;git pull&lt;/code&gt; do this automatically:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git config &lt;span class="nt"&gt;--global&lt;/span&gt; pull.rebase &lt;span class="nb"&gt;true&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;⚠️ Note: rebasing a branch you have already pushed and that a teammate has pulled means force-pushing over shared history. Their next &lt;code&gt;git pull&lt;/code&gt; brings back your old commits alongside your rewritten ones, and now the branch has both. The recovery is &lt;code&gt;git pull --rebase&lt;/code&gt;, which uses patch-id checksums to spot and drop the duplicates, but it is friction nobody wanted.&lt;/p&gt;

&lt;p&gt;The safe boundary is simple. Rebase commits that have never left your machine, or that are pushed but that nobody has based work on. For the mechanics of interactive rebase, stash, and cherry-pick, see &lt;a href="https://devtoolhub.com/advanced-git-rebase-stash-cherry-pick/" rel="noopener noreferrer"&gt;advanced Git rebase, stash, and cherry-pick&lt;/a&gt;. To tidy a messy branch into one commit before review, &lt;a href="https://devtoolhub.com/squashing-commits-in-a-git-feature-branch-a-cleaner-git-history/" rel="noopener noreferrer"&gt;squash the feature branch&lt;/a&gt; instead of rebasing interactively. When a rebase or merge does collide, &lt;a href="https://devtoolhub.com/how-to-handle-merge-conflicts-in-git/" rel="noopener noreferrer"&gt;resolving merge conflicts in Git&lt;/a&gt; walks through the markers and the fix.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Verdict:&lt;/strong&gt; rebase before the first push, merge after. The pull-request merge button is a merge, and that is correct — leave it alone.&lt;/p&gt;

&lt;h2&gt;
  
  
  What are Git hooks actually good for?
&lt;/h2&gt;

&lt;p&gt;Git hooks run your own scripts automatically at points in the Git lifecycle: before a commit, before a push, after a merge. The useful ones are &lt;code&gt;pre-commit&lt;/code&gt; for linting and formatting, &lt;code&gt;commit-msg&lt;/code&gt; for enforcing a message format, and &lt;code&gt;pre-push&lt;/code&gt; for a fast smoke test. One catch decides how you use them. Hooks live in &lt;code&gt;.git/hooks&lt;/code&gt; and are never copied on clone, so a hook on your machine does nothing for anyone else.&lt;/p&gt;

&lt;p&gt;When you run &lt;code&gt;git init&lt;/code&gt;, Git fills &lt;code&gt;.git/hooks&lt;/code&gt; with example scripts that end in &lt;code&gt;.sample&lt;/code&gt;. Drop the suffix and make the file executable to activate one. Any language works. If the script exits non-zero, Git aborts the operation — a failing &lt;code&gt;pre-commit&lt;/code&gt; stops the commit, a failing &lt;code&gt;pre-push&lt;/code&gt; stops the push. Anyone can skip a client-side hook with &lt;code&gt;git commit --no-verify&lt;/code&gt;, so a hook is a convenience, not a gate.&lt;/p&gt;

&lt;p&gt;Because hooks are not version-controlled, teams share them one of two ways. First, you can commit hook scripts to a tracked folder and point Git at it with &lt;code&gt;git config core.hooksPath .githooks&lt;/code&gt;. Or you can use a manager like &lt;code&gt;pre-commit&lt;/code&gt; or Husky that wires itself in during a setup step. Either way, keep hook work fast. A &lt;code&gt;pre-commit&lt;/code&gt; that runs the full test suite trains people to use &lt;code&gt;--no-verify&lt;/code&gt; every time.&lt;/p&gt;

&lt;p&gt;Server-side hooks (&lt;code&gt;pre-receive&lt;/code&gt;, &lt;code&gt;update&lt;/code&gt;, &lt;code&gt;post-receive&lt;/code&gt;) run on the remote and cannot be bypassed, but on GitHub you do not manage those directly — branch protection and Actions cover the same ground. For working examples of each client-side hook, see &lt;a href="https://devtoolhub.com/git-hooks-automate-workflow-examples/" rel="noopener noreferrer"&gt;Git hooks with real scripts&lt;/a&gt;, and check the &lt;a href="https://git-scm.com/book/en/v2/Customizing-Git-Git-Hooks" rel="noopener noreferrer"&gt;official Git hooks reference&lt;/a&gt; for the full list and their arguments.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Verdict:&lt;/strong&gt; use hooks for fast local checks only, and manage them with a tool so the whole team gets them. Real enforcement belongs in GitHub Actions and branch protection, not a hook that &lt;code&gt;--no-verify&lt;/code&gt; skips.&lt;/p&gt;

&lt;h2&gt;
  
  
  How are GitHub Actions workflows structured?
&lt;/h2&gt;

&lt;p&gt;A workflow is a YAML file in &lt;code&gt;.github/workflows/&lt;/code&gt;. It needs three pieces: &lt;code&gt;on:&lt;/code&gt; for what triggers it, &lt;code&gt;jobs:&lt;/code&gt; for the units of work, and inside each job, &lt;code&gt;runs-on:&lt;/code&gt; for the runner and &lt;code&gt;steps:&lt;/code&gt; for the commands. Jobs run in parallel by default. Add &lt;code&gt;needs:&lt;/code&gt; to make one job wait for another. Everything else — &lt;code&gt;env:&lt;/code&gt;, &lt;code&gt;permissions:&lt;/code&gt;, &lt;code&gt;concurrency:&lt;/code&gt;, &lt;code&gt;strategy.matrix:&lt;/code&gt; — is optional tuning.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;CI&lt;/span&gt;
&lt;span class="na"&gt;on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;push&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;branches&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;main&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
  &lt;span class="na"&gt;pull_request&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;

&lt;span class="na"&gt;permissions&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;contents&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;read&lt;/span&gt;

&lt;span class="na"&gt;jobs&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;test&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;runs-on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ubuntu-latest&lt;/span&gt;
    &lt;span class="na"&gt;timeout-minutes&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;15&lt;/span&gt;
    &lt;span class="na"&gt;steps&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;actions/checkout@v4&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;actions/setup-node@v4&lt;/span&gt;
        &lt;span class="na"&gt;with&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;node-version&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;22&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;npm ci&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;npm test&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The triggers you will actually use are &lt;code&gt;push&lt;/code&gt;, &lt;code&gt;pull_request&lt;/code&gt;, &lt;code&gt;schedule&lt;/code&gt; (POSIX cron, minimum five-minute intervals), &lt;code&gt;workflow_dispatch&lt;/code&gt; for a manual run, and &lt;code&gt;workflow_call&lt;/code&gt; to be invoked by another workflow. Branch and path filters narrow it further, so a docs-only change need not run the full build.&lt;/p&gt;

&lt;p&gt;Two tuning keys matter early. &lt;code&gt;concurrency&lt;/code&gt; groups runs so a new push cancels the in-progress one for the same branch:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;concurrency&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;group&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${{ github.workflow }}-${{ github.ref }}&lt;/span&gt;
  &lt;span class="na"&gt;cancel-in-progress&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;timeout-minutes&lt;/code&gt; caps a job. Without it, the platform still kills the job at six hours, but by then you have burned a runner for nothing. Environment variables follow a precedence order: a step-level &lt;code&gt;env&lt;/code&gt; beats a job-level one, which beats the workflow level.&lt;/p&gt;

&lt;p&gt;For the full syntax with every key explained, see the &lt;a href="https://devtoolhub.com/github-actions-workflow-yaml-guide/" rel="noopener noreferrer"&gt;workflow YAML guide&lt;/a&gt;. If you have never built one, &lt;a href="https://devtoolhub.com/github-actions-first-cicd-pipeline/" rel="noopener noreferrer"&gt;your first CI/CD pipeline&lt;/a&gt; starts from an empty repo, and &lt;a href="https://devtoolhub.com/introduction-to-github-actions-automating-your-workflow/" rel="noopener noreferrer"&gt;an introduction to automating your workflow&lt;/a&gt; covers the concepts. Real patterns live in &lt;a href="https://devtoolhub.com/github-actions-use-cases-examples/" rel="noopener noreferrer"&gt;common use cases and examples&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Verdict:&lt;/strong&gt; keep one workflow per purpose — CI, release, scheduled maintenance — not one file with twenty conditional jobs. Small files are far easier to reason about when a run fails at 2 a.m.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reusable workflows, composite actions, or a matrix: which one?
&lt;/h2&gt;

&lt;p&gt;Use a matrix when you run the same job across variations, like Node 18, 20, and 22, or three operating systems. A composite action fits a sequence of steps you repeat inside one repository, such as checkout, language setup, and cache restore. Reach for a reusable workflow when whole jobs need to be shared across many repositories. They solve different problems and often show up together.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Mechanism&lt;/th&gt;
&lt;th&gt;Scope&lt;/th&gt;
&lt;th&gt;What it shares&lt;/th&gt;
&lt;th&gt;Called with&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;strategy.matrix&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;One job, many variants&lt;/td&gt;
&lt;td&gt;Nothing — same steps, different inputs&lt;/td&gt;
&lt;td&gt;Built into the job&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Composite action&lt;/td&gt;
&lt;td&gt;Within or across repos&lt;/td&gt;
&lt;td&gt;A sequence of steps&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;uses:&lt;/code&gt; at the step level&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reusable workflow&lt;/td&gt;
&lt;td&gt;Across repos&lt;/td&gt;
&lt;td&gt;Whole jobs, with inputs and secrets&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;uses:&lt;/code&gt; at the job level, &lt;code&gt;secrets: inherit&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A matrix is the simplest and the one to try first:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;strategy&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;matrix&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;node&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;18&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;20&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;22&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That runs the job three times in parallel. Because a matrix caps at 256 jobs per workflow run, a two-dimensional matrix across many values adds up faster than people expect.&lt;/p&gt;

&lt;p&gt;Composite actions bundle steps into an &lt;code&gt;action.yml&lt;/code&gt; file so a repo can call &lt;code&gt;uses: ./.github/actions/setup&lt;/code&gt; instead of repeating six lines. The reusable workflow is a full workflow with &lt;code&gt;on: workflow_call&lt;/code&gt;, typed &lt;code&gt;inputs&lt;/code&gt;, and either explicit &lt;code&gt;secrets&lt;/code&gt; or &lt;code&gt;secrets: inherit&lt;/code&gt;. It is how a platform team standardizes CI across forty services. The &lt;a href="https://devtoolhub.com/github-actions-reusable-workflows-composite-actions/" rel="noopener noreferrer"&gt;reusable workflows and composite actions guide&lt;/a&gt; has working examples of both.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Verdict:&lt;/strong&gt; reach for a matrix first. Build a composite action or a reusable workflow only once you are copy-pasting the same block into a third place.&lt;/p&gt;

&lt;h2&gt;
  
  
  How should GitHub Actions handle secrets and cloud auth?
&lt;/h2&gt;

&lt;p&gt;Stop storing long-lived cloud keys as repository secrets. Use OpenID Connect (OIDC): the workflow requests a short-lived token from AWS, Azure, or Google Cloud at run time, scoped to that repository and branch. For everything else, keep the &lt;code&gt;GITHUB_TOKEN&lt;/code&gt; read-only by default and raise permissions per job. Never pass a secret as a command-line argument, because another job on the same runner can read it with &lt;code&gt;ps x -w&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;GitHub's own hardening guide is direct about the risks. It says automatic redaction "is not guaranteed," so mask anything sensitive that is not already a GitHub secret with the &lt;code&gt;::add-mask::&lt;/code&gt; workflow command. It also says "Never use structured data as a secret" — a JSON or YAML blob lowers the odds that redaction catches every piece of it. Create one secret per value.&lt;/p&gt;

&lt;p&gt;The &lt;code&gt;permissions&lt;/code&gt; key is the single highest-value line in most workflows:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;permissions&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;contents&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;read&lt;/span&gt;
  &lt;span class="na"&gt;id-token&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;write&lt;/span&gt;   &lt;span class="c1"&gt;# only in the job that needs OIDC&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;GitHub flipped the &lt;code&gt;GITHUB_TOKEN&lt;/code&gt; default to read-only for new repositories in 2023. But organizations created before February 2023 may still default to read-write. So check your organization's Actions settings and set it explicitly.&lt;/p&gt;

&lt;p&gt;OIDC removes the stored key entirely. The workflow proves its identity to the cloud provider, which hands back a token that expires in minutes. GitHub's docs recommend it directly for any workflow that deploys to a cloud provider or uses HashiCorp Vault. One caveat: "Support for custom claims for OIDC is unavailable in AWS." So scope AWS trust policies on the standard subject claim.&lt;/p&gt;

&lt;p&gt;For the day-to-day rules, see &lt;a href="https://devtoolhub.com/github-actions-secrets-security-best-practices/" rel="noopener noreferrer"&gt;managing workflow secrets safely&lt;/a&gt; and the wider &lt;a href="https://devtoolhub.com/github-actions-security-best-practices/" rel="noopener noreferrer"&gt;security best practices checklist&lt;/a&gt;. The &lt;a href="https://docs.github.com/en/actions/security-for-github-actions/security-guides/security-hardening-for-github-actions" rel="noopener noreferrer"&gt;official security-hardening docs&lt;/a&gt; are the source for all of the above.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Verdict:&lt;/strong&gt; OIDC for cloud, a read-only &lt;code&gt;GITHUB_TOKEN&lt;/code&gt; everywhere, and third-party actions pinned to a commit SHA. The section on what breaks explains why that last one is not optional.&lt;/p&gt;

&lt;h2&gt;
  
  
  How does GitHub Actions caching work, and what does it cost?
&lt;/h2&gt;

&lt;p&gt;The &lt;code&gt;actions/cache&lt;/code&gt; action stores a directory keyed by a hash, usually of your lockfile, and restores it on the next run. Each repository gets 10 GB of cache. GitHub deletes any cache not accessed in seven days. Once a repo passes 10 GB, it evicts entries "in order of last access date, from oldest to most recent." A pull request can restore caches from its own branch, the default branch, and its base branch. It cannot read caches from sibling or child branches.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;actions/cache@v4&lt;/span&gt;
  &lt;span class="na"&gt;with&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;path&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;~/.npm&lt;/span&gt;
    &lt;span class="na"&gt;key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;npm-${{ hashFiles('package-lock.json') }}&lt;/span&gt;
    &lt;span class="na"&gt;restore-keys&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;|&lt;/span&gt;
      &lt;span class="s"&gt;npm-&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The action looks for an exact &lt;code&gt;key&lt;/code&gt; match first, then partial matches, then each &lt;code&gt;restore-keys&lt;/code&gt; prefix in turn. A &lt;code&gt;key&lt;/code&gt; can be up to 512 characters. The cache is best-effort: a miss should slow the build, never break it.&lt;/p&gt;

&lt;p&gt;⚠️ Note: caches are not private within a repo. The docs are explicit: "Anyone who can open a pull request against your repository can read the contents of caches in the base branch." Do not cache credentials, tokens, or anything derived from a secret.&lt;/p&gt;

&lt;p&gt;For key design and layered caching strategies, see &lt;a href="https://devtoolhub.com/github-actions-caching-performance-optimization/" rel="noopener noreferrer"&gt;caching and performance tuning&lt;/a&gt;. The &lt;a href="https://docs.github.com/en/actions/writing-workflows/choosing-what-your-workflow-does/caching-dependencies-to-speed-up-workflows" rel="noopener noreferrer"&gt;official caching docs&lt;/a&gt; cover cross-OS archives and rate limits.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Verdict:&lt;/strong&gt; cache dependency downloads, never build output you cannot trust to be stale. If a cache miss breaks your build, the build was already broken.&lt;/p&gt;

&lt;h2&gt;
  
  
  How do you enforce that tests pass and the right people review?
&lt;/h2&gt;

&lt;p&gt;Enforcement lives in branch protection, not the workflow file. In the protected branch's rule, mark specific status checks as required, and the merge button stays disabled until they pass. Add a &lt;code&gt;CODEOWNERS&lt;/code&gt; file to require a named team's review on the paths they own. For a monorepo, path filters plus per-path approval rules stop a frontend change from needing the database team's sign-off.&lt;/p&gt;

&lt;p&gt;Required status checks are matched by job name. That creates a quiet trap. Rename a job in the workflow and the required check no longer matches, so the rule silently stops enforcing anything. So keep job names stable, or update the branch rule in the same change.&lt;/p&gt;

&lt;p&gt;A &lt;code&gt;CODEOWNERS&lt;/code&gt; file at &lt;code&gt;.github/CODEOWNERS&lt;/code&gt; maps globs to owners:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight codeowners"&gt;&lt;code&gt;&lt;span class="n"&gt;/api/&lt;/span&gt;&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="nf"&gt;@org/backend&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="n"&gt;/web/&lt;/span&gt;&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="nf"&gt;@org/frontend&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="n"&gt;*.tf&lt;/span&gt;&lt;span class="w"&gt;         &lt;/span&gt;&lt;span class="nf"&gt;@org/platform&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With "Require review from Code Owners" enabled, a pull request touching &lt;code&gt;/api/&lt;/code&gt; needs a &lt;code&gt;@org/backend&lt;/code&gt; approval before merge. The &lt;a href="https://devtoolhub.com/github-codeowners-permissions-best-practices/" rel="noopener noreferrer"&gt;CODEOWNERS and review permissions guide&lt;/a&gt; covers the syntax edge cases, and &lt;a href="https://devtoolhub.com/enforce-path-based-approvals-github-actions/" rel="noopener noreferrer"&gt;path-based approval rules&lt;/a&gt; show how to scope checks per directory.&lt;/p&gt;

&lt;p&gt;Two more pieces round this out. The &lt;a href="https://devtoolhub.com/github-actions-testing-unit-tests/" rel="noopener noreferrer"&gt;guide to running unit tests in CI&lt;/a&gt; covers wiring the test job so its result is a usable status check. And &lt;a href="https://devtoolhub.com/github-templates-pull-requests-issues-discussions/" rel="noopener noreferrer"&gt;pull request and issue templates&lt;/a&gt; give reviewers a checklist, so "looks fine" reviews get rarer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Verdict:&lt;/strong&gt; a test that is not a required status check is a suggestion. Wire the check into branch protection, or accept that people will merge red.&lt;/p&gt;

&lt;h2&gt;
  
  
  How do GitHub Actions deployment environments work?
&lt;/h2&gt;

&lt;p&gt;An environment is a named deployment target — &lt;code&gt;staging&lt;/code&gt;, &lt;code&gt;production&lt;/code&gt; — with its own secrets and protection rules. A job that sets &lt;code&gt;environment: production&lt;/code&gt; pauses until its rules pass: required reviewers approve, a wait timer elapses, or the branch matches an allowed pattern. Environment secrets are visible only to jobs targeting that environment, which keeps production credentials out of every other run.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;deploy&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;runs-on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ubuntu-latest&lt;/span&gt;
  &lt;span class="na"&gt;environment&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;production&lt;/span&gt;
  &lt;span class="na"&gt;concurrency&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;production&lt;/span&gt;
  &lt;span class="na"&gt;steps&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;./deploy.sh&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;environment&lt;/code&gt; line does two things: it gates the job behind the environment's rules, and it scopes which secrets the job can read. The &lt;code&gt;concurrency: production&lt;/code&gt; line stops two deploys running at once, so a fast follow-up merge queues behind the current release instead of racing it.&lt;/p&gt;

&lt;p&gt;Rolling, blue-green, and canary deploys are not features you toggle. You build them by calling your platform's CLI or API across environment-scoped jobs, with the environment gate providing the human checkpoint. Tag-based releases fit here too — cutting a release from a &lt;a href="https://devtoolhub.com/git-tags-releases-best-practices/" rel="noopener noreferrer"&gt;Git tag&lt;/a&gt; is a clean trigger for a production workflow. The &lt;a href="https://devtoolhub.com/github-actions-deployment-strategies-environments/" rel="noopener noreferrer"&gt;deployment strategies and environments guide&lt;/a&gt; walks through each pattern, and &lt;a href="https://devtoolhub.com/github-actions-deployment-automating-releases/" rel="noopener noreferrer"&gt;automating releases&lt;/a&gt; covers changelog and versioning steps.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Verdict:&lt;/strong&gt; put every deploy behind an environment with at least one required reviewer for production. It is the cheapest safety net Actions gives you.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually breaks GitHub Actions in production?
&lt;/h2&gt;

&lt;p&gt;Four failure modes show up again and again: a compromised third-party action leaking secrets, a &lt;code&gt;pull_request_target&lt;/code&gt; workflow running attacker code, caches vanishing mid-sprint, and jobs silently hitting a platform limit. None of these appear in a tutorial. All of them have cost real teams real incidents.&lt;/p&gt;

&lt;h3&gt;
  
  
  A compromised action dumps your secrets into the logs
&lt;/h3&gt;

&lt;p&gt;In March 2025, the widely used &lt;code&gt;tj-actions/changed-files&lt;/code&gt; action was compromised (CVE-2025-30066). An attacker repointed every tag from &lt;code&gt;v1&lt;/code&gt; through &lt;code&gt;v45.0.7&lt;/code&gt; to a single malicious commit, &lt;code&gt;0e58ed8&lt;/code&gt;. The injected code read secrets out of the runner's memory and printed them into the workflow log, which is public on any public repository. More than 23,000 repositories ran the poisoned version before it was caught. The fix shipped as &lt;code&gt;v46.0.1&lt;/code&gt;, and CISA advised every affected project to rotate every credential the workflow could touch during the window.&lt;/p&gt;

&lt;p&gt;The defense is pinning. A tag is a movable pointer; a commit SHA is not.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# vulnerable — a tag can be repointed&lt;/span&gt;
&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;tj-actions/changed-files@v45&lt;/span&gt;

&lt;span class="c1"&gt;# safe — an immutable reference&lt;/span&gt;
&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;tj-actions/changed-files@a5b3c1d2e4f5...&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;GitHub's hardening guide says a full-length commit SHA "is currently the only way to use an action as an immutable release." Turn on Dependabot so a known-bad version gets flagged, and audit that the SHA belongs to the action's real repository, not a fork.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;pull_request_target&lt;/code&gt; runs code from the fork with your token
&lt;/h3&gt;

&lt;p&gt;The &lt;code&gt;pull_request_target&lt;/code&gt; trigger runs with your repository's secrets and a writable &lt;code&gt;GITHUB_TOKEN&lt;/code&gt;, but in the base repository's context, so it can comment on or label a pull request from a fork. The danger is checking out the pull request's head and then running anything from it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# dangerous pattern&lt;/span&gt;
&lt;span class="na"&gt;on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;pull_request_target&lt;/span&gt;
&lt;span class="na"&gt;jobs&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;build&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;steps&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;actions/checkout@v4&lt;/span&gt;
        &lt;span class="na"&gt;with&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;ref&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${{ github.event.pull_request.head.sha }}&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;npm ci &amp;amp;&amp;amp; npm run build&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That &lt;code&gt;npm ci&lt;/code&gt; runs &lt;code&gt;postinstall&lt;/code&gt; scripts from an attacker's branch, with your secrets in scope. Security researchers used exactly this pattern to reach remote code execution in workflows at Microsoft, Google, and Nvidia. The rule from GitHub's docs: workflows using these triggers "must not explicitly check out untrusted code." Use plain &lt;code&gt;pull_request&lt;/code&gt; for anything that runs a contributor's code — it has no secrets and a read-only token — and move any privileged follow-up into a separate &lt;code&gt;workflow_run&lt;/code&gt; workflow.&lt;/p&gt;

&lt;h3&gt;
  
  
  Your cache disappears the week you need it
&lt;/h3&gt;

&lt;p&gt;The seven-day eviction plus the 10 GB cap means a cache you depend on can vanish after a quiet week, and a busy monorepo can evict its own caches within a day as new entries push old ones out. The symptom is CI time doubling overnight with no code change and no alert. Check the repository's cache list when build times jump:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;gh cache list &lt;span class="nt"&gt;--repo&lt;/span&gt; owner/name
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Scope keys tightly so unrelated changes do not invalidate everything, and design the build so a cold cache is slow but still green.&lt;/p&gt;

&lt;h3&gt;
  
  
  A job hits a limit and just stops
&lt;/h3&gt;

&lt;p&gt;These limits have no loud error. A job is killed at six hours of execution time. Queued jobs are cancelled after 24 hours waiting for a runner. Matrix builds cap at 256 jobs per run. Concurrent jobs cap at 20 on Free, 40 on Pro, 60 on Team, and 500 on Enterprise, so a large matrix on a Free plan queues behind itself. The &lt;code&gt;GITHUB_TOKEN&lt;/code&gt; also gets 1,000 API requests per hour per repository, so a workflow that loops over the API can start returning 403s late in the run.&lt;/p&gt;

&lt;p&gt;Set an explicit &lt;code&gt;timeout-minutes&lt;/code&gt; well under six hours so a hung job fails fast, keep matrices small, and know your plan's concurrency number before you fan out. The &lt;a href="https://devtoolhub.com/github-actions-monitoring-debugging-guide/" rel="noopener noreferrer"&gt;monitoring and debugging guide&lt;/a&gt; covers reading run logs when a job ends without an obvious reason. For the repository side of the same problem, see &lt;a href="https://devtoolhub.com/git-security-best-practices/" rel="noopener noreferrer"&gt;Git security best practices&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Verdict:&lt;/strong&gt; pin actions to SHAs, keep untrusted code out of privileged triggers, and assume every cache and every runner is about to disappear. Design for that and GitHub Actions is boringly reliable.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Q: Do I need GitHub Actions if I already use Git?&lt;/strong&gt;&lt;br&gt;
A: Git tracks code history; GitHub Actions automates what happens to that code. You can use Git alone, but you lose automated tests on every pull request, one-click deploys, and required checks that block bad merges. For a solo project it is optional. For a team, that automation is what keeps &lt;code&gt;main&lt;/code&gt; releasable.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Is rebase or merge better for a team?&lt;/strong&gt;&lt;br&gt;
A: Both, at different times. Rebase your own local commits to tidy them before the first push. Merge branches once other people have pulled them. The pull-request merge button is a merge and should stay that way. Never rebase a branch someone else has already based work on.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: How do I stop a third-party GitHub Action from stealing secrets?&lt;/strong&gt;&lt;br&gt;
A: Pin it to a full commit SHA, not a tag. A tag can be repointed to malicious code, as the &lt;code&gt;tj-actions/changed-files&lt;/code&gt; compromise showed in 2025. Set the &lt;code&gt;GITHUB_TOKEN&lt;/code&gt; to read-only, grant write access per job, and use OIDC instead of stored cloud keys. Turn on Dependabot to flag known-bad versions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Why did my GitHub Actions build suddenly get slower?&lt;/strong&gt;&lt;br&gt;
A: The most common cause is cache eviction. GitHub deletes caches not used in seven days and evicts the oldest once a repository passes 10 GB. A quiet week or a busy monorepo can wipe the cache you rely on, doubling build time with no code change and no alert. Check &lt;code&gt;gh cache list&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Where should I enforce that tests pass before a merge?&lt;/strong&gt;&lt;br&gt;
A: In branch protection, not the workflow file. Mark the test job as a required status check on the protected branch, and the merge button stays disabled until it passes. A workflow that runs tests but is not wired into branch protection does nothing to stop a red merge.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quick Summary:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;GitHub Flow — one &lt;code&gt;main&lt;/code&gt;, short-lived feature branches, merge via pull request — is the right default. Move to trunk-based only with a fast, trusted test suite.&lt;/li&gt;
&lt;li&gt;Rebase local commits before the first push; merge anything already shared. The Pro Git rule: never rebase commits others may have based work on.&lt;/li&gt;
&lt;li&gt;Pin third-party actions to a full commit SHA. The &lt;code&gt;tj-actions/changed-files&lt;/code&gt; compromise (CVE-2025-30066, March 2025) repointed every tag to secret-stealing code across 23,000+ repositories.&lt;/li&gt;
&lt;li&gt;GitHub Actions caches are 10 GB per repository and evicted after seven days unused. A build that slows overnight with no code change is usually a lost cache.&lt;/li&gt;
&lt;li&gt;Real enforcement lives in branch protection. A test that is not a required status check will not stop a red merge, and deploys belong behind an environment with a required reviewer.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Working through one of these decisions now? Start with the branching model — &lt;a href="https://devtoolhub.com/git-workflows-gitflow-githubflow-trunk-based/" rel="noopener noreferrer"&gt;the three-model breakdown&lt;/a&gt; has the commands for each. Then wire your test job into branch protection before you touch anything else.&lt;/p&gt;

</description>
      <category>githubactions</category>
      <category>git</category>
      <category>cicd</category>
      <category>devops</category>
    </item>
    <item>
      <title>GitHub Actions Self-Hosted Runners: Fix Before Sept 25</title>
      <dc:creator>Amaresh Pelleti</dc:creator>
      <pubDate>Thu, 27 Aug 2026 10:53:57 +0000</pubDate>
      <link>https://dev.to/amareswer/github-actions-self-hosted-runners-fix-before-sept-25-4fp4</link>
      <guid>https://dev.to/amareswer/github-actions-self-hosted-runners-fix-before-sept-25-4fp4</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;Originally published on &lt;a href="https://devtoolhub.com/github-actions-self-hosted-runner-enforcement/" rel="noopener noreferrer"&gt;DevToolHub&lt;/a&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If you run GitHub Actions self-hosted runners, the enforcement clock isn't approaching — it's already running. &lt;a href="https://github.blog/changelog/2026-06-12-github-actions-minimum-version-enforcement-timeline-for-self-hosted-runners/" rel="noopener noreferrer"&gt;GitHub's minimum-version enforcement&lt;/a&gt; blocks registration below version &lt;code&gt;2.329.0&lt;/code&gt;. It's been fully in effect for GitHub Enterprise Cloud with data residency since July 31, 2026, and for the rest of github.com the brownout windows started August 24 — full enforcement lands September 25.&lt;/p&gt;

&lt;p&gt;The practical takeaway isn't "update your runners." It's that GitHub already tried this once and had to pull it days before the deadline, and the fix has a gap the changelog doesn't spell out: auto-update can't save a runner that's already below the line, because the version check happens before auto-update ever gets a chance to run.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Breaks If You Don't Upgrade Your GitHub Actions Self-Hosted Runners
&lt;/h2&gt;

&lt;p&gt;Two separate requirements kick in. Runners need version &lt;code&gt;2.329.0&lt;/code&gt; or later just to register at all. Once registered, GitHub raises the bar again — runners must install new releases within 30 days of publication, or they stop picking up jobs even though they stay registered.&lt;/p&gt;

&lt;p&gt;Each deadline comes with brownout windows — scheduled periods where GitHub temporarily blocks old-version registration so you can catch problems before enforcement is permanent. For September 25, those windows are already running: August 24 was the first, with more on August 31, September 2, 7, 9, 11, 14, 16, and 18 (11:00 AM–3:00 PM ET each). Miss the deadline entirely and the failure mode is blunt. New runners fail to register. Existing runners stop picking up jobs, and workflows targeting them sit queued or fail outright. Separately, when GitHub ships a critical security update to the runner, job queuing pauses on runners that haven't applied it.&lt;/p&gt;

&lt;p&gt;[IMAGE: articles/images/2026-08-25-github-actions-self-hosted-runner-enforcement-diagram.png | alt: "registration flow for github actions self-hosted runners hitting the minimum version gate"]&lt;/p&gt;

&lt;p&gt;If you're new to running your own runners at all, the &lt;a href="https://devtoolhub.com/github-actions-first-cicd-pipeline/" rel="noopener noreferrer"&gt;guide to your first GitHub Actions CI/CD pipeline&lt;/a&gt; covers the baseline setup this enforcement sits on top of.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Enforcement Got Paused Once Already
&lt;/h2&gt;

&lt;p&gt;This isn't GitHub's first attempt. Enforcement was originally slated for early 2026, &lt;a href="https://github.blog/changelog/2026-02-05-github-actions-self-hosted-runner-minimum-version-enforcement-extended/" rel="noopener noreferrer"&gt;pushed to March 16&lt;/a&gt;, and then &lt;a href="https://github.blog/changelog/2026-03-13-self-hosted-runner-minimum-version-enforcement-paused/" rel="noopener noreferrer"&gt;paused three days before that deadline&lt;/a&gt;. GitHub's pause notice gave no technical reason — just wanting "a smooth transition" — but third-party reporting tied it to kernel compatibility failures &lt;code&gt;v2.329.0&lt;/code&gt; exposed in legacy cgroup environments, the kind of failure that breaks CI pipelines overnight. Three months later, GitHub came back with the current July 31 / September 25 timeline.&lt;/p&gt;

&lt;p&gt;That history matters for one reason. &lt;code&gt;2.329.0&lt;/code&gt; is the floor GitHub checks at registration, not a recommendation. Pull the current runner release instead of the minimum version number quoted in the changelog — newer releases carry every fix shipped since the pause, and the 30-day rule means you need to stay near-current anyway.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Registration Trap: Why Auto-Update Won't Save an Old Runner
&lt;/h2&gt;

&lt;p&gt;Here's the gotcha GitHub's own docs don't call out directly. Auto-update on a self-hosted runner only kicks in &lt;em&gt;after&lt;/em&gt; the runner successfully registers with &lt;code&gt;./config.sh&lt;/code&gt;. If a runner is already below the minimum version, registration gets rejected first. Auto-update never gets a chance to pull a newer release.&lt;/p&gt;

&lt;p&gt;That's a real chicken-and-egg problem. It hits teams provisioning runners from a golden image or a container that hasn't been rebuilt recently. &lt;a href="https://github.com/orgs/community/discussions/182046" rel="noopener noreferrer"&gt;A GitHub community discussion&lt;/a&gt; confirms the fix isn't waiting for auto-update. Bake a current runner tarball into the image or startup script &lt;em&gt;before&lt;/em&gt; &lt;code&gt;config.sh&lt;/code&gt; runs. Auto-update keeps it current from there.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# check what version a running self-hosted runner is on&lt;/span&gt;
&lt;span class="nb"&gt;cat &lt;/span&gt;_diag/&lt;span class="k"&gt;*&lt;/span&gt;.log | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-i&lt;/span&gt; &lt;span class="s2"&gt;"Runner Version"&lt;/span&gt;

&lt;span class="c"&gt;# disable auto-update explicitly (you'll then own manual updates&lt;/span&gt;
&lt;span class="c"&gt;# within 30 days of each new release, per GitHub's policy)&lt;/span&gt;
./config.sh &lt;span class="nt"&gt;--disableupdate&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;⚠️ &lt;strong&gt;Important:&lt;/strong&gt; if you use &lt;code&gt;--disableupdate&lt;/code&gt;, you're committing to manually updating that runner within 30 days of every new release GitHub ships — not just the current one. Most teams are better off leaving auto-update on and fixing the registration-time version at the image level instead.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to Check and Upgrade GitHub Actions Self-Hosted Runners
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Inventory your runner images and startup scripts.&lt;/strong&gt; Anywhere a VM image, container image, or Kubernetes manifest pins a runner tarball URL or version string, that's a place a current runner release needs to land before September 25.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Confirm the version already registered.&lt;/strong&gt; GitHub's Actions settings page lists each runner's version under Settings → Actions → Runners; cross-reference against anything still on &lt;code&gt;2.328.x&lt;/code&gt; or earlier.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rebuild images with the current runner release&lt;/strong&gt;, not the bare minimum. The reported cgroup issues with &lt;code&gt;v2.329.0&lt;/code&gt; are reason enough to stay off the exact floor version.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Re-run registration in a non-production environment first&lt;/strong&gt; if you manage runners through Terraform, Pulumi, or a custom operator — the registration handshake is where the version gate lives, and that's the step most likely to silently fail in automation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Leave auto-update enabled&lt;/strong&gt; on any runner you don't have a strict reason to pin, so this doesn't recur at the next enforcement wave.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Once your runners are current, the &lt;a href="https://devtoolhub.com/github-actions-caching-performance-optimization/" rel="noopener noreferrer"&gt;caching and performance guide&lt;/a&gt; and the &lt;a href="https://devtoolhub.com/github-actions-security-best-practices/" rel="noopener noreferrer"&gt;security best practices guide&lt;/a&gt; are worth a pass too — a version audit is a natural time to also check for other drift. And if you want visibility into runner health going forward instead of finding out at the next deadline, the &lt;a href="https://devtoolhub.com/github-actions-monitoring-debugging-guide/" rel="noopener noreferrer"&gt;monitoring and debugging guide&lt;/a&gt; covers exactly that.&lt;/p&gt;

&lt;h2&gt;
  
  
  Upgrade Now or Wait?
&lt;/h2&gt;

&lt;p&gt;Upgrade now, in a staging environment — the brownout windows running through September 18 are exactly the feedback loop to use. GitHub's first enforcement attempt got paused amid reports of &lt;code&gt;v2.329.0&lt;/code&gt; breaking legacy cgroup setups, so treat this as a real change to test, not a checkbox. Because the failure mode is silent queuing rather than a loud error, teams that wait until the deadline usually find out from a stalled deploy, not a warning.&lt;/p&gt;

&lt;p&gt;If your runners already sit on GitHub-managed VM images or an actively maintained container base, this is likely a non-event; auto-update has probably already carried you past &lt;code&gt;2.329.0&lt;/code&gt;. The risk concentrates in golden images, air-gapped runners, and anything provisioned by infrastructure code that hasn't been touched since early 2026.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Q: What's the actual minimum GitHub Actions runner version required?&lt;/strong&gt;&lt;br&gt;
A: &lt;code&gt;2.329.0&lt;/code&gt; for registration. But pull the current release instead of the bare minimum — reported compatibility issues with &lt;code&gt;2.329.0&lt;/code&gt; on legacy cgroup setups were tied to GitHub's paused first enforcement attempt, and the 30-day rule means you need to stay near-current anyway.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Will my existing, already-registered runners just stop working on September 25?&lt;/strong&gt;&lt;br&gt;
A: Not immediately for registration — that gate only applies to new or re-registering runners. But the separate 30-day job-execution rule means an already-registered runner that falls too far behind on releases will stop picking up jobs, even without re-registering.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Does this affect GitHub Enterprise Server customers?&lt;/strong&gt;&lt;br&gt;
A: No. The enforcement applies to self-hosted runners on github.com — including GitHub Enterprise Cloud and its data-residency variant. GitHub says Enterprise Server isn't impacted at this time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Can I just disable auto-update and update manually on my own schedule?&lt;/strong&gt;&lt;br&gt;
A: Yes, with &lt;code&gt;./config.sh --disableupdate&lt;/code&gt;, but you then own updating within 30 days of every new release GitHub ships, not a schedule you control. For most teams, leaving auto-update on and fixing the version at the image level is simpler.&lt;/p&gt;

&lt;h2&gt;
  
  
  Quick Summary
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Enforcement has been live for GitHub Enterprise Cloud data-residency orgs since July 31, 2026; for the rest of github.com, brownout windows are running now ahead of full enforcement September 25&lt;/li&gt;
&lt;li&gt;Minimum version for registration is &lt;code&gt;2.329.0&lt;/code&gt;, but pull the current release — reported kernel compatibility problems with the bare minimum on legacy cgroup setups were tied to GitHub pausing its first enforcement attempt in March&lt;/li&gt;
&lt;li&gt;Auto-update can't rescue a runner that's already below the minimum, because registration is rejected before auto-update gets a chance to run&lt;/li&gt;
&lt;li&gt;Fix it at the image or startup-script level, not by waiting for the runner to self-update&lt;/li&gt;
&lt;li&gt;GitHub Enterprise Server customers are not impacted at this time, per GitHub's changelog&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Rebuild your GitHub Actions self-hosted runners with the current release now, register one in staging during a brownout window, and confirm the version shown in GitHub's Actions settings page before September 25 turns this into a production incident.&lt;/p&gt;

</description>
      <category>githubactions</category>
      <category>cicd</category>
      <category>selfhostedrunners</category>
      <category>devops</category>
    </item>
    <item>
      <title>Git 2.55: History Fixup, Rust by Default, Safer Checkouts</title>
      <dc:creator>Amaresh Pelleti</dc:creator>
      <pubDate>Tue, 25 Aug 2026 10:43:55 +0000</pubDate>
      <link>https://dev.to/amareswer/git-255-history-fixup-rust-by-default-safer-checkouts-2lbb</link>
      <guid>https://dev.to/amareswer/git-255-history-fixup-rust-by-default-safer-checkouts-2lbb</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;Originally published on &lt;a href="https://devtoolhub.com/git-2-55-new-features/" rel="noopener noreferrer"&gt;DevToolHub&lt;/a&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Git 2.55 landed on June 29, 2026, and the headline change isn't a new command — it's a build requirement. Starting with this release, building Git from source requires a Rust toolchain unless you explicitly opt out. Alongside that, Git 2.55 Git ships &lt;code&gt;git history fixup&lt;/code&gt; for editing an old commit without a full interactive rebase, a safer &lt;code&gt;git checkout -m&lt;/code&gt; that won't leave you with a half-resolved mess, and native &lt;code&gt;fsmonitor&lt;/code&gt; support on Linux.&lt;/p&gt;

&lt;p&gt;Here's exactly what changed, the commands you'll actually use, and what breaks if you build Git from source in CI.&lt;/p&gt;

&lt;h2&gt;
  
  
  git history fixup: Editing an Old Commit Without Rebase
&lt;/h2&gt;

&lt;p&gt;Interactive rebase has always been the tool for folding a small fix into an earlier commit, but it means dropping into an editor and picking through a to-do list even for a one-line change. &lt;code&gt;git history fixup&lt;/code&gt; skips that:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git add foo.c
git &lt;span class="nb"&gt;history &lt;/span&gt;fixup abcdef01
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Stage your change normally, then point &lt;code&gt;git history fixup&lt;/code&gt; at the target commit. It folds the staged change into that commit and replays every commit after it on top — the target commit keeps its original message and authorship unless you pass &lt;code&gt;--reedit-message&lt;/code&gt;. The command is intentionally conservative: it requires a clean working tree and aborts outright if replaying the later commits would produce a conflict, rather than leaving you halfway through a rebase.&lt;/p&gt;

&lt;p&gt;If you're already comfortable with &lt;code&gt;git rebase --autosquash&lt;/code&gt; and &lt;code&gt;fixup!&lt;/code&gt; commit prefixes, this covers the same use case with one command instead of a two-step commit-then-rebase dance. For anyone who's avoided &lt;a href="https://devtoolhub.com/advanced-git-rebase-stash-cherry-pick/" rel="noopener noreferrer"&gt;rebase, stash, and cherry-pick&lt;/a&gt; because the interactive editor felt like overkill for a single-line fix, this is the more direct path.&lt;/p&gt;

&lt;h2&gt;
  
  
  Safer git checkout -m Behavior
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;git checkout -m &amp;lt;branch&amp;gt;&lt;/code&gt; re-applies your local changes on top of the branch you're switching to, merging them in. Before Git 2.55, if that merge hit a conflict, you got exactly one shot to resolve it on the spot — mess it up and you were stuck with a half-resolved working tree.&lt;/p&gt;

&lt;p&gt;Git 2.55 changes the mechanics: &lt;code&gt;checkout -m&lt;/code&gt; now uses an autostash internally, so your local changes land in a stash entry instead of a conflicted working tree. You can resolve the conflict immediately, or walk away and come back to it later with &lt;code&gt;git stash pop&lt;/code&gt; when you're ready. If you've ever hit a &lt;a href="https://devtoolhub.com/how-to-handle-merge-conflicts-in-git/" rel="noopener noreferrer"&gt;merge conflict&lt;/a&gt; mid-checkout and had to untangle it under pressure, this removes the time pressure entirely.&lt;/p&gt;

&lt;h2&gt;
  
  
  Rust Now Required to Build Git From Source
&lt;/h2&gt;

&lt;p&gt;This is the change that affects the most people indirectly, even if most Git users will never notice it. As of Git 2.55, the Rust compiler is required to build Git from source, unless you explicitly disable it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Makefile-based builds&lt;/span&gt;
make &lt;span class="nv"&gt;NO_RUST&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;YesPlease

&lt;span class="c"&gt;# Meson-based builds&lt;/span&gt;
meson configure &lt;span class="nt"&gt;-Drust&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;disabled
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you install Git through your OS package manager or a prebuilt binary, this changes nothing for you — the maintainers handle the build. It matters if your CI pipeline compiles Git from source as part of a custom image or hardened build process. Any pipeline like that needs a Rust toolchain available now, or it needs the &lt;code&gt;NO_RUST&lt;/code&gt; flag set explicitly, or the build breaks the next time it runs.&lt;/p&gt;

&lt;p&gt;This is part of a longer-term push by the Git project to replace selected components with Rust for memory safety, the same direction several other systems-level open source projects have taken. Don't expect Rust to become mandatory with no opt-out in the very next release, but treat this as the direction things are heading and plan your build pipeline accordingly.&lt;/p&gt;

&lt;h2&gt;
  
  
  fsmonitor Now Works on Linux
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;fsmonitor&lt;/code&gt; speeds up status checks on large repositories by watching the filesystem in the background instead of walking the whole working tree on every &lt;code&gt;git status&lt;/code&gt;. It's had Windows and macOS implementations for a while — Git 2.55 adds a Linux implementation, built on &lt;code&gt;inotify&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The daemon doesn't need elevated privileges to run, but it does need one watch per directory, so very large repositories may need a higher &lt;code&gt;fs.inotify.max_user_watches&lt;/code&gt; limit than the Linux default. Network-mounted repositories stay opt-in for fsmonitor, matching the behavior on the other platforms. If you manage a monorepo where &lt;code&gt;git status&lt;/code&gt; has always felt sluggish, this is worth enabling.&lt;/p&gt;

&lt;p&gt;[IMAGE: articles/images/2026-07-24-git-2-55-new-features-diagram.png | alt: "how git history fixup folds a staged change into an earlier commit"]&lt;/p&gt;

&lt;h2&gt;
  
  
  Git 2.55 Parallel Hooks
&lt;/h2&gt;

&lt;p&gt;Hooks configured through Git's config system — as opposed to plain scripts dropped in &lt;code&gt;.git/hooks&lt;/code&gt; — can now run in parallel instead of one after another. Turn it on per hook:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git config hook.pre-commit.parallel &lt;span class="nb"&gt;true&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Control concurrency globally with &lt;code&gt;hook.jobs&lt;/code&gt;, per-event with &lt;code&gt;hook.&amp;lt;event&amp;gt;.jobs&lt;/code&gt;, or per-invocation with &lt;code&gt;-j&lt;/code&gt;. Hooks that depend on shared state still run serially — Git doesn't parallelize anything that could race. If you're using &lt;a href="https://devtoolhub.com/git-hooks-automate-workflow-examples/" rel="noopener noreferrer"&gt;Git hooks to automate parts of your workflow&lt;/a&gt; and have several independent checks running on every commit, this can meaningfully cut down the wait.&lt;/p&gt;

&lt;h2&gt;
  
  
  Git 2.55 Remote-Group Push
&lt;/h2&gt;

&lt;p&gt;You can now push to multiple remotes in one command by defining a remote group:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git config remotes.publish &lt;span class="s2"&gt;"github gitlab mirror"&lt;/span&gt;
git push publish main
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;git push publish main&lt;/code&gt; pushes sequentially to every remote listed under &lt;code&gt;remotes.publish&lt;/code&gt;. One limitation worth knowing: &lt;code&gt;--atomic&lt;/code&gt; isn't supported for grouped pushes, since atomicity can't be guaranteed across multiple independent connections. If a push to one remote fails partway through, the others may already have succeeded.&lt;/p&gt;

&lt;h2&gt;
  
  
  Should You Upgrade to Git 2.55 Now
&lt;/h2&gt;

&lt;p&gt;For most people, yes — none of the changes in this release are disruptive if you're using Git normally through a package manager or prebuilt binary. The one group that needs to act before upgrading: anyone with a CI pipeline or Docker image that compiles Git from source. Add a Rust toolchain to that build, or set &lt;code&gt;NO_RUST=YesPlease&lt;/code&gt; explicitly, before you pull in 2.55.&lt;/p&gt;

&lt;p&gt;Everyone else gets a genuinely useful set of additions — &lt;code&gt;git history fixup&lt;/code&gt; for quick historical edits, a safer &lt;code&gt;checkout -m&lt;/code&gt;, and fsmonitor finally working on Linux — without anything to migrate or rewrite. If you want the full technical rationale behind each change, &lt;a href="https://github.blog/open-source/git/highlights-from-git-2-55/" rel="noopener noreferrer"&gt;GitHub's release highlights&lt;/a&gt; and &lt;a href="https://about.gitlab.com/blog/whats-new-in-git-2-55-0/" rel="noopener noreferrer"&gt;GitLab's breakdown&lt;/a&gt; both cover the implementation details this article doesn't.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Q: Do I need Rust installed to use Git 2.55?&lt;/strong&gt;&lt;br&gt;
A: Only if you're building Git from source. If you install it through a package manager, Homebrew, or a prebuilt binary, this change is invisible to you.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Does git history fixup replace git rebase --autosquash?&lt;/strong&gt;&lt;br&gt;
A: Not entirely — it covers the common case of folding one staged change into an earlier commit in a single step. Complex interactive rebases with reordering or multiple fixups still need the full &lt;code&gt;git rebase -i&lt;/code&gt; workflow.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Will git checkout -m break my existing scripts if they expect the old conflict behavior?&lt;/strong&gt;&lt;br&gt;
A: If a script parses the exact output of a conflicted &lt;code&gt;checkout -m&lt;/code&gt; and expects the working tree to be left in a conflicted state, yes, that output changes. Scripts that just check the exit code are unaffected.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Can I use remote-group push with --force?&lt;/strong&gt;&lt;br&gt;
A: Yes, &lt;code&gt;--force&lt;/code&gt; works with grouped pushes. It's specifically &lt;code&gt;--atomic&lt;/code&gt; that's unsupported, because Git can't guarantee all-or-nothing delivery across multiple separate remote connections.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quick Summary:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Git 2.55 requires Rust to build from source now, unless you set &lt;code&gt;NO_RUST=YesPlease&lt;/code&gt; or &lt;code&gt;-Drust=disabled&lt;/code&gt; explicitly — end users installing prebuilt Git are unaffected&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;git history fixup &amp;lt;commit&amp;gt;&lt;/code&gt; folds a staged change into an older commit and replays descendants, without a full interactive rebase&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;git checkout -m&lt;/code&gt; now autostashes conflicts instead of leaving a half-resolved working tree&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;fsmonitor&lt;/code&gt; finally has a Linux implementation, using &lt;code&gt;inotify&lt;/code&gt; to speed up &lt;code&gt;git status&lt;/code&gt; on large repos&lt;/li&gt;
&lt;li&gt;Parallel hooks and remote-group push (&lt;code&gt;git push publish main&lt;/code&gt;) round out the release&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That's the full Git 2.55 picture: if you maintain a CI image that compiles Git from source, check that build today — that's the one place this release can break something without warning.&lt;/p&gt;

</description>
      <category>git</category>
      <category>versioncontrol</category>
      <category>devtools</category>
      <category>cli</category>
    </item>
    <item>
      <title>PostgreSQL 18: The 6 Features Worth Upgrading For</title>
      <dc:creator>Amaresh Pelleti</dc:creator>
      <pubDate>Mon, 24 Aug 2026 22:25:28 +0000</pubDate>
      <link>https://dev.to/amareswer/postgresql-18-the-6-features-worth-upgrading-for-if</link>
      <guid>https://dev.to/amareswer/postgresql-18-the-6-features-worth-upgrading-for-if</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;Originally published on &lt;a href="https://devtoolhub.com/postgresql-18-new-features/" rel="noopener noreferrer"&gt;DevToolHub&lt;/a&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;a href="https://www.postgresql.org/about/news/postgresql-18-released-3142/" rel="noopener noreferrer"&gt;PostgreSQL 18 is out&lt;/a&gt;, and it's a bigger release than the version bump suggests. The headline feature is a new asynchronous I/O subsystem that can cut read latency by up to 3x on the right workload. But the release also ships virtual generated columns as the new default, a &lt;code&gt;uuidv7()&lt;/code&gt; function for sortable primary keys, and skip scan support for indexes that used to sit unused half the time.&lt;/p&gt;

&lt;p&gt;Here's what's actually worth upgrading for, what changed under the hood, and what to check before you run &lt;code&gt;pg_upgrade&lt;/code&gt; on production. The full technical detail lives in the &lt;a href="https://www.postgresql.org/docs/current/release-18.html" rel="noopener noreferrer"&gt;official PostgreSQL 18 release notes&lt;/a&gt; — this covers the parts that change how you design and operate a database day to day.&lt;/p&gt;

&lt;h2&gt;
  
  
  PostgreSQL 18's Async I/O: Up to 3x Faster Reads
&lt;/h2&gt;

&lt;p&gt;The AIO (asynchronous I/O) subsystem is the biggest architectural change in this release. Instead of issuing one disk read and waiting for it before starting the next, PostgreSQL 18 can queue multiple read requests at once. Sequential scans, bitmap heap scans, and vacuum operations all benefit directly.&lt;/p&gt;

&lt;p&gt;Turn it on with a single setting:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SHOW&lt;/span&gt; &lt;span class="n"&gt;io_method&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two more settings control the behavior: &lt;code&gt;io_combine_limit&lt;/code&gt; and &lt;code&gt;io_max_combine_limit&lt;/code&gt; cap how many reads get batched together. If you're on a system without native &lt;code&gt;fadvise()&lt;/code&gt; support, &lt;code&gt;effective_io_concurrency&lt;/code&gt; and &lt;code&gt;maintenance_io_concurrency&lt;/code&gt; now accept values above zero too, which wasn't possible before.&lt;/p&gt;

&lt;p&gt;Want to see it working? Query the new &lt;code&gt;pg_aios&lt;/code&gt; view — it shows the file handles currently in flight for async reads. If it's empty on a read-heavy workload, &lt;code&gt;io_method&lt;/code&gt; probably isn't set the way you think it is.&lt;/p&gt;

&lt;p&gt;[IMAGE: articles/images/2026-07-24-postgresql-18-new-features-diagram.png | alt: "sequential reads compared to the new asynchronous I/O subsystem"]&lt;/p&gt;

&lt;p&gt;Storage-bound workloads see the biggest jump. If your bottleneck is CPU, not disk, don't expect PostgreSQL 18 to feel dramatically different — &lt;a href="https://devtoolhub.com/postgresql-performance-tuning-with-pg_stat_statements/" rel="noopener noreferrer"&gt;profile your actual queries&lt;/a&gt; with &lt;code&gt;pg_stat_statements&lt;/code&gt; before assuming AIO will fix a slow endpoint.&lt;/p&gt;

&lt;h2&gt;
  
  
  Virtual Generated Columns Are Now the Default
&lt;/h2&gt;

&lt;p&gt;Generated columns compute a value from other columns automatically. Before PostgreSQL 18, that computation always happened at write time, and the result got stored on disk. Now, virtual is the default:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;orders&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="nb"&gt;INT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;price&lt;/span&gt; &lt;span class="nb"&gt;NUMERIC&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;quantity&lt;/span&gt; &lt;span class="nb"&gt;INT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;total&lt;/span&gt; &lt;span class="nb"&gt;NUMERIC&lt;/span&gt; &lt;span class="k"&gt;GENERATED&lt;/span&gt; &lt;span class="n"&gt;ALWAYS&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;price&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;quantity&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="n"&gt;VIRTUAL&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Virtual columns compute the value at read time instead of write time — nothing gets stored. That saves disk space and write overhead, but it costs a bit of CPU on every read. If you need the old write-time behavior for a hot read path, say so explicitly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;orders&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="nb"&gt;INT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;price&lt;/span&gt; &lt;span class="nb"&gt;NUMERIC&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;quantity&lt;/span&gt; &lt;span class="nb"&gt;INT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;total&lt;/span&gt; &lt;span class="nb"&gt;NUMERIC&lt;/span&gt; &lt;span class="k"&gt;GENERATED&lt;/span&gt; &lt;span class="n"&gt;ALWAYS&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;price&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;quantity&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="n"&gt;STORED&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Existing STORED columns from earlier versions keep working exactly as before — this only changes the default for new tables.&lt;/p&gt;

&lt;h2&gt;
  
  
  uuidv7(): Sortable UUIDs for Primary Keys
&lt;/h2&gt;

&lt;p&gt;Random UUIDs (&lt;code&gt;uuidv4()&lt;/code&gt;) have always been bad for index locality — every insert lands in a random spot in the B-tree, which fragments the index over time. PostgreSQL 18 ships &lt;code&gt;uuidv7()&lt;/code&gt;, which encodes a timestamp into the leading bits:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;uuidv7&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

&lt;span class="k"&gt;ALTER&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;events&lt;/span&gt; &lt;span class="k"&gt;ADD&lt;/span&gt; &lt;span class="k"&gt;COLUMN&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="n"&gt;uuid&lt;/span&gt; &lt;span class="k"&gt;DEFAULT&lt;/span&gt; &lt;span class="n"&gt;uuidv7&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;PRIMARY&lt;/span&gt; &lt;span class="k"&gt;KEY&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Because the timestamp sits at the front of the value, new rows insert close together in index order, the same way an auto-incrementing integer would. You get the collision-safety of a UUID without the write-amplification problem that's plagued &lt;code&gt;uuidv4()&lt;/code&gt; primary keys for years. If you're designing a new schema, this is the default worth reaching for.&lt;/p&gt;

&lt;h2&gt;
  
  
  Skip Scan: Multicolumn Indexes That Actually Get Used
&lt;/h2&gt;

&lt;p&gt;Multicolumn B-tree indexes have a well-known limitation: PostgreSQL could only use them efficiently if your query filtered on the leading column. An index on &lt;code&gt;(tenant_id, status, created_at)&lt;/code&gt; was mostly useless for a query that filtered on &lt;code&gt;status&lt;/code&gt; alone.&lt;/p&gt;

&lt;p&gt;Skip scan changes that. PostgreSQL 18 can now use a multicolumn index even when the leading column has no restriction, by internally probing each distinct value of that column:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;INDEX&lt;/span&gt; &lt;span class="n"&gt;idx_orders&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;orders&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;tenant_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;status&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;created_at&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="c1"&gt;-- Now benefits from the index above, even without a tenant_id filter&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;orders&lt;/span&gt; &lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;status&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'pending'&lt;/span&gt; &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="n"&gt;created_at&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;now&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;interval&lt;/span&gt; &lt;span class="s1"&gt;'1 day'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It's not free — skip scan works best when the leading column has a small number of distinct values. If &lt;code&gt;tenant_id&lt;/code&gt; has millions of distinct values, the planner will likely still choose a sequential scan. Run &lt;code&gt;EXPLAIN ANALYZE&lt;/code&gt; before and after upgrading to confirm the planner actually picks it up for your specific queries.&lt;/p&gt;

&lt;h2&gt;
  
  
  OAuth Authentication Support
&lt;/h2&gt;

&lt;p&gt;PostgreSQL 18 adds native OAuth token authentication, configured in &lt;code&gt;pg_hba.conf&lt;/code&gt; like any other auth method:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;host    database    user    address    oauth
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You'll need to build with &lt;code&gt;--with-libcurl&lt;/code&gt; and load a token validation library via &lt;code&gt;oauth_validator_libraries&lt;/code&gt;. This matters if your org is trying to get off long-lived database passwords and onto short-lived tokens tied to an identity provider — previously that meant a third-party proxy in front of Postgres. Now it's built in.&lt;/p&gt;

&lt;h2&gt;
  
  
  Upgrading to PostgreSQL 18 Without Losing Planner Statistics
&lt;/h2&gt;

&lt;p&gt;The upgrade pain point that's kept people on old major versions isn't the schema — it's the hours-long window where a freshly upgraded cluster runs on empty planner statistics until &lt;code&gt;ANALYZE&lt;/code&gt; catches up. PostgreSQL 18's &lt;code&gt;pg_upgrade&lt;/code&gt; now preserves those statistics by default:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pg_upgrade &lt;span class="nt"&gt;-d&lt;/span&gt; /path/to/old_data &lt;span class="nt"&gt;-D&lt;/span&gt; /path/to/new_data
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Extended statistics (the kind built with &lt;code&gt;CREATE STATISTICS&lt;/code&gt;) aren't carried over — you'll still need to rebuild those manually. But regular column statistics survive the jump, so query plans stay sane immediately after cutover instead of degrading until autovacuum's analyze catches up. If you'd rather start clean, &lt;code&gt;--no-statistics&lt;/code&gt; disables the behavior.&lt;/p&gt;

&lt;h2&gt;
  
  
  Breaking Changes to Check Before You Upgrade
&lt;/h2&gt;

&lt;p&gt;A few defaults changed in ways that can bite you mid-migration:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Data checksums are now on by default&lt;/strong&gt; in &lt;code&gt;initdb&lt;/code&gt;. &lt;code&gt;pg_upgrade&lt;/code&gt; requires matching checksum settings between the old and new cluster, so mismatched checksum config is the first thing to check if the upgrade fails immediately.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;VACUUM and ANALYZE now process partitioned table children by default.&lt;/strong&gt; If you were relying on the old skip-children behavior, add &lt;code&gt;ONLY&lt;/code&gt; to keep it: &lt;code&gt;VACUUM ONLY parent_table&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MD5 password authentication is deprecated&lt;/strong&gt; — not removed yet, but &lt;code&gt;CREATE ROLE&lt;/code&gt; and &lt;code&gt;ALTER ROLE&lt;/code&gt; now emit a warning when you set one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Full-text search now follows the cluster's default collation provider&lt;/strong&gt; instead of always using libc. If your cluster runs a non-libc provider, reindex your full-text and &lt;code&gt;pg_trgm&lt;/code&gt; indexes after upgrading.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of these block an upgrade on their own, but any one of them can produce a confusing error if you're not expecting it — &lt;a href="https://pgpedia.info/postgresql-versions/postgresql-18.html" rel="noopener noreferrer"&gt;pgpedia's PostgreSQL 18 page&lt;/a&gt; tracks the full list if you want to check something not covered here first. If you hit something not covered here, our &lt;a href="https://devtoolhub.com/postgresql-troubleshooting-guide/" rel="noopener noreferrer"&gt;PostgreSQL troubleshooting guide&lt;/a&gt; walks through the most common post-upgrade failures.&lt;/p&gt;

&lt;h2&gt;
  
  
  Patch to the Latest Minor Version Before You Deploy
&lt;/h2&gt;

&lt;p&gt;Don't install 18.0 straight off the release notes. PostgreSQL has shipped &lt;a href="https://www.postgresql.org/support/security/" rel="noopener noreferrer"&gt;18.1 through 18.6&lt;/a&gt; since the initial release, fixing 46 CVEs along the way — several of them CVSS 8.8 remote-code-execution-class bugs. That includes &lt;a href="https://www.postgresql.org/support/security/CVE-2026-14676/" rel="noopener noreferrer"&gt;CVE-2026-14676&lt;/a&gt;, a heap buffer overflow in &lt;code&gt;pg_stat_statements&lt;/code&gt; itself, the exact extension this article points you to for profiling queries before you chase an AIO-related performance win. Pull the current minor (18.6 as of this writing) before you touch production, and set a reminder to keep pulling new minors — none of this is a one-time patch.&lt;/p&gt;

&lt;h2&gt;
  
  
  Should You Upgrade to PostgreSQL 18 Now
&lt;/h2&gt;

&lt;p&gt;If you're running a workload where disk I/O is the bottleneck, the async I/O subsystem alone justifies testing PostgreSQL 18 in staging. Combined with statistics-preserving upgrades, the operational risk of the jump itself is lower than past major version upgrades. PostgreSQL 19 is in beta (Beta 3 as of this writing) and not recommended for production, so there's no reason to wait on 18 if you're evaluating a new deployment now.&lt;/p&gt;

&lt;p&gt;If your database runs on Kubernetes, check &lt;a href="https://devtoolhub.com/postgresql-on-kubernetes-cloudnativepg/" rel="noopener noreferrer"&gt;how CloudNativePG handles major version upgrades&lt;/a&gt; before scheduling this — the sequencing matters more in an operator-managed cluster than a standalone install. And if you're running partitioned tables at scale, revisit your &lt;a href="https://devtoolhub.com/postgresql-partitioning-ultimate-guide/" rel="noopener noreferrer"&gt;partitioning strategy&lt;/a&gt; against the new VACUUM defaults before you cut over, since the children-by-default change affects partition maintenance directly.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Q: Do I need to change my application code to benefit from async I/O in PostgreSQL 18?&lt;/strong&gt;&lt;br&gt;
A: No. AIO works at the storage engine level. You get the benefit automatically once &lt;code&gt;io_method&lt;/code&gt; is configured, with no query or schema changes required.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Will my existing STORED generated columns break after upgrading to PostgreSQL 18?&lt;/strong&gt;&lt;br&gt;
A: No. Existing STORED columns keep working exactly as before. The new VIRTUAL default only applies to columns you create after upgrading, unless you explicitly write STORED.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Is &lt;code&gt;uuidv7()&lt;/code&gt; a drop-in replacement for &lt;code&gt;uuidv4()&lt;/code&gt;?&lt;/strong&gt;&lt;br&gt;
A: For new tables, yes — swap the default and you get sortable inserts. For existing tables with &lt;code&gt;uuidv4()&lt;/code&gt; primary keys already in production, migrating is a separate project since existing values won't retroactively sort.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: How long does upgrading to PostgreSQL 18 actually take with statistics preservation?&lt;/strong&gt;&lt;br&gt;
A: It depends on database size, but the statistics-preservation feature removes the multi-hour "cold cache, bad plans" window that used to follow a &lt;code&gt;pg_upgrade&lt;/code&gt;. The physical upgrade time itself is unchanged — you're saving the recovery period after it, not the upgrade itself.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Is it safe to run PostgreSQL 18.0, or should I be on a specific minor version?&lt;/strong&gt;&lt;br&gt;
A: Don't run 18.0 in production. PostgreSQL has shipped 18.1 through 18.6 since the initial release, fixing 46 CVEs, several of them CVSS 8.8 remote-code-execution bugs — including one in &lt;code&gt;pg_stat_statements&lt;/code&gt; itself. Always deploy the current minor release and keep pulling new ones as they ship.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quick Summary:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;PostgreSQL 18's async I/O subsystem can cut read latency up to 3x on storage-bound workloads, with no query changes required&lt;/li&gt;
&lt;li&gt;Virtual generated columns are now the default — computed at read time instead of write time, use STORED explicitly if you need the old behavior&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;uuidv7()&lt;/code&gt; gives you sortable, collision-safe UUIDs without the index fragmentation &lt;code&gt;uuidv4()&lt;/code&gt; causes&lt;/li&gt;
&lt;li&gt;Skip scan makes multicolumn indexes usable even when queries don't filter on the leading column&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;pg_upgrade&lt;/code&gt; now preserves planner statistics by default, removing the post-upgrade performance dip&lt;/li&gt;
&lt;li&gt;Don't deploy 18.0 as-is — 18.1 through 18.6 fixed 46 CVEs, including a critical RCE in &lt;code&gt;pg_stat_statements&lt;/code&gt;; always pull the current minor&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Test PostgreSQL 18 against your actual query patterns in staging before committing to a production upgrade — &lt;code&gt;EXPLAIN ANALYZE&lt;/code&gt; your slowest queries first and last, since the AIO and skip scan gains vary a lot by workload shape.&lt;/p&gt;

</description>
      <category>postgres</category>
      <category>database</category>
      <category>sql</category>
      <category>backend</category>
    </item>
  </channel>
</rss>
