<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Viswaretas Kotra</title>
    <description>The latest articles on DEV Community by Viswaretas Kotra (@viswaretas_kotra).</description>
    <link>https://dev.to/viswaretas_kotra</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4172307%2F52f79e2f-4ce1-40f8-bb27-0a62c752e52c.png</url>
      <title>DEV Community: Viswaretas Kotra</title>
      <link>https://dev.to/viswaretas_kotra</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/viswaretas_kotra"/>
    <language>en</language>
    <item>
      <title>GPU cloud pricing, October 2026: what an H100, H200, B200 or B300 costs per hour</title>
      <dc:creator>Viswaretas Kotra</dc:creator>
      <pubDate>Fri, 09 Oct 2026 02:08:27 +0000</pubDate>
      <link>https://dev.to/viswaretas_kotra/gpu-cloud-pricing-october-2026-what-an-h100-h200-b200-or-b300-costs-per-hour-5a25</link>
      <guid>https://dev.to/viswaretas_kotra/gpu-cloud-pricing-october-2026-what-an-h100-h200-b200-or-b300-costs-per-hour-5a25</guid>
      <description>&lt;p&gt;&lt;strong&gt;Short answer (Oct 8, 2026):&lt;/strong&gt; on &lt;a href="https://www.nodus-compute.ai/pricing/" rel="noopener noreferrer"&gt;Nodus&lt;/a&gt;, an H100 SXM 80 GB is $2.60/hr, H200 SXM $3.59/hr, B200 $5.50/hr and B300 $6.94/hr on demand. Consumer and prosumer cards go much lower: an RTX 4090 is $0.29/hr and an RTX 3090 is $0.18/hr. Modal's listed rates for the same data center GPUs run from 2% higher (B300) to about 2.5x higher (L40S) depending on the card.&lt;/p&gt;

&lt;p&gt;Prices below are per GPU-hour, USD, taken from the public Nodus pricing page on October 8, 2026, which lists Modal's published rate beside each one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Data center GPUs
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;GPU&lt;/th&gt;
&lt;th&gt;VRAM&lt;/th&gt;
&lt;th&gt;Nodus&lt;/th&gt;
&lt;th&gt;Modal&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;B300&lt;/td&gt;
&lt;td&gt;288 GB&lt;/td&gt;
&lt;td&gt;$6.943&lt;/td&gt;
&lt;td&gt;$7.099&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;B200&lt;/td&gt;
&lt;td&gt;180 GB&lt;/td&gt;
&lt;td&gt;$5.500&lt;/td&gt;
&lt;td&gt;$6.250&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;H200 SXM&lt;/td&gt;
&lt;td&gt;141 GB&lt;/td&gt;
&lt;td&gt;$3.593&lt;/td&gt;
&lt;td&gt;$4.540&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;H100 SXM&lt;/td&gt;
&lt;td&gt;80 GB&lt;/td&gt;
&lt;td&gt;$2.600&lt;/td&gt;
&lt;td&gt;$3.949&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A100&lt;/td&gt;
&lt;td&gt;80 GB&lt;/td&gt;
&lt;td&gt;$1.090&lt;/td&gt;
&lt;td&gt;$2.498&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A100&lt;/td&gt;
&lt;td&gt;40 GB&lt;/td&gt;
&lt;td&gt;$0.904&lt;/td&gt;
&lt;td&gt;$2.099&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;RTX PRO 6000&lt;/td&gt;
&lt;td&gt;96 GB&lt;/td&gt;
&lt;td&gt;$1.350&lt;/td&gt;
&lt;td&gt;$3.031&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;L40S&lt;/td&gt;
&lt;td&gt;48 GB&lt;/td&gt;
&lt;td&gt;$0.793&lt;/td&gt;
&lt;td&gt;$1.951&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;L4&lt;/td&gt;
&lt;td&gt;24 GB&lt;/td&gt;
&lt;td&gt;$0.430&lt;/td&gt;
&lt;td&gt;$0.799&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Budget and prosumer GPUs
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;GPU&lt;/th&gt;
&lt;th&gt;VRAM&lt;/th&gt;
&lt;th&gt;Nodus&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;RTX 5090&lt;/td&gt;
&lt;td&gt;32 GB&lt;/td&gt;
&lt;td&gt;$0.480&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;RTX 4090&lt;/td&gt;
&lt;td&gt;24 GB&lt;/td&gt;
&lt;td&gt;$0.290&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;RTX 3090&lt;/td&gt;
&lt;td&gt;24 GB&lt;/td&gt;
&lt;td&gt;$0.180&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;RTX 6000 Ada&lt;/td&gt;
&lt;td&gt;48 GB&lt;/td&gt;
&lt;td&gt;$0.743&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;RTX A6000&lt;/td&gt;
&lt;td&gt;48 GB&lt;/td&gt;
&lt;td&gt;$0.333&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;RTX A4000&lt;/td&gt;
&lt;td&gt;16 GB&lt;/td&gt;
&lt;td&gt;$0.173&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Which one should you rent?
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Fine-tuning a 7B to 14B model with LoRA:&lt;/strong&gt; a 48 GB card (L40S, A6000) or a 24 GB card with QLoRA. A100 80 GB if you want full fine-tunes of small models.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Full fine-tunes and pretraining at scale:&lt;/strong&gt; H100 or H200. H200's 141 GB helps with long context and larger batch sizes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Serving frontier-size open models:&lt;/strong&gt; B200 or B300. B300's 288 GB per GPU means fewer GPUs per replica, which often lowers cost per token even though the hourly rate is higher.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Experiments, evals and small jobs:&lt;/strong&gt; RTX 4090 / 5090 at under $0.50/hr are hard to beat.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What the hourly rate does not tell you
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Boot and teardown.&lt;/strong&gt; Some platforms bill from machine creation to deletion. On Nodus the bill is split into &lt;code&gt;Boot&lt;/code&gt;, &lt;code&gt;Running&lt;/code&gt; and &lt;code&gt;Teardown&lt;/code&gt; segments so you can see it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Price changes mid-run.&lt;/strong&gt; Nodus freezes the rate when the machine is chosen for the whole run.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Interruptions.&lt;/strong&gt; Spot capacity is cheaper, but only if your job checkpoints. Otherwise one preemption can wipe out the savings. See our &lt;a href="https://www.nodus-compute.ai/docs/guides/checkpoints/" rel="noopener noreferrer"&gt;preemption guide&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reserved terms.&lt;/strong&gt; For multi-month commitments on B200/B300, reserved pricing is quoted separately and depends on term length, location and interconnect. Nodus handles these through &lt;a href="https://www.nodus-compute.ai/reserved-compute/" rel="noopener noreferrer"&gt;Reserved Compute&lt;/a&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;What is the cheapest H100 per hour in October 2026?&lt;/strong&gt; On Nodus, $2.60/hr for an H100 SXM 80 GB, on demand.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How much does a B300 cost per hour?&lt;/strong&gt; $6.94/hr on demand on Nodus as of Oct 8, 2026. Reserved terms are quoted per request.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is there a free tier?&lt;/strong&gt; New Nodus accounts get a $30 starter grant valid for 30 days.&lt;/p&gt;

&lt;p&gt;We refresh this table monthly. Live prices: &lt;a href="https://www.nodus-compute.ai/pricing/" rel="noopener noreferrer"&gt;nodus-compute.ai/pricing&lt;/a&gt;&lt;/p&gt;

</description>
      <category>gpu</category>
      <category>cloud</category>
      <category>machinelearning</category>
      <category>nvidia</category>
    </item>
    <item>
      <title>How to give Claude Code, Codex or Cursor access to GPUs (MCP setup)</title>
      <dc:creator>Viswaretas Kotra</dc:creator>
      <pubDate>Fri, 09 Oct 2026 02:00:11 +0000</pubDate>
      <link>https://dev.to/viswaretas_kotra/how-to-give-claude-code-codex-or-cursor-access-to-gpus-mcp-setup-573o</link>
      <guid>https://dev.to/viswaretas_kotra/how-to-give-claude-code-codex-or-cursor-access-to-gpus-mcp-setup-573o</guid>
      <description>&lt;p&gt;&lt;strong&gt;Short answer:&lt;/strong&gt; connect your coding agent to a GPU platform's MCP server. The agent can then estimate, launch, watch and fetch results from GPU jobs from the same chat where it wrote the code. With &lt;a href="https://www.nodus-compute.ai/" rel="noopener noreferrer"&gt;Nodus&lt;/a&gt; that is one command per client, and every paid action shows a dry run and cost estimate before it runs.&lt;/p&gt;

&lt;p&gt;Coding agents are great at writing a training script and terrible at the part after: finding a GPU, shipping code to it, watching logs, copying results back. MCP closes that loop.&lt;/p&gt;

&lt;h2&gt;
  
  
  Claude Code
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;claude mcp add &lt;span class="nt"&gt;--scope&lt;/span&gt; user &lt;span class="nt"&gt;--transport&lt;/span&gt; http nodus https://api-next.nodus-compute.ai/mcp
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then run &lt;code&gt;/mcp&lt;/code&gt;, pick &lt;code&gt;nodus&lt;/code&gt;, and sign in in your browser.&lt;/p&gt;

&lt;h2&gt;
  
  
  Codex
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;codex mcp add nodus &lt;span class="nt"&gt;--url&lt;/span&gt; https://api-next.nodus-compute.ai/mcp
codex mcp login nodus
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Cursor
&lt;/h2&gt;

&lt;p&gt;Add this to &lt;code&gt;~/.cursor/mcp.json&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"mcpServers"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"nodus"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"http"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"url"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://api-next.nodus-compute.ai/mcp"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Any other MCP client
&lt;/h2&gt;

&lt;p&gt;Add a remote HTTP MCP server with URL &lt;code&gt;https://api-next.nodus-compute.ai/mcp&lt;/code&gt; and follow the sign-in prompt. Prefer a local server? Install the CLI and run &lt;code&gt;nodus mcp install&lt;/code&gt;, which configures the agents on your machine.&lt;/p&gt;

&lt;h2&gt;
  
  
  Verify without spending money
&lt;/h2&gt;

&lt;p&gt;Ask the agent: "List my Nodus jobs." That is read only. A good first prompt for any agent that can read URLs:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Read https://nodus-compute.ai/connect.md and help me connect Nodus to this agent.
Verify setup by listing my jobs. Do not start paid compute.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  What the agent can actually do
&lt;/h2&gt;

&lt;p&gt;The server exposes generic tools over every resource: &lt;code&gt;get&lt;/code&gt;, &lt;code&gt;describe&lt;/code&gt;, &lt;code&gt;logs&lt;/code&gt;, &lt;code&gt;estimate&lt;/code&gt;, &lt;code&gt;apply&lt;/code&gt;, &lt;code&gt;exec&lt;/code&gt; and more. In practice that means prompts like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;"Run train.py on a 24 GB GPU with a $5 cap and stream the logs."&lt;/li&gt;
&lt;li&gt;"Why did job/train fail?"&lt;/li&gt;
&lt;li&gt;"Fine-tune Qwen3 0.6B with LoRA on chats.jsonl and download the adapter."&lt;/li&gt;
&lt;li&gt;"What did my sandboxes cost this week?"&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Guardrails that matter when an agent holds your credit card
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Writes are dry-run first.&lt;/strong&gt; The agent sees the object as it would be created plus the cost estimate, and it only runs once confirmed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Same limits as you.&lt;/strong&gt; The agent uses your account, projects, budgets and spending caps. A project budget stops it from overspending even if it gets creative.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hard caps per job.&lt;/strong&gt; &lt;code&gt;maxCostUSD&lt;/code&gt; stops a job gracefully, with a checkpoint, before it passes the cap.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Frozen rates.&lt;/strong&gt; The hourly rate is fixed when the machine is chosen and holds for the whole run.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Do I need a separate "agents" product to use Claude Code with GPUs?&lt;/strong&gt; No. MCP is enough. Nodus Agents is a different feature for running agent conversations inside Nodus.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Which GPU does the agent get?&lt;/strong&gt; By default the cheapest offering that fits the request and can start now, within your limits. You can ask for a specific type like H100 or B200.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What does it cost to try?&lt;/strong&gt; New accounts get a $30 starter grant, no card needed for the first runs.&lt;/p&gt;

&lt;p&gt;Setup for every client: &lt;a href="https://www.nodus-compute.ai/connect/" rel="noopener noreferrer"&gt;nodus-compute.ai/connect&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>mcp</category>
      <category>claude</category>
      <category>gpu</category>
    </item>
    <item>
      <title>How to make GPU training survive spot preemption (without babysitting it)</title>
      <dc:creator>Viswaretas Kotra</dc:creator>
      <pubDate>Fri, 09 Oct 2026 01:59:34 +0000</pubDate>
      <link>https://dev.to/viswaretas_kotra/how-to-make-gpu-training-survive-spot-preemption-without-babysitting-it-44k6</link>
      <guid>https://dev.to/viswaretas_kotra/how-to-make-gpu-training-survive-spot-preemption-without-babysitting-it-44k6</guid>
      <description>&lt;p&gt;&lt;strong&gt;Short answer:&lt;/strong&gt; save your model, optimizer and step counter to one directory, write those files atomically, load them on startup, and run on a platform that copies that directory before the machine disappears and restores it on the next machine. Do that and a preemption costs you minutes of work, not the whole run.&lt;/p&gt;

&lt;p&gt;Spot and interruptible GPUs are often a fraction of on-demand prices. The catch is that the provider can take the machine back. Here is the pattern we use at &lt;a href="https://www.nodus-compute.ai/" rel="noopener noreferrer"&gt;Nodus&lt;/a&gt;, and it works anywhere.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Put all recovery state in one directory
&lt;/h2&gt;

&lt;p&gt;Everything you need to continue goes in one place: weights, optimizer state, LR scheduler, RNG state, data loader position, current step. Results you want to download (final model, eval reports) go somewhere else. Mixing the two makes checkpoints big and slow.&lt;/p&gt;

&lt;p&gt;On Nodus that directory is &lt;code&gt;/nodus/state&lt;/code&gt; (also in &lt;code&gt;$NODUS_STATE_DIR&lt;/code&gt;), and outputs go to &lt;code&gt;/nodus/outputs&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Write atomically
&lt;/h2&gt;

&lt;p&gt;A checkpoint taken while you are halfway through writing a file restores a broken state. Write to a temp file, then rename it over the old one. Rename is atomic on POSIX filesystems.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pathlib&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Path&lt;/span&gt;

&lt;span class="n"&gt;STATE_DIR&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Path&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;NODUS_STATE_DIR&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/nodus/state&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="n"&gt;STATE&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;STATE_DIR&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;progress.json&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;save&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;step&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;opt&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;STATE_DIR&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;mkdir&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;parents&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;exist_ok&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;tmp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;STATE_DIR&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ckpt.pt.tmp&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;torch&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;save&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;state_dict&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;opt&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;opt&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;state_dict&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;step&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;step&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="n"&gt;tmp&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;tmp&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;STATE_DIR&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ckpt.pt&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  3. Resume on startup
&lt;/h2&gt;

&lt;p&gt;Checkpoints restore files, not process memory. Your program starts from the top, so it has to check for saved state:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;start&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
&lt;span class="n"&gt;ckpt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;STATE_DIR&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ckpt.pt&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;ckpt&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;exists&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="n"&gt;s&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;torch&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;load&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ckpt&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;load_state_dict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]);&lt;/span&gt; &lt;span class="n"&gt;opt&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;load_state_dict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;opt&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]);&lt;/span&gt; &lt;span class="n"&gt;start&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;step&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;resumed from step &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;start&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;step&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;start&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;total_steps&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="bp"&gt;...&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  4. Pick a checkpoint cadence
&lt;/h2&gt;

&lt;p&gt;Too often and you waste GPU time writing files. Too rarely and each preemption throws away a lot of work. A practical rule: aim for at least four checkpoints per expected run, and keep checkpointing under about 10% of runtime. Capacity that is rarely interrupted can be checkpointed less often.&lt;/p&gt;

&lt;p&gt;Nodus does this math for you with &lt;code&gt;interval: auto&lt;/code&gt;, using how often that capacity actually gets interrupted and how long your saves take.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Save when the machine is about to go away
&lt;/h2&gt;

&lt;p&gt;Most providers give a short reclaim notice. Use it. On Nodus your program can subscribe to checkpoint requests over a local socket and acknowledge when files are consistent:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;nodus&lt;/span&gt;  &lt;span class="c1"&gt;# pip install nodus-compute; no-ops outside Nodus
&lt;/span&gt;
&lt;span class="n"&gt;nodus&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;checkpoint&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;on_request&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;lambda&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;save&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;step&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;opt&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When the capacity gives notice, Nodus sends an urgent request, takes the checkpoint after your ack, and prepares a replacement machine at the same time. The replacement only starts once the old one is provably gone, so two machines never write the same state.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. Hugging Face Trainer users
&lt;/h2&gt;

&lt;p&gt;Trainer already saves and resumes. Point &lt;code&gt;output_dir&lt;/code&gt; at the state directory, set &lt;code&gt;save_steps&lt;/code&gt; and &lt;code&gt;save_total_limit=2&lt;/code&gt;, and call &lt;code&gt;trainer.train(resume_from_checkpoint=True)&lt;/code&gt; when &lt;code&gt;NODUS_RESTORED=1&lt;/code&gt;. Or set &lt;code&gt;integration: HFTrainer&lt;/code&gt; and Nodus registers the save-on-request callback for you.&lt;/p&gt;

&lt;h2&gt;
  
  
  Running it
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;nodus-compute
nodus login
nodus run &lt;span class="nt"&gt;--gpu&lt;/span&gt; H100 &lt;span class="nt"&gt;--interruptible&lt;/span&gt; &lt;span class="nt"&gt;--checkpoint&lt;/span&gt; /nodus/state &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="nt"&gt;--&lt;/span&gt; python train.py
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;nodus describe job/&amp;lt;name&amp;gt;&lt;/code&gt; shows each attempt, why it ended (for example &lt;code&gt;Preempted&lt;/code&gt;) and the latest checkpoint. New accounts get a $30 starter grant, so you can test a preemption-safe run without a card.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Does checkpointing restore GPU memory?&lt;/strong&gt; No. It restores files. Your code reloads its own state.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What if a checkpoint is empty?&lt;/strong&gt; An empty state directory never replaces an earlier useful checkpoint, so a crash before the first save cannot erase progress.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How many times will it retry?&lt;/strong&gt; By default up to 8 recoveries (3 for distributed jobs). Two attempts in a row with no progress stops the run with &lt;code&gt;NoProgress&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does this work for multi-node?&lt;/strong&gt; Yes, in beta. Each rank writes its shard with &lt;code&gt;torch.distributed.checkpoint&lt;/code&gt; and rank 0 commits it.&lt;/p&gt;

&lt;p&gt;Full guide: &lt;a href="https://www.nodus-compute.ai/docs/guides/checkpoints/" rel="noopener noreferrer"&gt;nodus-compute.ai/docs/guides/checkpoints&lt;/a&gt;&lt;/p&gt;

</description>
      <category>machinelearning</category>
      <category>pytorch</category>
      <category>gpu</category>
      <category>mlops</category>
    </item>
  </channel>
</rss>
