<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Sumukh Shenoy</title>
    <description>The latest articles on DEV Community by Sumukh Shenoy (@sumukh_dev).</description>
    <link>https://dev.to/sumukh_dev</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4119060%2F29c4889f-a897-4e59-a3e8-28bea0b43dc9.png</url>
      <title>DEV Community: Sumukh Shenoy</title>
      <link>https://dev.to/sumukh_dev</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/sumukh_dev"/>
    <language>en</language>
    <item>
      <title>Why GPU Availability Is Still the Biggest Bottleneck in ML Infra</title>
      <dc:creator>Sumukh Shenoy</dc:creator>
      <pubDate>Tue, 22 Sep 2026 08:52:52 +0000</pubDate>
      <link>https://dev.to/sumukh_dev/why-gpu-availability-is-still-the-biggest-bottleneck-in-ml-infra-2bf0</link>
      <guid>https://dev.to/sumukh_dev/why-gpu-availability-is-still-the-biggest-bottleneck-in-ml-infra-2bf0</guid>
      <description>&lt;p&gt;If you've ever had a training job ready to go and then sat there refreshing your cloud console because there's no GPU capacity available in your region, you already know this problem intimately. It's 2026, GPUs are everywhere in the headlines, and yet "I need a GPU right now" is still one of the most common blockers ML teams run into.&lt;br&gt;
This isn't a rant about hardware shortages (though that's part of it). It's a look at why this keeps happening structurally, and what teams actually do to work around it.&lt;/p&gt;

&lt;p&gt;**&lt;/p&gt;

&lt;h2&gt;
  
  
  Why "just spin up a GPU instance" doesn't work the way it used to
&lt;/h2&gt;

&lt;p&gt;**&lt;br&gt;
A few years ago, GPU access was mostly a capacity problem — not enough chips being manufactured relative to demand. That's improved, but a new, quieter problem has taken its place: capacity is available, just not where and when you need it.&lt;/p&gt;

&lt;p&gt;A few things are going on at once:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Regional fragmentation.&lt;/strong&gt; A specific GPU SKU (say, an A100 or H100) might be well-stocked in one region and completely unavailable in another. Teams end up either waiting, or provisioning in a region that adds latency to the rest of their stack.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Instance-type lock-in&lt;/strong&gt;. Reserved capacity and committed-use discounts are great for cost, but they quietly reduce your flexibility. If your reserved pool is full, you're back to on-demand pricing and on-demand availability — which is exactly when availability is worst, because everyone else is in the same boat.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Burst demand mismatches supply&lt;/strong&gt;. ML workloads are spiky by nature — a training run needs 8 GPUs for 6 hours, then nothing for two weeks. Most cloud capacity planning (yours and the provider's) is built around steadier baselines, not spikes.&lt;/p&gt;

&lt;p&gt;**&lt;/p&gt;

&lt;h2&gt;
  
  
  The three ways teams actually deal with this today
&lt;/h2&gt;

&lt;p&gt;**&lt;br&gt;
&lt;strong&gt;1. Multi-region, multi-provider fallback&lt;/strong&gt;&lt;br&gt;
Instead of hardcoding a single region/provider, teams build a fallback chain: try region A, if unavailable try region B, if that fails try a secondary provider entirely. This works, but it's operational overhead — someone has to maintain that logic, and it usually means holding accounts/credentials with more than one provider just for this contingency.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Reserving capacity you don't need yet&lt;/strong&gt;&lt;br&gt;
Some teams over-provision reserved instances specifically to guarantee availability during predictable busy periods (e.g., before a product launch, or a known quarterly retraining cycle). This solves availability at the cost of paying for idle capacity most of the time — trading a scheduling problem for a cost problem.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Queue-and-wait systems&lt;/strong&gt;&lt;br&gt;
Rather than failing immediately when capacity isn't available, some internal tooling queues the job and retries on a backoff schedule until capacity frees up. This is cheap to build but can silently turn a "5 minute job" into a "we don't know when this will run" job, which is its own kind of pain for anyone waiting on results.&lt;/p&gt;

&lt;p&gt;**&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually helps, in practice
&lt;/h2&gt;

&lt;p&gt;**&lt;br&gt;
A few patterns that genuinely reduce the pain, based on what's worked for teams dealing with this regularly:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Decouple job submission from job execution.&lt;/strong&gt; If your workflow can tolerate a queue (many training jobs can), build that in from day one rather than bolting it on after the third failed terraform apply.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Track availability, not just price&lt;/strong&gt;. Most cost dashboards show you what you're spending. Far fewer show you how often your preferred instance type was actually available when you needed it. That data is what tells you whether your fallback strategy is actually working or just theoretical.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Right-size instead of over-provisioning by default&lt;/strong&gt;. A lot of "no GPUs available" pain is self-inflicted by always requesting the biggest SKU out of habit, when a smaller one (or a fractional/shared GPU) would do the job and has a much deeper availability pool.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Separate your interruptible and non-interruptible workloads.&lt;/strong&gt; Spot/preemptible capacity is far more available, but only usable if your workload can checkpoint and resume. If it can't yet, that's often a higher-leverage engineering investment than chasing more reserved capacity.&lt;/p&gt;

&lt;p&gt;**&lt;/p&gt;

&lt;h2&gt;
  
  
  The uncomfortable truth
&lt;/h2&gt;

&lt;p&gt;**&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;None of this fully "solves" GPU availability&lt;/strong&gt;. It's a structural supply/demand issue, and any individual team's workarounds are just that: workarounds. But the difference between teams that get blocked for days and teams that barely notice usually isn't luck. It's whether they built for this problem before it happened, instead of improvising a fallback plan mid-incident.&lt;/p&gt;

&lt;p&gt;If you're dealing with this right now , genuinely curious what's worked (or hasn't) for you. Multi-cloud fallback? Aggressive checkpointing? Something else entirely?&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;I like helping engineering teams eliminate cloud bill surprises, fix database bottlenecks, and scale dedicated bare metal and GPU infrastructure alongside &lt;a href="https://racko.ai/" rel="noopener noreferrer"&gt;Racko&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;

</description>
      <category>machinelearning</category>
      <category>infrastructure</category>
      <category>gpu</category>
      <category>devops</category>
    </item>
    <item>
      <title>Migrating from Shared Hosting to a VPS: A Practical Guide</title>
      <dc:creator>Sumukh Shenoy</dc:creator>
      <pubDate>Thu, 17 Sep 2026 05:07:49 +0000</pubDate>
      <link>https://dev.to/sumukh_dev/migrating-from-shared-hosting-to-a-vps-a-practical-guide-47a5</link>
      <guid>https://dev.to/sumukh_dev/migrating-from-shared-hosting-to-a-vps-a-practical-guide-47a5</guid>
      <description>&lt;p&gt;There's a specific moment every growing website hits. The site that used to load fine now feels sluggish during busy hours. Support tickets mention timeouts. You check your shared hosting dashboard and see you're sharing a server with hundreds of other sites, and you have no real control over any of it.&lt;/p&gt;

&lt;p&gt;That's usually the sign it's time for a VPS. Here's how to actually make that move without breaking anything.&lt;/p&gt;

&lt;p&gt;**&lt;/p&gt;

&lt;h2&gt;
  
  
  Why shared hosting stops working
&lt;/h2&gt;

&lt;p&gt;**&lt;br&gt;
Shared hosting is cheap because you're splitting one server's resources — CPU, memory, disk I/O — across many accounts. It works fine at low traffic. The problem is you have no guaranteed resources. If a neighboring site on the same server gets a traffic spike, your site can slow down too, and there's nothing you can do about it.&lt;/p&gt;

&lt;p&gt;A VPS (Virtual Private Server) gives you dedicated, guaranteed resources — your own slice of CPU and memory that nobody else touches. It's not a full dedicated server, but it's a real step up in both performance and control.&lt;/p&gt;

&lt;p&gt;**&lt;/p&gt;
&lt;h2&gt;
  
  
  Step 1: Pick the right VPS size
&lt;/h2&gt;

&lt;p&gt;**&lt;br&gt;
Don't guess. Check your current shared hosting's resource usage if the dashboard shows it — CPU, memory, disk. As a rough starting point for a small-to-medium site:&lt;br&gt;
• 1-2 vCPU, 2-4GB RAM — small sites, low-to-moderate traffic, a typical WordPress or small app&lt;br&gt;
• 4 vCPU, 8GB RAM — growing sites, e-commerce, anything with a database doing real work&lt;br&gt;
• 8+ vCPU, 16GB+ RAM — high traffic, multiple services running together&lt;/p&gt;

&lt;p&gt;You can usually resize later, so don't over-engineer this upfront — start close to your actual needs.&lt;/p&gt;

&lt;p&gt;**&lt;/p&gt;
&lt;h2&gt;
  
  
  Step 2: Set up the server before you touch DNS
&lt;/h2&gt;

&lt;p&gt;**&lt;/p&gt;

&lt;p&gt;This is the part people rush, and it's the part that causes downtime. Get everything working on the new VPS first, fully tested, before pointing your domain at it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Basic setup checklist&lt;/span&gt;
&lt;span class="nb"&gt;sudo &lt;/span&gt;apt update &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;sudo &lt;/span&gt;apt upgrade &lt;span class="nt"&gt;-y&lt;/span&gt;
&lt;span class="nb"&gt;sudo &lt;/span&gt;apt &lt;span class="nb"&gt;install &lt;/span&gt;nginx mysql-server php-fpm &lt;span class="nt"&gt;-y&lt;/span&gt;
&lt;span class="nb"&gt;sudo &lt;/span&gt;ufw allow &lt;span class="s1"&gt;'Nginx Full'&lt;/span&gt;
&lt;span class="nb"&gt;sudo &lt;/span&gt;ufw allow OpenSSH
&lt;span class="nb"&gt;sudo &lt;/span&gt;ufw &lt;span class="nb"&gt;enable&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Install your actual stack (whatever it is — WordPress, Node, Django, whatever your site runs on), migrate your database, migrate your files, and test everything using the server's IP address directly before touching DNS at all.&lt;/p&gt;

&lt;p&gt;**&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3: Migrate your database properly
&lt;/h2&gt;

&lt;p&gt;**&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# On the old server&lt;/span&gt;
mysqldump &lt;span class="nt"&gt;-u&lt;/span&gt; username &lt;span class="nt"&gt;-p&lt;/span&gt; database_name &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; backup.sql

&lt;span class="c"&gt;# On the new VPS&lt;/span&gt;
mysql &lt;span class="nt"&gt;-u&lt;/span&gt; username &lt;span class="nt"&gt;-p&lt;/span&gt; database_name &amp;lt; backup.sql
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Actually check the data after importing — row counts, a few spot checks on real records. Don't assume it worked just because the command didn't error out.&lt;/p&gt;

&lt;p&gt;**&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 4: Test everything before the DNS switch
&lt;/h2&gt;

&lt;p&gt;**&lt;/p&gt;

&lt;p&gt;Edit your local machine's hosts file to point your domain at the new server's IP, just for your own testing — this doesn't affect anyone else and lets you browse the site "live" on the new server before it's actually live for the world.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight conf"&gt;&lt;code&gt;&lt;span class="c"&gt;# /etc/hosts (on your own machine, not the server)
&lt;/span&gt;&lt;span class="n"&gt;YOUR_NEW_VPS_IP&lt;/span&gt; &lt;span class="n"&gt;yourdomain&lt;/span&gt;.&lt;span class="n"&gt;com&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Click through the whole site. Test forms, checkout flows, logins — anything critical. This is your last chance to catch problems before real traffic hits the new server.&lt;/p&gt;

&lt;p&gt;**&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 5: Lower your DNS TTL in advance
&lt;/h2&gt;

&lt;p&gt;**&lt;/p&gt;

&lt;p&gt;A day or two before the actual switch, lower your DNS TTL (time-to-live) to something like 300 seconds. This means when you do switch, the change propagates fast instead of some visitors hitting the old server for up to 24-48 hours after you've already moved.&lt;/p&gt;

&lt;p&gt;**&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 6: Make the switch, then watch closely
&lt;/h2&gt;

&lt;p&gt;**&lt;/p&gt;

&lt;p&gt;Update your DNS A record to point at the new VPS IP. Keep the old server running for at least 24-48 hours as a safety net — don't cancel your old hosting the same day you switch, no matter how confident you are.&lt;br&gt;
Watch your error logs and uptime closely for the first day. Small issues (a missing PHP extension, a permissions problem) are much easier to catch and fix in the first few hours than after they've been quietly broken for a week.&lt;/p&gt;

&lt;p&gt;**&lt;/p&gt;

&lt;h2&gt;
  
  
  The part people forget
&lt;/h2&gt;

&lt;p&gt;**&lt;/p&gt;

&lt;p&gt;Once everything's stable, actually decommission the old hosting — don't just let it sit there costing money out of habit. And keep a backup of the final state of the old server for a few weeks, just in case something surfaces later that you didn't test for.&lt;/p&gt;

&lt;p&gt;Anyone else gone through this migration recently? Curious what broke that you didn't expect.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;I like helping engineering teams eliminate cloud bill surprises, fix database bottlenecks, and scale dedicated bare metal and GPU infrastructure alongside &lt;a href="https://racko.ai/" rel="noopener noreferrer"&gt;Racko_&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;

</description>
      <category>webdev</category>
      <category>devops</category>
      <category>vps</category>
      <category>host</category>
    </item>
    <item>
      <title>CUDA Cores vs Tensor Cores Explained</title>
      <dc:creator>Sumukh Shenoy</dc:creator>
      <pubDate>Mon, 14 Sep 2026 05:32:36 +0000</pubDate>
      <link>https://dev.to/sumukh_dev/cuda-cores-vs-tensor-cores-explained-o4h</link>
      <guid>https://dev.to/sumukh_dev/cuda-cores-vs-tensor-cores-explained-o4h</guid>
      <description>&lt;p&gt;If you've shopped for a GPU for ML work, you've seen both numbers on the spec sheet — CUDA cores and Tensor cores — usually with zero explanation of what either actually does for you.&lt;/p&gt;

&lt;p&gt;**&lt;/p&gt;

&lt;h2&gt;
  
  
  CUDA cores: the all-purpose workers
&lt;/h2&gt;

&lt;p&gt;**&lt;/p&gt;

&lt;p&gt;CUDA cores are general-purpose processing units. Think of them as thousands of small workers, each capable of doing a simple math operation, all working in parallel. This is what makes GPUs good at graphics in the first place — and it's also why they turned out to be great at deep learning, since neural networks are basically millions of small, repetitive math operations that can run at the same time. &lt;/p&gt;

&lt;p&gt;More CUDA cores generally means more raw parallel processing power. They handle pretty much anything — traditional deep learning operations, general compute, graphics rendering.&lt;/p&gt;

&lt;p&gt;**&lt;/p&gt;

&lt;h2&gt;
  
  
  Tensor cores: built for one specific job
&lt;/h2&gt;

&lt;p&gt;**&lt;/p&gt;

&lt;p&gt;Tensor cores are a different kind of core entirely — specialized hardware built specifically for one operation: &lt;strong&gt;matrix multiplication&lt;/strong&gt; and &lt;strong&gt;accumulation&lt;/strong&gt;, done in mixed precision. That sounds narrow, but matrix multiplication is the core operation in deep learning. Most of what happens during training and inference boils down to multiplying huge matrices together. &lt;/p&gt;

&lt;p&gt;Because Tensor cores are purpose-built for exactly this, they do it dramatically faster than CUDA cores doing the same job. We're talking multiple times the throughput for the specific operations they're designed for.&lt;/p&gt;

&lt;p&gt;**&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this actually matters for you
&lt;/h2&gt;

&lt;p&gt;**&lt;/p&gt;

&lt;p&gt;Here's the practical part: if your GPU has Tensor cores, but your code doesn't use mixed precision training, you're leaving most of that speed on the table. Tensor cores need FP16 or BF16 (mixed precision) to actually kick in — running everything in standard FP32 mostly ignores them. &lt;/p&gt;

&lt;p&gt;This is an easy win most people skip. Turning on mixed precision is often one flag:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;torch.cuda.amp&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;autocast&lt;/span&gt; 



&lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nf"&gt;autocast&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt; 

    &lt;span class="n"&gt;output&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;model&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;input&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; 

    &lt;span class="n"&gt;loss&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;loss_fn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;output&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;target&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Same model, same GPU, often noticeably faster training — just because you let the Tensor cores actually do their job. &lt;/p&gt;

&lt;p&gt;**&lt;/p&gt;

&lt;h2&gt;
  
  
  So which one matters more when choosing a GPU?
&lt;/h2&gt;

&lt;p&gt;**&lt;/p&gt;

&lt;p&gt;For deep learning specifically, Tensor core count and generation usually matters more than raw CUDA core count. A GPU with fewer CUDA cores but modern Tensor cores will often beat an older GPU with more CUDA cores but no (or older) Tensor cores, for training and inference workloads. &lt;/p&gt;

&lt;p&gt;For anything outside deep learning — general compute, rendering, tasks that aren't matrix-multiplication-heavy — CUDA core count is still the more relevant number.&lt;/p&gt;

&lt;p&gt;**&lt;/p&gt;

&lt;h2&gt;
  
  
  The short version
&lt;/h2&gt;

&lt;p&gt;**&lt;br&gt;
CUDA cores: general-purpose, handle anything. &lt;/p&gt;

&lt;p&gt;Tensor cores: specialized, dramatically faster at the one operation deep learning cares about most — but only if your code actually uses mixed precision to trigger them. &lt;/p&gt;

&lt;p&gt;If you're training and not using mixed precision yet, that's probably the single easiest performance win sitting on the table right now.&lt;/p&gt;

&lt;p&gt;.......................... &lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;_I like helping engineering teams eliminate cloud bill surprises, fix database bottlenecks, and scale dedicated bare metal and GPU infrastructure alongside &lt;a href="https://racko.ai/" rel="noopener noreferrer"&gt;Racko_&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;

</description>
      <category>gpu</category>
      <category>machinelearning</category>
      <category>deeplearning</category>
      <category>beginners</category>
    </item>
  </channel>
</rss>
