<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Vishal Chandak</title>
    <description>The latest articles on DEV Community by Vishal Chandak (@chandakvishal).</description>
    <link>https://dev.to/chandakvishal</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F199244%2F45c48884-25d7-4147-b482-6a1f1c8c9f0c.png</url>
      <title>DEV Community: Vishal Chandak</title>
      <link>https://dev.to/chandakvishal</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/chandakvishal"/>
    <language>en</language>
    <item>
      <title>What if corporate bureaucracy and data center networking suffered from the exact same problem?
 AWS just slashed 69% of their networking hardware by flattening their data centers and embracing quasi-random graph math.
 
Read about how it really works here.</title>
      <dc:creator>Vishal Chandak</dc:creator>
      <pubDate>Sat, 22 Aug 2026 15:18:55 +0000</pubDate>
      <link>https://dev.to/chandakvishal/what-if-corporate-bureaucracy-and-data-center-networking-suffered-from-the-exact-same-problem-2gil</link>
      <guid>https://dev.to/chandakvishal/what-if-corporate-bureaucracy-and-data-center-networking-suffered-from-the-exact-same-problem-2gil</guid>
      <description>&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/chandakvishal/why-aws-is-tearing-down-the-tree-4em7" class="crayons-story__hidden-navigation-link"&gt;Why AWS Is Tearing Down the "Tree"?&lt;/a&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;

          &lt;a href="/chandakvishal" class="crayons-avatar  crayons-avatar--l  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F199244%2F45c48884-25d7-4147-b482-6a1f1c8c9f0c.png" alt="chandakvishal profile" class="crayons-avatar__image"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/chandakvishal" class="crayons-story__secondary fw-medium m:hidden"&gt;
              Vishal Chandak
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                Vishal Chandak
                
                
              
              &lt;div id="story-author-preview-content-4247423" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/chandakvishal" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F199244%2F45c48884-25d7-4147-b482-6a1f1c8c9f0c.png" class="crayons-avatar__image" alt=""&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;Vishal Chandak&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

          &lt;/div&gt;
          &lt;a href="https://dev.to/chandakvishal/why-aws-is-tearing-down-the-tree-4em7" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Jul 27&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/chandakvishal/why-aws-is-tearing-down-the-tree-4em7" id="article-link-4247423"&gt;
          Why AWS Is Tearing Down the "Tree"?
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/aws"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;aws&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/networktopology"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;networktopology&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/networking"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;networking&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/cloudcomputing"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;cloudcomputing&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
          &lt;a href="https://dev.to/chandakvishal/why-aws-is-tearing-down-the-tree-4em7" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left"&gt;
            &lt;div class="multiple_reactions_aggregate"&gt;
              &lt;span class="multiple_reactions_icons_container"&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/fire-f60e7a582391810302117f987b22a8ef04a2fe0df7e3258a5f49332df1cec71e.svg" width="18" height="18"&gt;
                  &lt;/span&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/raised-hands-74b2099fd66a39f2d7eed9305ee0f4553df0eb7b4f11b01b6b1b499973048fe5.svg" width="18" height="18"&gt;
                  &lt;/span&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/sparkle-heart-5f9bee3767e18deb1bb725290cb151c25234768a0e9a2bd39370c382d02920cf.svg" width="18" height="18"&gt;
                  &lt;/span&gt;
              &lt;/span&gt;
              &lt;span class="aggregate_reactions_counter"&gt;3&lt;span class="hidden s:inline"&gt;&amp;nbsp;reactions&lt;/span&gt;&lt;/span&gt;
            &lt;/div&gt;
          &lt;/a&gt;
            &lt;a href="https://dev.to/chandakvishal/why-aws-is-tearing-down-the-tree-4em7#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              &lt;span class="hidden s:inline"&gt;Add&amp;nbsp;Comment&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            5 min read
          &lt;/small&gt;
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


</description>
    </item>
    <item>
      <title>From Massive to Miniature: How Small Language Models Are Engineered</title>
      <dc:creator>Vishal Chandak</dc:creator>
      <pubDate>Sat, 22 Aug 2026 14:36:47 +0000</pubDate>
      <link>https://dev.to/chandakvishal/from-massive-to-miniature-how-small-language-models-are-engineered-223i</link>
      <guid>https://dev.to/chandakvishal/from-massive-to-miniature-how-small-language-models-are-engineered-223i</guid>
      <description>&lt;p&gt;In Part 1, we started with a simple question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Does every AI task need the biggest model we can afford?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Often, the answer is no.&lt;/p&gt;

&lt;p&gt;A model that is considerably smaller than a frontier model can be the better choice for a well-defined workload—especially when latency, cost, privacy, or on-device execution matters.&lt;/p&gt;

&lt;p&gt;But that immediately creates a harder question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;How do we build a smaller model that is still good enough?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This is where the conversation moves from strategy to engineering.&lt;/p&gt;

&lt;p&gt;You cannot take a 400-billion-parameter model, delete 99% of its parameters, and expect the remaining billion parameters to magically retain everything the original model knew.&lt;/p&gt;

&lt;p&gt;Model compression is not a file-size problem.&lt;/p&gt;

&lt;p&gt;It is a &lt;strong&gt;capability-preservation problem&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;And there are several different tools for attacking it.&lt;/p&gt;

&lt;p&gt;The most important ones are:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;Knowledge distillation&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Quantization&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Pruning&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Fine-tuning&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Parameter-efficient fine-tuning, especially LoRA and QLoRA&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;They sound similar because they all help make AI systems more efficient.&lt;/p&gt;

&lt;p&gt;Under the hood, however, they do very different things.&lt;/p&gt;




&lt;h3&gt;
  
  
  First, let's separate two problems
&lt;/h3&gt;

&lt;p&gt;Before getting into the techniques, there is an important distinction.&lt;/p&gt;

&lt;p&gt;Suppose you have a 7B model and your application needs a 3B model.&lt;/p&gt;

&lt;p&gt;There are actually two questions:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How do I make the model cheaper to run?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;and:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How do I make the model better at my particular task?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Quantization primarily attacks the first problem.&lt;/p&gt;

&lt;p&gt;Fine-tuning primarily attacks the second.&lt;/p&gt;

&lt;p&gt;Distillation can address both, because it can transfer useful behavior from a larger teacher into a smaller student.&lt;/p&gt;

&lt;p&gt;This distinction is useful because you can combine these techniques.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;            Large Teacher
                  |
                  | Distillation
                  v
           Small Base Model
                  |
                  | Fine-tuning / LoRA
                  v
          Domain Specialist
                  |
                  | Quantization
                  v
         Efficient Deployment

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is not one optimization.&lt;/p&gt;

&lt;p&gt;It's a pipeline.&lt;/p&gt;




&lt;h3&gt;
  
  
  1. Knowledge Distillation: Teach the Small Model
&lt;/h3&gt;

&lt;p&gt;Let's start with the technique most closely associated with the idea of transferring capability from a larger model.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Knowledge distillation&lt;/strong&gt; uses a larger model—the &lt;strong&gt;teacher&lt;/strong&gt; —to help train a smaller &lt;strong&gt;student&lt;/strong&gt; model.&lt;/p&gt;

&lt;p&gt;The concept predates today's LLMs. Hinton, Vinyals, and Dean described a method for transferring knowledge from an ensemble of models into a smaller model that is easier to deploy.&lt;/p&gt;

&lt;p&gt;The basic idea is beautifully simple.&lt;/p&gt;

&lt;p&gt;Instead of asking the student to learn only from the original training labels, we also let it learn from the behavior of the teacher.&lt;/p&gt;

&lt;p&gt;Conceptually:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;        Input
          |
    +-----+------+
    | |
    v v
 Teacher Student
    | |
    v v
Teacher output Student output
    | |
    +-----+------+
          |
          v
       Loss
          |
          v
   Update Student

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The teacher already contains useful information about the task.&lt;/p&gt;

&lt;p&gt;The student learns to approximate that behavior.&lt;/p&gt;




&lt;h3&gt;
  
  
  Hard targets vs. soft targets
&lt;/h3&gt;

&lt;p&gt;This is one of the most important ideas in traditional knowledge distillation.&lt;/p&gt;

&lt;p&gt;Imagine a classification problem with three classes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Cat 0.92

Dog 0.07

Rabbit 0.01

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The hard label might simply be:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;Cat&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;The hard label tells the student which answer is correct.&lt;/p&gt;

&lt;p&gt;The soft distribution tells it something more subtle:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"This is overwhelmingly a cat, but it has some resemblance to a dog and very little resemblance to a rabbit."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That additional structure can contain useful information.&lt;/p&gt;

&lt;p&gt;The original distillation work showed how these &lt;strong&gt;soft targets&lt;/strong&gt; can transfer information from a larger model into a smaller one.&lt;/p&gt;

&lt;p&gt;For modern language models, the picture becomes more complicated because the output is a sequence of tokens rather than a single class.&lt;/p&gt;

&lt;p&gt;But the principle remains:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Don't just teach the student the answer. Teach it something about how the teacher behaves.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h3&gt;
  
  
  Distillation for LLMs is more complicated
&lt;/h3&gt;

&lt;p&gt;With a language model, the teacher might generate:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"The crash is most likely caused by an invalid pointer dereference in the native library."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A straightforward approach is to train the student to reproduce that response.&lt;/p&gt;

&lt;p&gt;But there are many ways to distill an LLM.&lt;/p&gt;

&lt;p&gt;You can transfer:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;generated answers,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;token-level probabilities,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;reasoning-oriented examples,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;task-specific demonstrations,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;intermediate representations,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;preferences,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;or specialized behavior.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Modern LLM distillation research has become a large field in its own right, with different approaches targeting algorithms, skills, and domain specialization.&lt;/p&gt;

&lt;p&gt;This gives us an important correction to a common oversimplification:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Distillation isn't simply "copy the big model into the small model."&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It is a family of training techniques for transferring useful behavior.&lt;/p&gt;




&lt;h3&gt;
  
  
  Why distillation is so attractive for SLMs
&lt;/h3&gt;

&lt;p&gt;Imagine a large model is excellent at a particular task but far too expensive to deploy at scale.&lt;/p&gt;

&lt;p&gt;You can use that model offline as a teacher.&lt;/p&gt;

&lt;p&gt;Generate high-quality examples.&lt;/p&gt;

&lt;p&gt;Train a smaller model on those examples.&lt;/p&gt;

&lt;p&gt;Then deploy the smaller model for the high-volume workload.&lt;/p&gt;

&lt;p&gt;The expensive teacher doesn't necessarily have to serve every production request.&lt;/p&gt;

&lt;p&gt;That gives us an architecture like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;          EXPENSIVE / OFFLINE

         +---------------+
         | Large Teacher |
         +-------+-------+
                 |
         Generate examples
                 |
                 v
         +---------------+
         | Training Data |
         +-------+-------+
                 |
                 v
         +---------------+
         | Small Student |
         +-------+-------+
                 |
                 v

          CHEAP / ONLINE

         Millions of requests
                 |
                 v
            Small Model

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This separation between &lt;strong&gt;expensive capability acquisition&lt;/strong&gt; and &lt;strong&gt;cheap production inference&lt;/strong&gt; is one of the reasons distillation is so interesting for SLMs.&lt;/p&gt;




&lt;h3&gt;
  
  
  But distillation has a catch
&lt;/h3&gt;

&lt;p&gt;A student cannot learn what the teacher never demonstrates.&lt;/p&gt;

&lt;p&gt;If your distillation dataset contains only easy questions, the student may become excellent at easy questions and terrible at edge cases.&lt;/p&gt;

&lt;p&gt;If the teacher itself makes systematic mistakes, those mistakes can be transferred too.&lt;/p&gt;

&lt;p&gt;And if the student is substantially smaller than the teacher, there is a limit to how much information it can absorb.&lt;/p&gt;

&lt;p&gt;So a good distillation pipeline isn't:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Teacher → dump outputs → train student → ship.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It's closer to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Teacher
   |
   v
Generate candidate data
   |
   v
Filter / validate
   |
   v
Balance easy + hard examples
   |
   v
Train student
   |
   v
Evaluate against real workloads
   |
   +----&amp;gt; Fail? ----&amp;gt; Improve dataset
   |
   v
Deploy

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The dataset becomes part of the engineering&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flm9rb80sfgpkwe8iddn1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flm9rb80sfgpkwe8iddn1.png" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Quantization: Fewer Bits, Less Baggage
&lt;/h3&gt;

&lt;p&gt;Distillation changes the model itself.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quantization attacks representation.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Neural-network parameters are stored as numerical values.&lt;/p&gt;

&lt;p&gt;A model might commonly use formats such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;FP32 — 32-bit floating point&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;FP16 — 16-bit floating point&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;BF16 — 16-bit brain floating point&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;INT8 — 8-bit integer&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;INT4 — 4-bit integer&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The basic intuition is straightforward:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;If we can represent the model's numbers using fewer bits, we can reduce its memory footprint&lt;/strong&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;For example, ignoring overhead and implementation details, storing a billion parameters at 16 bits requires roughly 2 GB just for the weights.&lt;/p&gt;

&lt;p&gt;At 8 bits, it is roughly 1 GB.&lt;/p&gt;

&lt;p&gt;At 4 bits, roughly 0.5 GB.&lt;/p&gt;

&lt;p&gt;Those are simplified calculations, because real systems contain additional metadata, scaling factors, buffers, KV cache, runtime overhead, and other components.&lt;/p&gt;

&lt;p&gt;But the intuition is important.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The numerical representation matters&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Quantization isn't just "round the numbers"
&lt;/h3&gt;

&lt;p&gt;A naive approach would be:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;FP16 value
    |
    v
Round it
    |
    v
INT4 value

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;But neural networks contain distributions that aren't always friendly to naive quantization.&lt;/p&gt;

&lt;p&gt;Some values matter disproportionately.&lt;/p&gt;

&lt;p&gt;That means good quantization methods use scaling, calibration, mixed precision, or other techniques to preserve important information.&lt;/p&gt;

&lt;p&gt;LLM.int8(), for example, demonstrated an approach that handled outlier features separately while performing the majority of multiplication in 8-bit precision. The authors reported that this allowed large models to be run with substantially lower memory requirements without the performance degradation they observed from simpler approaches.&lt;/p&gt;

&lt;p&gt;The broader lesson is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Quantization is an optimization problem, not simply a bit-counting exercise.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkd7nfz04g2fvnkxnt93z.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkd7nfz04g2fvnkxnt93z.png" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h3&gt;
  
  
  What quantization buys you
&lt;/h3&gt;

&lt;p&gt;Depending on the model, hardware, and runtime, quantization can provide:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;lower model memory,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;lower bandwidth requirements,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;the ability to run models on smaller GPUs,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;improved feasibility for CPU or edge inference,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;and potentially higher throughput.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But there are trade-offs.&lt;/p&gt;

&lt;p&gt;Quantization can affect:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;accuracy,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;perplexity,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;reasoning performance,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;generation quality,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;and sometimes latency.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And the trade-off isn't identical for every model.&lt;/p&gt;

&lt;p&gt;A 4-bit model isn't automatically "better" than a 8-bit model.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;You need to measure.&lt;/strong&gt;&lt;/p&gt;




&lt;h3&gt;
  
  
  3. Pruning: Remove What You Don't Need
&lt;/h3&gt;

&lt;p&gt;Quantization changes how parameters are represented.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Pruning tries to remove parameters or structures altogether&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The intuition is similar to trimming a tree.&lt;/p&gt;

&lt;p&gt;Some branches contribute more than others.&lt;/p&gt;

&lt;p&gt;If certain weights contribute very little to the final behavior, perhaps they can be removed.&lt;/p&gt;

&lt;p&gt;There are several forms of pruning, including:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;unstructured pruning,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;structured pruning,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;neuron pruning,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;head pruning,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;layer pruning,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;and other architecture-specific approaches.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The trade-off is that sparsity is only useful if the hardware and runtime can exploit it.&lt;/p&gt;

&lt;p&gt;Imagine removing 50% of the weights but still performing essentially the same dense matrix multiplication.&lt;/p&gt;

&lt;p&gt;You may have a smaller file.&lt;/p&gt;

&lt;p&gt;You may not have a faster model.&lt;/p&gt;

&lt;p&gt;This is an important engineering distinction:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Compression does not automatically translate into acceleration&lt;/strong&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A technique can reduce storage while providing little real-world latency benefit.&lt;/p&gt;




&lt;h3&gt;
  
  
  The three techniques so far
&lt;/h3&gt;

&lt;p&gt;At this point, we can summarize the difference:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Technique&lt;/th&gt;
&lt;th&gt;What changes?&lt;/th&gt;
&lt;th&gt;Primary goal&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Distillation&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Training behavior&lt;/td&gt;
&lt;td&gt;Transfer capability&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Quantization&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Numerical representation&lt;/td&gt;
&lt;td&gt;Reduce memory/compute cost&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Pruning&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Model structure&lt;/td&gt;
&lt;td&gt;Remove unnecessary computation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Fine-tuning&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Model behavior&lt;/td&gt;
&lt;td&gt;Specialize for a task&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;LoRA / QLoRA&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Trainable parameters&lt;/td&gt;
&lt;td&gt;Make adaptation cheaper&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;And these techniques can be combined.&lt;/p&gt;

&lt;p&gt;That's where things get powerful.&lt;/p&gt;




&lt;h3&gt;
  
  
  4. Fine-tuning: Make the Model Care About Your Problem
&lt;/h3&gt;

&lt;p&gt;Pre-trained language models are generalists.&lt;/p&gt;

&lt;p&gt;They have learned from enormous and diverse datasets.&lt;/p&gt;

&lt;p&gt;But your application probably doesn't need a generalist.&lt;/p&gt;

&lt;p&gt;It needs something specific.&lt;/p&gt;

&lt;p&gt;Imagine a device diagnostics application.&lt;/p&gt;

&lt;p&gt;The model doesn't need to be an expert in Shakespeare.&lt;/p&gt;

&lt;p&gt;It needs to understand:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;crash signatures,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;error messages,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;device metadata,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;component names,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;severity levels,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;and your organization's troubleshooting vocabulary.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Fine-tuning allows us to adapt a pre-trained model to a particular task or domain.&lt;/p&gt;

&lt;p&gt;Instead of starting from zero:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Random weights
      |
      v
Train enormous model
      |
      v
General-purpose model

we start from an existing model:

Pre-trained model
       |
       v
Task-specific data
       |
       v
Fine-tuned specialist

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is dramatically more practical.&lt;/p&gt;

&lt;p&gt;But traditional fine-tuning still has a problem.&lt;/p&gt;

&lt;p&gt;You have to update the model's parameters.&lt;/p&gt;

&lt;p&gt;For a large model, that can be expensive.&lt;/p&gt;

&lt;p&gt;That's where parameter-efficient fine-tuning enters.&lt;/p&gt;




&lt;h3&gt;
  
  
  5. LoRA: Don't Rewrite the Whole Model
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;LoRA—Low-Rank Adaptation of Large Language Models—takes a clever approach.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Instead of updating all of the original model weights, LoRA freezes the pre-trained model and introduces small trainable matrices into the model's layers.&lt;/p&gt;

&lt;p&gt;Conceptually:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;          Original Model
         +--------------+
         | Frozen |

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Input ------&amp;gt;| Weights |----+ +--------------+ | +----&amp;gt; Output +--------------+ | Input ------&amp;gt;| LoRA Adapter |----+ | Trainable | +--------------+&lt;/p&gt;

&lt;p&gt;The base model stays frozen.&lt;/p&gt;

&lt;p&gt;The adapter learns the task-specific modification.&lt;/p&gt;

&lt;p&gt;The original LoRA paper demonstrated that this can dramatically reduce the number of trainable parameters and memory requirements compared with full fine-tuning while achieving comparable or better quality on the tasks they evaluated.&lt;/p&gt;

&lt;p&gt;This changes the economics of specialization.&lt;/p&gt;

&lt;p&gt;Instead of storing a complete copy of a model for every task, you can conceptually maintain:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;             Base Model
                 |
    +------------+------------+
    | | |
    v v v
 Adapter A Adapter B Adapter C
 Medical Support Coding

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The same base model can therefore support multiple specialized behaviors.&lt;/p&gt;




&lt;h3&gt;
  
  
  6. QLoRA: Quantization Meets LoRA
&lt;/h3&gt;

&lt;p&gt;Now combine two ideas.&lt;/p&gt;

&lt;p&gt;LoRA says:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Don't update the entire model.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Quantization says:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Don't store the model using unnecessarily high numerical precision.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;QLoRA combines these ideas.&lt;/p&gt;

&lt;p&gt;The base model is loaded in a quantized representation, while LoRA adapters are trained on top of it.&lt;/p&gt;

&lt;p&gt;The QLoRA paper demonstrated that this approach could reduce memory requirements enough to fine-tune a 65B-parameter model on a single 48 GB GPU while maintaining the authors' reported 16-bit fine-tuning performance on their evaluated setup.&lt;/p&gt;

&lt;p&gt;QLoRA introduced several components, including:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;4-bit NormalFloat (NF4),&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;double quantization,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;paged optimizers,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;and LoRA adapters.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The important architectural idea is simpler than the terminology:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Keep the expensive base model compressed and frozen; learn a small amount of task-specific information on top&lt;/strong&gt;.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h3&gt;
  
  
  This is where SLM engineering gets interesting
&lt;/h3&gt;

&lt;p&gt;Now we can combine the techniques.&lt;/p&gt;

&lt;p&gt;Suppose you want a specialized model for a production application.&lt;/p&gt;

&lt;p&gt;A possible pipeline is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;            Large Teacher
                 |
                 | Distillation
                 v
          Small Base Model
                 |
                 | QLoRA / Fine-tuning
                 v
         Domain Specialist
                 |
                 | Quantization
                 v
         Deployment Model
                 |
         +-------+-------+
         | |
         v v
      Cloud Edge

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Notice what happened.&lt;/p&gt;

&lt;p&gt;We didn't simply "make an LLM smaller."&lt;/p&gt;

&lt;p&gt;We built a specialized model optimized for a workload.&lt;/p&gt;

&lt;p&gt;That is a fundamentally different mindset.&lt;/p&gt;




&lt;h3&gt;
  
  
  The hidden trade-off: every optimization changes something
&lt;/h3&gt;

&lt;p&gt;There is a tendency in AI discussions to present optimization techniques as if they are free.&lt;/p&gt;

&lt;p&gt;They're not.&lt;/p&gt;

&lt;p&gt;Every technique introduces a trade-off.&lt;/p&gt;

&lt;h6&gt;
  
  
  &lt;strong&gt;Distillation&lt;/strong&gt;
&lt;/h6&gt;

&lt;p&gt;Can reduce model size while transferring useful behavior.&lt;/p&gt;

&lt;p&gt;But the student can lose capabilities the teacher had, particularly outside the distilled distribution.&lt;/p&gt;

&lt;h6&gt;
  
  
  &lt;strong&gt;Quantization&lt;/strong&gt;
&lt;/h6&gt;

&lt;p&gt;Can dramatically reduce memory requirements.&lt;/p&gt;

&lt;p&gt;But aggressive quantization can affect model quality, and the impact varies by model and task.&lt;/p&gt;

&lt;h6&gt;
  
  
  &lt;strong&gt;Pruning&lt;/strong&gt;
&lt;/h6&gt;

&lt;p&gt;Can reduce the number of parameters or operations.&lt;/p&gt;

&lt;p&gt;But irregular sparsity may not translate into real speedups on the target hardware.&lt;/p&gt;

&lt;h6&gt;
  
  
  &lt;strong&gt;Fine-tuning&lt;/strong&gt;
&lt;/h6&gt;

&lt;p&gt;Can dramatically improve performance on a domain.&lt;/p&gt;

&lt;p&gt;But a poorly designed dataset can cause overfitting, unwanted behavior, or loss of general capabilities.&lt;/p&gt;

&lt;h6&gt;
  
  
  &lt;strong&gt;LoRA&lt;/strong&gt;
&lt;/h6&gt;

&lt;p&gt;Can make specialization much cheaper.&lt;/p&gt;

&lt;p&gt;But the adapter still depends on the underlying base model, and the chosen rank, target modules, training data, and task determine how effective it is.&lt;/p&gt;

&lt;h6&gt;
  
  
  &lt;strong&gt;QLoRA&lt;/strong&gt;
&lt;/h6&gt;

&lt;p&gt;Can make fine-tuning much more memory-efficient.&lt;/p&gt;

&lt;p&gt;But quantized training introduces its own numerical and implementation considerations.&lt;/p&gt;

&lt;p&gt;There is no magic compression button.&lt;/p&gt;




&lt;h3&gt;
  
  
  The benchmark trap
&lt;/h3&gt;

&lt;p&gt;This is where experienced engineers should be particularly skeptical.&lt;/p&gt;

&lt;p&gt;Suppose someone tells you:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Our 3B model is almost as good as a 70B model."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The next question should be:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;At what?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A model can perform extremely well on one benchmark while failing badly on another.&lt;/p&gt;

&lt;p&gt;Even worse, a benchmark may not resemble your production workload.&lt;/p&gt;

&lt;p&gt;Consider a model used for structured extraction.&lt;/p&gt;

&lt;p&gt;The benchmark might measure semantic accuracy.&lt;/p&gt;

&lt;p&gt;Your application might require:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"severity"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"critical"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"component"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"camera"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"confidence"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.93&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the model instead produces:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;The severity appears to be critical and the affected componentis probably the camera. I would estimate confidence at around 93%.&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;a human might consider that a good answer.&lt;/p&gt;

&lt;p&gt;Your parser might consider it a complete failure.&lt;/p&gt;

&lt;p&gt;This is why &lt;strong&gt;application-level evaluation&lt;/strong&gt; matters.&lt;/p&gt;

&lt;p&gt;For an SLM, you should measure things like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;task accuracy,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;structured-output validity,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;latency,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;memory usage,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;throughput,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;energy consumption where relevant,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;failure rate,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;escalation rate,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;and recovery behavior.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The best model is the one that performs well across the metrics your application actually cares about.&lt;/p&gt;




&lt;h3&gt;
  
  
  A production SLM is more than a model file
&lt;/h3&gt;

&lt;p&gt;This is perhaps the most important engineering lesson from Part 2.&lt;/p&gt;

&lt;p&gt;When you deploy an SLM, you aren't deploying:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"model.bin"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;You're deploying a system.&lt;/p&gt;

&lt;p&gt;That system may include:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                User Request
                     |
                     v
             +---------------+
             | Preprocessor |
             +-------+-------+
                     |
                     v
             +---------------+
             | SLM |
             +-------+-------+
                     |
                     v
             +---------------+
             | Schema |
             | Validation |
             +-------+-------+
                     |
          +----------+----------+
          | |
        Valid Invalid
          | |
          v v
       Accept Retry /
                            Escalate
                                |
                                v
                              LLM

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is especially important for on-device applications.&lt;/p&gt;

&lt;p&gt;A model might produce a semantically correct answer but violate the application's output contract.&lt;/p&gt;

&lt;p&gt;Your runtime needs to handle that.&lt;/p&gt;

&lt;p&gt;A model might run beautifully for a 2K-token context and then exhaust memory at 16K.&lt;/p&gt;

&lt;p&gt;Your application needs to handle that.&lt;/p&gt;

&lt;p&gt;A quantized model might be fast on one device and slower on another because the runtime doesn't have optimized kernels.&lt;/p&gt;

&lt;p&gt;Your deployment system needs to handle that.&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;model is only one component&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6y5m2tnocpvlom1u2f5q.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6y5m2tnocpvlom1u2f5q.png" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h3&gt;
  
  
  So, how do you actually choose?
&lt;/h3&gt;

&lt;p&gt;There isn't one universally optimal compression strategy.&lt;/p&gt;

&lt;p&gt;Instead, start with the deployment constraint.&lt;/p&gt;

&lt;h6&gt;
  
  
  If memory is the primary problem
&lt;/h6&gt;

&lt;p&gt;Start by investigating quantization.&lt;/p&gt;

&lt;h6&gt;
  
  
  If task performance is the primary problem
&lt;/h6&gt;

&lt;p&gt;Investigate fine-tuning or distillation.&lt;/p&gt;

&lt;h6&gt;
  
  
  If training cost is the primary problem
&lt;/h6&gt;

&lt;p&gt;Investigate LoRA/QLoRA and parameter-efficient methods.&lt;/p&gt;

&lt;h6&gt;
  
  
  If the model contains unnecessary capacity
&lt;/h6&gt;

&lt;p&gt;Investigate distillation or pruning.&lt;/p&gt;

&lt;h6&gt;
  
  
  If the device is extremely constrained
&lt;/h6&gt;

&lt;p&gt;Combine techniques.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Distill → specialize → quantize → benchmark on the actual device&lt;/code&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;And don't forget the final step.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Benchmark on the actual device&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;A model that looks fantastic on an A100 benchmark may behave very differently on a phone.&lt;/p&gt;




&lt;h3&gt;
  
  
  The real optimization target
&lt;/h3&gt;

&lt;p&gt;At the beginning of this article, we asked:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;How do we make a model smaller without losing everything useful?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The answer isn't a single algorithm.&lt;/p&gt;

&lt;p&gt;It's a sequence of trade-offs.&lt;/p&gt;

&lt;p&gt;Distillation asks:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What knowledge can we transfer?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Quantization asks:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;How precisely do we need to represent it?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Pruning asks:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What computation can we remove?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Fine-tuning asks:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What behavior does this application actually need?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;LoRA asks:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;How little of the model do we need to change?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;QLoRA asks:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Can we do that while keeping the base model heavily compressed?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Together, these techniques let us move from a general-purpose model toward a model that is &lt;strong&gt;smaller, more specialized, and easier to deploy&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;But that still leaves one major problem.&lt;/p&gt;

&lt;p&gt;We've optimized the model.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;We haven't yet optimized the system.&lt;/strong&gt;&lt;/p&gt;




&lt;h3&gt;
  
  
  From model optimization to system optimization
&lt;/h3&gt;

&lt;p&gt;Imagine we have three models:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Small model → Fast, cheap, limited reasoning
Medium model → Balanced
Large model → Expensive, powerful, broad reasoning

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Which one should receive the next request?&lt;/p&gt;

&lt;p&gt;If we always choose the large model, we've thrown away much of the benefit of SLMs.&lt;/p&gt;

&lt;p&gt;If we always choose the small model, we'll eventually encounter tasks it can't handle.&lt;/p&gt;

&lt;p&gt;The interesting solution is to make the system decide.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                   Request
                      |
                      v
                +-----------+
                | Router |
                +-----+-----+
                      |
         +------------+------------+
         | | |
         v v v
       Small Medium Large
       Model Model Model

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now the optimization problem becomes much more interesting.&lt;/p&gt;

&lt;p&gt;We're no longer asking:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;"How do I make one model do everything?"&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;We're asking:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;"How do I use the right amount of intelligence for every request?"&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is the bridge between Part 2 and Part 3.&lt;/p&gt;

&lt;p&gt;And it leads us to the final—and perhaps most important—idea in this series:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The future of efficient AI may not be a smaller model. It may be a system that knows when to use a small model.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h3&gt;
  
  
  What's next?
&lt;/h3&gt;

&lt;p&gt;In Part 3, we move from &lt;strong&gt;model engineering to system orchestration&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;We'll look at:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;intelligent model routing;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;SLM + LLM hybrid architectures;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;confidence-based escalation;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;agents and tool use;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;local versus cloud execution;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;measuring energy and water consumption;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;the difference between model efficiency and system efficiency;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;and whether an SLM-first architecture actually makes AI more sustainable.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Because "small model = green AI" is an appealing story.&lt;/p&gt;

&lt;p&gt;But the real story is much more complicated.&lt;/p&gt;

&lt;p&gt;And much more interesting.&lt;/p&gt;




&lt;h3&gt;
  
  
  Part 3: The Orchestrated Future
&lt;/h3&gt;

&lt;p&gt;&lt;em&gt;When the smartest AI system isn't the one with the smartest model—but the one that knows which model to use.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>smalllanguagemodel</category>
      <category>artificialintelligen</category>
      <category>machinelearning</category>
      <category>llm</category>
    </item>
    <item>
      <title>The Small Language Model Revolution: Why Fit Beats Force</title>
      <dc:creator>Vishal Chandak</dc:creator>
      <pubDate>Tue, 11 Aug 2026 19:53:22 +0000</pubDate>
      <link>https://dev.to/chandakvishal/the-small-language-model-revolution-why-fit-beats-force-1a05</link>
      <guid>https://dev.to/chandakvishal/the-small-language-model-revolution-why-fit-beats-force-1a05</guid>
      <description>&lt;p&gt;For the last few years, AI engineering has operated under a surprisingly simple assumption:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Bigger models are better models.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;More parameters. More training data. More GPUs. More compute.&lt;/p&gt;

&lt;p&gt;And to be fair, that strategy has worked remarkably well. Scaling has produced huge improvements in language understanding, coding, reasoning, multimodal capabilities, and general-purpose AI.&lt;/p&gt;

&lt;p&gt;But engineering is rarely about maximizing one metric.&lt;/p&gt;

&lt;p&gt;Eventually, the question changes from:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Can the model solve this?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;to:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"What does it cost us to make the model solve this?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That's where Small Language Models become interesting.&lt;/p&gt;




&lt;h3&gt;
  
  
  The problem with using a giant model for everything
&lt;/h3&gt;

&lt;p&gt;Imagine a production system processing millions of AI requests every day.&lt;/p&gt;

&lt;p&gt;A request might ask the model to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;classify a support ticket;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;extract a few fields from a document;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;identify the language of a message;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;summarize a paragraph;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;detect a known failure pattern in a device log.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These are useful AI workloads.&lt;/p&gt;

&lt;p&gt;But they aren't all difficult reasoning problems.&lt;/p&gt;

&lt;p&gt;If every request is sent to the largest model available, the architecture is effectively saying:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Every problem deserves our most expensive reasoning engine."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That can work.&lt;/p&gt;

&lt;p&gt;It can also be a terrible optimization strategy.&lt;/p&gt;

&lt;p&gt;A better question is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;What is the smallest model that can solve this task reliably?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That question is at the heart of the SLM approach.&lt;/p&gt;

&lt;p&gt;The goal isn't to prove that a 1B model is "as intelligent" as a 100B model. It isn't.&lt;/p&gt;

&lt;p&gt;The goal is to recognize that &lt;strong&gt;model capability exists on a spectrum, while application requirements are usually much narrower.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A support-ticket classifier doesn't need to write a novel.&lt;/p&gt;

&lt;p&gt;A document extractor doesn't need to solve an Olympiad problem.&lt;/p&gt;

&lt;p&gt;A device assistant doesn't necessarily need the world's broadest knowledge.&lt;/p&gt;

&lt;p&gt;If a smaller model can reliably do the job, using a larger one may simply be unnecessary.&lt;/p&gt;




&lt;h3&gt;
  
  
  So, what exactly is an SLM?
&lt;/h3&gt;

&lt;p&gt;Here's where things get slightly messy.&lt;/p&gt;

&lt;p&gt;There isn't one universally accepted parameter-count boundary that separates an SLM from an LLM. Different researchers and vendors use different definitions. Some emphasize parameter count. Others focus on memory, latency, deployment environment, or computational constraints.&lt;/p&gt;

&lt;p&gt;For engineers, I find a capability-and-deployment definition more useful:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;&lt;em&gt;An SLM is a language model designed to provide useful language capabilities within a substantially smaller computational and memory footprint than large general-purpose models.&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The important part is not the exact number of parameters. Because parameter count alone doesn't tell you how practical a model is.&lt;/p&gt;

&lt;p&gt;Runtime memory depends on much more:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;parameter precision;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;architecture;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;context length;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;KV-cache size;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;batch size;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;inference runtime;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;hardware.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That becomes especially important when the target isn't a GPU cluster.&lt;/p&gt;

&lt;p&gt;It's a phone.&lt;/p&gt;




&lt;h3&gt;
  
  
  Small doesn't necessarily mean weak
&lt;/h3&gt;

&lt;p&gt;Reducing model size doesn't necessarily mean randomly throwing away capability.&lt;/p&gt;

&lt;p&gt;There are several techniques for improving the efficiency of language models.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Knowledge distillation&lt;/strong&gt; transfers useful behavior from a stronger teacher model into a smaller student.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quantization&lt;/strong&gt; represents model parameters using fewer bits, reducing memory requirements and potentially improving inference efficiency.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Pruning&lt;/strong&gt; removes parameters or structures that contribute less to computation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fine-tuning and parameter-efficient methods&lt;/strong&gt; specialize an existing model for a particular domain or workflow.&lt;/p&gt;

&lt;p&gt;None of these magically turns a small model into a frontier model.&lt;/p&gt;

&lt;p&gt;What they can do is improve the amount of useful capability we get for a given resource budget.&lt;/p&gt;

&lt;p&gt;And that is a much more interesting metric.&lt;/p&gt;




&lt;h3&gt;
  
  
  The device changes the equation
&lt;/h3&gt;

&lt;p&gt;The difference becomes especially interesting when inference moves from the data center to the device.&lt;/p&gt;

&lt;p&gt;A phone has a finite amount of:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;RAM;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;compute;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;battery;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;storage;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;memory bandwidth;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;thermal capacity.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And there is another memory consumer that developers sometimes underestimate:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;the KV cache.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;As context grows, the KV cache grows too. A model that comfortably fits into memory with a short prompt can behave very differently when an application starts processing long conversations or documents.&lt;/p&gt;

&lt;p&gt;So the engineering question isn't simply:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"&lt;strong&gt;&lt;em&gt;Can I fit the model on the phone?&lt;/em&gt;&lt;/strong&gt;"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It's:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"&lt;strong&gt;&lt;em&gt;Can I fit the entire inference workload on the phone while keeping the application responsive?&lt;/em&gt;&lt;/strong&gt;"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A 2026 practitioner case study on integrating Qwen3 0.6B and Gemma 4 E2B into a production Android word-guessing game illustrates this very well. The authors encountered output-format violations, constraint violations, context degradation, latency problems, and model-selection instability. The final architecture deliberately reduced the amount of work delegated to the model and added deterministic fallbacks.&lt;/p&gt;

&lt;p&gt;That gives us an important engineering lesson:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Putting a model on a device is an application-engineering problem, not just a model-download problem.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi3bzbodbuqn3ez3fuhtl.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi3bzbodbuqn3ez3fuhtl.png" alt="SLM Meme" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Current&lt;/strong&gt; &lt;strong&gt;Scenario&lt;/strong&gt;:&lt;/p&gt;

&lt;p&gt;Use case: "We need to classify ERROR vs WARNING."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Engineering Team:&lt;/strong&gt; Let's deploy the frontier model.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h3&gt;
  
  
  Three reasons engineers should care
&lt;/h3&gt;

&lt;p&gt;The appeal of SLMs isn't simply that they're smaller.&lt;/p&gt;

&lt;p&gt;Their smaller footprint can change three important characteristics of an AI application: &lt;strong&gt;cost, latency, and data locality.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A fourth benefit—&lt;strong&gt;deployment flexibility&lt;/strong&gt;—often follows from those three.&lt;/p&gt;




&lt;h4&gt;
  
  
  1. Cost
&lt;/h4&gt;

&lt;p&gt;Inference requires compute, and compute costs money.&lt;/p&gt;

&lt;p&gt;At low request volumes, the difference between models may not matter much.&lt;/p&gt;

&lt;p&gt;At very high volumes, however, even modest differences in per-request compute can become significant.&lt;/p&gt;

&lt;p&gt;Suppose an application receives 100 million requests.&lt;/p&gt;

&lt;p&gt;If a large fraction of those requests can be handled accurately by a smaller model, there is little reason to automatically send all 100 million to the most expensive model available.&lt;/p&gt;

&lt;p&gt;Instead, expensive inference can be reserved for requests that actually need it.&lt;/p&gt;

&lt;p&gt;This leads to an architectural principle we'll keep returning to:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Use expensive intelligence where expensive intelligence is actually necessar&lt;/strong&gt;y.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That doesn't mean "always use the cheapest model."&lt;/p&gt;

&lt;p&gt;A cheap model that fails frequently can become expensive once retries, fallbacks, downstream failures, and human review are included.&lt;/p&gt;

&lt;p&gt;So a more useful metric is: &lt;strong&gt;Cost per successful task.&lt;/strong&gt;&lt;/p&gt;




&lt;h4&gt;
  
  
  2. Latency
&lt;/h4&gt;

&lt;p&gt;Latency is another reason smaller or local models can be attractive.&lt;/p&gt;

&lt;p&gt;A cloud request can involve network transfer, queuing, inference scheduling, generation, and response transmission.&lt;/p&gt;

&lt;p&gt;A local model changes that equation.&lt;/p&gt;

&lt;p&gt;It moves computation closer to the user.&lt;/p&gt;

&lt;p&gt;For mobile assistants, interactive applications, robotics, device diagnostics, and other latency-sensitive workloads, that architectural difference can matter.&lt;/p&gt;

&lt;p&gt;But we should avoid simplistic statements such as:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;"SLMs are always X times faster."&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;They aren't.&lt;/p&gt;

&lt;p&gt;Actual latency depends on:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;model architecture;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;quantization;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;hardware;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;inference runtime;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;context length;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;concurrency;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;workload.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The defensible statement is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;&lt;em&gt;Smaller models expand the range of workloads for which local and edge inference becomes practical&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h4&gt;
  
  
  3. Privacy and data locality
&lt;/h4&gt;

&lt;p&gt;Now consider an application processing sensitive information.&lt;/p&gt;

&lt;p&gt;Perhaps it is analyzing:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;proprietary source code;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;internal documents;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;customer information;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;device telemetry;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;industrial data;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;confidential material.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A cloud architecture requires that information to cross a network boundary.&lt;/p&gt;

&lt;p&gt;That doesn't automatically make the architecture insecure. Cloud AI systems can have strong security controls.&lt;/p&gt;

&lt;p&gt;But it introduces another boundary that needs to be governed.&lt;/p&gt;

&lt;p&gt;A local model offers another option:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;&lt;em&gt;Bring the model to the data instead of bringing the data to the model&lt;/em&gt;&lt;/strong&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Again, this isn't a magic security solution.&lt;/p&gt;

&lt;p&gt;A local model doesn't protect an insecure application.&lt;/p&gt;

&lt;p&gt;But it can change the threat model and reduce the amount of information that needs to leave the device.&lt;/p&gt;




&lt;h3&gt;
  
  
  The scalpel and the Swiss Army knife
&lt;/h3&gt;

&lt;p&gt;This is where the distinction between SLMs and large general-purpose models becomes useful.&lt;/p&gt;

&lt;p&gt;Think of a large general-purpose model as a &lt;strong&gt;Swiss Army knife&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;It has a huge range of capabilities and is useful when the shape of the problem is unknown.&lt;/p&gt;

&lt;p&gt;An SLM is more like a &lt;strong&gt;specialized tool&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;It may have considerably less general capability, but if the task falls within its intended operating range, it can be a much more efficient choice.&lt;/p&gt;

&lt;p&gt;A larger model makes sense when:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;the problem is genuinely open-ended;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;complex multi-step reasoning is required;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;the request combines many capabilities;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;the cost of an incorrect answer justifies additional model capability.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A smaller model becomes attractive when:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;the task is well defined;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;the workload is repetitive;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;request volume is high;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;latency matters;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;privacy or data locality matters;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;the task can be evaluated automatically;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;the model can be specialized.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The question therefore isn't:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;&lt;em&gt;"Which model is smarter?"&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It's:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;&lt;em&gt;"Which model is sufficient?"&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h3&gt;
  
  
  But there is a catch: the capability cliff
&lt;/h3&gt;

&lt;p&gt;We also need to resist the hype.&lt;/p&gt;

&lt;p&gt;SLMs aren't miniature versions of frontier models with exactly the same capabilities.&lt;/p&gt;

&lt;p&gt;Smaller models can be remarkably capable on constrained or specialized workloads. But reducing model capacity can affect performance on unfamiliar problems, complex reasoning, and tasks requiring broad knowledge.&lt;/p&gt;

&lt;p&gt;Compare:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Classify this support ticket."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;with:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Read these conflicting reports, identify the hidden assumption, construct a counterexample, and explain why the conclusion doesn't follow."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Both requests involve language.&lt;/p&gt;

&lt;p&gt;The second requires substantially more reasoning.&lt;/p&gt;

&lt;p&gt;That's why benchmark results need context.&lt;/p&gt;

&lt;p&gt;A small model can perform extremely well on a particular benchmark and still be a poor choice for your workload.&lt;/p&gt;

&lt;p&gt;The right question isn't:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"What's this model's benchmark score?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It's:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;"How does this model perform on the tasks my application actually needs?"&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h3&gt;
  
  
  Don't choose one model. Build a system.
&lt;/h3&gt;

&lt;p&gt;This is where things get really interesting.&lt;/p&gt;

&lt;p&gt;The future doesn't have to look like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                Every Request
                     |
                     v
                 +-----------+
                |    LLM    |
                +-----------+
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Instead, introduce a routing layer:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                User Request
                     |
                     v
               +-----------+
               |   Router  |
               +-----+-----+
                     |
            +--------+--------+
            |                 |
            v                 v
       +---------+       +---------+
       |   SLM   |       |   LLM   |
       +---------+       +---------+
       Routine work     Complex work
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A routine classification task might remain local.&lt;/p&gt;

&lt;p&gt;A complicated reasoning problem can be escalated.&lt;/p&gt;

&lt;p&gt;A privacy-sensitive request might stay on-device.&lt;/p&gt;

&lt;p&gt;A request that fails validation can be retried or sent to a stronger model.&lt;/p&gt;

&lt;p&gt;This is more powerful than simply choosing "the best model."&lt;/p&gt;

&lt;p&gt;You're building a &lt;strong&gt;model hierarchy.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;And that idea becomes central in Part 3.&lt;/p&gt;




&lt;h3&gt;
  
  
  What about sustainability?
&lt;/h3&gt;

&lt;p&gt;AI systems ultimately consume physical resources.&lt;/p&gt;

&lt;p&gt;Behind an inference request are processors, memory, networking, cooling, electricity, and physical infrastructure.&lt;/p&gt;

&lt;p&gt;It's tempting to jump from that observation to:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;"Small models are green."&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That's too simplistic.&lt;/p&gt;

&lt;p&gt;A smaller model may require less computation for a comparable workload, but environmental impact depends on the complete system.&lt;/p&gt;

&lt;p&gt;We need to consider:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;model architecture;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;precision;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;hardware;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;utilization;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;context length;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;generated tokens;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;retries;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;routing;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;data-center efficiency;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;electricity source.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A small model that fails repeatedly and requires escalation may not be more efficient than a larger model that succeeds on the first attempt.&lt;/p&gt;

&lt;p&gt;The environmental question therefore isn't simply:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"How much energy does this model use?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It's:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;"How much useful work does the complete system deliver per unit of resource consumption?"&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;We'll return to that question in Part 3 of this series.&lt;/p&gt;




&lt;h3&gt;
  
  
  The principle of least intelligence
&lt;/h3&gt;

&lt;p&gt;There is a broader architectural principle hiding underneath all of this:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Give each task the minimum model capability required to solve it reliably.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Notice the important word:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;reliably&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Sometimes a deterministic program is better than an SLM.&lt;/p&gt;

&lt;p&gt;Sometimes an SLM is better than a large model.&lt;/p&gt;

&lt;p&gt;Sometimes the large model is exactly what the task requires.&lt;/p&gt;

&lt;p&gt;The engineering challenge is knowing which is which.&lt;/p&gt;

&lt;p&gt;That means evaluating models against real workloads, not just parameter counts or headline benchmarks.&lt;/p&gt;

&lt;p&gt;It also means building systems that can recover when the smaller model isn't good enough.&lt;/p&gt;

&lt;p&gt;That might involve:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;schema validation;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;confidence estimation;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;retries;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;deterministic post-processing;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;fallback models;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;human review;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;escalation.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The future of SLMs isn't just about making models smaller.&lt;/p&gt;

&lt;p&gt;It's about making &lt;strong&gt;systems smarter about when and where intelligence is used.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  What's next?
&lt;/h2&gt;

&lt;p&gt;We've answered the why.&lt;/p&gt;

&lt;p&gt;But there's an obvious question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;&lt;em&gt;How do you actually make a model smaller without throwing away everything that makes it useful?&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That's where things get interesting.&lt;/p&gt;

&lt;p&gt;In &lt;strong&gt;Part 2 of this series&lt;/strong&gt;, we'll open the machine shop.&lt;/p&gt;

&lt;p&gt;We'll look at &lt;strong&gt;knowledge distillation, quantization, pruning, fine-tuning, LoRA, and QLoRA&lt;/strong&gt;—and, more importantly, understand what each technique actually changes.&lt;/p&gt;

&lt;p&gt;We'll also look at a question that's easy to gloss over:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;When you make a model smaller, what capability did you actually lose?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If this helped you understand why SLMs matter, follow me for Part 2.&lt;/p&gt;

&lt;p&gt;In a couple of days, we'll move from the architecture diagram to the actual engineering behind the model.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Part 2: From Massive to Miniature — How We Build Small Language Models.- Coming Soon&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
    </item>
    <item>
      <title>Why AWS Is Tearing Down the "Tree"?</title>
      <dc:creator>Vishal Chandak</dc:creator>
      <pubDate>Mon, 27 Jul 2026 18:40:24 +0000</pubDate>
      <link>https://dev.to/chandakvishal/why-aws-is-tearing-down-the-tree-4em7</link>
      <guid>https://dev.to/chandakvishal/why-aws-is-tearing-down-the-tree-4em7</guid>
      <description>&lt;p&gt;Amazon CEO Andy Jassy has been on a highly publicized crusade lately to crush corporate bureaucracy. His goal? Flatten the org chart, remove the layers of middle management that slow things down, and let individual teams move faster.&lt;/p&gt;

&lt;p&gt;​As it turns out, AWS’s network engineers took his mandate incredibly literally—they just applied it to their data centers.&lt;/p&gt;

&lt;p&gt;​For thirty years, the networking industry has been stuck with a rigid hierarchy known as the "fat-tree." In this model, data center architecture looks exactly like an aging, bloated corporate org chart. To move a packet of data from one server rack to its neighbor, the data can’t just walk across the hallway. It has to climb the corporate ladder—from Top-of-Rack (ToR) switches, up to aggregation layers, and finally to the "executive" spine switches—before trickling back down to its destination.&lt;/p&gt;

&lt;p&gt;​Every layer in this network acts exactly like a middle manager: it adds latency, drains power, and serves as a potential bottleneck. While the fat-tree has been the workhorse of the cloud, AWS is officially tearing it down. They are replacing it with a radically flat alternative called the Resilient Network Graph (RNG). By taking quasi-random graph theory out of academia and into hyperscale production, AWS has figured out how to commoditize randomness.&lt;/p&gt;

&lt;p&gt;​Here is how AWS is slashing their hardware footprint while actually making the network faster.&lt;/p&gt;

&lt;h3&gt;
  
  
  ​Takeaway 1: Why "Random" Beats "Ordered"
&lt;/h3&gt;

&lt;p&gt;​In traditional networking, order equals efficiency. But RNG proves the exact opposite: randomly connecting routers is mathematically superior to stacking them in a neat hierarchy.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjvxbn6c18s7gdtkfp0d4.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjvxbn6c18s7gdtkfp0d4.jpg" width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;​In a traditional fat-tree, traffic is limited to strict pathways. Think of it like forcing all highway traffic through a few toll booths; it creates "small cuts," or bottlenecks where upper-layer links are jammed while other routes sit completely idle. Because of this, a fat-tree can strand up to 60% of its capacity.&lt;/p&gt;

&lt;p&gt;​In an RNG fabric, those toll booths don't exist. Every node has high bandwidth to all other nodes. This is powered by the &lt;em&gt;Expansion Property&lt;/em&gt; of random graphs, creating "capacity fungibility"—a fancy way of saying no bandwidth ever goes to waste.&lt;/p&gt;

&lt;p&gt;​So why didn't we do this decades ago? Because it looked terrible on paper. As Giacomo Bernardi, AWS Principal Applied Scientist, noted:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;​*“It was typical for academia. Everybody's excited, but then the real world hits.”*&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The real world meant an unmanageable hairball of physical cables and routing memory requirements that no standard switch could handle. Until now.&lt;/p&gt;

&lt;h3&gt;
  
  
  ​Takeaway 2: The "Magic Number" 69%
&lt;/h3&gt;

&lt;p&gt;​By firing the "middle managers" (removing the aggregation and spine layers entirely), AWS is pushing more data using drastically less metal. This shift to RNG is one of the most massive infrastructure optimizations in cloud history.&lt;/p&gt;

&lt;p&gt;​The impact breaks down into four staggering metrics:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;​ &lt;strong&gt;69% fewer networking devices&lt;/strong&gt; (routers and switches).&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;​ &lt;strong&gt;Up to 33% higher throughput&lt;/strong&gt; compared to traditional fat-trees.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;​ &lt;strong&gt;40% less power consumption&lt;/strong&gt; for network equipment.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;​ &lt;strong&gt;9% to 45% in cost savings&lt;/strong&gt; , depending on the setup.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;​Beyond the balance sheet, eliminating thousands of power-hungry routers directly shrinks the carbon footprint of AWS's global grids. After all, the greenest energy is the power you never have to use in the first place.&lt;/p&gt;

&lt;h3&gt;
  
  
  ​Takeaway 3: The ShuffleBox &lt;em&gt;(Taming the Spaghetti Monster)&lt;/em&gt;
&lt;/h3&gt;

&lt;p&gt;​The biggest barrier to flat networks has always been the physical wiring. Connecting routers hundreds of meters apart in a truly random pattern usually results in a logistical nightmare affectionately known as the "spaghetti monster."&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F426lk8l9qz8kgo0jtvvm.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F426lk8l9qz8kgo0jtvvm.gif" width="402" height="235"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;AWS solved this with a secret weapon called the &lt;strong&gt;ShuffleBox&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;​Think of the ShuffleBox as a master illusionist. It’s a passive optical device that internally scrambles fiber wiring into a deterministic, random pattern. To the network, the topology looks wonderfully chaotic and random; to the technician on the floor, the cabling looks perfectly structured, neat, and maintainable. Because it’s completely passive, it requires no power and has no active parts that can fail.&lt;/p&gt;

&lt;p&gt;​Crucially, the ShuffleBox fixes the expansion problem. In older random graphs, adding a new server rack meant unplugging live cables, risking a total outage. With ShuffleBoxes, technicians can plug in new racks just like they always have, growing the network without disrupting live traffic.&lt;/p&gt;

&lt;p&gt;🎥 Watch: &lt;a href="https://youtube.com/shorts/N9saEhXqF9g?si=sybAbKhzv-MKhpIy" rel="noopener noreferrer"&gt;How AWS Tames the Data Center "Spaghetti Monster" 📦⚡&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  ​Takeaway 4: Spraypoint Routing &lt;em&gt;(Taking the Side Streets)&lt;/em&gt;
&lt;/h3&gt;

&lt;p&gt;​To navigate a structure-less network, AWS had to throw out the traditional networking rulebook. Standard routing logic uses too much memory—scaling it up would require 20 to 80 times more memory than commodity switches actually have.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxkptmv6y6c2xvam7c7x6.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxkptmv6y6c2xvam7c7x6.jpg" width="599" height="627"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;​Instead, AWS built a lightweight protocol called &lt;strong&gt;Spraypoint&lt;/strong&gt;. It works on a simple "Spray and Point" strategy:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;​ &lt;strong&gt;Spray:&lt;/strong&gt; The source router essentially blasts the data out to all its immediate neighbors.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;​ &lt;strong&gt;Point:&lt;/strong&gt; The packets bounce through "rings" of waypoints surrounding the destination, guiding the data in from every possible angle.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;​It sounds counterintuitive. Why take a longer, zig-zagging path? Because in a flat network, taking a dozen different side streets is actually faster than getting stuck in a traffic jam on the main corporate highway.&lt;/p&gt;

&lt;p&gt;🎥 &lt;a href="https://youtube.com/shorts/acU4tG4Ew8E?si=409GJcbUPrjhrUO3" rel="noopener noreferrer"&gt;Watch: SprayPoint Explained&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  ​Takeaway 5: Surviving the Blast Radius
&lt;/h3&gt;

&lt;p&gt;​The best argument for RNG brings us right back to Andy Jassy’s war on corporate structure: removing single points of failure.&lt;/p&gt;

&lt;p&gt;​In a hierarchical fat-tree, if a critical "VP" spine router fails, it can wipe out half the capacity for massive regions of a data center. The blast radius is catastrophic.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F216sgyujlej7qqikkim3.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F216sgyujlej7qqikkim3.jpg" width="512" height="512"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;In an RNG topology, because there is no "top" of the network, a failure is barely a blip. If 1% of the routers fail, you only lose 1% of your capacity.&lt;/p&gt;

&lt;p&gt;​The network degrades proportionally, making outages virtually invisible to the end user. As Matt Rehder, VP of Network Engineering, put it:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;​“You turn it on, it functions. It's not something we want our customers to think about at all.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  ​The Future of the Flat Fabric
&lt;/h3&gt;

&lt;p&gt;​As of April 2026, this isn't just an ambitious science experiment. RNG is the global default for all new general-compute AWS data center builds. The network "tree" has been officially chopped down for standard workloads.&lt;/p&gt;

&lt;p&gt;​The final boss is specialized AI training. Because massive AI clusters require highly synchronized, centralized traffic patterns, they still rely on AWS's structured UltraServer architecture. But as RNG research evolves to handle these coordinated workloads, even those AI islands might eventually get flattened.&lt;/p&gt;

&lt;p&gt;​AWS has proven that chaos—when properly managed—scales better than order. The question now is whether competitors will follow them into the mesh, or remain stubbornly stuck in the trees while AWS enjoys a 69% hardware advantage.&lt;/p&gt;

</description>
      <category>aws</category>
      <category>networktopology</category>
      <category>networking</category>
      <category>cloudcomputing</category>
    </item>
    <item>
      <title>Serverless vs Managed Servers</title>
      <dc:creator>Vishal Chandak</dc:creator>
      <pubDate>Mon, 14 Dec 2020 05:56:43 +0000</pubDate>
      <link>https://dev.to/chandakvishal/serverless-vs-managed-servers-1752</link>
      <guid>https://dev.to/chandakvishal/serverless-vs-managed-servers-1752</guid>
      <description>&lt;p&gt;Serverless seems really nice to begin with but does it make your bills a lot more heavier as you grow?&lt;/p&gt;

&lt;p&gt;Uing AWS Lambda when the traffic received is around 100 transactions per seconds works fine and costs nominal but if the Transactions per seconds were to grow ten fold (or more) is Lambda still a viable option?&lt;/p&gt;

&lt;p&gt;Along with Lamba, the Cloud watch costs extra to house the logs this is not an issue with managed servers. &lt;/p&gt;

</description>
      <category>discuss</category>
      <category>serverless</category>
      <category>aws</category>
    </item>
    <item>
      <title>Demystifying Machine Learning</title>
      <dc:creator>Vishal Chandak</dc:creator>
      <pubDate>Thu, 29 Oct 2020 07:26:51 +0000</pubDate>
      <link>https://dev.to/chandakvishal/demystifying-machine-learning-1m03</link>
      <guid>https://dev.to/chandakvishal/demystifying-machine-learning-1m03</guid>
      <description>&lt;p&gt;Even though a quick google search can give you plenty of “answers” for this question, it appears that they have actually not been satisfactory enough. The fact that you are reading this post should be enough to back my claim but if you are still not convinced, take a look at the google trends report for the term “Machine Learning”.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://res.cloudinary.com/practicaldev/image/fetch/s--TLJ1MhFZ--/c_limit%2Cf_auto%2Cfl_progressive%2Cq_auto%2Cw_880/https://dev-to-uploads.s3.amazonaws.com/i/lzvyoi9sfp0ubmjidx8a.png" class="article-body-image-wrapper"&gt;&lt;img src="https://res.cloudinary.com/practicaldev/image/fetch/s--TLJ1MhFZ--/c_limit%2Cf_auto%2Cfl_progressive%2Cq_auto%2Cw_880/https://dev-to-uploads.s3.amazonaws.com/i/lzvyoi9sfp0ubmjidx8a.png" alt="Alt Text" width="880" height="269"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;So what is Machine Learning? Let’s find out!&lt;/p&gt;

&lt;p&gt;Before you go any further, I should warn you that this post is intended for people with little or no idea about the topic and thus has a super-simplified view of what Machine learning is.&lt;/p&gt;

&lt;p&gt;To put it in Layman’s terms— Machine Learning is the ability to let machine form an equation to correctly solve any problem with little or no human intervention.&lt;/p&gt;

&lt;p&gt;With Machine Learning, instead of programming the machine to do something, we let the machine learn to find a solution to the problem by itself.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://res.cloudinary.com/practicaldev/image/fetch/s--daEvihKp--/c_limit%2Cf_auto%2Cfl_progressive%2Cq_auto%2Cw_880/https://dev-to-uploads.s3.amazonaws.com/i/inf4h17ltelxx43vyn64.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://res.cloudinary.com/practicaldev/image/fetch/s--daEvihKp--/c_limit%2Cf_auto%2Cfl_progressive%2Cq_auto%2Cw_880/https://dev-to-uploads.s3.amazonaws.com/i/inf4h17ltelxx43vyn64.jpeg" alt="Machine Learning Training Phase" title="Machine Learning Training Phase" width="880" height="439"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://res.cloudinary.com/practicaldev/image/fetch/s--lV7hzyC0--/c_limit%2Cf_auto%2Cfl_progressive%2Cq_auto%2Cw_880/https://dev-to-uploads.s3.amazonaws.com/i/hkta7dbybz10z4po0u0s.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://res.cloudinary.com/practicaldev/image/fetch/s--lV7hzyC0--/c_limit%2Cf_auto%2Cfl_progressive%2Cq_auto%2Cw_880/https://dev-to-uploads.s3.amazonaws.com/i/hkta7dbybz10z4po0u0s.jpeg" alt="Alt Text" width="623" height="459"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://res.cloudinary.com/practicaldev/image/fetch/s--yIsbRHrU--/c_limit%2Cf_auto%2Cfl_progressive%2Cq_auto%2Cw_880/https://dev-to-uploads.s3.amazonaws.com/i/0pymkf16s5cyu2fpybgg.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://res.cloudinary.com/practicaldev/image/fetch/s--yIsbRHrU--/c_limit%2Cf_auto%2Cfl_progressive%2Cq_auto%2Cw_880/https://dev-to-uploads.s3.amazonaws.com/i/0pymkf16s5cyu2fpybgg.jpeg" alt="Machine Learning in action" title="Machine Learning in action" width="880" height="528"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  “It sounds cool, but how is that it even possible.” — Your Brain
&lt;/h2&gt;

&lt;p&gt;Hold your horses, we will get to that question in a minute. Before that, let’s understand how is it different than the “usual” way of solving a problem. &lt;/p&gt;

&lt;p&gt;To understand this analogy better, let’s take a simple real-world example of — “Classifying shapes”.&lt;/p&gt;




&lt;h2&gt;
  
  
  Problem:
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Given an image of a particular shape, let’s find out if the image is for a Square or Triangle.&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Non Machine Learning Approach
&lt;/h3&gt;

&lt;p&gt;The pseudo-code (fancy term programmers use to define “steps”) for that would look like:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Take the input Image&lt;/li&gt;
&lt;li&gt;Find the number of corners in the image. This operation would basically include: (Don’t worry if you don’t understand all the terms below, neither do I)

&lt;ul&gt;
&lt;li&gt;Convert the image from RGB (i.e. Colored Image) to Grayscale (i.e. B/w)&lt;/li&gt;
&lt;li&gt;Smooth-en out the image&lt;/li&gt;
&lt;li&gt;Use thresholding to convert the image to Bi-Level Binary Image&lt;/li&gt;
&lt;li&gt;Use an edge-detector algorithm to get the list of edges using pixel gradients&lt;/li&gt;
&lt;li&gt;Use contours to get the list of corners in the image&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Write an if-else block to return the shape of the object based on the number of corners returned&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Here’s a small visualization to help you better visualize what the above steps might look like:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://res.cloudinary.com/practicaldev/image/fetch/s--fejA4QIb--/c_limit%2Cf_auto%2Cfl_progressive%2Cq_auto%2Cw_880/https://dev-to-uploads.s3.amazonaws.com/i/74l3zn93hsezqyynlbni.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://res.cloudinary.com/practicaldev/image/fetch/s--fejA4QIb--/c_limit%2Cf_auto%2Cfl_progressive%2Cq_auto%2Cw_880/https://dev-to-uploads.s3.amazonaws.com/i/74l3zn93hsezqyynlbni.jpeg" alt="Alt Text" width="880" height="540"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Huh! You must be thinking — That seems a lot of work for such a trivial task. You are not entirely wrong but this has been the general approach to solve such problems until a few years back. Even though this approach “worked”, it had a lot of issues. For starters, it only worked well with perfectly captured images i.e. pictures taken under proper light and clear background which means that it will have a high percentage of failure in real-world scenarios.&lt;/p&gt;

&lt;h3&gt;
  
  
  Machine Learning Approach to the rescue
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://res.cloudinary.com/practicaldev/image/fetch/s--yx_mGoYE--/c_limit%2Cf_auto%2Cfl_progressive%2Cq_66%2Cw_880/https://dev-to-uploads.s3.amazonaws.com/i/spnpw565rbebkut2bygl.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://res.cloudinary.com/practicaldev/image/fetch/s--yx_mGoYE--/c_limit%2Cf_auto%2Cfl_progressive%2Cq_66%2Cw_880/https://dev-to-uploads.s3.amazonaws.com/i/spnpw565rbebkut2bygl.gif" alt="Alt Text" width="500" height="281"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Let’s see what the same would look like with Machine Learning:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Gather some images for different shapes — squares and triangles.&lt;/li&gt;
&lt;li&gt;Feed it to a ML Algorithm (more technical term would be a Convolutional Neural Network)&lt;/li&gt;
&lt;li&gt;Run the image to be tested on the above trained model&lt;/li&gt;
&lt;li&gt;Sit back and relax!&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Now you must be thinking &lt;strong&gt;“Ok. That seems to be awfully easy. But how does that even work?”&lt;/strong&gt; Let me answer that question with a question of my own — &lt;strong&gt;How would you teach a child a difference between a square and a triangle?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;You won’t start by telling him about the number of corners or that the angles between them should be 90; Rather you would show him some images and help him relate that one of them is called a square and other is called a triangle.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Exactly! This is the way Machine Learning works too — based on feedback from the label provided with the data.&lt;/strong&gt;&lt;/p&gt;




&lt;p&gt;The below image might help understand it better:&lt;br&gt;
&lt;a href="https://res.cloudinary.com/practicaldev/image/fetch/s--Usk-Yo4G--/c_limit%2Cf_auto%2Cfl_progressive%2Cq_auto%2Cw_880/https://dev-to-uploads.s3.amazonaws.com/i/6rngxb0m6vkbtc67zcg1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://res.cloudinary.com/practicaldev/image/fetch/s--Usk-Yo4G--/c_limit%2Cf_auto%2Cfl_progressive%2Cq_auto%2Cw_880/https://dev-to-uploads.s3.amazonaws.com/i/6rngxb0m6vkbtc67zcg1.png" alt="Alt Text" width="880" height="724"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://res.cloudinary.com/practicaldev/image/fetch/s--E8N1g1dE--/c_limit%2Cf_auto%2Cfl_progressive%2Cq_auto%2Cw_880/https://dev-to-uploads.s3.amazonaws.com/i/3fmidy6ofos9cbe8u9f9.png" class="article-body-image-wrapper"&gt;&lt;img src="https://res.cloudinary.com/practicaldev/image/fetch/s--E8N1g1dE--/c_limit%2Cf_auto%2Cfl_progressive%2Cq_auto%2Cw_880/https://dev-to-uploads.s3.amazonaws.com/i/3fmidy6ofos9cbe8u9f9.png" alt="Alt Text" width="880" height="685"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Let’s take off the training wheels and see how far our newly trained child can go:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://res.cloudinary.com/practicaldev/image/fetch/s--vSchBu24--/c_limit%2Cf_auto%2Cfl_progressive%2Cq_auto%2Cw_880/https://dev-to-uploads.s3.amazonaws.com/i/xvvasjhbk5mtpim9e6ko.png" class="article-body-image-wrapper"&gt;&lt;img src="https://res.cloudinary.com/practicaldev/image/fetch/s--vSchBu24--/c_limit%2Cf_auto%2Cfl_progressive%2Cq_auto%2Cw_880/https://dev-to-uploads.s3.amazonaws.com/i/xvvasjhbk5mtpim9e6ko.png" alt="Alt Text" width="880" height="200"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://res.cloudinary.com/practicaldev/image/fetch/s--sTyIftKx--/c_limit%2Cf_auto%2Cfl_progressive%2Cq_66%2Cw_880/https://dev-to-uploads.s3.amazonaws.com/i/hasps4nw0ll89cvfcum5.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://res.cloudinary.com/practicaldev/image/fetch/s--sTyIftKx--/c_limit%2Cf_auto%2Cfl_progressive%2Cq_66%2Cw_880/https://dev-to-uploads.s3.amazonaws.com/i/hasps4nw0ll89cvfcum5.gif" alt="Alt Text" width="480" height="270"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
   “All that’s cool but I don’t need a machine to categorize shapes.” — Almost everyone
&lt;/h2&gt;

&lt;p&gt;To give a little more perspective on its potential - Machine Learning is the magic behind how Google lets you find your favorite cat photos (Who doesn’t love cats!) on the internet. If you look closely, almost everything you interact with on a day-to-day basis has machine learning weaved into it — Amazon product recommendations, Netflix movie recommendations, predicting stock prices, Alexa Voice, Sales forecasting, Automatic subtitle generation, Language translation and a lot more. The details pertaining to those is outside the scope of this post.&lt;br&gt;
&lt;strong&gt;The major advantages with the Machine Learning approach is that:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;You don’t need any domain knowledge to get started&lt;/li&gt;
&lt;li&gt;The same classifier which can recognize shapes can also recognize numerical digits, cats vs dogs etc. since the Machine trains itself based on the input data&lt;/li&gt;
&lt;li&gt;Can be continuously improved based on its own performance&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The few disadvantages with Machine Learning are:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Needs a lot of training data for better predictions&lt;/li&gt;
&lt;li&gt;Needs a lot of processing power for training the model and deriving inferences&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;However, the advent of the information age, providing us with humongous amount of data, and the drastic reduction in CPU and GPU costs over the years have helped solve both the above problems. Thus, it makes sense now more than ever, to have a generalized solution for such cases rather than a problem-specific solution.&lt;/p&gt;

&lt;p&gt;That’s where Machine Learning comes into play. Sounds really cool, doesn’t it? If it’s so fascinating, why isn’t everyone trying to get their hands dirty with Machine Learning?&lt;/p&gt;

&lt;p&gt;I will be addressing this question in my next post. Stay Tuned!&lt;/p&gt;

&lt;p&gt;Until then, That’s all folks!&lt;/p&gt;

</description>
      <category>machinelearning</category>
      <category>ai</category>
      <category>technology</category>
    </item>
  </channel>
</rss>
