<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: mohit modi</title>
    <description>The latest articles on DEV Community by mohit modi (@mohit_modi_e86a932fb11e61).</description>
    <link>https://dev.to/mohit_modi_e86a932fb11e61</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4084180%2F94579c93-e39f-44b7-80a1-945f0f0c0475.jpg</url>
      <title>DEV Community: mohit modi</title>
      <link>https://dev.to/mohit_modi_e86a932fb11e61</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/mohit_modi_e86a932fb11e61"/>
    <language>en</language>
    <item>
      <title>Understanding Principal Component Analysis</title>
      <dc:creator>mohit modi</dc:creator>
      <pubDate>Tue, 01 Sep 2026 03:52:54 +0000</pubDate>
      <link>https://dev.to/mohit_modi_e86a932fb11e61/understanding-principal-component-analysis-3n84</link>
      <guid>https://dev.to/mohit_modi_e86a932fb11e61/understanding-principal-component-analysis-3n84</guid>
      <description>&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/mohit_modi_e86a932fb11e61/what-pca-is-actually-doing-to-your-data-4a77" class="crayons-story__hidden-navigation-link"&gt;What PCA is actually doing to your data&lt;/a&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;

          &lt;a href="/mohit_modi_e86a932fb11e61" class="crayons-avatar  crayons-avatar--l  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4084180%2F94579c93-e39f-44b7-80a1-945f0f0c0475.jpg" alt="mohit_modi_e86a932fb11e61 profile" class="crayons-avatar__image" width="96" height="96"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/mohit_modi_e86a932fb11e61" class="crayons-story__secondary fw-medium m:hidden"&gt;
              mohit modi
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                mohit modi
                
                
              
              &lt;div id="story-author-preview-content-4541777" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/mohit_modi_e86a932fb11e61" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4084180%2F94579c93-e39f-44b7-80a1-945f0f0c0475.jpg" class="crayons-avatar__image" alt="" width="96" height="96"&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;mohit modi&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

          &lt;/div&gt;
          &lt;a href="https://dev.to/mohit_modi_e86a932fb11e61/what-pca-is-actually-doing-to-your-data-4a77" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Sep 1&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/mohit_modi_e86a932fb11e61/what-pca-is-actually-doing-to-your-data-4a77" id="article-link-4541777"&gt;
          What PCA is actually doing to your data
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag crayons-tag--filled  " href="/t/discuss"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;discuss&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/machinelearning"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;machinelearning&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/datascience"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;datascience&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/learning"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;learning&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
            &lt;a href="https://dev.to/mohit_modi_e86a932fb11e61/what-pca-is-actually-doing-to-your-data-4a77#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              &lt;span class="hidden s:inline"&gt;Add&amp;nbsp;Comment&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            5 min read
          &lt;/small&gt;
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


</description>
    </item>
    <item>
      <title>Understanding Principal Component Analysis</title>
      <dc:creator>mohit modi</dc:creator>
      <pubDate>Tue, 01 Sep 2026 03:52:54 +0000</pubDate>
      <link>https://dev.to/mohit_modi_e86a932fb11e61/understanding-principal-component-analysis-3ep4</link>
      <guid>https://dev.to/mohit_modi_e86a932fb11e61/understanding-principal-component-analysis-3ep4</guid>
      <description>&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/mohit_modi_e86a932fb11e61/what-pca-is-actually-doing-to-your-data-4a77" class="crayons-story__hidden-navigation-link"&gt;What PCA is actually doing to your data&lt;/a&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;

          &lt;a href="/mohit_modi_e86a932fb11e61" class="crayons-avatar  crayons-avatar--l  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4084180%2F94579c93-e39f-44b7-80a1-945f0f0c0475.jpg" alt="mohit_modi_e86a932fb11e61 profile" class="crayons-avatar__image" width="96" height="96"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/mohit_modi_e86a932fb11e61" class="crayons-story__secondary fw-medium m:hidden"&gt;
              mohit modi
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                mohit modi
                
                
              
              &lt;div id="story-author-preview-content-4541777" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/mohit_modi_e86a932fb11e61" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4084180%2F94579c93-e39f-44b7-80a1-945f0f0c0475.jpg" class="crayons-avatar__image" alt="" width="96" height="96"&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;mohit modi&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

          &lt;/div&gt;
          &lt;a href="https://dev.to/mohit_modi_e86a932fb11e61/what-pca-is-actually-doing-to-your-data-4a77" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Sep 1&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/mohit_modi_e86a932fb11e61/what-pca-is-actually-doing-to-your-data-4a77" id="article-link-4541777"&gt;
          What PCA is actually doing to your data
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag crayons-tag--filled  " href="/t/discuss"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;discuss&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/machinelearning"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;machinelearning&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/datascience"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;datascience&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/learning"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;learning&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
            &lt;a href="https://dev.to/mohit_modi_e86a932fb11e61/what-pca-is-actually-doing-to-your-data-4a77#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              &lt;span class="hidden s:inline"&gt;Add&amp;nbsp;Comment&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            5 min read
          &lt;/small&gt;
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


</description>
    </item>
    <item>
      <title>What PCA is actually doing to your data</title>
      <dc:creator>mohit modi</dc:creator>
      <pubDate>Tue, 01 Sep 2026 03:31:17 +0000</pubDate>
      <link>https://dev.to/mohit_modi_e86a932fb11e61/what-pca-is-actually-doing-to-your-data-4a77</link>
      <guid>https://dev.to/mohit_modi_e86a932fb11e61/what-pca-is-actually-doing-to-your-data-4a77</guid>
      <description>&lt;p&gt;Most explanations of PCA stop at "it reduces dimensions." That is true, and it is roughly as useful as saying a regression "fits a line." It tells you what happens without telling you what the method is doing or when it will fail you.&lt;/p&gt;

&lt;p&gt;Here is the whole idea in one sentence: &lt;strong&gt;PCA finds the directions your data varies in most, by taking the eigenvectors of its covariance matrix.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Everything else is detail. But that sentence only helps if you know what a covariance matrix is and what an eigenvector does, so let's build it up — and if the prior question is &lt;a href="https://dev.to/library/principal-component-analysis/dimensionality-reduction-motivation"&gt;why you'd reduce dimensions at all&lt;/a&gt;, start there.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with the covariance matrix
&lt;/h2&gt;

&lt;p&gt;If you centre your data — subtract the mean from every column — the &lt;a href="https://dev.to/library/correlation-covariance/understanding-covariance"&gt;covariance matrix&lt;/a&gt; is:&lt;/p&gt;

&lt;p&gt;Sigma = 1/(n-1).X.X_t&lt;/p&gt;

&lt;p&gt;For p features you get a pxp matrix. The diagonal holds each feature's variance. Everything off the diagonal holds the covariance between a pair of features: how much they move together.&lt;/p&gt;

&lt;p&gt;That matrix is the entire input to PCA. Not the raw data — the covariance structure. Which means PCA only ever sees how your features vary and co-vary, and is blind to anything else about them.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the eigenvectors are doing
&lt;/h2&gt;

&lt;p&gt;For a square matrix Sigma, an eigenvector v is a direction that the matrix does not rotate:&lt;/p&gt;

&lt;p&gt;Sigma v = lambda v&lt;/p&gt;

&lt;p&gt;Multiplying by Sigma only stretches v, by a factor of lambda. Every other direction gets turned.&lt;/p&gt;

&lt;p&gt;When Sigma is a covariance matrix, those special directions have a concrete meaning: &lt;strong&gt;they are the axes along which the data varies independently&lt;/strong&gt;, and lambda is &lt;em&gt;how much&lt;/em&gt; variance lies along each one.&lt;/p&gt;

&lt;p&gt;So sorting the eigenvectors by their eigenvalues sorts the directions of your data from "most spread out" to "least." The first eigenvector is the first principal component. It is the single direction that captures more of your data's variance than any other.&lt;/p&gt;

&lt;p&gt;That is the whole method. Compute the covariance matrix, take its eigenvectors, order them by eigenvalue, keep the top few. The &lt;a href="https://dev.to/library/principal-component-analysis/eigen-decomposition-principal-directions"&gt;full derivation&lt;/a&gt; is worth walking through once.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the components are uncorrelated
&lt;/h2&gt;

&lt;p&gt;A covariance matrix is symmetric, and symmetric matrices have orthogonal eigenvectors. Every principal component is at right angles to every other.&lt;/p&gt;

&lt;p&gt;Orthogonal directions have zero covariance. So your new features are, by construction, &lt;strong&gt;&lt;a href="https://dev.to/library/correlation-covariance/understanding-correlation"&gt;completely uncorrelated&lt;/a&gt; with each other&lt;/strong&gt; — a property none of your original features are likely to have.&lt;/p&gt;

&lt;p&gt;Hold onto that. It is the source of PCA's most underrated use, which we'll come to.&lt;/p&gt;

&lt;h2&gt;
  
  
  Choosing how many components to keep
&lt;/h2&gt;

&lt;p&gt;Each eigenvalue is the variance along its component, so the fraction of total variance a component explains is just its share of the sum:&lt;/p&gt;

&lt;p&gt;explained variance ratio (i) = lambda(i)/ sigma of lambda(j)&lt;/p&gt;

&lt;p&gt;Plot the cumulative version and keep however many components clear the bar you care about — 90%, 95%, whatever the application justifies. There is no correct threshold, which is worth saying plainly rather than pretending an elbow plot settles it. More on &lt;a href="https://dev.to/library/principal-component-analysis/determining-optimal-components"&gt;choosing the number of components&lt;/a&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;sklearn.decomposition&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;PCA&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;sklearn.preprocessing&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;StandardScaler&lt;/span&gt;

&lt;span class="n"&gt;X_scaled&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;StandardScaler&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;fit_transform&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;   &lt;span class="c1"&gt;# not optional — see below
&lt;/span&gt;&lt;span class="n"&gt;pca&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;PCA&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;fit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X_scaled&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;pca&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;explained_variance_ratio_&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;cumsum&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  PCA and feature selection
&lt;/h2&gt;

&lt;p&gt;This is where terminology causes real confusion, so let's be exact.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;PCA is not feature selection. It is feature extraction.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;em&gt;Selection&lt;/em&gt; keeps a subset of your original columns. You start with 50 features, you end with 12 of the original 50, and you can still say "this model uses income and tenure."&lt;/li&gt;
&lt;li&gt;
&lt;em&gt;Extraction&lt;/em&gt; builds new features out of the old ones. PC1 is not one of your columns; it is a weighted blend of &lt;strong&gt;all&lt;/strong&gt; of them — something like 0.4xIncome + 0.3xtenure - 0.2xage + ..&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So PCA never selects anything. It replaces your feature set with a smaller, rotated one. If someone tells you they "used PCA for feature selection," they mean one of the three things below.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Dimensionality reduction before modelling
&lt;/h3&gt;

&lt;p&gt;The standard use. Fifty correlated features become ten components carrying 95% of the variance, and the model trains on those. It serves the same purpose as selection — fewer inputs, less overfitting, faster training — while doing something mathematically different.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Reading the loadings to inform real selection
&lt;/h3&gt;

&lt;p&gt;Each component has loadings: the weight of every original feature within it. If one feature dominates the top components, that is evidence it carries a lot of the variance, and you might keep it in a genuine selection step.&lt;/p&gt;

&lt;p&gt;This is legitimate but weaker than it looks. High loading means high &lt;em&gt;variance contribution&lt;/em&gt;, not high &lt;em&gt;predictive value&lt;/em&gt; — and those are not the same thing. See &lt;a href="https://dev.to/library/principal-component-analysis/interpreting-component-loadings"&gt;interpreting component loadings&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Removing multicollinearity
&lt;/h3&gt;

&lt;p&gt;The one people underuse. Because the components are orthogonal, they have &lt;strong&gt;zero correlation with each other by construction&lt;/strong&gt;. Regressing on principal components instead of raw features makes multicollinearity structurally impossible.&lt;/p&gt;

&lt;p&gt;If you have ever watched a VIF climb past 10 and had to decide which of two near-identical predictors to drop, PCA sidesteps that decision entirely — it is one of the few honest answers to that &lt;a href="https://dev.to/library/linear-regression/assumptions-linear-regression"&gt;violated regression assumption&lt;/a&gt;, rather than a coin toss between two predictors.&lt;/p&gt;

&lt;h3&gt;
  
  
  Advantages
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;No multicollinearity&lt;/strong&gt;, guaranteed by orthogonality — not something you check for afterwards.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fewer inputs&lt;/strong&gt; means less overfitting and faster training, particularly when p approaches n.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Noise reduction.&lt;/strong&gt; Low-variance components are often mostly measurement noise; dropping them can improve signal.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It keeps information that selection throws away.&lt;/strong&gt; Dropping a column discards everything in it. PCA compresses all fifty features into ten components, so a feature's contribution survives even when it isn't individually important.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It's unsupervised&lt;/strong&gt;, so it can run before you have labels, and it cannot leak your target into the transformation.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  The costs, which are not small
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;You lose interpretability.&lt;/strong&gt; PC1 is a blend of everything. You can no longer say "a year of tenure is worth £400." For anything that needs explaining to a regulator, a clinician, or a stakeholder, that alone can rule PCA out.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;PCA maximises variance, not predictive power.&lt;/strong&gt; It never sees your target. The direction your data varies in most is not necessarily the direction that predicts y — a low-variance component can easily be the one that matters, and dropping it because it explains 2% of variance can quietly cost you the model.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It only finds linear structure.&lt;/strong&gt; Data on a curved manifold won't be captured well. That's what kernel PCA, t-SNE and UMAP are for.&lt;/p&gt;

&lt;h2&gt;
  
  
  The mistake almost everyone makes once
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Standardise your features first.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;PCA maximises variance, and variance has units. A feature measured in rupees will have a variance millions of times larger than one measured in years — so the first principal component will point almost exactly along the rupee axis, and you will have discovered nothing except which column has the biggest numbers.&lt;/p&gt;

&lt;p&gt;Scaling puts every feature on equal footing so the components reflect structure rather than measurement units. The only time you skip it is when your features are already in the same units and their relative scales are meaningful — the rest of the &lt;a href="https://dev.to/library/principal-component-analysis/data-preprocessing-pca"&gt;preprocessing PCA expects&lt;/a&gt; is short but non-negotiable.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where to go from here
&lt;/h2&gt;

&lt;p&gt;PCA is the point where linear algebra stops being abstract. Eigenvectors go from a homework exercise to the thing deciding which ten of your fifty columns survive.&lt;/p&gt;

&lt;p&gt;If any of the pieces felt shaky, fix them before you use this in anger — the covariance structure and the eigen-decomposition do all the work here, and everything else is bookkeeping:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://dev.to/library/principal-component-analysis"&gt;Principal Component Analysis&lt;/a&gt; — the full topic, free to read&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://dev.to/library/correlation-covariance/understanding-covariance"&gt;Understanding covariance&lt;/a&gt; — the matrix PCA actually consumes&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://dev.to/library/correlation-covariance/understanding-correlation"&gt;Understanding correlation&lt;/a&gt; — why orthogonal means uncorrelated&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://dev.to/library/linear-regression/assumptions-linear-regression"&gt;Linear regression assumptions&lt;/a&gt; — the multicollinearity problem PCA dissolves&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>machinelearning</category>
      <category>datascience</category>
      <category>discuss</category>
      <category>learning</category>
    </item>
    <item>
      <title>Correlation Is Pairwise. Multicollinearity Isn't.</title>
      <dc:creator>mohit modi</dc:creator>
      <pubDate>Sat, 22 Aug 2026 04:58:52 +0000</pubDate>
      <link>https://dev.to/mohit_modi_e86a932fb11e61/correlation-is-pairwise-multicollinearity-isnt-2970</link>
      <guid>https://dev.to/mohit_modi_e86a932fb11e61/correlation-is-pairwise-multicollinearity-isnt-2970</guid>
      <description>&lt;p&gt;Almost every regression project I've seen in the last decade starts the same way. Load the data, &lt;code&gt;df.corr()&lt;/code&gt;, plot the heatmap, scan for red squares, drop one variable from every pair above 0.8. Then we move on, feeling like we've handled multicollinearity.&lt;/p&gt;

&lt;p&gt;We haven't. We've handled the version of it that happens to show up in pairs.&lt;/p&gt;

&lt;h2&gt;
  
  
  The three-line example that should end the ritual
&lt;/h2&gt;

&lt;p&gt;Take two independent features and construct a third as their sum:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;numpy&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;pandas&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;pd&lt;/span&gt;

&lt;span class="n"&gt;rng&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;random&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;default_rng&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;42&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;x1&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;rng&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;normal&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;10_000&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;x2&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;rng&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;normal&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;10_000&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;x3&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;x1&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;x2&lt;/span&gt;

&lt;span class="n"&gt;pd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;DataFrame&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;x1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;x1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;x2&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;x2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;x3&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;x3&lt;/span&gt;&lt;span class="p"&gt;}).&lt;/span&gt;&lt;span class="nf"&gt;corr&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;round&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The correlation matrix shows roughly 0.7 between x1 and x3, roughly 0.7 between x2 and x3, and approximately 0 between x1 and x2.&lt;/p&gt;

&lt;p&gt;Every pair passes a 0.8 threshold. Every pair passes a 0.75 threshold. And yet &lt;code&gt;x3&lt;/code&gt; is &lt;em&gt;perfectly&lt;/em&gt; determined by the other two. The design matrix is singular. There is no unique solution for the coefficients — your software will either return garbage or silently regularise its way out of the problem.&lt;/p&gt;

&lt;p&gt;The heatmap saw nothing, because there was nothing to see &lt;em&gt;pairwise&lt;/em&gt;. Correlation is a two-variable statistic. Multicollinearity is a property of the entire design matrix. Those are different questions, and the heatmap only answers the first one.&lt;/p&gt;

&lt;p&gt;This isn't a contrived edge case. It's every ratio you've engineered, every "total" column that's the sum of its parts, every set of channel spends that add up to a budget, every one-hot encoding you forgot to drop a level from. In practice, near-dependence across three or four features is far more common than a single scary pair — and it is exactly what the heatmap is blind to.&lt;/p&gt;

&lt;h2&gt;
  
  
  What VIF actually measures
&lt;/h2&gt;

&lt;p&gt;The Variance Inflation Factor asks the right question. For each feature, regress it on &lt;strong&gt;all&lt;/strong&gt; the other features and take:&lt;/p&gt;

&lt;p&gt;VIF = 1/(1 - R^2)&lt;/p&gt;

&lt;p&gt;That R^2 is multivariate by construction. In the example above, regressing &lt;code&gt;x3&lt;/code&gt; on &lt;code&gt;x1&lt;/code&gt; and &lt;code&gt;x2&lt;/code&gt; gives R^2 = 1, so the VIF is infinite. The heatmap said fine; the VIF says the model is unidentifiable.&lt;/p&gt;

&lt;p&gt;The name is literal, which is the part people miss. A VIF of 10 means the variance of that coefficient estimate is ten times what it would be if the feature were orthogonal to the others — so the standard error is sqrt(10) ~ 3.2 times larger. That's the whole mechanism. Your coefficient isn't wrong on average; it's just so unstable that it's useless. Resample the data and it swings, sometimes across zero.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;statsmodels.stats.outliers_influence&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;variance_inflation_factor&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;statsmodels.api&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;sm&lt;/span&gt;

&lt;span class="n"&gt;X&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;sm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add_constant&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;pd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;DataFrame&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;x1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;x1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;x2&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;x2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;x3&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;x3&lt;/span&gt;&lt;span class="p"&gt;}))&lt;/span&gt;
&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nf"&gt;variance_inflation_factor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;values&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;shape&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;])]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A note on the thresholds, in the spirit of this whole post: VIF &amp;gt; 5 and VIF &amp;gt; 10 are conventions, not findings. They're no more principled than the 0.8 on the heatmap. Use them to rank features by severity, not as a pass/fail gate.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part nobody says out loud: this is an inference problem, not a prediction problem
&lt;/h2&gt;

&lt;p&gt;If you only care about predictions, multicollinearity mostly doesn't matter. Correlated features don't bias the fitted values, don't inflate test error, and don't need to be removed. A gradient boosting model on collinear features will predict perfectly well.&lt;/p&gt;

&lt;p&gt;Multicollinearity only hurts when you need to &lt;strong&gt;read the coefficients&lt;/strong&gt; — when the model is going to answer "how much should we spend on this channel" or "what happens if we raise price by 5%." That's where unstable coefficients become bad decisions.&lt;/p&gt;

&lt;p&gt;So the first question isn't "which features are correlated." It's "am I predicting or explaining?" Half the arguments about multicollinearity are people answering different questions and not realising it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why I sometimes drop a variable that correlates strongly with the target
&lt;/h2&gt;

&lt;p&gt;Building Marketing Mix Models, I started the way everyone does: rank variables by correlation with sales, keep the strong ones. It worked, in the sense that the model fit beautifully. It also quietly pushed every media driver out of the model.&lt;/p&gt;

&lt;p&gt;The pattern was consistent. Distribution, seasonality, holiday flags — these sat at 0.8 and above. The media variables I actually needed to report on sat between 0.4 and 0.6. Ranked on correlation, the media never stood a chance.&lt;/p&gt;

&lt;p&gt;The problem is that in any real market, promotions, media and seasonal demand all move together. December sales are up. So is TV spend, so is display, so is the holiday flag. Every one of those variables is competing to explain the same peak, and the one that wins the competition is the one most tightly coupled to it — which is almost always the variable you can't control.&lt;/p&gt;

&lt;p&gt;So I dropped variables with high correlation to the target, deliberately, and accepted a worse fit. Because a marketing mix model whose coefficients say "sales are seasonal" is not a model. Nobody can act on it. The entire point of the exercise was the media coefficients, and the high-correlation variables were crowding out the only part anyone was going to read.&lt;/p&gt;

&lt;p&gt;This is the inference-versus-prediction distinction from earlier, arriving with a budget attached. If I were forecasting sales, keeping seasonality and dropping media would be the right call — better fit, better forecast, done. But the model existed to answer "what should we spend on this channel next quarter," and that question is answered by coefficients, not by fit.&lt;/p&gt;

&lt;p&gt;What's left is a judgement call that no diagnostic makes for you: balancing model fit against business expectation and explainability. The correlation matrix can't help here. It ranks variables by their relationship to the target and has nothing at all to say about which relationships you need to be reportable.&lt;/p&gt;

&lt;p&gt;The general principle I've settled on: &lt;strong&gt;feature selection is a question about what the model is for, not just what's in the data.&lt;/strong&gt; A variable that improves fit but destabilises the coefficient you're going to act on is a bad trade, regardless of what the heatmap says.&lt;/p&gt;

&lt;h2&gt;
  
  
  The mirror-image mistake: dropping features with low target correlation
&lt;/h2&gt;

&lt;p&gt;While we're dismantling this ritual, the other half of it deserves the same treatment. People keep the features that correlate strongly with the target and cut the ones that don't.&lt;/p&gt;

&lt;p&gt;But a feature can have near-zero correlation with the target and still be one of the most important variables in the model. These are suppressor variables: they don't predict the target directly, they explain away noise in &lt;em&gt;another&lt;/em&gt; predictor, sharpening its signal. Cut it on a univariate screen and the model gets worse.&lt;/p&gt;

&lt;p&gt;The pattern is the same error in both directions — using a pairwise statistic to make a multivariate decision.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where eigenvalues come in
&lt;/h2&gt;

&lt;p&gt;If you want the honest, single-number version of "is my design matrix in trouble," it's the condition number: the ratio of the largest to the smallest singular value of your scaled feature matrix. Values above ~30 are the usual flag.&lt;/p&gt;

&lt;p&gt;That's not a coincidence — it's the same underlying idea. Near-dependence among your features means the matrix X^T X has eigenvalues close to zero, and inverting something with a near-zero eigenvalue is what blows the coefficient variances up in the first place. VIF and the condition number are two views of one geometric fact: your features don't span the space you think they do.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd actually do instead
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Decide first whether you're predicting or explaining. If predicting, most of this is moot.&lt;/li&gt;
&lt;li&gt;Skip the heatmap as a decision tool. Keep it for exploration — it's genuinely useful for spotting data errors and leakage.&lt;/li&gt;
&lt;li&gt;Compute VIF on the full feature set. Rank, don't threshold.&lt;/li&gt;
&lt;li&gt;Check the condition number as a whole-matrix sanity check.&lt;/li&gt;
&lt;li&gt;For anything with a high VIF, ask &lt;em&gt;why&lt;/em&gt; it's dependent. Engineered ratio? Sum of components? Dummy trap? The fix is usually structural, not deletion.&lt;/li&gt;
&lt;li&gt;Only then decide what to drop — and decide it against what the model is for.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;None of this is more work than making the heatmap. It's just less satisfying, because there's no red square to point at.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Want the underlying concepts, properly? Explore &lt;a href="https://www.bitelrn.com/library/correlation-covariance" rel="noopener noreferrer"&gt;Correlation &amp;amp; Covariance&lt;/a&gt;, &lt;a href="https://www.bitelrn.com/library/linear-regression" rel="noopener noreferrer"&gt;Linear Regression&lt;/a&gt;, and &lt;a href="https://www.bitelrn.com/library/linear-algebra-eigenvalues" rel="noopener noreferrer"&gt;Eigenvalues&lt;/a&gt; in the Bitelrn Open Library — no sign-up.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;--&lt;/p&gt;

</description>
      <category>datascience</category>
      <category>statistics</category>
      <category>python</category>
      <category>machinelearning</category>
    </item>
  </channel>
</rss>
