<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Sevan Safarian</title>
    <description>The latest articles on DEV Community by Sevan Safarian (@sevan_safarian_555dc84374).</description>
    <link>https://dev.to/sevan_safarian_555dc84374</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F2134386%2F551695dc-7c92-4253-9353-79f51bead58a.png</url>
      <title>DEV Community: Sevan Safarian</title>
      <link>https://dev.to/sevan_safarian_555dc84374</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/sevan_safarian_555dc84374"/>
    <language>en</language>
    <item>
      <title>What a vision model can (and can't) tell you from a selfie</title>
      <dc:creator>Sevan Safarian</dc:creator>
      <pubDate>Fri, 09 Oct 2026 13:17:44 +0000</pubDate>
      <link>https://dev.to/sevan_safarian_555dc84374/what-a-vision-model-can-and-cant-tell-you-from-a-selfie-3noj</link>
      <guid>https://dev.to/sevan_safarian_555dc84374/what-a-vision-model-can-and-cant-tell-you-from-a-selfie-3noj</guid>
      <description>&lt;p&gt;A few weeks ago I added a seasonal colour analysis to &lt;a href="https://fring.sevan.zone/en" rel="noopener noreferrer"&gt;Fring&lt;/a&gt;, the free wardrobe app I build on my own: you upload a selfie, you get your colour season, your undertone and six colours that suit you. It looks like a small feature. It taught me four things about putting a vision model in front of users.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Let the model judge, never let it invent
&lt;/h2&gt;

&lt;p&gt;The tempting design is to ask the model for everything: season, undertone, and a palette. But a generated palette would change from one run to the next — plausible colours every time, and nothing the rest of the app could rely on.&lt;/p&gt;

&lt;p&gt;So the model answers a much smaller question:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Judge the face only: skin undertone, hair and eye colour, contrast between them.
Classify into one of the four colour seasons:
spring = warm and light/bright, summer = cool and soft/light,
autumn = warm and deep/muted, winter = cool and deep/high contrast.
If no human face is clearly visible, answer null for both fields.
Respond with ONLY a JSON object of the shape
{"season":"spring|summer|autumn|winter","undertone":"warm|cool|neutral"}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The palette is a fixed table on the server, six colours per season. The model classifies; the code decides what a classification means. Anything outside the allowed values is treated as "no readable face" and the user is asked for a better photo, instead of getting a confident wrong answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Four seasons, not twelve
&lt;/h2&gt;

&lt;p&gt;Colour analysts often work with twelve seasons (soft autumn, cool summer, deep winter…). A phone selfie under random indoor light cannot carry that level of detail reliably: white balance alone moves a face from "warm" to "neutral". So the app sticks to four seasons and three undertones, and says plainly that it's an estimate to check against a test with real fabric. A feature that is honest about its precision is more useful than one that pretends.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Features die silently
&lt;/h2&gt;

&lt;p&gt;For a while the analysis called an old vision microservice. When the AI moved to a single GPU at home (a used Tesla P40 running Qwen3-VL-8B through llama.cpp's server), that service stopped running — and the colour analysis broke with no visible error, just "service unavailable" for every user. I only found it by walking through every feature by hand. Since then, every model call goes through one client module, so a backend change can't quietly orphan a feature again.&lt;/p&gt;

&lt;p&gt;The other side of running the model at home: the selfie is processed in memory and never stored, and it never leaves the machine.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. A result is only worth something if it's used
&lt;/h2&gt;

&lt;p&gt;A palette you look at once is a curiosity. In Fring the saved colours feed the outfit suggestions (outfits in your colours score higher) and the shopping assistant flags a colour outside your palette before you buy. That's where the feature earns its place: not in the result screen, but in the decisions it nudges afterwards.&lt;/p&gt;

&lt;p&gt;If you're curious about the colour theory side, I wrote a longer guide: &lt;a href="https://fring.sevan.zone/en/guides/seasonal-color-analysis" rel="noopener noreferrer"&gt;seasonal color analysis, the 4 and 12 seasons and a free at-home test&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Written with AI assistance.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>computervision</category>
      <category>webdev</category>
      <category>showdev</category>
    </item>
    <item>
      <title>On a Tesla P40, loading a diffusion model in fp16 made it 14x slower than fp32</title>
      <dc:creator>Sevan Safarian</dc:creator>
      <pubDate>Fri, 09 Oct 2026 08:11:08 +0000</pubDate>
      <link>https://dev.to/sevan_safarian_555dc84374/on-a-tesla-p40-loading-a-diffusion-model-in-fp16-made-it-14x-slower-than-fp32-3fa7</link>
      <guid>https://dev.to/sevan_safarian_555dc84374/on-a-tesla-p40-loading-a-diffusion-model-in-fp16-made-it-14x-slower-than-fp32-3fa7</guid>
      <description>&lt;p&gt;I run the AI for a small side project, &lt;a href="https://fring.sevan.zone/en?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=p40" rel="noopener noreferrer"&gt;Fring&lt;/a&gt;, a free wardrobe app, on a second-hand Tesla P40 in my homelab. The virtual try-on uses Leffa, a diffusion model. For a month every render took about 21 minutes, and I assumed that was just what a 2016 card could do.&lt;/p&gt;

&lt;p&gt;It wasn't. The model was loaded in fp16, because "use half precision to save VRAM" is the default advice everywhere. On the P40 (GP102, compute capability 6.1), fp16 runs at roughly 1/64 of fp32 throughput. Same inputs, same seed:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;s / diffusion step&lt;/th&gt;
&lt;th&gt;VRAM&lt;/th&gt;
&lt;th&gt;20-step render&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;fp16&lt;/td&gt;
&lt;td&gt;63.8&lt;/td&gt;
&lt;td&gt;6.6 GB&lt;/td&gt;
&lt;td&gt;1241 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;fp32&lt;/td&gt;
&lt;td&gt;4.7&lt;/td&gt;
&lt;td&gt;9.8 GB&lt;/td&gt;
&lt;td&gt;85 s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The images were practically identical (mean difference 0.77/255). The card never needed the VRAM savings: the model used 6.4 GB out of 23.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two other traps from the same month
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;The Dockerfile installed a CPU-only torch wheel, inherited from an older 4 GB card. &lt;code&gt;nvidia-smi&lt;/code&gt; saw the GPU while &lt;code&gt;torch.cuda.is_available()&lt;/code&gt; returned &lt;code&gt;False&lt;/code&gt;, and renders took hours. Check inside the container, not on the host.&lt;/li&gt;
&lt;li&gt;An unpinned &lt;code&gt;mediapipe&lt;/code&gt; upgrade removed &lt;code&gt;mediapipe.solutions&lt;/code&gt;. The health check only tested the import, so it kept reporting pose estimation as healthy while the service silently fell back to a centred paste.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What I changed
&lt;/h2&gt;

&lt;p&gt;The service now picks fp16 only when the GPU actually accelerates it (Volta and newer, or P100), or when fp32 would not fit in VRAM. An environment variable can force either.&lt;/p&gt;

&lt;p&gt;The general lesson for old datacenter cards: measure before applying the usual advice.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://fring.sevan.zone/en/guides/tesla-p40-fp16-vs-fp32" rel="noopener noreferrer"&gt;Fring's guides&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>gpu</category>
      <category>machinelearning</category>
      <category>docker</category>
      <category>homelab</category>
    </item>
  </channel>
</rss>
