<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: RV42</title>
    <description>The latest articles on DEV Community by RV42 (@richael_42).</description>
    <link>https://dev.to/richael_42</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4096610%2F5443aff6-9426-4869-a468-4e642143220e.jpg</url>
      <title>DEV Community: RV42</title>
      <link>https://dev.to/richael_42</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/richael_42"/>
    <language>en</language>
    <item>
      <title>9 Triton kernels in 9 weeks: what PyTorch taught me by winning</title>
      <dc:creator>RV42</dc:creator>
      <pubDate>Thu, 08 Oct 2026 17:32:02 +0000</pubDate>
      <link>https://dev.to/richael_42/9-triton-kernels-in-9-weeks-what-pytorch-taught-me-by-winning-3jim</link>
      <guid>https://dev.to/richael_42/9-triton-kernels-in-9-weeks-what-pytorch-taught-me-by-winning-3jim</guid>
      <description>&lt;p&gt;I'm Richael (dh8116), a Year 11 student in Auckland. Since late July I've written one Triton kernel a week and benchmarked each against PyTorch, publishing the numbers even when PyTorch won, which was often. Here's what nine weeks taught me, kernel by kernel.&lt;/p&gt;

&lt;p&gt;The ones that won&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fused softmax (week 2).&lt;/strong&gt; Naive softmax makes three passes over memory. Fusing them into one took it from ~55 GB/s to ~230 GB/s, a 4x jump, and it held a tighter line than PyTorch's own kernel. Lesson one: on a memory-bound op, the passes you don't make are the speedup.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fused cross-entropy (week 7).&lt;/strong&gt; Forward and backward in one pass, with the gradient written straight over the logits buffer, so PyTorch's second [N, V] tensor never exists. On a T4 at vocab 131,072 in fp16: 15.90 ms vs 24.03 ms (1.51x), and 1.67x less peak memory.&lt;/p&gt;

&lt;p&gt;The first version of this benchmark said 2.5x. That was wrong: I was comparing my fused kernel against an unfused PyTorch baseline. Comparing fused against fused is how I caught myself handicapping the baseline.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;RoPE (week 6).&lt;/strong&gt; Not the speedup, the elegance: backward reuses the exact same kernel as forward. You just flip the sign on sin.&lt;/p&gt;

&lt;h2&gt;
  
  
  The ones that lost, and why that was the useful part
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Flash attention (week 3).&lt;/strong&gt; Correct, but about 25x slower than PyTorch on my T4.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Matmul (week 4).&lt;/strong&gt; The output matched, but throughput sat flat around 1 TFLOPS while cuBLAS climbed to 38. Reading the compiled PTX showed zero tensor-core instructions: the compiler never emitted them for my layout.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fused linear + cross-entropy (week 8).&lt;/strong&gt; Chunking the lm_head projection into the loss means the [batch, vocab] logits never exist at all. That's about 10x less activation memory, and it was slower at every single size. Sometimes memory is the thing you're buying, and you pay for it in time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fused SwiGLU MLP (week 9).&lt;/strong&gt; Level with torch.compile on time. The win is memory: it holds two big activation tensors between forward and backward, where eager holds four. A compiler will fuse for you, but it won't throw an activation away.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd tell anyone starting
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Benchmark fused against fused.&lt;/strong&gt; An unfused baseline makes every kernel look like a win.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Read the PTX when the numbers are flat.&lt;/strong&gt; The matmul answer was sitting in the compiled output the whole time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Memory is a result too.&lt;/strong&gt; Two of my best kernels "lost" on time and won on memory, and that trade decides batch size.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Publish the losses.&lt;/strong&gt; They taught me more, and they're the posts people actually argue with.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Code: &lt;a href="https://github.com/dh8116/triton-kernels" rel="noopener noreferrer"&gt;https://github.com/dh8116/triton-kernels&lt;/a&gt;&lt;br&gt;
Every write-up, with methods: &lt;a href="https://dh8116.github.io/blog" rel="noopener noreferrer"&gt;https://dh8116.github.io/blog&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Next on my list is revisiting flash attention on Ampere, because the T4 result deserves a second look.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>coding</category>
    </item>
    <item>
      <title>I fine-tuned Qwen3-14B for my AI companion app for $1.30. The hard part wasn't the model</title>
      <dc:creator>RV42</dc:creator>
      <pubDate>Sat, 03 Oct 2026 09:29:26 +0000</pubDate>
      <link>https://dev.to/richael_42/i-fine-tuned-qwen3-14b-for-my-ai-companion-app-for-130-the-hard-part-wasnt-the-model-56f5</link>
      <guid>https://dev.to/richael_42/i-fine-tuned-qwen3-14b-for-my-ai-companion-app-for-130-the-hard-part-wasnt-the-model-56f5</guid>
      <description>&lt;p&gt;I built Soulor (soulor.app) independently as a 16y/o high school student. It helps reading social situations with multiple perspectives, and helps users rehearse hard conversations by simulating others with analyses and suggestions, there are also companions that knows the full context if users choose to share that can chat with them naturally. It can be both serious and entertaining, by switching different world views, and it fully respects users' privacy by keeping all texts encrypted. This post is about the model underneath: a LoRA fine-tune of Qwen3-14B, served with vLLM on Modal, and the four things that mattered more than I expected.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. The first adapter failed on format, not voice
&lt;/h2&gt;

&lt;p&gt;The app doesn't want free text from the model. Every reply comes back in a JSON envelope the frontend parses: the message, plus fields the app acts on. My first adapter learned the persona perfectly and couldn't produce that envelope at all. It sounded right, and the app couldn't use any of it.&lt;/p&gt;

&lt;p&gt;The fix was the training data, not the hyperparameters. I rewrote the corpus so every example was written &lt;em&gt;inside&lt;/em&gt; the envelope, the same shape production asks for. The retrained adapter scores &lt;strong&gt;15/15 on companion chat and 15/15 on feed comments&lt;/strong&gt; at production prompt size.&lt;/p&gt;

&lt;p&gt;Training cost about &lt;strong&gt;$1.30&lt;/strong&gt; and took half an hour.&lt;/p&gt;

&lt;p&gt;Lesson: if your app parses the model's output, the output format is part of the behaviour you're fine-tuning, so train on it.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Latency was memory bandwidth, not a vLLM flag
&lt;/h2&gt;

&lt;p&gt;Replies took 7.9s. I spent a while on serving flags before measuring properly. Decoding a 14B model is bound by memory bandwidth, so each token means reading the weights again, and no flag changes how fast the card can read them. Moving to an H100 took it to &lt;strong&gt;2.0s&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The hosted Qwen fallback had a different problem. With thinking enabled it took 14.6s, and the reasoning was eating the token budget and truncating replies. Turning &lt;code&gt;enable_thinking&lt;/code&gt; off gave &lt;strong&gt;5.7s&lt;/strong&gt; and complete answers.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Scale-to-zero needs a gate in front of it
&lt;/h2&gt;

&lt;p&gt;The Modal endpoint scales to zero, which is what keeps it affordable. A cold start takes &lt;strong&gt;108s&lt;/strong&gt;, though, and nobody waits 108s for a chat reply.&lt;/p&gt;

&lt;p&gt;So the gateway asks one question before every request: is a container warm? If yes, the fine-tune goes first in the provider chain. If not, it skips straight to the hosted models. The probe doubles as the warm-up, so the first message of a conversation is answered by a hosted model and the rest can go to the fine-tune. The voice can shift once, early, but I don't pay for an idle GPU between conversations.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Two bug shapes caused four different-looking outages
&lt;/h2&gt;

&lt;p&gt;Signups leaving unreachable accounts, cleared chats coming back, companions forgetting the last few turns, and chat hanging for a minute all looked unrelated. They came from just two patterns:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A write fired without &lt;code&gt;await&lt;/code&gt;, with its error swallowed.&lt;/strong&gt; On serverless the function often returns before the write happens, so the work silently never runs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A network call with no timeout.&lt;/strong&gt; A fallback chain can fall through an error, but it can't fall through a hang.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Both are now fixed where they can't be forgotten: writes are awaited, and every provider call has a timeout.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Simulation Mode does with all this
&lt;/h2&gt;

&lt;p&gt;The same engine runs the companion and the panel. In Simulation Mode, one situation goes to five analysts: an Optimist, a Cynic, a Mentor, a Status Observer and a Gossiper. Because it's the same app, the panel already knows the context you gave your companion. You can rehearse a hard conversation before having it for real, then switch back mid-conversation without starting over.&lt;/p&gt;

&lt;p&gt;Try Soulor: &lt;strong&gt;&lt;a href="https://soulor.app/" rel="noopener noreferrer"&gt;https://soulor.app/&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I'm happy to go deeper on any of this in the comments: the envelope corpus, the warm gate, or the eval harness. More of what I build, including a weekly Triton kernel series benchmarked against PyTorch, is at &lt;a href="https://dh8116.github.io/" rel="noopener noreferrer"&gt;https://dh8116.github.io/&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>llm</category>
      <category>showdev</category>
    </item>
  </channel>
</rss>
