<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: marcelotaparelli</title>
    <description>The latest articles on DEV Community by marcelotaparelli (@marcelotaparelli).</description>
    <link>https://dev.to/marcelotaparelli</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1388551%2F8e0f7cab-1676-4dc0-b3cb-f7074e3ce0b7.png</url>
      <title>DEV Community: marcelotaparelli</title>
      <link>https://dev.to/marcelotaparelli</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/marcelotaparelli"/>
    <language>en</language>
    <item>
      <title>Jev 1.13 in Practice: I Tested a Decision Model Against Rules and a Local LLM</title>
      <dc:creator>marcelotaparelli</dc:creator>
      <pubDate>Wed, 23 Sep 2026 02:33:26 +0000</pubDate>
      <link>https://dev.to/marcelotaparelli/jev-113-in-practice-i-tested-a-decision-model-against-rules-and-a-local-llm-3p1c</link>
      <guid>https://dev.to/marcelotaparelli/jev-113-in-practice-i-tested-a-decision-model-against-rules-and-a-local-llm-3p1c</guid>
      <description>&lt;p&gt;Jev 1.13 got every category classification right in my held-out set: &lt;strong&gt;1.0000 category accuracy&lt;/strong&gt;. The same experiment measured 0.9857 priority accuracy and 0.9571 risk accuracy. Those are results on 70 synthetic tickets; they do not prove that Jev is superior in production.&lt;/p&gt;

&lt;p&gt;I decided to test Jev because I already had two reference points in &lt;code&gt;ops-triage-ai&lt;/code&gt;: a deterministic rules baseline and a classifier using local Ollama. The question was whether a third paradigm — a typed probabilistic decision model — would produce different results on the same taxonomy.&lt;/p&gt;

&lt;h2&gt;
  
  
  What changes compared with a chat LLM
&lt;/h2&gt;

&lt;p&gt;I did not use Jev as a free-form conversation and then interpret prose afterward. A separate adapter sent decisions through the OpenRouter Decisions API using &lt;code&gt;typesafe/jev-1.13&lt;/code&gt;, with typed responses and probability distributions. All 73 API responses during the experiment resolved to &lt;code&gt;typesafe/jev-1.13-20260917&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The adapter implements the &lt;code&gt;TriageClassifier&lt;/code&gt; interface, but remained separate from the Ollama classifier and &lt;code&gt;HybridPolicy&lt;/code&gt;. I used Bun native fetch and added no dependencies. Jev was not integrated into the application flow and did not replace the local model.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I froze the experiment
&lt;/h2&gt;

&lt;p&gt;I used the same taxonomy and the same 70 synthetic labels from the historical held-out benchmark. Development calls happened before the freeze. In commit &lt;a href="https://github.com/marcelotaparelli/ops-triage-ai/tree/4c41e0b" rel="noopener noreferrer"&gt;&lt;code&gt;4c41e0b&lt;/code&gt;&lt;/a&gt;, I froze the configuration and evaluation, then ran the held-out set once for the official result.&lt;/p&gt;

&lt;p&gt;That separation matters: the final set was not used to tune the model and then measure that tuning on the same set. It is still a small synthetic sample, but the protocol makes clear what was frozen and which run produced the numbers.&lt;/p&gt;

&lt;h2&gt;
  
  
  Results
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Deterministic&lt;/th&gt;
&lt;th&gt;Ollama&lt;/th&gt;
&lt;th&gt;Jev 1.13&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Category accuracy&lt;/td&gt;
&lt;td&gt;0.8286&lt;/td&gt;
&lt;td&gt;0.9571&lt;/td&gt;
&lt;td&gt;1.0000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Category macro-F1&lt;/td&gt;
&lt;td&gt;0.8512&lt;/td&gt;
&lt;td&gt;0.9550&lt;/td&gt;
&lt;td&gt;1.0000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Priority accuracy&lt;/td&gt;
&lt;td&gt;0.9000&lt;/td&gt;
&lt;td&gt;0.9143&lt;/td&gt;
&lt;td&gt;0.9857&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Risk accuracy&lt;/td&gt;
&lt;td&gt;0.9571&lt;/td&gt;
&lt;td&gt;0.9143&lt;/td&gt;
&lt;td&gt;0.9571&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HIGH/CRITICAL priority recall&lt;/td&gt;
&lt;td&gt;0.7857&lt;/td&gt;
&lt;td&gt;1.0000&lt;/td&gt;
&lt;td&gt;1.0000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HIGH risk recall&lt;/td&gt;
&lt;td&gt;0.5714&lt;/td&gt;
&lt;td&gt;0.7143&lt;/td&gt;
&lt;td&gt;0.8571&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Jev got the full category, priority, and risk tuple right in &lt;strong&gt;66 of 70 cases (0.9429)&lt;/strong&gt;. The historical benchmark did not emit standalone exact-tuple accuracy for the deterministic classifier or Ollama. I therefore do not compare 66/70 with per-field metrics or the hybrid path's exact tuple: they measure different things.&lt;/p&gt;

&lt;h2&gt;
  
  
  Measured cost and latency
&lt;/h2&gt;

&lt;p&gt;Across Jev's 70 standalone calls, mean latency was &lt;strong&gt;569.4 ms&lt;/strong&gt;, with p50 at &lt;strong&gt;545.5 ms&lt;/strong&gt;, p95 at &lt;strong&gt;712.2 ms&lt;/strong&gt;, and a maximum of &lt;strong&gt;1,142.2 ms&lt;/strong&gt;. The report recorded &lt;strong&gt;90,229 input tokens&lt;/strong&gt;, a total reported cost of &lt;strong&gt;US$ 0.003789618&lt;/strong&gt;, and zero API or schema failures.&lt;/p&gt;

&lt;p&gt;That cost is the reported amount for this OpenRouter run. The historical Ollama benchmark did not emit cost or standalone latency. Its published 6–7 seconds describe the full hybrid path; that is not an equivalent comparison with a standalone Jev call.&lt;/p&gt;

&lt;h2&gt;
  
  
  Confidence is not semantic correctness
&lt;/h2&gt;

&lt;p&gt;Jev preserved probability distributions, continuous confidence, and top-1/top-2 margins. That gives more information for evaluating a decision, but it does not guarantee the decision is right: one incorrect HIGH risk classification had &lt;strong&gt;0.86 confidence&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;For this observed dataset, exploratory calibration analysis reported Brier scores of 0.0005 for category, 0.0386 for priority, and 0.0800 for risk; top-1 ECE was 0.0054, 0.0329, and 0.0226, respectively. These are statistics over 70 synthetic tickets, with few examples in each range. They do not show that probabilities are calibrated for real tickets.&lt;/p&gt;

&lt;h2&gt;
  
  
  Selective automation: an exploratory signal
&lt;/h2&gt;

&lt;p&gt;When requiring all three confidence values to reach a threshold of 0.80, &lt;strong&gt;54 of 70 tickets&lt;/strong&gt; remained covered, and all covered cases had the observed tuple correct. At 0.90, &lt;strong&gt;51 of 70&lt;/strong&gt; remained, also with 100% observed exact accuracy.&lt;/p&gt;

&lt;p&gt;This does not define a production threshold. The denominator is small, the examples are synthetic, and the result does not establish calibration or performance beyond this set. It is a hypothesis to test with representative data and separately defined review criteria.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this benchmark does not prove
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;It does not estimate performance in production or on real tickets.&lt;/li&gt;
&lt;li&gt;One official run does not measure run-to-run variance.&lt;/li&gt;
&lt;li&gt;70 synthetic examples do not represent the full domain distribution.&lt;/li&gt;
&lt;li&gt;The result does not establish that Jev is generally superior to Ollama or rules.&lt;/li&gt;
&lt;li&gt;Exploratory calibration metrics do not validate confidence for real decisions.&lt;/li&gt;
&lt;li&gt;I did not compare Ollama standalone cost or latency because the historical benchmark did not emit them.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What I would do architecturally
&lt;/h2&gt;

&lt;p&gt;I would keep the current decisions: the deterministic baseline and Ollama remain in the existing evaluation and hybrid policy; Jev stays benchmark/evaluation only. Before any integration, I would expand and diversify the labeled data, repeat the frozen evaluation to measure variance, and define how human review and severity errors factor into acceptance criteria.&lt;/p&gt;

&lt;p&gt;The useful result is not “Jev won.” A third paradigm produced different measurable signals in a reproducible run with explicit limits — material for guiding the next evaluation, not for justifying a production replacement.&lt;/p&gt;

&lt;h2&gt;
  
  
  Evidence
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://marcelotaparelli.com.br/en/projects/ops-triage-ai/" rel="noopener noreferrer"&gt;Ops Triage AI case&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Repository: &lt;a href="https://github.com/marcelotaparelli/ops-triage-ai" rel="noopener noreferrer"&gt;ops-triage-ai on GitHub&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Jev 1.13 held-out report at the freeze commit: &lt;a href="https://github.com/marcelotaparelli/ops-triage-ai/blob/4c41e0b/docs/evaluation/jev-1.13-held-out-4c41e0b.md" rel="noopener noreferrer"&gt;jev-1.13-held-out-4c41e0b.md&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Run JSON artifact: &lt;a href="https://github.com/marcelotaparelli/ops-triage-ai/blob/4c41e0b/artifacts/jev-1.13-held-out-4c41e0b.json" rel="noopener noreferrer"&gt;jev-1.13-held-out-4c41e0b.json&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>machinelearning</category>
      <category>programming</category>
    </item>
    <item>
      <title>A Valid JWT Does Not Mean Authorized Access</title>
      <dc:creator>marcelotaparelli</dc:creator>
      <pubDate>Sat, 19 Sep 2026 21:12:04 +0000</pubDate>
      <link>https://dev.to/marcelotaparelli/a-valid-jwt-does-not-mean-authorized-access-1m38</link>
      <guid>https://dev.to/marcelotaparelli/a-valid-jwt-does-not-mean-authorized-access-1m38</guid>
      <description>&lt;p&gt;A security review of Salus brought me back to a question that remains after authentication:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can this user access this specific resource?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Consider this scenario:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;User A creates patient X.&lt;/li&gt;
&lt;li&gt;User B signs in and receives a valid JWT.&lt;/li&gt;
&lt;li&gt;B calls &lt;code&gt;GET /patients/X&lt;/code&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If the application looks up the patient by &lt;code&gt;patientId&lt;/code&gt; alone, B may receive a patient belonging to A. B is authenticated, but is not authorized to access that object. This is &lt;strong&gt;Broken Object Level Authorization (BOLA)&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Authorization must reach the query
&lt;/h2&gt;

&lt;p&gt;A valid JWT establishes who made the request. The access rule still needs to use that identity to scope the resource lookup:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;patientId + currentUser.id → authorized resource
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Instead of:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;findById(patientId)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;the operation could conceptually use:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;findByIdAndOwner(patientId, currentUser.id)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The method name is incidental. The access boundary belongs in the rule and the lookup, including reads, updates, and deletes.&lt;/p&gt;

&lt;h2&gt;
  
  
  What should the API return for someone else's patient?
&lt;/h2&gt;

&lt;p&gt;A &lt;code&gt;403 Forbidden&lt;/code&gt; response can tell B that patient X exists. When even that fact is sensitive, &lt;code&gt;404 Not Found&lt;/code&gt; may be preferable: X is outside the set of resources B is allowed to see.&lt;/p&gt;

&lt;p&gt;The choice depends on the API's policy and should be applied consistently. It makes HTTP responses less useful for probing which IDs exist.&lt;/p&gt;

&lt;h2&gt;
  
  
  Security as testable behavior
&lt;/h2&gt;

&lt;p&gt;The path does not end at &lt;code&gt;login → JWT → middleware&lt;/code&gt;. It continues through the operation on each object:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;identity → authentication → authorization → ownership → resource access
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is the property that became a Salus security regression test:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User A creates patient X.
User B tries to access X.
Expected result: 404.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The hardening work reproduced this scenario as a failing test (RED): User A creates X, User B receives a valid JWT and tries to access it. Ownership was then introduced, operations were scoped by &lt;code&gt;ownerId&lt;/code&gt;, and cross-user access began returning 404 in this context. The regression test now passes (GREEN), alongside JWT verification on patient routes.&lt;/p&gt;

&lt;p&gt;TDD and security regression tests protect this behavior as the system evolves.&lt;/p&gt;

&lt;p&gt;Security should be verifiable behavior, not just configuration.&lt;/p&gt;

&lt;p&gt;Project: &lt;a href="https://github.com/marcelotaparelli/salus" rel="noopener noreferrer"&gt;Salus on GitHub&lt;/a&gt;&lt;/p&gt;

</description>
      <category>security</category>
      <category>api</category>
      <category>backend</category>
      <category>typescript</category>
    </item>
    <item>
      <title>Evals: I Stopped Asking Whether the LLM “Looks Good” and Started Measuring</title>
      <dc:creator>marcelotaparelli</dc:creator>
      <pubDate>Wed, 16 Sep 2026 16:03:28 +0000</pubDate>
      <link>https://dev.to/marcelotaparelli/evals-i-stopped-asking-whether-the-llm-looks-good-and-started-measuring-536b</link>
      <guid>https://dev.to/marcelotaparelli/evals-i-stopped-asking-whether-the-llm-looks-good-and-started-measuring-536b</guid>
      <description>&lt;p&gt;When I started working with LLMs, one of the hardest questions looked deceptively simple:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;how do I know the model is actually getting better?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Running a few examples by hand and thinking "that answer looks good" works at first.&lt;/p&gt;

&lt;p&gt;But it does not scale.&lt;/p&gt;

&lt;p&gt;And, more importantly, it produces no evidence.&lt;/p&gt;

&lt;p&gt;That is when I started to understand the role of &lt;strong&gt;evals&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is an eval?
&lt;/h2&gt;

&lt;p&gt;An eval is a structured way of testing the behavior of an AI system.&lt;/p&gt;

&lt;p&gt;The idea is fairly simple:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;input
+
produced answer
+
expected answer
+
metric
=
evaluation
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Instead of subjectively asking whether an answer turned out well, you define up front what you expect and measure the distance between actual behavior and expected behavior.&lt;/p&gt;

&lt;p&gt;An article by Martin Fowler on &lt;a href="https://martinfowler.com/articles/gen-ai-patterns/" rel="noopener noreferrer"&gt;GenAI patterns&lt;/a&gt; describes this kind of mechanism as &lt;strong&gt;scoring and judging&lt;/strong&gt;: the model's output goes through a scorer that produces metrics or feedback about the result.&lt;/p&gt;

&lt;p&gt;That is exactly the principle I applied in my operational triage project.&lt;/p&gt;

&lt;h2&gt;
  
  
  Before the LLM, I wrote code
&lt;/h2&gt;

&lt;p&gt;In &lt;code&gt;ops-triage-ai&lt;/code&gt;, the system receives tickets and has to determine things like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;category;&lt;/li&gt;
&lt;li&gt;priority;&lt;/li&gt;
&lt;li&gt;risk;&lt;/li&gt;
&lt;li&gt;suggested team.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Before putting an LLM on the problem, I implemented a deterministic classifier with hand-written rules in TypeScript.&lt;/p&gt;

&lt;p&gt;Simplified:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ticket
   ↓
deterministic rules
   ↓
classification
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That created something extremely valuable:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;a baseline.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I now had a concrete implementation to compare any AI-based solution against.&lt;/p&gt;

&lt;h2&gt;
  
  
  Then I wrote down expected answers
&lt;/h2&gt;

&lt;p&gt;I set aside a collection of tickets and defined in advance what the correct classification for each one should be.&lt;/p&gt;

&lt;p&gt;Conceptually something like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"ticket"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"title"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Production API unavailable"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"description"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Users cannot access the service"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"expected"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"category"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"INCIDENT"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"priority"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"CRITICAL"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"risk"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"HIGH"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The classifier receives the ticket.&lt;/p&gt;

&lt;p&gt;Its output is compared against &lt;code&gt;expected&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;And then we compute metrics.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;               ┌──────────────────┐
ticket ───────►│    classifier    │
               └────────┬─────────┘
                        │
                        ▼
                 predicted output
                        │
expected output ────────┤
                        ▼
                    scorer
                        │
                        ▼
                    metrics
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That scorer can be plain code.&lt;/p&gt;

&lt;p&gt;It does not need to be another LLM.&lt;/p&gt;

&lt;h2&gt;
  
  
  My code became part of the experiment
&lt;/h2&gt;

&lt;p&gt;That is where I found the most interesting idea.&lt;/p&gt;

&lt;p&gt;The deterministic code I would normally write to solve the problem also became an &lt;strong&gt;experimental reference&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;I could run:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;dataset
   ├── deterministic classifier
   └── LLM classifier
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and compare both on exactly the same examples.&lt;/p&gt;

&lt;p&gt;On the project's final held-out set of 70 synthetic tickets, for instance:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Category accuracy

Deterministic: 82.9%
LLM:           95.7%
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For &lt;code&gt;HIGH/CRITICAL&lt;/code&gt; priority recall:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Deterministic: 78.6%
LLM:           100%
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;But something even more important happened.&lt;/p&gt;

&lt;p&gt;The model &lt;strong&gt;did not win everywhere&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;On overall risk classification:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Deterministic: 95.7%
LLM:            91.4%
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Without an eval, it would have been easy to look at a few good LLM answers and conclude the model was simply better.&lt;/p&gt;

&lt;p&gt;The metrics told a more interesting story.&lt;/p&gt;

&lt;h2&gt;
  
  
  And then the eval started shaping the architecture
&lt;/h2&gt;

&lt;p&gt;At that point the question stopped being:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;How do I make the LLM replace my rules?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;and became:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Where does each approach work best?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That pushed the project toward a hybrid architecture.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                 ticket
                   │
          ┌────────┴────────┐
          ▼                 ▼
 deterministic           LLM
 classifier           classifier
          │                 │
          └────────┬────────┘
                   ▼
              hybrid policy
                   │
             ┌─────┴─────┐
             ▼           ▼
         decision    human review
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The baseline stopped being just an old version of the system.&lt;/p&gt;

&lt;p&gt;It started serving as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a reference;&lt;/li&gt;
&lt;li&gt;a divergence signal;&lt;/li&gt;
&lt;li&gt;a fallback;&lt;/li&gt;
&lt;li&gt;a component of the human-review policy.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The eval did not just measure the architecture.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It helped determine the architecture.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  It also changes how you develop with LLMs
&lt;/h2&gt;

&lt;p&gt;Without structured evaluation, the loop tends to look like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;change the prompt
↓
run a few examples
↓
looks better
↓
deploy
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With evals:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;change prompt/model/policy
↓
run the dataset
↓
measure results
↓
compare against baseline
↓
analyze regressions
↓
decide
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That difference looks small.&lt;/p&gt;

&lt;p&gt;But it is the difference between experimenting and simply trusting an impression.&lt;/p&gt;

&lt;h2&gt;
  
  
  Not every metric needs to come from an LLM
&lt;/h2&gt;

&lt;p&gt;There is a lot of discussion about &lt;strong&gt;LLM-as-a-judge&lt;/strong&gt;, where another model grades the produced answer.&lt;/p&gt;

&lt;p&gt;That is useful when the criteria are subjective, such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;clarity;&lt;/li&gt;
&lt;li&gt;relevance;&lt;/li&gt;
&lt;li&gt;coherence;&lt;/li&gt;
&lt;li&gt;quality of an open-ended answer.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But when there is a verifiable answer, traditional code is usually simpler.&lt;/p&gt;

&lt;p&gt;In my case:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nx"&gt;predicted&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;category&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="nx"&gt;expected&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;category&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;already answers an important question.&lt;/p&gt;

&lt;p&gt;The evaluation tool should be proportional to the problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  The main takeaway
&lt;/h2&gt;

&lt;p&gt;I used to think of evals as something that happens after building an AI system.&lt;/p&gt;

&lt;p&gt;I see it differently now.&lt;/p&gt;

&lt;p&gt;The eval is part of development itself.&lt;/p&gt;

&lt;p&gt;It helps answer:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;did the model improve?
where did it get worse?
by how much?
in which cases?
compared to what?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And perhaps the most important question:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;does this improvement actually justify putting the LLM in this part of the system?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Without that, it is easy to build a convincing demo.&lt;/p&gt;

&lt;p&gt;With it, engineering starts to show up.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep exploring
&lt;/h2&gt;

&lt;p&gt;The full details — official metrics, trade-offs, and limitations — are in the &lt;a href="https://marcelotaparelli.com.br/en/projects/ops-triage-ai/" rel="noopener noreferrer"&gt;project case&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Project: &lt;a href="https://github.com/marcelotaparelli/ops-triage-ai" rel="noopener noreferrer"&gt;ops-triage-ai on GitHub&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>machinelearning</category>
      <category>programming</category>
    </item>
    <item>
      <title>The LLM Didn't Win Everywhere — and That's What Made the Project Interesting</title>
      <dc:creator>marcelotaparelli</dc:creator>
      <pubDate>Tue, 15 Sep 2026 09:33:21 +0000</pubDate>
      <link>https://dev.to/marcelotaparelli/the-llm-didnt-win-everywhere-and-thats-what-made-the-project-interesting-a1a</link>
      <guid>https://dev.to/marcelotaparelli/the-llm-didnt-win-everywhere-and-thats-what-made-the-project-interesting-a1a</guid>
      <description>&lt;p&gt;I closed an important stage of &lt;code&gt;ops-triage-ai&lt;/code&gt;, an operational triage&lt;br&gt;
system that combines a deterministic baseline, a local LLM, and a hybrid&lt;br&gt;
policy with human review.&lt;/p&gt;

&lt;p&gt;The most interesting result was not simply "the LLM was better."&lt;/p&gt;

&lt;h2&gt;
  
  
  What the benchmark measured
&lt;/h2&gt;

&lt;p&gt;Evaluation on a frozen held-out set of 70 synthetic tickets, in a single&lt;br&gt;
run. Deterministic baseline versus the local LLM:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Category accuracy: 82.9% → 95.7%&lt;/li&gt;
&lt;li&gt;HIGH/CRITICAL priority recall: 78.6% → 100%&lt;/li&gt;
&lt;li&gt;HIGH risk recall: 57.1% → 71.4%&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And the point that matters: the deterministic baseline still won on&lt;br&gt;
overall risk accuracy — 95.7% against the LLM's 91.4%.&lt;/p&gt;

&lt;p&gt;The LLM did not win everywhere. The regression is published alongside the&lt;br&gt;
gains, with nothing smoothed over.&lt;/p&gt;

&lt;h2&gt;
  
  
  The question changed
&lt;/h2&gt;

&lt;p&gt;Instead of:&lt;/p&gt;

&lt;p&gt;"How do we replace rules with AI?"&lt;/p&gt;

&lt;p&gt;the architecture answers:&lt;/p&gt;

&lt;p&gt;"How do we combine different behaviors in a way that is safe, observable,&lt;br&gt;
and auditable?"&lt;/p&gt;

&lt;h2&gt;
  
  
  The final architecture
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Deterministic baseline as reference and fallback;&lt;/li&gt;
&lt;li&gt;Local LLM for semantic interpretation;&lt;/li&gt;
&lt;li&gt;Hybrid policy to decide when human review is required;&lt;/li&gt;
&lt;li&gt;Audit trail preserving ticket, run, decision, and feedback.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In the final benchmark:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;51.4% of cases were routed to human review;&lt;/li&gt;
&lt;li&gt;the classifiers disagreed on 32.9% of tickets;&lt;/li&gt;
&lt;li&gt;there was no fallback and no LLM failure during the official run.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Review rate describes operations here — how many cases asked for review —&lt;br&gt;
not review precision/recall. The held-out set has no independent ground&lt;br&gt;
truth for that.&lt;/p&gt;

&lt;h2&gt;
  
  
  Accepted limits
&lt;/h2&gt;

&lt;p&gt;Synthetic dataset, in English, with 70 examples. A single official run,&lt;br&gt;
with no measurement of LLM variance. The system has not been validated&lt;br&gt;
under real production traffic. Hybrid latency sits around 6 to 7 seconds:&lt;br&gt;
fine for asynchronous triage, not for a critical synchronous path. The&lt;br&gt;
full details — official metrics, trade-offs, and the 13 published&lt;br&gt;
limitations — are in the &lt;a href="https://marcelotaparelli.com.br/en/projects/ops-triage-ai/" rel="noopener noreferrer"&gt;project case&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The takeaway
&lt;/h2&gt;

&lt;p&gt;A useful AI system does not need to trust the model blindly.&lt;/p&gt;

&lt;p&gt;Sometimes the best architecture comes precisely from understanding where&lt;br&gt;
the model is better — and where it is not.&lt;/p&gt;

&lt;p&gt;Project: &lt;a href="https://github.com/marcelotaparelli/ops-triage-ai" rel="noopener noreferrer"&gt;https://github.com/marcelotaparelli/ops-triage-ai&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>machinelearning</category>
      <category>showdev</category>
    </item>
    <item>
      <title>Why I Built My Portfolio with Bun + Astro + MDX Instead of a More Complex Stack</title>
      <dc:creator>marcelotaparelli</dc:creator>
      <pubDate>Sun, 13 Sep 2026 13:05:47 +0000</pubDate>
      <link>https://dev.to/marcelotaparelli/why-i-built-my-portfolio-with-bun-astro-mdx-instead-of-a-more-complex-stack-2cmd</link>
      <guid>https://dev.to/marcelotaparelli/why-i-built-my-portfolio-with-bun-astro-mdx-instead-of-a-more-complex-stack-2cmd</guid>
      <description>&lt;p&gt;Choosing a stack is a product decision before it is a technical one. This&lt;br&gt;
article documents why this portfolio runs on Bun, Astro, and MDX — and,&lt;br&gt;
more importantly, the reasoning that led me to turn down more sophisticated&lt;br&gt;
alternatives. Nothing here is theory: every claim points to something&lt;br&gt;
verifiable in this repository.&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem
&lt;/h2&gt;

&lt;p&gt;I needed a portfolio that could carry professional authority as a Software&lt;br&gt;
Engineer going deeper into Applied AI Engineering. The requirements were&lt;br&gt;
concrete: bilingual content (Portuguese and English), project cases and&lt;br&gt;
articles with their own identity, solid SEO, high performance, genuine&lt;br&gt;
accessibility, and maintenance simple enough that I would never have to&lt;br&gt;
think about it twice. No backend, no platform team, no infrastructure&lt;br&gt;
budget — just me, the code, and the content.&lt;/p&gt;

&lt;h2&gt;
  
  
  The question behind the architecture
&lt;/h2&gt;

&lt;p&gt;The guiding question was simple: what is the simplest solution capable of&lt;br&gt;
solving this problem? Not "which stack is trending," and not "which one&lt;br&gt;
shows off the most skill." Technology chosen for hype often charges&lt;br&gt;
interest in complexity. Every layer I added had to justify itself against this specific&lt;br&gt;
problem — a mostly static content site maintained by one person.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Bun
&lt;/h2&gt;

&lt;p&gt;Bun is this project's toolchain: package manager, runner for the&lt;br&gt;
development scripts, and runtime for the validation tooling. The dev&lt;br&gt;
server, build, artifact validation, asset measurement, and release-check&lt;br&gt;
scripts are TypeScript executed directly by Bun, with no separate&lt;br&gt;
transpilation layer for tooling. I make no comparative claims I haven't&lt;br&gt;
measured here — the decision was about structural simplicity: one tool&lt;br&gt;
where there would otherwise be several. What is observable is objective&lt;br&gt;
but contextual: in this implementation, on the local environment, the full&lt;br&gt;
build renders every page in about two seconds.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Astro
&lt;/h2&gt;

&lt;p&gt;This site's content is predominantly static, so HTML is generated at build&lt;br&gt;
time and served as files — no application server, no database, no&lt;br&gt;
per-request runtime. Astro was chosen for exactly that model: it lets me&lt;br&gt;
write composable editorial content while shipping almost no JavaScript to&lt;br&gt;
the browser. The script embedded in the home page is 277 bytes (198&lt;br&gt;
gzipped) and exists for exactly one job: the mobile menu. Zero bytes of&lt;br&gt;
client-side framework — measured, not estimated. React was never installed&lt;br&gt;
because no problem required React. This model also simplifies SEO: every&lt;br&gt;
page arrives as complete HTML, with its own canonical and alternates.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why MDX
&lt;/h2&gt;

&lt;p&gt;Articles and cases live versioned alongside the code, in the same&lt;br&gt;
repository, under the same review process. There is no CMS or backend to&lt;br&gt;
maintain, update, or pay for — and no content outside version control. MDX&lt;br&gt;
gives the content structure (schema-validated frontmatter requiring a&lt;br&gt;
bilingual pair, slug, category, and review status), which makes every piece&lt;br&gt;
of writing an artifact as reviewable as any other code. Publishing means&lt;br&gt;
merging and building; rolling back means reverting a commit.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why not Next.js, SPA, or CMS
&lt;/h2&gt;

&lt;p&gt;Not because they are bad technologies — they are excellent for the right&lt;br&gt;
problems. A full-stack framework would solve problems I don't have:&lt;br&gt;
per-request rendering, API routes, global client state. A client-rendered&lt;br&gt;
SPA would add JavaScript runtime and browser state to a site whose main&lt;br&gt;
job is delivering ready-to-read content. A CMS would trade versioned files for an external dependency with logins,&lt;br&gt;
backups, and a bill. None of that complexity was justified by the current&lt;br&gt;
problem, so none of it got in.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bilingualism as an architecture requirement
&lt;/h2&gt;

&lt;p&gt;Bilingualism is not a plugin bolted on afterward — it shaped the&lt;br&gt;
architecture. Portuguese lives at &lt;code&gt;/&lt;/code&gt;, English at &lt;code&gt;/en/&lt;/code&gt;, and every&lt;br&gt;
published item is required to exist in both languages: the publication&lt;br&gt;
model rejects at build time anything published without its reviewed&lt;br&gt;
counterpart. Every page carries its own canonical and three alternates&lt;br&gt;
(current language, opposite language, and x-default), and the artifact&lt;br&gt;
validator checks the pairs across every generated page. Incomplete content&lt;br&gt;
simply never reaches production: drafts are excluded from the production&lt;br&gt;
build and included only in the preview build. The &lt;a href="https://marcelotaparelli.com.br/en/projects/" rel="noopener noreferrer"&gt;projects page&lt;/a&gt;&lt;br&gt;
and the &lt;a href="https://marcelotaparelli.com.br/en/about/" rel="noopener noreferrer"&gt;about page&lt;/a&gt; exist in both languages because the&lt;br&gt;
system would accept nothing less.&lt;/p&gt;

&lt;h2&gt;
  
  
  Quality as part of the product
&lt;/h2&gt;

&lt;p&gt;Here, quality is not a promise — it is a pipeline. Every change goes&lt;br&gt;
through strict typechecking (zero errors), linting, verified formatting,&lt;br&gt;
unit tests for the publication model (4/4), thirteen browser tests&lt;br&gt;
(reciprocal PT/EN navigation, keyboard menu with Escape, no-JavaScript&lt;br&gt;
navigation, localized 404, explicit external links, 320px reflow with&lt;br&gt;
automated WCAG checks), and&lt;br&gt;
artifact validation in both build modes. If something breaks, the build&lt;br&gt;
says so — before any human needs to check.&lt;/p&gt;

&lt;p&gt;CI/CD pipeline with GitHub Actions covering type checking, linting, unit&lt;br&gt;
tests, Playwright E2E tests, artifact validation, and production builds.&lt;br&gt;
Delivery follows a Continuous Delivery model: the same CI-approved&lt;br&gt;
artifact is published after explicit manual approval, without rebuilding,&lt;br&gt;
to a dedicated production branch and then deployed to Hostinger.&lt;/p&gt;

&lt;h2&gt;
  
  
  Measured performance
&lt;/h2&gt;

&lt;p&gt;What follows are laboratory measurements, never real-user data. I ran&lt;br&gt;
Lighthouse 13.4.1 in headless Chromium, mobile simulation (412×823&lt;br&gt;
viewport, simulated network and CPU throttling), three runs per page,&lt;br&gt;
serving the local static build — methodology recorded in&lt;br&gt;
&lt;code&gt;scripts/lighthouse.ts&lt;/code&gt;, results in &lt;code&gt;reports/lighthouse-summary.json&lt;/code&gt;. On&lt;br&gt;
the home page: performance 100 and accessibility 100 across all three&lt;br&gt;
runs, LCP between roughly 1.5 s and 1.7 s, CLS around 0.0006, zero TBT,&lt;br&gt;
and about 93.5 KB of initial transfer. The asset budget&lt;br&gt;
(&lt;code&gt;reports/assets.json&lt;/code&gt;) shows where the lightness comes from: roughly&lt;br&gt;
6 KB of gzipped CSS, about 0.2 KB of inline JavaScript, and fonts totaling&lt;br&gt;
around 48 KB. Lab numbers inform decisions; they prove nothing about real&lt;br&gt;
user experience, and I would never present them as such.&lt;/p&gt;

&lt;h2&gt;
  
  
  A bug worth finding
&lt;/h2&gt;

&lt;p&gt;The artifact validator flagged a broken link on the Portuguese 404 page:&lt;br&gt;
&lt;code&gt;/404/&lt;/code&gt;. The cause was in the header's language switcher, which used&lt;br&gt;
&lt;code&gt;Astro.url.pathname&lt;/code&gt; as the current-language link — and with trailing&lt;br&gt;
slashes always on, that pathname renders as &lt;code&gt;/404/&lt;/code&gt;. But the actual&lt;br&gt;
artifact Astro generates for &lt;code&gt;404.astro&lt;/code&gt; is &lt;code&gt;404.html&lt;/code&gt;, a special 404&lt;br&gt;
document for static hosting; the &lt;code&gt;/404/&lt;/code&gt; route never existed. The fix went&lt;br&gt;
into the routing model (the 404 page now links its canonical&lt;br&gt;
&lt;code&gt;/404.html&lt;/code&gt; and &lt;code&gt;/en/404/&lt;/code&gt; paths), not into the test. Weakening the&lt;br&gt;
validator would have hidden the symptom and destroyed its value. The&lt;br&gt;
episode became a working rule: fix the publication model, never work&lt;br&gt;
around the validator.&lt;/p&gt;

&lt;h2&gt;
  
  
  Trade-offs
&lt;/h2&gt;

&lt;p&gt;Every choice has a cost, and these are mine, accepted: no backend, no&lt;br&gt;
database, no CMS — any new content requires a commit, a build, and a&lt;br&gt;
deploy. No React at launch — richer interactivity in the future will&lt;br&gt;
require revisiting the decision. No analytics — there is no real-usage&lt;br&gt;
telemetry, so the lab metrics above are the ceiling of what I can claim&lt;br&gt;
today. And bilingualism costs double review: every piece must exist, make&lt;br&gt;
sense, and be approved in both languages. A deliberate cost, because&lt;br&gt;
international reach and consistency across languages are part of the&lt;br&gt;
product I want to build.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Mature engineering is not about choosing the most sophisticated stack —&lt;br&gt;
it is about choosing complexity proportional to the problem. This&lt;br&gt;
portfolio could have been a full-stack monolith, an SPA, or a CMS&lt;br&gt;
instance; those approaches could work, but they would introduce&lt;br&gt;
capabilities and operational costs this project's requirements never asked&lt;br&gt;
for. Technology starts from the human&lt;br&gt;
problem, not from the code: understand what needs to change, choose&lt;br&gt;
deliberately, and examine the outcome with evidence. FROM REAL PROBLEMS TO&lt;br&gt;
INTELLIGENT PRODUCTS — including when building the showcase itself.&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>astro</category>
      <category>bunjs</category>
      <category>showdev</category>
    </item>
    <item>
      <title>⚡ C# Tip: ToCharArray() vs ToArray() — Why It Matters More Than You Think</title>
      <dc:creator>marcelotaparelli</dc:creator>
      <pubDate>Fri, 04 Jul 2025 18:29:21 +0000</pubDate>
      <link>https://dev.to/marcelotaparelli/c-tip-tochararray-vs-toarray-why-it-matters-more-than-you-think-56ba</link>
      <guid>https://dev.to/marcelotaparelli/c-tip-tochararray-vs-toarray-why-it-matters-more-than-you-think-56ba</guid>
      <description>&lt;p&gt;As developers, we often write code that “just works” — but sometimes, small choices can have a big impact on performance.&lt;/p&gt;

&lt;p&gt;One example? The subtle but meaningful difference between ToCharArray() and ToArray() when working with strings in C#. Let’s break it down 👇&lt;/p&gt;



&lt;h2&gt;
  
  
  🧩 What Each One Does
&lt;/h2&gt;

&lt;p&gt;✅ &lt;strong&gt;ToCharArray()&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Native method on the string class.&lt;/li&gt;
&lt;li&gt;Returns a char[].&lt;/li&gt;
&lt;li&gt;Fast and direct.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;🚧 &lt;strong&gt;ToArray()&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Extension method from LINQ (System.Linq).&lt;/li&gt;
&lt;li&gt;Treats the string as IEnumerable.&lt;/li&gt;
&lt;li&gt;Adds abstraction and intermediate steps.&lt;/li&gt;
&lt;/ul&gt;



&lt;h2&gt;
  
  
  🔍 The Overhead Problem
&lt;/h2&gt;

&lt;p&gt;Calling ToArray() on a string might seem harmless, but it invokes LINQ under the hood. That means:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Extra allocations&lt;/li&gt;
&lt;li&gt;Added abstraction&lt;/li&gt;
&lt;li&gt;Unnecessary iteration&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Meanwhile, ToCharArray() skips all that and copies the characters directly — with less memory and better performance. In tight loops, I've seen it run 2–3x faster. That adds up.&lt;/p&gt;



&lt;h2&gt;
  
  
  🧪 When Should You Use Each?
&lt;/h2&gt;

&lt;p&gt;Let’s keep it simple:&lt;/p&gt;

&lt;p&gt;👉 Use ToCharArray() when you're working directly with strings — parsing, looping, checking characters, etc.&lt;br&gt;
👉 Use ToArray() when you're working with LINQ and want to convert the result of a query to an array.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Example:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;// Using LINQ to filter and convert to array
var expensive = products
 .Where(p =&amp;gt; p.Price &amp;gt; 100)
 .ToArray(); // 👍 This is a perfect case for ToArray()
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;h2&gt;
  
  
  💡 Final Thoughts
&lt;/h2&gt;

&lt;p&gt;Is this a micro-optimization? Maybe.&lt;br&gt;
But understanding how your tools work under the hood can make you a better developer — and sometimes, a faster one too.&lt;/p&gt;



&lt;p&gt;If this helped or surprised you, give it a share so others don’t fall into the same trap 😄&lt;/p&gt;

&lt;p&gt;I’d love to hear your thoughts — have you encountered other subtle performance gotchas in C#?&lt;/p&gt;

</description>
      <category>csharp</category>
      <category>dotnet</category>
      <category>performance</category>
      <category>programmingtips</category>
    </item>
  </channel>
</rss>
