<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: özkan pakdil</title>
    <description>The latest articles on DEV Community by özkan pakdil (@ozkanpakdil).</description>
    <link>https://dev.to/ozkanpakdil</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F302594%2Fe6338387-2374-431d-8e72-c364cdcad84a.jpg</url>
      <title>DEV Community: özkan pakdil</title>
      <link>https://dev.to/ozkanpakdil</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/ozkanpakdil"/>
    <language>en</language>
    <item>
      <title>MariaDB vs MySQL vs PostgreSQL: an mpazari Benchmark</title>
      <dc:creator>özkan pakdil</dc:creator>
      <pubDate>Wed, 23 Sep 2026 00:00:00 +0000</pubDate>
      <link>https://dev.to/ozkanpakdil/mariadb-vs-mysql-vs-postgresql-an-mpazari-benchmark-4kp7</link>
      <guid>https://dev.to/ozkanpakdil/mariadb-vs-mysql-vs-postgresql-an-mpazari-benchmark-4kp7</guid>
      <description>&lt;p&gt;I recently compared MariaDB, MySQL, and PostgreSQL using the real&lt;a href="https://www.mpazari.com" rel="noopener noreferrer"&gt;mpazari.com&lt;/a&gt; application. The goal was not to produce a synthetic database benchmark. I wanted to answer a more practical question: &lt;strong&gt;does MariaDB really have an advantage when an application sends many simple queries?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This question came up after reading an argument that PostgreSQL is obviously faster than MySQL, but that MariaDB should be faster for applications with many simple queries, while PostgreSQL is probably better for complex joins. The suggestion was that a MariaDB versus PostgreSQL comparison would be more interesting than a PostgreSQL versus MySQL comparison.&lt;/p&gt;

&lt;p&gt;I had already migrated mpazari.com from MySQL 8 to PostgreSQL 12. That migration made the server feel much lighter, but it did not answer the MariaDB part of the question. So I restored the same data into MariaDB and MySQL and compared those results with the earlier PostgreSQL run.&lt;/p&gt;

&lt;h2&gt;
  
  
  The application
&lt;/h2&gt;

&lt;p&gt;mpazari.com is a Turkish motorcycle classifieds site. The application is a Spring Boot 4 / Java 25 GraalVM native image with hand-written SQL. The home page is not a single complicated query; it performs several small reads for listings, taxonomy data, counts, and footer information.&lt;/p&gt;

&lt;p&gt;The database is modest by production standards, but it is real application data rather than a generated benchmark dataset:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;31 tables&lt;/li&gt;
&lt;li&gt;20,750 rows in &lt;code&gt;motor_ilanlar&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;approximately 111 MB of data used by the tested workload&lt;/li&gt;
&lt;li&gt;roughly five simple database queries per home-page request&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That makes this a useful test for this particular workload, not a universal ranking of database engines.&lt;/p&gt;

&lt;h2&gt;
  
  
  Test setup
&lt;/h2&gt;

&lt;p&gt;I used the same Hetzner server for both runs: 8 cores, 32 GB of RAM, and Ubuntu 20.04. The system was otherwise idle, with the load average below 1 before the tests.&lt;/p&gt;

&lt;p&gt;The important comparison details were:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The same pre-cutover dump was restored into MariaDB and MySQL: 31 tables and 20,750&lt;code&gt;motor_ilanlar&lt;/code&gt; rows.&lt;/li&gt;
&lt;li&gt;The same application build and runtime settings were used for the MariaDB and MySQL runs.&lt;/li&gt;
&lt;li&gt;The application was freshly restarted before each main run.&lt;/li&gt;
&lt;li&gt;MariaDB and MySQL listened only on &lt;code&gt;127.0.0.1&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;MySQL ran in Docker with &lt;code&gt;--network host&lt;/code&gt;, so there was no NAT or proxy hop.&lt;/li&gt;
&lt;li&gt;MariaDB used port 3306 and MySQL used port 3307.&lt;/li&gt;
&lt;li&gt;k6 ramped to 50 virtual users for one minute, held that load for five minutes, and ramped down for one minute. Each iteration loaded the home page and slept for one second.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The PostgreSQL numbers come from an earlier same-box benchmark against the live stack. That stack speaks PostgreSQL-dialect SQL, and its PostgreSQL instance hosts several sites’ databases, so the PostgreSQL result is useful context rather than a perfectly engine-pure comparison.&lt;/p&gt;

&lt;p&gt;The k6 thresholds were &lt;code&gt;p(95) &amp;lt; 200 ms&lt;/code&gt; and &lt;code&gt;p(99) &amp;lt; 500 ms&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Main benchmark
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;MariaDB 10.3.39&lt;/th&gt;
&lt;th&gt;MySQL 8.0.42&lt;/th&gt;
&lt;th&gt;PostgreSQL 12*&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Average&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;125.52 ms&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;132.40 ms&lt;/td&gt;
&lt;td&gt;156.97 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;p(95)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;129.73 ms&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;139.43 ms&lt;/td&gt;
&lt;td&gt;171.86 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;p(99)&lt;/td&gt;
&lt;td&gt;182.13 ms&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;150.07 ms&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;177.16 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Maximum&lt;/td&gt;
&lt;td&gt;229.31 ms&lt;/td&gt;
&lt;td&gt;243.90 ms&lt;/td&gt;
&lt;td&gt;287.17 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Requests&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;16,025&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;15,930&lt;/td&gt;
&lt;td&gt;15,590&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Failed requests&lt;/td&gt;
&lt;td&gt;0%&lt;/td&gt;
&lt;td&gt;0%&lt;/td&gt;
&lt;td&gt;0%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Database RSS peak&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;141 MB&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;411 MB&lt;/td&gt;
&lt;td&gt;1.7 GB**&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Load average, 1-minute average / peak&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.88 / 1.34&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1.27 / 2.01&lt;/td&gt;
&lt;td&gt;–&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;MariaDB was about 5% faster than MySQL on average and about 7% faster at p(95). PostgreSQL was about 25% slower than MariaDB on average in the contextual comparison. The MariaDB and MySQL difference is real, but not dramatic. All three results passed the latency thresholds comfortably.&lt;/p&gt;

&lt;p&gt;The more interesting result is the tail: MySQL’s p(99) was 150 ms, compared with 182 ms for MariaDB, while PostgreSQL measured 177 ms. MariaDB won the average and p(95), but MySQL handled the slowest one percent of requests more consistently in this run.&lt;/p&gt;

&lt;p&gt;* PostgreSQL was measured in the earlier live-stack run, with different SQL text and a shared cluster.&lt;/p&gt;

&lt;p&gt;** The production PostgreSQL instance hosts several sites’ databases, so its CPU and memory cannot be attributed to this workload alone.&lt;/p&gt;

&lt;h2&gt;
  
  
  A smaller CPU comparison
&lt;/h2&gt;

&lt;p&gt;I also ran a smaller 15-VU test for two and a half minutes. The databases were equally warm when this test started.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;MariaDB 10.3.39&lt;/th&gt;
&lt;th&gt;MySQL 8.0.42&lt;/th&gt;
&lt;th&gt;MySQL 8.0.42 (&lt;code&gt;performance_schema=OFF&lt;/code&gt;)&lt;/th&gt;
&lt;th&gt;PostgreSQL 12*&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Average / p(95)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;125.12 / 130.26 ms&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;130.25 / 136.35 ms&lt;/td&gt;
&lt;td&gt;130.81 / 136.82 ms&lt;/td&gt;
&lt;td&gt;158.7 / 169.2 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Database CPU average / peak&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.2% / 1.0%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;6.6% / 9.0%&lt;/td&gt;
&lt;td&gt;8.3% / 10.0%&lt;/td&gt;
&lt;td&gt;n/a**&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Database RSS&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;141 MB&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;411 MB&lt;/td&gt;
&lt;td&gt;162 MB&lt;/td&gt;
&lt;td&gt;1.7 GB**&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;* PostgreSQL was measured in the earlier live-stack run.&lt;/p&gt;

&lt;p&gt;** PostgreSQL CPU and memory are not attributable to this workload alone because the instance is shared by several sites.&lt;/p&gt;

&lt;p&gt;For this workload, MariaDB used noticeably less CPU. The application sends many uncomplicated queries, and MariaDB’s executor appears to handle them with less overhead. This is the result that most closely matches the original claim.&lt;/p&gt;

&lt;p&gt;The CPU numbers should not be interpreted as a general statement that MySQL is expensive. The tested request rate was low and the database was mostly waiting. At this scale, a few percentage points of CPU do not affect the SLA, but they may matter on a smaller server or at a much higher request rate.&lt;/p&gt;

&lt;h2&gt;
  
  
  Memory and &lt;code&gt;performance_schema&lt;/code&gt;
&lt;/h2&gt;

&lt;p&gt;The first MySQL run used 411 MB of resident memory, compared with 141 MB for MariaDB. That looks like a large difference, but the defaults are not equivalent: MariaDB ships with&lt;code&gt;performance_schema&lt;/code&gt; disabled, while MySQL enables it by default.&lt;/p&gt;

&lt;p&gt;With &lt;code&gt;performance_schema&lt;/code&gt; disabled, MySQL used 162 MB. That is much closer to MariaDB’s 141 MB, so the original memory comparison was mostly a configuration difference rather than an inherent 900 MB advantage for one engine.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this does and does not prove
&lt;/h2&gt;

&lt;p&gt;This benchmark supports a narrow conclusion:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;For this application, with many simple queries and a small, fully resident dataset, MariaDB 10.3 was slightly faster on average and used less CPU than MySQL 8.0. The available PostgreSQL context was slower on this box, but it was not an engine-pure comparison.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It does not prove that MariaDB is always faster than MySQL or PostgreSQL. It also does not show how any engine would behave with large joins, analytical queries, writes under contention, replication, different indexes, or a much larger dataset.&lt;/p&gt;

&lt;p&gt;There is also a version caveat. MariaDB 10.3 is a 2019-generation release, while MySQL 8.0.42 is current. A comparison with MariaDB 10.6 or 11.x might produce different CPU and latency results. MySQL 8 features such as CTEs and window functions were not relevant to this workload.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;The simple-query prediction held for mpazari.com. MariaDB was approximately 5% faster than MySQL on average, 7% faster at p(95), and used much less CPU in the smaller comparison. The contextual PostgreSQL result was approximately 25% slower than MariaDB on average, although its different SQL dialect and shared cluster make that comparison directional rather than definitive. MySQL had the better p(99) in the main run, and its apparent memory disadvantage largely disappeared when &lt;code&gt;performance_schema&lt;/code&gt; was disabled.&lt;/p&gt;

&lt;p&gt;The differences are interesting, but they are not large enough to make the database engine the bottleneck here. Both databases passed the SLA easily. For this application, schema design, indexes, caching, connection handling, and the rest of the server stack are more important than choosing between MariaDB and MySQL based on a 5% average-latency difference.&lt;/p&gt;

&lt;p&gt;So the answer to the original question is: &lt;strong&gt;yes, MariaDB can be faster for many simple queries, and this real application showed that pattern. PostgreSQL was slower in the available contextual run, but that result needs a like-for-like rerun before making a strong engine-level claim. The MariaDB advantage over MySQL was small, workload-specific, and not enough to make MySQL an unreasonable choice.&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>database</category>
      <category>performance</category>
      <category>postgres</category>
      <category>sql</category>
    </item>
    <item>
      <title>From Spring Boot to Rust: Rewriting a Live Marketplace</title>
      <dc:creator>özkan pakdil</dc:creator>
      <pubDate>Mon, 21 Sep 2026 00:00:00 +0000</pubDate>
      <link>https://dev.to/ozkanpakdil/from-spring-boot-to-rust-rewriting-a-live-marketplace-bed</link>
      <guid>https://dev.to/ozkanpakdil/from-spring-boot-to-rust-rewriting-a-live-marketplace-bed</guid>
      <description>&lt;h2&gt;
  
  
  From Spring Boot to Rust: Rewriting a Live Marketplace
&lt;/h2&gt;

&lt;p&gt;mpazari.com spent years on a legacy ASP.NET stack — the &lt;code&gt;.aspx&lt;/code&gt; URLs still in the wild prove it. A recent Spring Boot rewrite replaced that engine, first as a JVM jar, then as a GraalVM native image to tame memory usage. The question that started this project: &lt;em&gt;if we rewrote the whole application in Rust, what would we gain, what would we break, and how do we prove we broke nothing?&lt;/em&gt; This post covers the rewrite, the parity hunting, the server-side build/deploy pipeline, and the load test comparing all three runtimes on the same production box.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Why, and what exactly
&lt;/h2&gt;

&lt;p&gt;The goal was not a rewrite “someday”. It was: &lt;strong&gt;same database, same URLs, same behavior — different engine.&lt;/strong&gt; Every page had to behave identically to the &lt;code&gt;aspx&lt;/code&gt; version:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;Java (old)&lt;/th&gt;
&lt;th&gt;Rust (new)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Framework&lt;/td&gt;
&lt;td&gt;Spring Boot / MVC + Thymeleaf&lt;/td&gt;
&lt;td&gt;warp 0.3 + hand-rolled filter chain&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Templates&lt;/td&gt;
&lt;td&gt;Thymeleaf (&lt;code&gt;th:*&lt;/code&gt; attributes)&lt;/td&gt;
&lt;td&gt;minijinja 2 (Jinja syntax)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DB access&lt;/td&gt;
&lt;td&gt;Spring JDBC / NamedParameterJdbcTemplate&lt;/td&gt;
&lt;td&gt;sqlx 0.8, plain SQL, same PostgreSQL schema&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sessions&lt;/td&gt;
&lt;td&gt;Spring Session&lt;/td&gt;
&lt;td&gt;stateless HMAC-SHA256 signed JSON cookie&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Deploy artifact&lt;/td&gt;
&lt;td&gt;GraalVM native image (112 MB) or fat jar&lt;/td&gt;
&lt;td&gt;single 21 MB binary&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  2. The load test: same box, three runtimes
&lt;/h2&gt;

&lt;p&gt;Method:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Same machine (Hetzner, Ubuntu 20.04), same PostgreSQL data, same Apache-less path — k6 maps &lt;code&gt;www.mpazari.com&lt;/code&gt; straight to the app port (&lt;code&gt;9100&lt;/code&gt;) via &lt;code&gt;/etc/hosts&lt;/code&gt; override, so the measurements are app-only, no proxy noise.&lt;/li&gt;
&lt;li&gt;k6 scenario: ramp to 50 VUs in 1m, hold 50 VUs for 5m, ramp down in 1m; each iteration fetches the home page and sleeps 1s.&lt;/li&gt;
&lt;li&gt;Thresholds: &lt;code&gt;p(95) &amp;lt; 200ms&lt;/code&gt;, &lt;code&gt;p(99) &amp;lt; 500ms&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Runtime&lt;/th&gt;
&lt;th&gt;avg&lt;/th&gt;
&lt;th&gt;p(95)&lt;/th&gt;
&lt;th&gt;p(99)&lt;/th&gt;
&lt;th&gt;max&lt;/th&gt;
&lt;th&gt;reqs&lt;/th&gt;
&lt;th&gt;failed&lt;/th&gt;
&lt;th&gt;RSS @ 50 VU&lt;/th&gt;
&lt;th&gt;Artifact&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GraalVM native&lt;/td&gt;
&lt;td&gt;158.58ms&lt;/td&gt;
&lt;td&gt;172.57ms&lt;/td&gt;
&lt;td&gt;178.64ms&lt;/td&gt;
&lt;td&gt;374.39ms&lt;/td&gt;
&lt;td&gt;15,570&lt;/td&gt;
&lt;td&gt;0%&lt;/td&gt;
&lt;td&gt;134–154 MB&lt;/td&gt;
&lt;td&gt;112 MB binary&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Spring Boot jar&lt;/td&gt;
&lt;td&gt;156.97ms&lt;/td&gt;
&lt;td&gt;171.86ms&lt;/td&gt;
&lt;td&gt;177.16ms&lt;/td&gt;
&lt;td&gt;287.17ms&lt;/td&gt;
&lt;td&gt;15,590&lt;/td&gt;
&lt;td&gt;0%&lt;/td&gt;
&lt;td&gt;477–949 MB&lt;/td&gt;
&lt;td&gt;32 MB jar&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Rust (warp + sqlx)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;160.89ms&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;179.45ms&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;191.91ms&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;351.72ms&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;15,545&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;20–40 MB&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;21 MB binary&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;All three runtimes pass both thresholds on the same day, on the same box, against the same database. Rust lands within ~7ms of the JVM’s p(95) — and serves the load from a 21 MB binary topping out at 40 MB RSS, while the JVM pays its heap for the same traffic.&lt;/p&gt;

&lt;h3&gt;
  
  
  2.1 The bug the load test caught: a &lt;code&gt;Lazy&lt;/code&gt; regex that compiled per request
&lt;/h3&gt;

&lt;p&gt;The first Rust run was a collapse: avg 757ms, p(95) 952ms, 10k reqs vs 15.5k for Java — and &lt;code&gt;top&lt;/code&gt; showed the Rust process eating &lt;strong&gt;623% CPU&lt;/strong&gt; for 28 req/s. Nothing about the workload explained it, so we profiled the live process (&lt;code&gt;perf record&lt;/code&gt; on the server, then the same repro locally with macOS &lt;code&gt;sample&lt;/code&gt;).&lt;/p&gt;

&lt;p&gt;The call graph said it all:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;regex_automata::dfa::determinize ... 596 samples
  once_cell::Lazy::get_or_init
    regex::Regex::new
      mpazari::util::taxonomy_slug
        mpazari::web::home::home

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;taxonomy_slug&lt;/code&gt; declared its regex &lt;strong&gt;inside the function&lt;/strong&gt; :&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight rust"&gt;&lt;code&gt;&lt;span class="k"&gt;pub&lt;/span&gt; &lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="nf"&gt;taxonomy_slug&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;String&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;x&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;transliterate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="nf"&gt;.to_lowercase&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
    &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Lazy&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;Regex&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nn"&gt;Lazy&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;new&lt;/span&gt;&lt;span class="p"&gt;(||&lt;/span&gt; &lt;span class="nn"&gt;Regex&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;new&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"[^a-z0-9]+"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="nf"&gt;.unwrap&lt;/span&gt;&lt;span class="p"&gt;());&lt;/span&gt; &lt;span class="c1"&gt;// ← per call!&lt;/span&gt;
    &lt;span class="o"&gt;...&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A &lt;code&gt;let re: Lazy&amp;lt;Regex&amp;gt;&lt;/code&gt; inside a function is not a cache — it is a fresh &lt;code&gt;Lazy&lt;/code&gt; per call, so every invocation re-compiled the regex, lazily building its full DFA. The home page renders ~183 brand slugs + 82 cities + 40 categories plus 39 listing cards — &lt;strong&gt;hundreds of regex compilations per request&lt;/strong&gt;. It looked like a &lt;code&gt;Lazy&lt;/code&gt; (static-shaped code), but it was a &lt;code&gt;let&lt;/code&gt;, so it behaved like &lt;code&gt;Regex::new&lt;/code&gt; on the hot path.&lt;/p&gt;

&lt;p&gt;Moving the regex to a module-level &lt;code&gt;static RE_NON_ALNUM: Lazy&amp;lt;Regex&amp;gt;&lt;/code&gt;:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;before&lt;/th&gt;
&lt;th&gt;after&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;home page (single request, prod)&lt;/td&gt;
&lt;td&gt;170–198ms&lt;/td&gt;
&lt;td&gt;30–40ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;k6 50 VUs&lt;/td&gt;
&lt;td&gt;avg 757ms, p(95) 952ms&lt;/td&gt;
&lt;td&gt;avg 161ms, p(95) 179ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;app CPU under load&lt;/td&gt;
&lt;td&gt;623%&lt;/td&gt;
&lt;td&gt;18%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;regex frames in sample&lt;/td&gt;
&lt;td&gt;~12% of samples&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The lesson generalizes: &lt;strong&gt;in Rust, a &lt;code&gt;Lazy&lt;/code&gt; is only a cache if it outlives the call.&lt;/strong&gt; &lt;code&gt;let re = Lazy::new(...)&lt;/code&gt; reads like &lt;code&gt;static RE: Lazy&amp;lt;Regex&amp;gt;&lt;/code&gt; and compiles silently — unit tests pass, pages render correctly, and only a production load test with a CPU counter exposes it.&lt;/p&gt;

&lt;p&gt;The second fix was Java-parity caching, found by reading &lt;code&gt;GlobalModelAdvice&lt;/code&gt;: Spring caches brands/cities/categories/counts in memory for 1h and the two footer COUNT(*)s for 10m — the Rust side re-queried &lt;code&gt;brand_hub&lt;/code&gt;/&lt;code&gt;city_hub&lt;/code&gt;/counts on every request. Porting the same TTLs (&lt;code&gt;TaxonomyCache&lt;/code&gt; + a 10-minute &lt;code&gt;CountsCache&lt;/code&gt;) brought per-request SQL to ~5 queries, matching the Java home handler.&lt;/p&gt;

&lt;h3&gt;
  
  
  2.2 Memory: the number the whole rewrite was chasing
&lt;/h3&gt;

&lt;p&gt;The RSS column in the main table comes from the same production box: swap the runtime, restart, run a separate 2.5-minute 50-VU k6 while sampling the process RSS every 15s. Idle values, for completeness:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Runtime&lt;/th&gt;
&lt;th&gt;RSS idle&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GraalVM native image&lt;/td&gt;
&lt;td&gt;~4 MB¹&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Spring Boot jar&lt;/td&gt;
&lt;td&gt;~477 MB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Rust (warp + sqlx)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;19 MB&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The JVM pays for its configured heap: &lt;code&gt;run-prod.sh&lt;/code&gt; starts it with &lt;code&gt;-Xms1g&lt;/code&gt;, so RSS begins near half a gigabyte and climbs toward ~1 GB as pages are touched under load. The native image is impressively compact at boot (~4 MB of resident pages before requests pull code and heap in) but settles around 140 MB serving traffic. Rust idles at 19 MB and tops out at 40 MB — roughly &lt;strong&gt;3.5× less than the native image and 24× less than the JVM&lt;/strong&gt; , while serving the same 50 VUs within a few ms of both.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. What we learned
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Parity testing is the whole project.&lt;/strong&gt; A rewrite is not a port; it is a proof. The test harness (31 acceptance + 150 e2e) is the actual deliverable — the Rust code is almost a byproduct of writing it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Silent emptiness is the enemy.&lt;/strong&gt; A missing context key renders an empty list, not an error. Every “empty page” report (&lt;code&gt;/markalar&lt;/code&gt;, &lt;code&gt;/sehirler&lt;/code&gt;, &lt;code&gt;/arama&lt;/code&gt;, favorites, the edit form) was a handler-template key mismatch, never a database problem.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Legacy URLs are a product feature.&lt;/strong&gt; A decade of &lt;code&gt;.aspx&lt;/code&gt; links, query-string shapes and redirect quirks live in Java’s &lt;code&gt;LegacyRedirectController&lt;/code&gt;. Rust reimplemented the full map — and Playwright tests every row of it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One binary named the same as the old one makes deploy a copy.&lt;/strong&gt; Zero ceremony: cp, restart, health sweep, rollback if unhealthy.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>architecture</category>
      <category>java</category>
      <category>performance</category>
      <category>rust</category>
    </item>
    <item>
      <title>From MySQL to PostgreSQL: a much lighter server</title>
      <dc:creator>özkan pakdil</dc:creator>
      <pubDate>Fri, 18 Sep 2026 00:00:00 +0000</pubDate>
      <link>https://dev.to/ozkanpakdil/from-mysql-to-postgresql-a-much-lighter-server-24i7</link>
      <guid>https://dev.to/ozkanpakdil/from-mysql-to-postgresql-a-much-lighter-server-24i7</guid>
      <description>&lt;p&gt;I recently migrated &lt;a href="https://www.mpazari.com" rel="noopener noreferrer"&gt;mpazari.com&lt;/a&gt;, a Turkish motorcycle classifieds site, from MySQL 8 to PostgreSQL 12.&lt;/p&gt;

&lt;p&gt;The application is a Spring Boot 4 / Java 25 GraalVM native image using hand-written SQL. The migration involved moving the existing database, adapting the MySQL-specific queries, and testing the application against a real production snapshot.&lt;/p&gt;

&lt;p&gt;The migration was completed with approximately &lt;strong&gt;three minutes of downtime&lt;/strong&gt;. The main steps were:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Create the PostgreSQL schema.&lt;/li&gt;
&lt;li&gt;Copy the MySQL data into PostgreSQL.&lt;/li&gt;
&lt;li&gt;Update the application queries for PostgreSQL.&lt;/li&gt;
&lt;li&gt;Run the test suite and verify the row counts.&lt;/li&gt;
&lt;li&gt;Stop the application, perform the final copy, switch the JDBC URL, and restart it.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;PostgreSQL’s stricter typing exposed a few old data and SQL issues that MySQL had silently accepted. The test suite helped find and fix those issues before the production cutover.&lt;/p&gt;

&lt;h2&gt;
  
  
  The unexpected result
&lt;/h2&gt;

&lt;p&gt;The most interesting result was the server load.&lt;/p&gt;

&lt;p&gt;With MySQL, &lt;code&gt;top&lt;/code&gt; commonly showed a load average of around &lt;strong&gt;1.5&lt;/strong&gt;. After switching to PostgreSQL, the same server and application workload usually showed a load average around &lt;strong&gt;0.3-0.5&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That is a significant improvement. The server feels much lighter, even though the database, application, indexes, and hardware are essentially the same.&lt;/p&gt;

&lt;p&gt;This is not a formal benchmark. I did not run a controlled performance test or collect latency percentiles before and after the migration. It is simply a real production observation from the server while handling the same site traffic.&lt;/p&gt;

&lt;p&gt;The database is relatively small, at about 131 MB, and both databases had the same application workload and indexes. So I cannot claim that PostgreSQL is universally faster than MySQL. For this particular application, however, PostgreSQL uses noticeably less server capacity.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;The migration reduced downtime and produced a much calmer server:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;MySQL: approximately &lt;strong&gt;1.5 load average&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;PostgreSQL: approximately &lt;strong&gt;0.3-0.5 load average&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The biggest lesson is simple: even when two databases support the same application, their behavior under a real workload can be surprisingly different.&lt;/p&gt;

</description>
      <category>backend</category>
      <category>database</category>
      <category>postgres</category>
      <category>sql</category>
    </item>
    <item>
      <title>Playwright Resource Monitor: making CI fail when your browser tabs burn CPU</title>
      <dc:creator>özkan pakdil</dc:creator>
      <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
      <link>https://dev.to/ozkanpakdil/playwright-resource-monitor-making-ci-fail-when-your-browser-tabs-burn-cpu-n8i</link>
      <guid>https://dev.to/ozkanpakdil/playwright-resource-monitor-making-ci-fail-when-your-browser-tabs-burn-cpu-n8i</guid>
      <description>&lt;p&gt;Green build, slow pages. That combination has annoyed me for years: Playwright happily reports 136/136 passing while the browser quietly burns 90% of one core on a badly-behaved tab, or the whole runner sits at 99% memory and everything gets &lt;em&gt;suspiciously&lt;/em&gt; slow. Test outcomes say nothing about &lt;strong&gt;how much resource the tests consumed&lt;/strong&gt;. Resource waste in CI is exactly the kind of thing you want to fail loudly, not discover from user complaints later.&lt;/p&gt;

&lt;p&gt;So I finally built it: &lt;a href="https://github.com/ozkanpakdil/github-playwright-monitor" rel="noopener noreferrer"&gt;playwright-resource-monitor&lt;/a&gt;, a GitHub Action that wraps your Playwright run, watches machine-wide CPU/memory &lt;em&gt;and&lt;/em&gt; the worst single browser tab, fails the build when configurable thresholds are crossed, and records every run’s peaks in the job summary plus a cross-run history.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it watches
&lt;/h2&gt;

&lt;p&gt;Two layers, four thresholds, each with its own meaning:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;What it measures&lt;/th&gt;
&lt;th&gt;Default&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Machine&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;CPU % across &lt;strong&gt;all cores combined&lt;/strong&gt; , memory % of the effective RAM limit&lt;/td&gt;
&lt;td&gt;70 / 70&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Tab&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;The &lt;strong&gt;worst single tab&lt;/strong&gt; (renderer process): CPU as % of &lt;strong&gt;one&lt;/strong&gt; core (yes, it can exceed 100%), memory % of the RAM limit&lt;/td&gt;
&lt;td&gt;70 / 70&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The per-tab layer is the one I actually care about. A machine-wide average on an 8-core runner looks serene even when one tab is melting. The average dilutes exactly the signal you want. “Worst single renderer process” is the meaningful question: &lt;em&gt;is any page we serve behaving badly?&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  How it sees your browser (the fun part)
&lt;/h2&gt;

&lt;p&gt;A GitHub Action is a separate process from your test run: it spawns &lt;code&gt;bun run test:e2e&lt;/code&gt; (or &lt;code&gt;npm run test&lt;/code&gt;) as a child and can’t reach inside Playwright to hook anything. That constraint shaped the whole design. On Linux, the answer is the good old process table:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;/proc/&amp;lt;pid&amp;gt;/cmdline&lt;/code&gt; finds Chromium processes and classifies them; &lt;code&gt;--type=renderer&lt;/code&gt; processes are &lt;strong&gt;tabs&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;/proc/&amp;lt;pid&amp;gt;/stat&lt;/code&gt; gives CPU time deltas between polls → the tab’s CPU consumed as &lt;strong&gt;% of one core&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;/proc/&amp;lt;pid&amp;gt;/status&lt;/code&gt; + a cgroup-aware limit check gives memory as &lt;strong&gt;% of RAM&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Zero configuration in your project. No reporter to add, no dependency, nothing. The action also polls for a Chrome DevTools Protocol port if you want exact per-tab labels, and it injects that port for you as an environment variable.&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;machine layer&lt;/strong&gt; is sampled by the action itself too (deltas from &lt;code&gt;/proc/stat&lt;/code&gt;, usage from &lt;code&gt;/proc/meminfo&lt;/code&gt;, respecting cgroup memory limits for containerized runners). And because the action owns this sampler, machine breaches are logged &lt;strong&gt;live during the run&lt;/strong&gt; as alert groups in the step log, not discovered after the fact:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;::start-group::RESOURCE ALERT #1 at 12:03:41 — worst tab: renderer pid 4242
::warning::Threshold breach: tab CPU 96.4% &amp;gt; 70% of one core
::endgroup::

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;blockquote&gt;
&lt;p&gt;Optional extra: if your project uses &lt;a href="https://github.com/cenfun/monocart-reporter" rel="noopener noreferrer"&gt;monocart-reporter&lt;/a&gt; with &lt;code&gt;json: true&lt;/code&gt;, the action automatically prefers its report and gets a machine-wide CPU/memory &lt;em&gt;timeline&lt;/em&gt; in an HTML report for free. But it’s optional. The built-in sampler covers enforcement on its own.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Using it
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;actions/checkout@v4&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;npm ci&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;npx playwright install chromium&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ozkanpakdil/github-playwright-monitor@v1&lt;/span&gt;
        &lt;span class="na"&gt;with&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;machine-cpu-threshold&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;80&lt;/span&gt; &lt;span class="c1"&gt;# % of all cores combined&lt;/span&gt;
          &lt;span class="na"&gt;machine-memory-threshold&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;70&lt;/span&gt; &lt;span class="c1"&gt;# % of effective RAM limit&lt;/span&gt;
          &lt;span class="na"&gt;tab-cpu-threshold&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;70&lt;/span&gt; &lt;span class="c1"&gt;# % of ONE core, worst single tab&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;actions/upload-artifact@v4&lt;/span&gt;
        &lt;span class="na"&gt;if&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;always()&lt;/span&gt;
        &lt;span class="na"&gt;with&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;resource-reports&lt;/span&gt;
          &lt;span class="na"&gt;path&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;resource-monitor/&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That’s the whole setup. Thresholds are strictly-greater-than, so 70 means “fail if any single tab holds more than 70% of one core”. Set per-tab high (like 300) if you only care about runaway tabs, and keep machine thresholds meaningful for your runner’s size.&lt;/p&gt;

&lt;p&gt;Every run appends a record to a history file, rendered into the job summary as a table (newest first, linked to each run):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"outcome"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"failed"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"machine"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"peakCpuPercent"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;66.3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"peakMemoryPercent"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;99.6&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"samples"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;45&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"tab"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"peakCpuPercent"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;84.1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"peakMemoryPercent"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;11.2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"samples"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;88&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"breached"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"machineMemory"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"tabCpu"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Runners are ephemeral, so keep the trend across runs with &lt;code&gt;actions/cache&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;actions/cache@v4&lt;/span&gt;
        &lt;span class="na"&gt;with&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;path&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;resource-monitor/history.json&lt;/span&gt;
          &lt;span class="na"&gt;key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;resource-history-${{ github.run_id }}&lt;/span&gt;
          &lt;span class="na"&gt;restore-keys&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;resource-history-&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The verdict rules
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;run-command fails&lt;/strong&gt; → the action fails; the test result is authoritative.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Any threshold breach with &lt;code&gt;fail-on-breach: true&lt;/code&gt;&lt;/strong&gt; (the default) → the action fails naming each breached layer with peaks vs thresholds. With &lt;code&gt;false&lt;/code&gt;, it only warns and stays green, which is handy for the first weeks on a new project while you calibrate.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Nothing to monitor&lt;/strong&gt; (no browser launched, no report) → the action stays &lt;em&gt;silent and green&lt;/em&gt;. Wrapping a command that happens to launch no browsers is not an error.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last one matters more than it sounds: the action just wraps a shell command, and “it did nothing weird to my run” is a hard requirement for a CI gate.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it taught me about thresholds
&lt;/h2&gt;

&lt;p&gt;I validated the whole loop on my own &lt;a href="https://github.com/ozkanpakdil/TrendCast" rel="noopener noreferrer"&gt;TrendCast&lt;/a&gt; repo (a browser extension with 136 Playwright tests). Machine memory read &lt;strong&gt;99.6% of RAM&lt;/strong&gt; and tripped the 70% threshold (on my laptop, not the runner). Machine memory % counts &lt;em&gt;the whole host&lt;/em&gt;, other apps included, and macOS especially reads high. On a dedicated GitHub runner it’s a healthy signal; on a shared machine, either size the machine thresholds for the runner or use &lt;code&gt;fail-on-breach: false&lt;/code&gt; locally. Per-tab numbers don’t have this problem: a tab’s RSS is the tab’s.&lt;/p&gt;

&lt;h2&gt;
  
  
  Testing before “publishing”
&lt;/h2&gt;

&lt;p&gt;One more thing I care about: you can consume the action by repo reference (&lt;code&gt;owner/repo@main&lt;/code&gt;) with &lt;strong&gt;no release required&lt;/strong&gt;. The Marketplace “publish” button only adds the listing (and asks for a tag when you’re ready). The repo’s CI runs a two-leg smoke matrix on every push: a “native” leg (no extra dependencies, asserting the built-in machine sampler and the &lt;code&gt;/proc&lt;/code&gt; tab scanner both produce data around a real browser test) and a “monocart” leg. So every push is a pre-release test of both monitoring paths. When it’s green, tag &lt;code&gt;v1&lt;/code&gt;, draft a release, tick &lt;em&gt;Publish to the GitHub Marketplace&lt;/em&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Limitations, honestly
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;The tab layer is Linux-only (process table). On macOS/Windows runners you’d need the CDP opt-in for tabs.&lt;/li&gt;
&lt;li&gt;Machine memory on shared hosts reads high (see above).&lt;/li&gt;
&lt;li&gt;Parallel CDP mode shares one debug port, so on Linux just use the &lt;code&gt;/proc&lt;/code&gt; scanner.&lt;/li&gt;
&lt;li&gt;Nothing to monitor → nothing enforced; the pass-through is intentional and quiet.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The code is small and readable: one sampler per data source, ~30 unit tests around the parsing and threshold math. Browse it at &lt;a href="https://github.com/ozkanpakdil/github-playwright-monitor" rel="noopener noreferrer"&gt;github.com/ozkanpakdil/github-playwright-monitor&lt;/a&gt;, or look at the &lt;a href="https://github.com/ozkanpakdil/TrendCast/blob/main/.github/workflows/ci.yml" rel="noopener noreferrer"&gt;TrendCast workflow&lt;/a&gt; for a real-world usage. If your green build has ever felt slower than it should, you now have a way to catch the culprit per-tab ,and to make the build fail loudly when a page burns your CI budget.&lt;/p&gt;

</description>
      <category>cicd</category>
      <category>monitoring</category>
      <category>performance</category>
      <category>testing</category>
    </item>
    <item>
      <title>How to add NVIDIA free models to VS Code</title>
      <dc:creator>özkan pakdil</dc:creator>
      <pubDate>Sun, 19 Jul 2026 00:00:00 +0000</pubDate>
      <link>https://dev.to/ozkanpakdil/how-to-add-nvidia-free-models-to-vs-code-3c01</link>
      <guid>https://dev.to/ozkanpakdil/how-to-add-nvidia-free-models-to-vs-code-3c01</guid>
      <description>&lt;p&gt;NVIDIA offers free access to their powerful Nemotron models through their &lt;a href="https://build.nvidia.com/" rel="noopener noreferrer"&gt;NVIDIA Build platform&lt;/a&gt;. In this post, I'll walk you through how to get a free API key and configure it in VS Code to use models like &lt;code&gt;nvidia/nemotron-3-ultra-550b-a55b&lt;/code&gt; directly in VS Code. Also you can choose any other model from &lt;a href="https://build.nvidia.com/models?orderBy=weightPopular%3ADESC&amp;amp;filters=nimType%3Anim_type_preview" rel="noopener noreferrer"&gt;here&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Prerequisites
&lt;/h3&gt;

&lt;p&gt;VScode needs to be above 1.27 it needs to have customendpoint supporting&lt;/p&gt;

&lt;h2&gt;
  
  
  Getting Your Free NVIDIA API Key
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Go to &lt;a href="https://build.nvidia.com/" rel="noopener noreferrer"&gt;build.nvidia.com&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Sign in or create an NVIDIA account&lt;/li&gt;
&lt;li&gt;Navigate to the model you want to use (e.g., &lt;a href="https://build.nvidia.com/nvidia/nemotron-3-ultra" rel="noopener noreferrer"&gt;Nemotron 3 Ultra&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;Click "Get API Key" button&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Note&lt;/strong&gt;: In my experience, the "Get API Key" button in the UI didn't work properly (it appeared to do nothing). If this happens to you, here's how to get your API key:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Open your browser's Developer Tools (F12) → Network tab&lt;/li&gt;
&lt;li&gt;Click the "Get API Key" button again&lt;/li&gt;
&lt;li&gt;Look for a POST request to &lt;code&gt;https://api.nvcf.nvidia.com/v2/nvcf/api-keys&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Click on that request and check the &lt;strong&gt;Response&lt;/strong&gt; tab - you'll find your API key there&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The API call looks like this (for reference - you don't need to run this manually):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# This is what the UI calls behind the scenes - for reference only&lt;/span&gt;
curl &lt;span class="s1"&gt;'https://api.nvcf.nvidia.com/v2/nvcf/api-keys'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s1"&gt;'Accept: */*'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s1"&gt;'Accept-Language: en-GB,en;q=0.9'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s1"&gt;'Accept-Encoding: gzip, deflate, br, zstd'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s1"&gt;'Referer: https://build.nvidia.com/'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s1"&gt;'content-type: application/json'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s1"&gt;'Origin: https://build.nvidia.com'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s1"&gt;'Connection: keep-alive'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s1"&gt;'Cookie: ...'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s1"&gt;'Sec-Fetch-Dest: empty'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s1"&gt;'Sec-Fetch-Mode: cors'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s1"&gt;'Sec-Fetch-Site: same-site'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--data-raw&lt;/span&gt; &lt;span class="s1"&gt;'{"name":"","type":"AI_PLAYGROUNDS_KEY","expiryDate":"2027-01-19T11:37:44Z","policies":[{"product":"nv-cloud-functions","scopes":["invoke_function"],"resources":[{"id":"*","type":"account-functions"}]}]}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The response will contain your API key:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"apiKey"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"keyId"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"1a9e12d7-b35f-4ed0-943c-3935c442a5c9"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"NVIDIABuild-Autogen-54"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"AI_PLAYGROUNDS_KEY"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"value"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"nvapi-..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"ownerId"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"status"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"ACTIVE"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"expiryDate"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2027-01-19T11:37:44.000Z"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"createdDate"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2026-07-19T11:39:27.197Z"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"policies"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"product"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"nv-cloud-functions"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"productDisplayName"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"NVIDIA Cloud Functions"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"resources"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
            &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"*"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
            &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"account-functions"&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
            &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"*"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
            &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"authorized-functions"&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"roles"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[],&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"scopes"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="s2"&gt;"invoke_function"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="s2"&gt;"list_functions"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="s2"&gt;"list_functions_details"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="s2"&gt;"queue_details"&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"product"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"ai-foundations"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"productDisplayName"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Public API Endpoints"&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"product"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"artifact-catalog"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"productDisplayName"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"NGC Catalog"&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"product"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"private-registry"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"productDisplayName"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"NVIDIA Private Registry"&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Important&lt;/strong&gt;: Save the &lt;code&gt;value&lt;/code&gt; field (the &lt;code&gt;nvapi-...&lt;/code&gt; key) - you'll need it for VS Code configuration.&lt;/p&gt;

&lt;h2&gt;
  
  
  Configuring VS Code
&lt;/h2&gt;

&lt;p&gt;The &lt;a href="https://code.visualstudio.com/docs/agent-customization/language-models" rel="noopener noreferrer"&gt;recommended way&lt;/a&gt; is to use the VS Code UI to add the custom model, which will properly encrypt and store your API key. &lt;strong&gt;Custom language models are configured in a separate file called &lt;code&gt;chatLanguageModels.json&lt;/code&gt;&lt;/strong&gt; (not in &lt;code&gt;settings.json&lt;/code&gt; under &lt;code&gt;chat.mcp.servers&lt;/code&gt; - that's for MCP servers, which is a different feature).&lt;/p&gt;

&lt;h3&gt;
  
  
  Using VS Code UI
&lt;/h3&gt;

&lt;p&gt;Check &lt;a href="https://youtu.be/VBSVSxs16_I?t=64" rel="noopener noreferrer"&gt;the video&lt;/a&gt; if you have not done before. Below are the steps&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Open VS Code Command Palette (Cmd/Ctrl + Shift + P)&lt;/li&gt;
&lt;li&gt;Search for "Chat: Manage Language Models" &lt;a href="https://code.visualstudio.com/docs/agent-customization/language-models#_manage-language-models" rel="noopener noreferrer"&gt;here the docs&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Click "Add Model" or "Add Custom Model"&lt;/li&gt;
&lt;li&gt;Fill in the details:

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Name&lt;/strong&gt;: &lt;code&gt;Nvidia&lt;/code&gt; (or any name you prefer)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Vendor&lt;/strong&gt;: &lt;code&gt;customendpoint&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;API Key&lt;/strong&gt;: Enter your &lt;code&gt;nvapi-...&lt;/code&gt; key (VS Code will encrypt and store it securely)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Model ID&lt;/strong&gt;: &lt;code&gt;nvidia/nemotron-3-ultra-550b-a55b&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Model Name&lt;/strong&gt;: &lt;code&gt;nvidia/nemotron-3-ultra-550b-a55b&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;URL&lt;/strong&gt;: &lt;code&gt;https://integrate.api.nvidia.com/v1/chat/completions&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tool Calling&lt;/strong&gt;: ✅ Enabled&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Max Input Tokens&lt;/strong&gt;: 262144&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Max Output Tokens&lt;/strong&gt;: 32000&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;in &lt;code&gt;chatLanguageModels.json&lt;/code&gt; will have json like below&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Nvidia"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"vendor"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"customendpoint"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"apiKey"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"${input:chat.lm.secret.6878c97f}"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"models"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
            &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
                &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"nvidia/nemotron-3-ultra-550b-a55b"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
                &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"nvidia/nemotron-3-ultra-550b-a55b"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
                &lt;/span&gt;&lt;span class="nl"&gt;"url"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://integrate.api.nvidia.com/v1/chat/completions"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
                &lt;/span&gt;&lt;span class="nl"&gt;"toolCalling"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
                &lt;/span&gt;&lt;span class="nl"&gt;"vision"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
                &lt;/span&gt;&lt;span class="nl"&gt;"maxInputTokens"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;262144&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
                &lt;/span&gt;&lt;span class="nl"&gt;"maxOutputTokens"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;32000&lt;/span&gt;&lt;span class="w"&gt;
            &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Using the Model in VS Code
&lt;/h2&gt;

&lt;p&gt;Once configured, you can use the model in VS Code Chat:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Open the Chat view (Ctrl+Alt+I / Cmd+Option+I)&lt;/li&gt;
&lt;li&gt;Click the model selector dropdown&lt;/li&gt;
&lt;li&gt;Select &lt;strong&gt;"Nvidia"&lt;/strong&gt; → &lt;strong&gt;"nvidia/nemotron-3-ultra-550b-a55b"&lt;/strong&gt; (or whatever name you gave the model)&lt;/li&gt;
&lt;li&gt;Start chatting!&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The model supports:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Tool calling&lt;/strong&gt; - Can use tools and functions&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Vision&lt;/strong&gt; - Can analyze images&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;128K context window&lt;/strong&gt; - Large context for long conversations&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;16K output tokens&lt;/strong&gt; - Long responses&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Model Details
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Property&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Model ID&lt;/td&gt;
&lt;td&gt;&lt;code&gt;nvidia/nemotron-3-ultra-550b-a55b&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Endpoint&lt;/td&gt;
&lt;td&gt;&lt;code&gt;https://integrate.api.nvidia.com/v1/chat/completions&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Context Window&lt;/td&gt;
&lt;td&gt;128,000 tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Max Output&lt;/td&gt;
&lt;td&gt;16,000 tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tool Calling&lt;/td&gt;
&lt;td&gt;✅ Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vision&lt;/td&gt;
&lt;td&gt;✅ Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Free Tier&lt;/td&gt;
&lt;td&gt;✅ Yes (with NVIDIA account)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Tips
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Rate Limits&lt;/strong&gt;: The free tier has rate limits. Check &lt;a href="https://build.nvidia.com/" rel="noopener noreferrer"&gt;NVIDIA's documentation&lt;/a&gt; for current limits.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;API Key Expiry&lt;/strong&gt;: The key expires on the date shown in the response (e.g., 2027-01-19). You'll need to generate a new one when it expires.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Multiple Models&lt;/strong&gt;: You can add multiple NVIDIA models to the same configuration by adding more entries to the &lt;code&gt;models&lt;/code&gt; array.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;VS Code Insiders&lt;/strong&gt;: If you're on VS Code Insiders, the settings path might be slightly different.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Resources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://build.nvidia.com/" rel="noopener noreferrer"&gt;NVIDIA Build Platform&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://build.nvidia.com/nvidia/nemotron-3-ultra" rel="noopener noreferrer"&gt;Nemotron 3 Ultra Model Card&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.nvidia.com/" rel="noopener noreferrer"&gt;NVIDIA API Documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://code.visualstudio.com/docs/copilot/custom-models" rel="noopener noreferrer"&gt;VS Code Copilot Custom Models Documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://build.nvidia.com/models?orderBy=weightPopular%3ADESC&amp;amp;filters=nimType%3Anim_type_preview" rel="noopener noreferrer"&gt;All free model from Nvidia&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Happy coding with NVIDIA's free models! 🚀&lt;/em&gt;&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Beyond Functionality: Why Code Hygiene is Your Project's Immune System</title>
      <dc:creator>özkan pakdil</dc:creator>
      <pubDate>Wed, 22 Apr 2026 00:00:00 +0000</pubDate>
      <link>https://dev.to/ozkanpakdil/beyond-functionality-why-code-hygiene-is-your-projects-immune-system-2lia</link>
      <guid>https://dev.to/ozkanpakdil/beyond-functionality-why-code-hygiene-is-your-projects-immune-system-2lia</guid>
      <description>&lt;p&gt;In the world of software engineering, we often obsess over the “Testing Pyramid.” We pour resources into unit tests, integration tests, and E2E suites. These are critical—they tell us that our features work as designed. But there’s a shadowy category of bugs that traditional tests often miss: the architectural “anti-patterns” and “API misuses” that don’t break functionality today but lead to system failures, memory leaks, or portability issues tomorrow.&lt;/p&gt;

&lt;p&gt;This is where &lt;strong&gt;Code Hygiene&lt;/strong&gt; comes in.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is Code Hygiene?
&lt;/h2&gt;

&lt;p&gt;Code Hygiene is the practice of using automated tools to enforce non-functional constraints, coding standards, and safety patterns across a codebase. Unlike a unit test that checks if &lt;code&gt;Add(1, 1) == 2&lt;/code&gt;, a hygiene scanner checks if you’re using an API in a way that’s technically “valid” but practically dangerous.&lt;/p&gt;

&lt;p&gt;In Computer Science, this is formally known as &lt;strong&gt;Static Program Analysis&lt;/strong&gt;. While general “linting” catches stylistic issues, Code Hygiene focuses on domain-specific safety and architectural integrity.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Evolution: From API Safety to Platform Portability
&lt;/h2&gt;

&lt;p&gt;One of the most powerful applications of code hygiene I’ve ever implemented involved managing a high-stakes support matrix. We had a system that executed various Linux commands across a wide range of distributions: Ubuntu, SUSE, and RedHat, all in different versions.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Problem: The Command Parameter Minefield
&lt;/h3&gt;

&lt;p&gt;As Linux distributions evolve, command-line parameters change. A flag that works on Ubuntu 20.04 might be deprecated on RedHat 9, or worse, behave subtly differently. Standard CI runs on a single OS wouldn’t catch these issues until a customer on a specific RedHat version hit a runtime error.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Solution: Static Analysis + Testcontainers
&lt;/h3&gt;

&lt;p&gt;To solve this, I developed a two-tier hygiene strategy:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;The Scanner:&lt;/strong&gt; We built a custom static analysis tool to crawl through the codebase and identify every single &lt;code&gt;exec&lt;/code&gt; or shell-out call.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The Matrix Validation:&lt;/strong&gt; Using &lt;strong&gt;Testcontainers with C#&lt;/strong&gt; , we spun up the exact versions of Ubuntu, SUSE, and RedHat defined in our support matrix. We then fed the commands discovered by our scanner into these containers to verify their exit codes and behavior.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This “Hygiene for Portability” transformed our release process. We stopped guessing if our commands were compatible and started &lt;em&gt;knowing&lt;/em&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Case Study: The “Order of Operations” Leak
&lt;/h2&gt;

&lt;p&gt;Hygiene isn’t just for external commands; it’s for internal APIs too. Recently, I dealt with a scenario where a database library had a specific requirement: &lt;code&gt;BeginTransaction()&lt;/code&gt; had to be called &lt;em&gt;after&lt;/em&gt; certain context configurations, but &lt;em&gt;before&lt;/em&gt; others.&lt;/p&gt;

&lt;p&gt;Syntactically, the code &lt;code&gt;db.WithContext(ctx).Begin()&lt;/code&gt; looked fine. It compiled. It even passed unit tests. However, in tests, it caused intermittent memory leaks because the context wasn’t being associated with the transaction object correctly.&lt;/p&gt;

&lt;p&gt;A simple &lt;strong&gt;AST (Abstract Syntax Tree)&lt;/strong&gt; scanner fixed this permanently. We wrote a rule that flags any instance where these calls are out of order. Institutional knowledge was turned into an automated gatekeeper.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why You Need a Hygiene Strategy
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Preventing “Silent” Failures:&lt;/strong&gt; Catch memory leaks, race conditions, and API misuses that don’t trigger a test failure but kill performance.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Institutional Memory:&lt;/strong&gt; When a team learns a hard lesson, a hygiene rule ensures that new developers (or your future self) don’t repeat the mistake.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Scalable Reviews:&lt;/strong&gt; Humans are bad at spotting subtle pattern errors in 1,000-line PRs. Computers are perfect at it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Platform Confidence:&lt;/strong&gt; In a world of multi-arch and multi-OS support, scanners can validate that your code respects the constraints of every target environment.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Testing tells you the code is &lt;strong&gt;right&lt;/strong&gt;. Hygiene tells you the code is &lt;strong&gt;healthy&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;By integrating custom static analysis and matrix-based validation into your workflow, you move from a reactive “hotfix” stance to a proactive “preventative” culture. It’s time to look beyond functionality and start caring about the structural health of your codebase.&lt;/p&gt;

</description>
      <category>architecture</category>
      <category>coding</category>
      <category>softwareengineering</category>
      <category>testing</category>
    </item>
    <item>
      <title>Accelerating LLMs on Debian 13: Setting up Vulkan for llama.cpp</title>
      <dc:creator>özkan pakdil</dc:creator>
      <pubDate>Sun, 22 Mar 2026 06:47:00 +0000</pubDate>
      <link>https://dev.to/ozkanpakdil/accelerating-llms-on-debian-13-setting-up-vulkan-for-llamacpp-f10</link>
      <guid>https://dev.to/ozkanpakdil/accelerating-llms-on-debian-13-setting-up-vulkan-for-llamacpp-f10</guid>
      <description>&lt;p&gt;After setting up CUDA on my other laptop, I moved to a different(older) machine that doesn’t have an NVIDIA GPU. This one is an everyday laptop with integrated Intel graphics, but that doesn’t mean we have to settle for slow CPU-only performance.&lt;/p&gt;

&lt;p&gt;On this machine, I switched to the &lt;strong&gt;Vulkan&lt;/strong&gt; backend for &lt;code&gt;llama.cpp&lt;/code&gt; and the results were even more dramatic than I expected.&lt;/p&gt;

&lt;h3&gt;
  
  
  Machine Hardware Info
&lt;/h3&gt;

&lt;p&gt;This laptop is running &lt;strong&gt;Debian 13 (Trixie/Sid)&lt;/strong&gt; with the following specs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;CPU&lt;/strong&gt; : Intel(R) Core(TM) i5-8250U @ 1.60GHz (4 Cores, 8 Threads)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;GPU&lt;/strong&gt; : Intel(R) UHD Graphics 620 (Integrated)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;RAM&lt;/strong&gt; : 8 GB&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;OS&lt;/strong&gt; : Debian GNU/Linux 13 (trixie)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Kernel&lt;/strong&gt; : 6.12.74+deb13+1-amd64&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  The Performance Gap: CPU vs. Vulkan
&lt;/h3&gt;

&lt;p&gt;I tested both the &lt;strong&gt;Qwen 3.5 2B&lt;/strong&gt; and the more capable &lt;strong&gt;Qwen 2.5 3B&lt;/strong&gt; models (GGUF format) to see how the integrated Intel GPU handles different LLM sizes.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Setup&lt;/th&gt;
&lt;th&gt;Response Time (Eval)&lt;/th&gt;
&lt;th&gt;Total Time&lt;/th&gt;
&lt;th&gt;Tokens/sec&lt;/th&gt;
&lt;th&gt;Notes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Qwen 3.5 2B&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;CPU Only&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;~7 minutes&lt;/strong&gt; (428s)&lt;/td&gt;
&lt;td&gt;431s&lt;/td&gt;
&lt;td&gt;2.32&lt;/td&gt;
&lt;td&gt;Purely on i5-8250U&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Qwen 3.5 2B&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Vulkan&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;14 seconds&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;21s&lt;/td&gt;
&lt;td&gt;6.07&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;30x improvement!&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Qwen 2.5 3B&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Vulkan&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;47 seconds&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;52s&lt;/td&gt;
&lt;td&gt;4.54&lt;/td&gt;
&lt;td&gt;More capable reasoning&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The &lt;strong&gt;14-second response&lt;/strong&gt; vs. &lt;strong&gt;7 minutes&lt;/strong&gt; on the 2B model is a game-changer, but the 3B model (answering “write me hello world in rust” in &lt;strong&gt;47 seconds&lt;/strong&gt; ) is the “sweet spot” for this machine. While the 2B model can be fully offloaded, the 3B model is too large to fit entirely in the GPU’s shared memory, but it still performs admirably.&lt;/p&gt;

&lt;h3&gt;
  
  
  Compiling llama.cpp with Vulkan
&lt;/h3&gt;

&lt;p&gt;Compiling for Vulkan on Debian is straightforward but requires the right development headers. It took me about &lt;strong&gt;10 minutes&lt;/strong&gt; to finish the compilation.&lt;/p&gt;

&lt;p&gt;First, ensure you have the Vulkan development packages:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;apt update
&lt;span class="nb"&gt;sudo &lt;/span&gt;apt &lt;span class="nb"&gt;install &lt;/span&gt;libvulkan-dev vulkan-tools

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then, compile &lt;code&gt;llama.cpp&lt;/code&gt; using CMake with the Vulkan option enabled:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight cmake"&gt;&lt;code&gt;cmake -B build -DGGML_VULKAN=ON
cmake --build build --config Release

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Running with Vulkan Acceleration
&lt;/h3&gt;

&lt;p&gt;Once compiled, you can run &lt;code&gt;llama-server&lt;/code&gt; (or &lt;code&gt;llama-cli&lt;/code&gt;). The server will automatically detect your Vulkan-compatible devices.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;./build/bin/llama-server &lt;span class="nt"&gt;-hf&lt;/span&gt; unsloth/Qwen3.5-2B-GGUF &lt;span class="nt"&gt;--jinja&lt;/span&gt; &lt;span class="nt"&gt;-c&lt;/span&gt; 4096 &lt;span class="nt"&gt;--host&lt;/span&gt; 127.0.0.1 &lt;span class="nt"&gt;--port&lt;/span&gt; 8033

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In the logs, you’ll see it picking up the Intel UHD Graphics. For the 2B model, I was able to offload all 25 layers:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ggml_vulkan: Found 1 Vulkan devices:
ggml_vulkan: 0 = Intel(R) UHD Graphics 620 (KBL GT2) (Intel open-source Mesa driver) | uma: 1 | fp16: 1 | ...
...
load_tensors: offloading 23 repeating layers to GPU
load_tensors: offloaded 25/25 layers to GPU

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Pushing the Limits: Qwen 2.5 3B
&lt;/h3&gt;

&lt;p&gt;When I moved to the &lt;strong&gt;3.4B parameter&lt;/strong&gt; model (&lt;code&gt;Qwen2.5-3B-Instruct-Q4_K_M&lt;/code&gt;), the memory management became more complex. The system had to balance between the GPU’s shared memory and the CPU:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;llama_params_fit_impl: projected to use 2283 MiB of device memory vs. 2200 MiB of free device memory
llama_params_fit_impl: cannot meet free memory target of 1024 MiB, need to reduce device memory by 1106 MiB
llama_params_fit_impl: filling dense layers back-to-front:
llama_params_fit_impl: - Vulkan0 (Intel(R) UHD Graphics 620 (KBL GT2)): 13 layers, 1137 MiB used, 1062 MiB free
...
load_tensors: offloaded 13/37 layers to GPU

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Even though it only offloaded &lt;strong&gt;13/37 layers&lt;/strong&gt; to the GPU, it still maintained a respectable &lt;strong&gt;4.54 tokens/sec&lt;/strong&gt;. This shows that even partial offloading on integrated graphics provides a significant boost over pure CPU execution for larger models.&lt;/p&gt;

&lt;h3&gt;
  
  
  Summary &amp;amp; Next Steps
&lt;/h3&gt;

&lt;p&gt;If you don’t have an NVIDIA card, don’t ignore your integrated GPU. Vulkan provides a fantastic alternative that works out-of-the-box on Debian with Intel and AMD hardware.&lt;/p&gt;

&lt;p&gt;My next target is to use &lt;strong&gt;Qwen on OpenClaw&lt;/strong&gt; to further explore local LLM capabilities. Stay tuned!&lt;/p&gt;

</description>
      <category>llama</category>
    </item>
    <item>
      <title>Accelerating LLMs on Debian 13: Setting up CUDA for llama.cpp</title>
      <dc:creator>özkan pakdil</dc:creator>
      <pubDate>Thu, 19 Mar 2026 22:35:00 +0000</pubDate>
      <link>https://dev.to/ozkanpakdil/accelerating-llms-on-debian-13-setting-up-cuda-for-llamacpp-22lb</link>
      <guid>https://dev.to/ozkanpakdil/accelerating-llms-on-debian-13-setting-up-cuda-for-llamacpp-22lb</guid>
      <description>&lt;p&gt;Setting up NVIDIA CUDA on Debian 13 (Trixie/Sid) to run Large Language Models (LLMs) can be a bit of a journey, especially if you’re transitioning from the default open-source drivers to the proprietary stack required for GPGPU workloads.&lt;/p&gt;

&lt;p&gt;Over the last few days, I’ve been working on getting &lt;code&gt;llama.cpp&lt;/code&gt; to run with CUDA on my laptop to see how much of a difference it makes compared to pure CPU execution.&lt;/p&gt;

&lt;p&gt;Initially, I tested a &lt;strong&gt;35B model on macOS&lt;/strong&gt; , where it was responding in about &lt;strong&gt;17 seconds&lt;/strong&gt;. When I moved that same 35B model to my old laptop running Debian 13 (on CPU), the response time plummeted to &lt;strong&gt;4 minutes and 30 seconds&lt;/strong&gt;. This massive gap was my main motivation to try and enable CUDA on the laptop’s GPU.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Problem: Nouveau vs. Proprietary Drivers
&lt;/h3&gt;

&lt;p&gt;By default, Debian might use the open-source &lt;code&gt;nouveau&lt;/code&gt; driver. While great for basic display tasks, it doesn’t support CUDA. To run &lt;code&gt;llama-server&lt;/code&gt; with GPU acceleration, you need the official NVIDIA drivers and the CUDA toolkit.&lt;/p&gt;

&lt;p&gt;I followed the &lt;a href="https://docs.nvidia.com/datacenter/tesla/driver-installation-guide/debian.html" rel="noopener noreferrer"&gt;NVIDIA Tesla Driver Installation Guide for Debian&lt;/a&gt;, which is a critical resource for getting the right packages.&lt;/p&gt;

&lt;p&gt;One specific hurdle with Secure Boot enabled was the need to trust the DKMS-generated keys:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;mokutil &lt;span class="nt"&gt;--import&lt;/span&gt; /var/lib/dkms/mok.pub

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After a reboot and enrolling the key in the MOK manager, the driver was finally active and recognized by the system.&lt;/p&gt;

&lt;h3&gt;
  
  
  Compiling llama.cpp with CUDA Support
&lt;/h3&gt;

&lt;p&gt;Once the drivers and &lt;code&gt;nvcc&lt;/code&gt; were ready, I recompiled &lt;code&gt;llama.cpp&lt;/code&gt; with CUDA enabled (see the &lt;a href="https://github.com/ggml-org/llama.cpp/blob/master/docs/build.md#cuda" rel="noopener noreferrer"&gt;official CUDA build documentation&lt;/a&gt; for more details).&lt;/p&gt;

&lt;p&gt;The compilation process is quite resource-intensive and took about &lt;strong&gt;15 minutes&lt;/strong&gt; on my laptop:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F0gz1nn2wivunyjusccbs.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F0gz1nn2wivunyjusccbs.png" alt="llama.cpp compilation" width="800" height="369"&gt;&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;CUDACXX&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;/usr/local/cuda/bin/nvcc
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;CUDA_HOME&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;/usr/local/cuda
cmake &lt;span class="nt"&gt;-B&lt;/span&gt; build &lt;span class="nt"&gt;-DGGML_CUDA&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;ON
cmake &lt;span class="nt"&gt;--build&lt;/span&gt; build &lt;span class="nt"&gt;--config&lt;/span&gt; Release

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  The VRAM Reality Check (OOM Errors)
&lt;/h3&gt;

&lt;p&gt;My laptop has an &lt;strong&gt;NVIDIA GeForce MX450&lt;/strong&gt; with &lt;strong&gt;2 GB of VRAM&lt;/strong&gt;. This is quite modest for modern LLMs.&lt;/p&gt;

&lt;p&gt;Initially, I tried running that 35B model that was so slow on the CPU:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;llama-server &lt;span class="nt"&gt;-hf&lt;/span&gt; unsloth/Qwen3.5-35B-A3B-GGUF &lt;span class="nt"&gt;--jinja&lt;/span&gt; &lt;span class="nt"&gt;-c&lt;/span&gt; 16384 &lt;span class="nt"&gt;--host&lt;/span&gt; 127.0.0.1 &lt;span class="nt"&gt;--port&lt;/span&gt; 8033

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It failed with a &lt;code&gt;cudaMalloc failed: out of memory&lt;/code&gt; error. Looking at the logs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Model parameters: ~857 MiB&lt;/li&gt;
&lt;li&gt;Context/CLIP buffers: ~899 MiB&lt;/li&gt;
&lt;li&gt;Total requested: &amp;gt; 1.7 GB&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;With the OS and display driver already taking up some of that 2 GB, there just wasn’t enough room. The 35B model was simply too large for this specific hardware’s VRAM. Even though CUDA would have been faster than the CPU-only 4.5 minutes, the hardware limit forced me to pivot.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Result: 2B Model Benchmark
&lt;/h3&gt;

&lt;p&gt;I switched to a smaller 2B model to stay within the VRAM limits. The results were impressive and clearly showed why we go through this trouble.&lt;/p&gt;

&lt;p&gt;Asking Qwen 2B to &lt;strong&gt;“write me hello world in rust”&lt;/strong&gt; :&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Setup&lt;/th&gt;
&lt;th&gt;Time to Complete&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;CPU Only&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1 minute 32 seconds&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;CUDA (GPU)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;24 seconds&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That’s nearly a &lt;strong&gt;4x speed improvement&lt;/strong&gt; on a entry-level mobile GPU!&lt;/p&gt;

&lt;h3&gt;
  
  
  Summary
&lt;/h3&gt;

&lt;p&gt;While the setup can be “so complicated” (dealing with drivers, Secure Boot, and compilation), the performance gains are undeniable. Even on a low-end GPU like the MX450, offloading the heavy lifting to CUDA makes the local LLM experience much more interactive.&lt;/p&gt;

&lt;h3&gt;
  
  
  Bonus: NVIDIA GPU Diagnostic Script
&lt;/h3&gt;

&lt;p&gt;To help troubleshoot my setup, I wrote a small script &lt;code&gt;nvidia_check_and_run.sh&lt;/code&gt; to verify the driver, kernel modules, and &lt;code&gt;llama.cpp&lt;/code&gt; support.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;#!/bin/bash&lt;/span&gt;

&lt;span class="c"&gt;# Configuration&lt;/span&gt;
&lt;span class="nv"&gt;LLAMA_PATH&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"/nix/store/wr7vi3957cx751la7q490h9v2m6q71fm-llama-cpp-8255/bin"&lt;/span&gt;
&lt;span class="nv"&gt;LLAMA_SERVER&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$LLAMA_PATH&lt;/span&gt;&lt;span class="s2"&gt;/llama-server"&lt;/span&gt;
&lt;span class="nv"&gt;LLAMA_BENCH&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$LLAMA_PATH&lt;/span&gt;&lt;span class="s2"&gt;/llama-bench"&lt;/span&gt;
&lt;span class="nv"&gt;LLAMA_CLI&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$LLAMA_PATH&lt;/span&gt;&lt;span class="s2"&gt;/llama-cli"&lt;/span&gt;

&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"--- NVIDIA GPU Diagnostic ---"&lt;/span&gt;

&lt;span class="c"&gt;# 1. Check for the NVIDIA device via PCI&lt;/span&gt;
&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"[1/4] Checking PCI devices for NVIDIA GPU..."&lt;/span&gt;
&lt;span class="k"&gt;if &lt;/span&gt;lspci | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-i&lt;/span&gt; nvidia&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
    &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;" - NVIDIA hardware detected via lspci."&lt;/span&gt;
&lt;span class="k"&gt;else
    &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;" - No NVIDIA hardware found on PCI bus."&lt;/span&gt;
&lt;span class="k"&gt;fi&lt;/span&gt;

&lt;span class="c"&gt;# 2. Check for the driver status&lt;/span&gt;
&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="nt"&gt;-e&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="s2"&gt;[2/4] Checking NVIDIA driver status..."&lt;/span&gt;
&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="nb"&gt;command&lt;/span&gt; &lt;span class="nt"&gt;-v&lt;/span&gt; nvidia-smi &amp;amp;&amp;gt; /dev/null&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
    &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;" - nvidia-smi found. Running..."&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="o"&gt;!&lt;/span&gt; nvidia-smi&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
        &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;" - CRITICAL: nvidia-smi failed. Kernel modules might not be loaded."&lt;/span&gt;
        &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;" - ACTION: Try running 'sudo modprobe nvidia' and then 'nvidia-smi' again."&lt;/span&gt;
    &lt;span class="k"&gt;fi
else
    &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;" - nvidia-smi NOT found. Driver might not be installed or active."&lt;/span&gt;
&lt;span class="k"&gt;fi&lt;/span&gt;

&lt;span class="c"&gt;# 3. Check for the kernel modules&lt;/span&gt;
&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="nt"&gt;-e&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="s2"&gt;[3/4] Checking for loaded NVIDIA kernel modules..."&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; /sbin/lsmod | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-i&lt;/span&gt; nvidia &amp;amp;&amp;gt; /dev/null&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
    &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;" - NVIDIA kernel modules are loaded."&lt;/span&gt;
    /sbin/lsmod | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-i&lt;/span&gt; nvidia
&lt;span class="k"&gt;else
    &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;" - CRITICAL: No NVIDIA kernel modules loaded."&lt;/span&gt;
    &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;" - ACTION: Run 'sudo modprobe nvidia' to load the driver."&lt;/span&gt;
&lt;span class="k"&gt;fi&lt;/span&gt;

&lt;span class="c"&gt;# 4. Check for llama.cpp device support&lt;/span&gt;
&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="nt"&gt;-e&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="s2"&gt;[4/4] Checking llama.cpp device support..."&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="o"&gt;[&lt;/span&gt;&lt;span class="nt"&gt;-f&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$LLAMA_CLI&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="o"&gt;]&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
    &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;" - Checking llama-cli with -ngl flag..."&lt;/span&gt;
    &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$LLAMA_CLI&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;-ngl&lt;/span&gt; 1 &lt;span class="nt"&gt;--version&lt;/span&gt;
&lt;span class="k"&gt;else
    &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;" - llama-cli not found at &lt;/span&gt;&lt;span class="nv"&gt;$LLAMA_CLI&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;span class="k"&gt;fi&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Running this script gave me a clear picture of what was missing:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="go"&gt;--- NVIDIA GPU Diagnostic ---
[1/4] Checking PCI devices for NVIDIA GPU...
0000:01:00.0 3D controller: NVIDIA Corporation TU117M [GeForce MX450] (rev a1)
  - NVIDIA hardware detected via lspci.

[2/4] Checking NVIDIA driver status...
  - nvidia-smi found. Running...
Fri Mar 20 01:55:51 2026       
+-----------------------------------------------------------------------------------------+
| NVIDIA-SMI 595.45.04 Driver Version: 595.45.04 CUDA Version: 13.2 |
+-----------------------------------------+------------------------+----------------------+
| GPU Name Persistence-M | Bus-Id Disp.A | Volatile Uncorr. ECC |
| Fan Temp Perf Pwr:Usage/Cap | Memory-Usage | GPU-Util Compute M. |
| | | MIG M. |
|=========================================+========================+======================|
| 0 NVIDIA GeForce MX450 On | 00000000:01:00.0 Off | N/A |
| N/A 53C P8 N/A / 5001W | 5MiB / 2048MiB | 0% Default |
| | | N/A |
+-----------------------------------------+------------------------+----------------------+

[3/4] Checking for loaded NVIDIA kernel modules...
  - NVIDIA kernel modules are loaded.
&lt;/span&gt;&lt;span class="c"&gt;...
&lt;/span&gt;&lt;span class="go"&gt;[4/4] Checking llama.cpp device support...
  - Checking llama-cli with -ngl flag...
warning: no usable GPU found, --gpu-layers option will be ignored
warning: one possible reason is that llama.cpp was compiled without GPU support
&lt;/span&gt;&lt;span class="c"&gt;...
&lt;/span&gt;&lt;span class="go"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you are on Debian 13 and want to try this, make sure you check your VRAM limits before picking a model, and don’t forget that &lt;code&gt;mokutil&lt;/code&gt; step if you have Secure Boot enabled!&lt;/p&gt;

</description>
      <category>ai</category>
      <category>linux</category>
      <category>llm</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>Tuning Podman on macOS to Match OrbStack Performance</title>
      <dc:creator>özkan pakdil</dc:creator>
      <pubDate>Sun, 08 Mar 2026 19:02:00 +0000</pubDate>
      <link>https://dev.to/ozkanpakdil/tuning-podman-on-macos-to-match-orbstack-performance-3f39</link>
      <guid>https://dev.to/ozkanpakdil/tuning-podman-on-macos-to-match-orbstack-performance-3f39</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Note:&lt;/strong&gt; These are suggested optimizations I have not personally tried yet. I’m blogging them as I’m planning to test them throughout this week.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;OrbStack is highly optimized for macOS, using a proprietary, high-performance networking stack and a custom VirtioFS implementation with aggressive caching. Podman, while being open-source and standard-compliant, can be tuned to significantly bridge the performance gap.&lt;/p&gt;

&lt;p&gt;The following plan outlines key areas where Podman’s performance can be improved on macOS:&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Enable Rosetta 2 for x86_64 Emulation (Apple Silicon only)
&lt;/h3&gt;

&lt;p&gt;If you are on an Apple Silicon (M1/M2/M3/M4) Mac, running x86_64 containers is often much slower than ARM64. Podman supports Apple’s native Rosetta 2 for Linux, which is substantially faster than QEMU-based emulation.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Check if Rosetta is enabled:&lt;/strong&gt; Run &lt;code&gt;podman machine inspect&lt;/code&gt; and look for &lt;code&gt;"Rosetta": true&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Enable Rosetta:&lt;/strong&gt; When creating a new machine, use the &lt;code&gt;--rosetta&lt;/code&gt; flag (if your Podman version and macOS version support it, typically macOS 13+):
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;podman machine init &lt;span class="nt"&gt;--rosetta&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Note: If you have an existing machine, you may need to recreate it to enable Rosetta.&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Optimize Volume Mounting (VirtioFS)
&lt;/h3&gt;

&lt;p&gt;Podman uses &lt;code&gt;virtiofs&lt;/code&gt; by default on macOS, which is the fastest way to share files between the host and the VM using Apple’s Virtualization.framework. However, file system I/O can still be a bottleneck.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Avoid Deeply Nested Mounts:&lt;/strong&gt; Minimize the number of files synced by mounting only the necessary sub-directories instead of the entire home directory.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Use Named Volumes:&lt;/strong&gt; For high-I/O workloads (like database storage or &lt;code&gt;node_modules&lt;/code&gt;), use named volumes instead of bind mounts. Named volumes reside within the VM’s disk image and operate at near-native speeds.
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;podman volume create my-data
podman run &lt;span class="nt"&gt;-v&lt;/span&gt; my-data:/app/data ...

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  3. Tuning Resource Allocation
&lt;/h3&gt;

&lt;p&gt;Ensure the Podman machine has sufficient resources. The default settings might be conservative for demanding workloads.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Increase CPUs and Memory:&lt;/strong&gt; Adjust the machine’s resources to match your workload.
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;podman machine &lt;span class="nb"&gt;set&lt;/span&gt; &lt;span class="nt"&gt;--cpus&lt;/span&gt; 4 &lt;span class="nt"&gt;--memory&lt;/span&gt; 8192

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;(Requires the machine to be stopped: &lt;code&gt;podman machine stop&lt;/code&gt;)&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Networking Performance (gvproxy)
&lt;/h3&gt;

&lt;p&gt;Podman uses &lt;code&gt;gvproxy&lt;/code&gt; for user-mode networking. This is often the primary reason OrbStack feels faster for network-heavy tasks, as OrbStack uses a more direct networking approach.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Reduce Network Hops:&lt;/strong&gt; If possible, avoid complex port mappings or heavy network traffic through the user-mode proxy.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MTU Tuning:&lt;/strong&gt; In some environments, increasing the MTU within the container can improve throughput, though this is dependent on the host’s network configuration.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  5. Experiment with the &lt;code&gt;libkrun&lt;/code&gt; Provider
&lt;/h3&gt;

&lt;p&gt;Podman on macOS supports multiple virtualization backends. While &lt;code&gt;applehv&lt;/code&gt; (default) is stable, &lt;code&gt;libkrun&lt;/code&gt; (based on &lt;code&gt;krun&lt;/code&gt;) can sometimes offer better performance for specific workloads, especially those involving GPU acceleration or specialized Virtio devices.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Try libkrun:&lt;/strong&gt; You can initialize a machine with the &lt;code&gt;libkrun&lt;/code&gt; provider:
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;podman machine init &lt;span class="nt"&gt;--provider&lt;/span&gt; libkrun

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  6. Summary of Recommended Configuration for Speed
&lt;/h3&gt;

&lt;p&gt;To get the best performance today, use the following initialization command (on Apple Silicon):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Stop and remove existing machine if necessary&lt;/span&gt;
podman machine stop
podman machine &lt;span class="nb"&gt;rm&lt;/span&gt;

&lt;span class="c"&gt;# Initialize with optimized settings&lt;/span&gt;
podman machine init &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--cpus&lt;/span&gt; 4 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--memory&lt;/span&gt; 8192 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--disk-size&lt;/span&gt; 50 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--rosetta&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--rootful&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;false&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;By applying these optimizations, Podman’s performance on macOS will be significantly closer to OrbStack, especially for CPU-intensive emulation and file-system heavy development workflows.&lt;/p&gt;

&lt;p&gt;Happy containerizing!&lt;/p&gt;

</description>
      <category>containers</category>
      <category>opensource</category>
      <category>performance</category>
      <category>tooling</category>
    </item>
    <item>
      <title>Atlassian MCP</title>
      <dc:creator>özkan pakdil</dc:creator>
      <pubDate>Sun, 08 Feb 2026 00:00:00 +0000</pubDate>
      <link>https://dev.to/ozkanpakdil/atlassian-mcp-229b</link>
      <guid>https://dev.to/ozkanpakdil/atlassian-mcp-229b</guid>
      <description>&lt;p&gt;I have been using &lt;a href="https://hub.docker.com/r/mcp/atlassian" rel="noopener noreferrer"&gt;Atlassian MCP&lt;/a&gt; with internal Confluence and Jira, and it has been wonderful.&lt;/p&gt;

&lt;p&gt;Finding internal information is often challenging and time-consuming. To be honest, searching through Jira or Confluence and locating the right information can be really difficult.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Create Jira and Confluence API tokens from your internal site profile page. For example: &lt;code&gt;https://internalconfluence.company.com/profile/personal&lt;/code&gt; for Confluence and &lt;code&gt;https://jira.company.com/secure/admin/CreateAPIToken!default.jspa&lt;/code&gt; for Jira. These URLs may vary depending on your setup.&lt;/li&gt;
&lt;li&gt;Create an &lt;code&gt;mcp.json&lt;/code&gt; file in the &lt;code&gt;.vscode&lt;/code&gt; folder for Visual Studio Code, or place this MCP configuration in the appropriate folder for your IDE of choice:
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"mcpServers"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"mcp-atlassian"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"command"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"uvx"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"args"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"mcp-atlassian"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"env"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"JIRA_URL"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://jira.company.com"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"JIRA_USERNAME"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"your.email@company.com"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"JIRA_API_TOKEN"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"your_api_token"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"CONFLUENCE_URL"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://internalconfluence.company.com/wiki"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"CONFLUENCE_USERNAME"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"your.email@company.com"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"CONFLUENCE_API_TOKEN"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"your_api_token"&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Remember to run podman-desktop or docker desktop. Because this MCP works as a docker container.&lt;/p&gt;

&lt;p&gt;After that, open GitHub Copilot in your IDE and instruct it to use the Atlassian MCP to search Confluence and Jira. This makes finding internal information incredibly easy—it goes through pages systematically and retrieves all the details you need.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>mcp</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Building a Real-time File I/O Heatmap with eBPF and Java 25</title>
      <dc:creator>özkan pakdil</dc:creator>
      <pubDate>Wed, 21 Jan 2026 05:52:00 +0000</pubDate>
      <link>https://dev.to/ozkanpakdil/building-a-real-time-file-io-heatmap-with-ebpf-and-java-25-2fnb</link>
      <guid>https://dev.to/ozkanpakdil/building-a-real-time-file-io-heatmap-with-ebpf-and-java-25-2fnb</guid>
      <description>&lt;p&gt;Have you ever wondered exactly which files are being hammered by your Linux system in real-time? While tools like &lt;code&gt;iotop&lt;/code&gt; or &lt;code&gt;lsof&lt;/code&gt; are great, sometimes you want something more visual, custom, and lightweight.&lt;/p&gt;

&lt;p&gt;In this post, I’ll walk you through how I built a &lt;strong&gt;Real-time File I/O Heatmap&lt;/strong&gt; using the power of &lt;strong&gt;eBPF&lt;/strong&gt; for data collection and &lt;strong&gt;Java 25&lt;/strong&gt; for a modern Terminal UI (TUI).&lt;/p&gt;

&lt;h3&gt;
  
  
  What is eBPF and Why Use It?
&lt;/h3&gt;

&lt;p&gt;eBPF (Extended Berkeley Packet Filter) is a revolutionary technology that allows you to run sandboxed programs in the Linux kernel without changing kernel source code or loading kernel modules.&lt;/p&gt;

&lt;p&gt;Think of it as &lt;strong&gt;JavaScript for the Kernel&lt;/strong&gt;. It allows you to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Observe&lt;/strong&gt; : Attach to almost any function in the kernel (kprobes) or userspace (uprobes).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Filter&lt;/strong&gt; : Process data efficiently at the source, inside the kernel.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Perform&lt;/strong&gt; : It’s extremely fast because it avoids expensive context switches between kernel and userspace for every event.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For this project, we use eBPF to hook into &lt;code&gt;vfs_read&lt;/code&gt; and &lt;code&gt;vfs_write&lt;/code&gt;, the gatekeepers of all filesystem activity in Linux.&lt;/p&gt;

&lt;p&gt;If you want to dive deeper into eBPF, I highly recommend reading &lt;a href="https://web.archive.org/web/20251129115431/https://cilium.isovalent.com/hubfs/Learning-eBPF%20-%20Full%20book.pdf" rel="noopener noreferrer"&gt;Learning eBPF by Liz Rice&lt;/a&gt;, it’s an excellent resource.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Architecture
&lt;/h3&gt;

&lt;p&gt;Our project consists of three main layers:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;The eBPF Program (C)&lt;/strong&gt;: Sits in the kernel, intercepts VFS calls, and aggregates stats (reads, writes, bytes) into a BPF Hash Map.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The Java Backend (Java 25 + JNA)&lt;/strong&gt;: Uses &lt;code&gt;libbpf&lt;/code&gt; via Java Native Access (JNA) to load the BPF program into the kernel and poll the maps.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The TUI (Lanterna)&lt;/strong&gt;: A Terminal User Interface that renders the data as a color-coded heatmap.&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  1. The Kernel Side: eBPF in C
&lt;/h3&gt;

&lt;p&gt;We use BPF CO-RE (Compile Once – Run Everywhere) to ensure our program works across different kernel versions without recompilation.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;SEC("kprobe/vfs_read")
int BPF_KPROBE(vfs_read, struct file *file, char *buf, size_t count) {
    // Extract filename from the file struct
    // Filter out non-file noise (sockets/pipes)
    // Update the BPF map with bytes read
    return 0;
}

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The magic happens in &lt;code&gt;file_heatmap.bpf.c&lt;/code&gt;, where we traverse the kernel’s &lt;code&gt;dentry&lt;/code&gt; structures to reconstruct partial file paths so we can actually see &lt;em&gt;what&lt;/em&gt; is being accessed.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. The Bridge: JNA and libbpf
&lt;/h3&gt;

&lt;p&gt;Interfacing Java with the kernel might sound scary, but &lt;code&gt;libbpf&lt;/code&gt; makes it manageable. We defined a JNA interface to map the C functions:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;public interface LibBpf extends Library {
    Pointer bpf_object__open(String path);
    int bpf_object__load(Pointer obj);
    int bpf_map_lookup_elem(int fd, Pointer key, Pointer value);
    // ...
}

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This allows our Java app to behave like a first-class Linux observability tool.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. The Frontend: A Terminal Heatmap
&lt;/h3&gt;

&lt;p&gt;Using the &lt;strong&gt;Lanterna&lt;/strong&gt; library, we created a TUI that updates every 2 seconds. The heatmap effect is achieved by calculating the “intensity” of I/O for each file and mapping it to a color gradient from white to red.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;float intensity = (float) currentVal / maxVal;
int green = (int) (255 * (1 - intensity));
int blue = (int) (255 * (1 - intensity));
tg.setBackgroundColor(new TextColor.RGB(255, green, blue));

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Challenges Overcome
&lt;/h3&gt;

&lt;p&gt;Building this wasn’t without its hurdles:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;SDKMAN &amp;amp; Sudo&lt;/strong&gt; : BPF requires root privileges, but &lt;code&gt;sudo&lt;/code&gt; often strips the environment variables (like &lt;code&gt;JAVA_HOME&lt;/code&gt;) set by SDKMAN. I solved this in the &lt;code&gt;Makefile&lt;/code&gt; by using absolute paths and passing environment variables explicitly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;BPF Verifier&lt;/strong&gt; : The kernel is very strict. Reconstructing file paths required careful use of &lt;code&gt;bpf_probe_read_kernel_str&lt;/code&gt; and &lt;code&gt;bpf_snprintf&lt;/code&gt; to keep the verifier happy.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;vmlinux.h Size&lt;/strong&gt; : The standard &lt;code&gt;vmlinux.h&lt;/code&gt; is over 2MB. I optimized this by using &lt;code&gt;bpftool gen min_core_btf&lt;/code&gt; to generate a &lt;strong&gt;minified&lt;/strong&gt; header (~2KB) containing only the types we actually use.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Noise Filtering&lt;/strong&gt; : Initially, the heatmap was flooded with TCP/UDP socket activity. Adding a filter for &lt;code&gt;S_IFREG&lt;/code&gt; (regular files) made the output much cleaner.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  How to Run It
&lt;/h3&gt;

&lt;p&gt;You can find the full source code for this project on GitHub: &lt;a href="https://github.com/ozkanpakdil/java-examlpes/tree/master/ebpf-file-heatmap" rel="noopener noreferrer"&gt;ozkanpakdil/java-examples/ebpf-file-heatmap&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If you’re on a Linux machine with &lt;code&gt;clang&lt;/code&gt;, &lt;code&gt;bpftool&lt;/code&gt;, and &lt;code&gt;maven&lt;/code&gt; installed, you can try it out:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;git clone https://github.com/ozkanpakdil/java-examples.git
cd java-examples/ebpf-file-heatmap
sudo make run

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Once it’s running, you can press &lt;strong&gt;1-5&lt;/strong&gt; to sort by different metrics (Reads, Writes, Bytes) and watch your system’s I/O come to life!&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/user-attachments/assets/8fee6e53-6a2a-41b9-aa70-33c2002c4dc2" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F27ibaxoiov6dug6cmbmg.png" alt="Image" width="800" height="680"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Conclusion
&lt;/h3&gt;

&lt;p&gt;Combining eBPF’s low-level performance with Java’s high-level productivity (and the latest features in Java 25!) is a powerful way to build Linux tooling. Whether you’re debugging a database or just curious about what your IDE is doing in the background, this heatmap gives you a unique window into your system.&lt;/p&gt;

&lt;p&gt;Happy hacking!&lt;/p&gt;

</description>
      <category>java</category>
      <category>linux</category>
      <category>monitoring</category>
      <category>showdev</category>
    </item>
    <item>
      <title>Bun Joins the Microservice Framework Benchmark: Surprisingly Fast JavaScript Runtime</title>
      <dc:creator>özkan pakdil</dc:creator>
      <pubDate>Sat, 10 Jan 2026 17:00:00 +0000</pubDate>
      <link>https://dev.to/ozkanpakdil/bun-joins-the-microservice-framework-benchmark-surprisingly-fast-javascript-runtime-119d</link>
      <guid>https://dev.to/ozkanpakdil/bun-joins-the-microservice-framework-benchmark-surprisingly-fast-javascript-runtime-119d</guid>
      <description>&lt;h2&gt;
  
  
  Introduction
&lt;/h2&gt;

&lt;p&gt;Today I’m excited to announce the addition of &lt;strong&gt;Bun&lt;/strong&gt; to our &lt;a href="https://ozkanpakdil.github.io/test-microservice-frameworks/" rel="noopener noreferrer"&gt;microservice framework benchmark suite&lt;/a&gt;. The results are nothing short of remarkable . Bun has proven to be one of the fastest runtimes in our entire test suite, competing directly with Rust frameworks!&lt;/p&gt;

&lt;h2&gt;
  
  
  What is Bun?
&lt;/h2&gt;

&lt;p&gt;Bun is a modern JavaScript runtime built from scratch using &lt;a href="https://ziglang.org/" rel="noopener noreferrer"&gt;Zig&lt;/a&gt; and &lt;a href="https://developer.apple.com/documentation/javascriptcore" rel="noopener noreferrer"&gt;JavaScriptCore&lt;/a&gt; (the engine that powers Safari). It’s designed to be a drop-in replacement for Node.js with a focus on:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Speed&lt;/strong&gt; - Native code execution and optimized I/O&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;TypeScript support&lt;/strong&gt; - First-class TypeScript without transpilation&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;All-in-one toolkit&lt;/strong&gt; - Runtime, bundler, test runner, and package manager&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Implementation Details
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Bun Version:&lt;/strong&gt; 1.3.5&lt;/p&gt;

&lt;p&gt;The implementation uses Bun’s native HTTP server API, which is incredibly simple and performant:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;const port = 8080;

const server = Bun.serve({
    port: port,
    fetch(req) {
        const url = new URL(req.url);

        if (url.pathname === "/hello") {
            const info = {
                name: "bun",
                releaseYear: new Date().getFullYear()
            };
            return new Response(JSON.stringify(info), {
                headers: { "Content-Type": "application/json" }
            });
        }

        return new Response("Not Found", { status: 404 });
    }
});

console.log(`Bun server started on port ${server.port}`);

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The build process is straightforward . Bun can compile TypeScript directly to a standalone executable:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;bun build --compile ./main.ts --outfile bun-demo

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Benchmark Results: The Numbers Speak
&lt;/h2&gt;

&lt;p&gt;Here are the complete benchmark results for Bun:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;---- Global Information --------------------------------------------------------
&amp;gt; request count 32000 (OK=32000 KO=0 )
&amp;gt; min response time 0 (OK=0 KO=- )
&amp;gt; max response time 569 (OK=569 KO=- )
&amp;gt; mean response time 157 (OK=157 KO=- )
&amp;gt; std deviation 115 (OK=115 KO=- )
&amp;gt; response time 50th percentile 148 (OK=148 KO=- )
&amp;gt; response time 75th percentile 208 (OK=208 KO=- )
&amp;gt; response time 95th percentile 403 (OK=402 KO=- )
&amp;gt; response time 99th percentile 483 (OK=483 KO=- )
&amp;gt; mean requests/sec 6400 (OK=6400 KO=- )

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Key highlights:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;157ms mean response time&lt;/strong&gt; : faster than Golang (227ms), .NET 9 AOT (255ms), and all JVM frameworks&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;6,400 requests/sec&lt;/strong&gt; : matching the throughput of Rust frameworks&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;0 failed requests&lt;/strong&gt; : 100% success rate under load (unlike Express.js which had 75% failure rate)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;569ms max response time&lt;/strong&gt; : excellent consistency&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Performance Comparison
&lt;/h2&gt;

&lt;p&gt;Let’s put Bun’s performance in perspective with the top performers:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Rank&lt;/th&gt;
&lt;th&gt;Technology&lt;/th&gt;
&lt;th&gt;Mean Response Time (ms)&lt;/th&gt;
&lt;th&gt;Requests/sec&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Rust (Warp)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;144&lt;/td&gt;
&lt;td&gt;6,400&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Rust (Actix)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;154&lt;/td&gt;
&lt;td&gt;6,400&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Rust (Axum)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;154&lt;/td&gt;
&lt;td&gt;6,400&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Bun&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;157&lt;/td&gt;
&lt;td&gt;6,400&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Rust (Rocket)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;238&lt;/td&gt;
&lt;td&gt;5,333&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Golang&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;227&lt;/td&gt;
&lt;td&gt;5,333&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;.NET 9 AOT&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;255&lt;/td&gt;
&lt;td&gt;5,333&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;.NET 7 AOT&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;284&lt;/td&gt;
&lt;td&gt;5,333&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;.NET 8 AOT&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;285&lt;/td&gt;
&lt;td&gt;5,333&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;GraalVM Micronaut&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;339&lt;/td&gt;
&lt;td&gt;5,333&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Bun vs Express.js: A JavaScript Runtime Showdown
&lt;/h2&gt;

&lt;p&gt;The comparison between Bun and Express.js (Node.js) is particularly striking:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Bun&lt;/th&gt;
&lt;th&gt;Express.js (Node.js)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Mean Response Time&lt;/td&gt;
&lt;td&gt;157ms&lt;/td&gt;
&lt;td&gt;815ms (3,247ms for OK requests)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Requests/sec&lt;/td&gt;
&lt;td&gt;6,400&lt;/td&gt;
&lt;td&gt;667 (successful only)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Failed Requests&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;24,000 (75% failure rate)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Max Response Time&lt;/td&gt;
&lt;td&gt;569ms&lt;/td&gt;
&lt;td&gt;10,719ms&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Bun is approximately &lt;strong&gt;5x faster&lt;/strong&gt; than Express.js in mean response time and handles &lt;strong&gt;~10x more successful requests per second&lt;/strong&gt;. Most importantly, Bun maintained 100% stability under load while Express.js struggled significantly.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why is Bun So Fast?
&lt;/h2&gt;

&lt;p&gt;Several factors contribute to Bun’s impressive performance:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;JavaScriptCore Engine&lt;/strong&gt; : Safari’s JS engine is highly optimized and often outperforms V8 in certain workloads&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Zig Implementation&lt;/strong&gt; : Low-level systems language with minimal overhead&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Native HTTP Server&lt;/strong&gt; : Built-in server implementation bypasses the overhead of frameworks like Express&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Optimized I/O&lt;/strong&gt; : Uses io_uring on Linux for efficient async I/O operations&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No Transpilation Overhead&lt;/strong&gt; : Native TypeScript execution&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Updated Performance Tiers
&lt;/h2&gt;

&lt;p&gt;With Bun’s addition, our performance tiers now look like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Performance Tiers:

🥇 TIER 1 (&amp;lt; 200ms): Rust frameworks, Bun
   - Native compilation or highly optimized runtimes
   - Minimal overhead, maximum throughput

🥈 TIER 2 (200-300ms): Golang, .NET AOT, GraalVM Native
   - Excellent performance with broader ecosystem

🥉 TIER 3 (300-600ms): GraalVM Java frameworks
   - Native compilation benefits for JVM

🏅 TIER 4 (&amp;gt; 600ms): JVM frameworks, Node.js/Express.js
   - Full-featured but with more overhead

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Bun’s benchmark results are genuinely surprising. A JavaScript/TypeScript runtime competing with Rust frameworks was not something I expected to see. Here are the key takeaways:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Bun is production-ready for high-performance workloads&lt;/strong&gt; : The 157ms mean response time and 0% failure rate prove it can handle serious traffic.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;JavaScript doesn’t have to be slow&lt;/strong&gt; : Bun demonstrates that with the right architecture, JavaScript can achieve near-native performance.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Consider Bun for new projects&lt;/strong&gt; : If you’re starting a new microservice and your team knows JavaScript/TypeScript, Bun offers an excellent balance of developer experience and performance.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;The gap between Bun and Node.js is massive&lt;/strong&gt; : If you’re currently using Express.js and need better performance, Bun is worth serious consideration.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;For the complete benchmark results including all frameworks, GraalVM native builds, and detailed statistics, check out the &lt;a href="https://ozkanpakdil.github.io/test-microservice-frameworks/posts/2026/2026-01-10-microservice-framework-test-25/" rel="noopener noreferrer"&gt;full benchmark report&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;a href="https://github.com/ozkanpakdil/test-microservice-frameworks" rel="noopener noreferrer"&gt;Source code for tests&lt;/a&gt; 👈 &lt;a href="https://github.com/ozkanpakdil/rust-examples" rel="noopener noreferrer"&gt;Rust examples&lt;/a&gt; 👈&lt;/p&gt;

</description>
      <category>javascript</category>
      <category>microservices</category>
      <category>performance</category>
    </item>
  </channel>
</rss>
