<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Jinxuan Che</title>
    <description>The latest articles on DEV Community by Jinxuan Che (@jinxuan_ai).</description>
    <link>https://dev.to/jinxuan_ai</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4104683%2Fa4ae9a31-6dd3-4d8e-98f3-5947e0695f7a.png</url>
      <title>DEV Community: Jinxuan Che</title>
      <link>https://dev.to/jinxuan_ai</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/jinxuan_ai"/>
    <language>en</language>
    <item>
      <title>Run a private OpenAI-compatible LLM endpoint on Apple Silicon with one Rust binary</title>
      <dc:creator>Jinxuan Che</dc:creator>
      <pubDate>Wed, 02 Sep 2026 03:12:46 +0000</pubDate>
      <link>https://dev.to/jinxuan_ai/run-a-private-openai-compatible-llm-endpoint-on-apple-silicon-with-one-rust-binary-29h0</link>
      <guid>https://dev.to/jinxuan_ai/run-a-private-openai-compatible-llm-endpoint-on-apple-silicon-with-one-rust-binary-29h0</guid>
      <description>&lt;p&gt;Disclosure: I maintain &lt;a href="https://github.com/sizzlecar/ferrum-infer-rs" rel="noopener noreferrer"&gt;Ferrum&lt;/a&gt;, an MIT-licensed local LLM inference server written in Rust.&lt;/p&gt;

&lt;p&gt;If you want a private OpenAI-compatible endpoint on an M-series Mac without setting up Python or a container, this is the shortest path I currently recommend.&lt;/p&gt;

&lt;p&gt;The model walkthrough below was tested with Ferrum v0.8.3 on an M1 Max. The one-command installation was checked with the published v0.8.8 macOS Metal build on September 8.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Install Ferrum
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-fsSL&lt;/span&gt; https://ferrum.pandaailabs.com/install.sh | sh
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The installer downloads the latest stable macOS Metal build, checks its SHA-256 checksum, and adds Ferrum to PATH. Open a new terminal before continuing. Run the same install command again to upgrade.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Run the preflight check
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ferrum doctor qwen3.5:4b-q4_k_m
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;ferrum doctor&lt;/code&gt; checks the local setup before inference. In one real install it caught missing Xcode Command Line Tools before the user reached a harder-to-diagnose failure.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Start an interactive chat
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The first run downloads about 2.55 GiB.&lt;/strong&gt; The terminal can look quiet during that download, so give it time before assuming it has hung.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ferrum run qwen3.5:4b-q4_k_m &lt;span class="nt"&gt;--disable-thinking&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Qwen3.5 emits verbose reasoning by default. &lt;code&gt;--disable-thinking&lt;/code&gt; gives a more conventional first-chat experience. Omit the flag when you want the reasoning behavior.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Start the OpenAI-compatible API
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ferrum serve &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--model&lt;/span&gt; qwen3.5:4b-q4_k_m &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--served-model-name&lt;/span&gt; ferrum &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--disable-thinking&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--port&lt;/span&gt; 8000
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In another terminal:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl http://localhost:8000/v1/chat/completions &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s1"&gt;'Content-Type: application/json'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "model": "ferrum",
    "messages": [{"role": "user", "content": "Reply exactly: ferrum-ok"}],
    "max_tokens": 32
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A working setup returns HTTP 200 with a non-empty assistant response.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I am looking for
&lt;/h2&gt;

&lt;p&gt;If you try it, I’d like to hear which model and client you connect, and whether you reached your first useful response. If something gets in the way, include your Mac model and the step so I can investigate.&lt;/p&gt;

&lt;p&gt;Repository and issue tracker: &lt;a href="https://github.com/sizzlecar/ferrum-infer-rs" rel="noopener noreferrer"&gt;github.com/sizzlecar/ferrum-infer-rs&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;This post was drafted with assistance from OpenAI Codex. The commands and stated behavior were checked against the published build; no comparative performance claim is being made here.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>showdev</category>
      <category>ai</category>
      <category>rust</category>
    </item>
  </channel>
</rss>
