<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: hui feng</title>
    <description>The latest articles on DEV Community by hui feng (@hui_feng_f2247629b1d2be00).</description>
    <link>https://dev.to/hui_feng_f2247629b1d2be00</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4173072%2F8303de09-b85f-43e9-adb3-a586ec4ba099.jpg</url>
      <title>DEV Community: hui feng</title>
      <link>https://dev.to/hui_feng_f2247629b1d2be00</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/hui_feng_f2247629b1d2be00"/>
    <language>en</language>
    <item>
      <title>Run Full Kimi K3 on a Single Machine with Deltafin: Rust-Powered Local Inference and OpenAI-Compatible API</title>
      <dc:creator>hui feng</dc:creator>
      <pubDate>Sat, 10 Oct 2026 06:58:28 +0000</pubDate>
      <link>https://dev.to/hui_feng_f2247629b1d2be00/run-full-kimi-k3-on-a-single-machine-with-deltafin-rust-powered-local-inference-and-97l</link>
      <guid>https://dev.to/hui_feng_f2247629b1d2be00/run-full-kimi-k3-on-a-single-machine-with-deltafin-rust-powered-local-inference-and-97l</guid>
      <description>&lt;h3&gt;
  
  
  TL;DR
&lt;/h3&gt;

&lt;p&gt;Deltafin is an ultra-fast, Rust-native inference engine that allows you to run the massive Kimi K3 model locally on a single machine without complex distributed cluster orchestration. By bundling low-overhead compute kernels with a drop-in OpenAI-compatible API server, it drastically slashes local deployment costs and lets you power local chat and autonomous coding agents instantly.&lt;/p&gt;




&lt;h3&gt;
  
  
  Key Features &amp;amp; Benchmarks
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Single-Device Execution:&lt;/strong&gt; Native Rust memory safety and aggressive offloading strategies squeeze Kimi K3 onto single-host architectures without requiring multi-node enterprise rigs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Drop-in OpenAI Compatible API:&lt;/strong&gt; Serves &lt;code&gt;/v1/chat/completions&lt;/code&gt; out of the box, integrating seamlessly with Cursor, Continue.dev, Cline, and agent frameworks like AutoGen or LangGraph.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Zero-Python Overhead:&lt;/strong&gt; Pure Rust runtime eliminates Python runtime bloat, GIL contention, and heavy dependency trees, leading to sub-millisecond server latency.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Massive Context Optimization:&lt;/strong&gt; Specially tuned attention and memory-mapped weight loading designed to leverage Kimi's deep-context retrieval capabilities efficiently.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Optimized for Coding Agents:&lt;/strong&gt; High sustained token throughput tailored specifically for continuous diff generation, codebase indexing, and multi-turn refactoring loops.&lt;/li&gt;
&lt;/ul&gt;




&lt;h3&gt;
  
  
  Quick Start
&lt;/h3&gt;

&lt;p&gt;Get Deltafin running on your machine in just a few commands:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# 1. Install Deltafin via Cargo or download pre-built binaries&lt;/span&gt;
cargo &lt;span class="nb"&gt;install &lt;/span&gt;deltafin

&lt;span class="c"&gt;# 2. Download and launch Kimi K3 with the built-in OpenAI-compatible API server&lt;/span&gt;
deltafin serve &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--model&lt;/span&gt; kimi-k3 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--host&lt;/span&gt; 0.0.0.0 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--port&lt;/span&gt; 8080 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--quant&lt;/span&gt; q4_k_m
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Once the server is up, test the endpoint with a standard OpenAI cURL request:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl http://localhost:8080/v1/chat/completions &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "model": "kimi-k3",
    "messages": [{"role": "user", "content": "Explain Rust lifetime bounds in one sentence."}],
    "temperature": 0.2
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h3&gt;
  
  
  Why It Matters
&lt;/h3&gt;

&lt;p&gt;Foundation models with long-context strengths like Kimi have historically required massive multi-GPU cloud instances or proprietary API contracts. Deltafin opens the door for:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Self-Hosted AI Engineers:&lt;/strong&gt; Run autonomous coding agents locally with zero data leaks and zero per-token inference bills.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Privacy-Constrained Teams:&lt;/strong&gt; Deploy state-of-the-art context reasoning entirely behind corporate firewalls.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Agent Infrastructure Builders:&lt;/strong&gt; Benefit from an ultra-lightweight Rust backend that doesn't waste precious VRAM on bloated runtime environments.&lt;/li&gt;
&lt;/ol&gt;




&lt;h3&gt;
  
  
  🛠️ Recommended AI Stack &amp;amp; Resources
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Cloud GPU Hosting:&lt;/strong&gt; Need raw power to scale model evaluations or host larger checkpoints? Spin up dedicated instances at low hourly rates on &lt;a href="https://runpod.io?ref=sr37lgmj" rel="noopener noreferrer"&gt;RunPod&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI Code Editor:&lt;/strong&gt; Supercharge your developer velocity and pair local model endpoints directly with &lt;a href="https://cursor.com" rel="noopener noreferrer"&gt;Cursor&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Production Database:&lt;/strong&gt; Power your retrieval-augmented workflows and structured data layers with &lt;a href="https://supabase.com" rel="noopener noreferrer"&gt;Supabase&lt;/a&gt; or &lt;a href="https://pinecone.io" rel="noopener noreferrer"&gt;Pinecone&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Global Dev Network:&lt;/strong&gt; Ensure lightning-fast Hugging Face downloads, ultra-low latency model syncs, and stable remote server administration with &lt;a href="https://wd-gold.net/aff.php?aff=13920" rel="noopener noreferrer"&gt;WD-Gold Network&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Newsletter CTA:&lt;/strong&gt; Subscribe to &lt;strong&gt;Local AI Daily&lt;/strong&gt; for curated breakdowns of the newest open-source runtimes, local quantization tools, and infrastructure updates delivered straight to your inbox.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sponsorship &amp;amp; Partnerships:&lt;/strong&gt; Want to showcase your AI runtime, model, or developer tool to thousands of active builders? Reach out at: &lt;a href="mailto:fengdahui195@gmail.com"&gt;fengdahui195@gmail.com&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>python</category>
      <category>localai</category>
    </item>
    <item>
      <title>Stop Renting Your Intelligence: Build Local-First RAG and AI Agents with AnythingLLM</title>
      <dc:creator>hui feng</dc:creator>
      <pubDate>Sat, 10 Oct 2026 00:35:53 +0000</pubDate>
      <link>https://dev.to/hui_feng_f2247629b1d2be00/stop-renting-your-intelligence-build-local-first-rag-and-ai-agents-with-anythingllm-26jf</link>
      <guid>https://dev.to/hui_feng_f2247629b1d2be00/stop-renting-your-intelligence-build-local-first-rag-and-ai-agents-with-anythingllm-26jf</guid>
      <description>&lt;h3&gt;
  
  
  TL;DR
&lt;/h3&gt;

&lt;p&gt;AnythingLLM replaces fragmented, subscription-based AI stacks with a turnkey, local-first platform for enterprise-grade RAG and autonomous agents. By unifying document processing, vector search, and model execution into a single installable package, it eliminates recurring SaaS fees and data leak risks. You can turn any local LLM into a multi-user, context-aware assistant in under five minutes.&lt;/p&gt;

&lt;h3&gt;
  
  
  Key Features &amp;amp; Benchmarks
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Universal Model &amp;amp; Vector Compatibility:&lt;/strong&gt; Connect effortlessly to local runtimes like Ollama, LM Studio, LocalAI, and vLLM, or swap between built-in LanceDB and enterprise vector stores (Chroma, Qdrant, Weaviate, Pinecone).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Zero-Code Multi-Modal Ingestion:&lt;/strong&gt; Native parsing and chunking for PDFs, office documents, YouTube transcripts, audio, and web crawls without writing custom LangChain pipelines.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Multi-Tenant Workspace Isolation:&lt;/strong&gt; Granular role-based access control (RBAC), multi-user workspaces, and session tracking built directly into the UI.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Local Autonomous Agents:&lt;/strong&gt; Configure custom agent skills, web searching, and multi-step tool execution running entirely within your self-hosted perimeter.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;High-Performance Footprint:&lt;/strong&gt; Lightweight Node.js/JavaScript architecture designed to run on consumer hardware, laptops, or low-cost bare-metal servers without heavy resource overhead.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Quick Start
&lt;/h3&gt;

&lt;p&gt;The fastest way to run AnythingLLM locally is via Docker:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Pull and run the official AnythingLLM container&lt;/span&gt;
docker run &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="nt"&gt;-p&lt;/span&gt; 3001:3001 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--name&lt;/span&gt; anything-llm &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-v&lt;/span&gt; anythingllm_storage:/app/server/storage &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-e&lt;/span&gt; &lt;span class="nv"&gt;STORAGE_DIR&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"/app/server/storage"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  mintplexlabs/anythingllm
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Open &lt;code&gt;http://localhost:3001&lt;/code&gt; in your browser, select your preferred local LLM provider (such as Ollama at &lt;code&gt;http://host.docker.internal:11434&lt;/code&gt;), drop in your documents, and start chatting.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why It Matters
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Privacy-Sensitive Engineering Teams:&lt;/strong&gt; Perfect for organizations bound by GDPR, HIPAA, or strict IP policies that prohibit sending proprietary source code or docs to third-party APIs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Homelab &amp;amp; Edge AI Enthusiasts:&lt;/strong&gt; Eliminates the complexity of wiring frontends, embedding pipelines, vector databases, and inference servers from scratch.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Internal Tool Builders:&lt;/strong&gt; Provides a complete, customizable platform with full REST API support to embed private RAG workflows into your existing internal applications.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Call To Action
&lt;/h3&gt;

&lt;p&gt;Ready to stay ahead of the self-hosted AI curve? Subscribe to &lt;strong&gt;Local AI &amp;amp; Infra Daily&lt;/strong&gt; for hands-on architectural deep dives, benchmarks, and daily open-source AI infrastructure updates delivered straight to your inbox.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>python</category>
      <category>localai</category>
    </item>
  </channel>
</rss>
