<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: SpicyAPI</title>
    <description>The latest articles on DEV Community by SpicyAPI (@spicyapi_ai).</description>
    <link>https://dev.to/spicyapi_ai</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4175210%2Fd9d3218b-e314-4ddb-8e42-8f9350e3f42d.png</url>
      <title>DEV Community: SpicyAPI</title>
      <link>https://dev.to/spicyapi_ai</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/spicyapi_ai"/>
    <language>en</language>
    <item>
      <title>How to run a Civitai or Hugging Face LoRA through an API (no GPU, no upload)</title>
      <dc:creator>SpicyAPI</dc:creator>
      <pubDate>Sat, 10 Oct 2026 12:10:07 +0000</pubDate>
      <link>https://dev.to/spicyapi_ai/how-to-run-a-civitai-or-hugging-face-lora-through-an-api-no-gpu-no-upload-5a43</link>
      <guid>https://dev.to/spicyapi_ai/how-to-run-a-civitai-or-hugging-face-lora-through-an-api-no-gpu-no-upload-5a43</guid>
      <description>&lt;p&gt;You found a LoRA you like on Civitai or Hugging Face. Running it usually means renting a GPU, installing ComfyUI or a diffusers script, downloading the base checkpoint and keeping all of it alive. If all you want is images or short clips from that LoRA inside your own app, that is a lot of machinery.&lt;/p&gt;

&lt;p&gt;This post shows the other route: pass the LoRA's download link in an API request and let a hosted model load it for that one task. The examples use &lt;a href="https://spicyapi.ai" rel="noopener noreferrer"&gt;SpicyAPI&lt;/a&gt;, which exposes a &lt;code&gt;loras&lt;/code&gt; field on 60+ image and video endpoints (FLUX, Z-Image, Qwen-Image, Wan 2.2, LTX-2, MiniMax H3 and more). The platform adds no content filter of its own, so what the LoRA and the base model can produce is what you get back. The same ideas apply to any host that takes a LoRA URL.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Pick a LoRA trained for the base you will call
&lt;/h2&gt;

&lt;p&gt;A LoRA only works on the base model it was trained on. A FLUX.1 [dev] LoRA does nothing useful on Z-Image Turbo, and most hosts will not tell you: the task succeeds and the LoRA is silently ignored. So start from the base:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;On Civitai, each file version lists its "Base Model" in the details panel.&lt;/li&gt;
&lt;li&gt;On Hugging Face, it is usually in the model card and the repo tags.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Also note the &lt;strong&gt;trigger word&lt;/strong&gt; and the &lt;strong&gt;recommended strength&lt;/strong&gt; from the card. You will need both.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://spicyapi.ai/lora" rel="noopener noreferrer"&gt;LoRA overview page&lt;/a&gt; lists which LoRA families map to which hosted base, so you can match them before you spend anything.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Get a direct link to the .safetensors file
&lt;/h2&gt;

&lt;p&gt;The API needs a URL that downloads the weights file itself, not the page about it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hugging Face&lt;/strong&gt; — use &lt;code&gt;resolve&lt;/code&gt;, not &lt;code&gt;blob&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;https://huggingface.co/OWNER/REPO/resolve/main/FILE.safetensors
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A &lt;code&gt;blob&lt;/code&gt; link returns an HTML page. Only public repos work; a gated repo answers the download with 401 and the task fails.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Civitai&lt;/strong&gt; — use the API download link with the &lt;em&gt;version&lt;/em&gt; ID (the number behind the Download button), not the model ID:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;https://civitai.com/api/download/models/VERSION_ID
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In our tests these links loaded without a token, including files Civitai marks as sign-in only. If you ever get a 401, append &lt;code&gt;?token=YOUR_CIVITAI_KEY&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Your own files&lt;/strong&gt; — any public HTTPS link to a &lt;code&gt;.safetensors&lt;/code&gt; file works, for example a file in your own S3 or R2 bucket.&lt;/p&gt;

&lt;p&gt;Nothing is uploaded ahead of time. The weights are fetched for the task, and there is no separate LoRA fee.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Send the request
&lt;/h2&gt;

&lt;p&gt;Every model has its own input schema, but the LoRA part looks the same: a &lt;code&gt;loras&lt;/code&gt; array of &lt;code&gt;{ path, scale }&lt;/code&gt; objects. Here is Z-Image Turbo LoRA, which takes up to three LoRAs at a scale from 0 to 4 (1 is neutral):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-X&lt;/span&gt; POST &lt;span class="s2"&gt;"https://api.spicyapi.ai/api/v1/jobs/createTask?wait=60"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$SPICY_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Idempotency-Key: lora-demo-001"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "model": "alibaba/z-image-turbo-lora/text-to-image",
    "input": {
      "prompt": "TRIGGERWORD, a woman reading by a rain-streaked window at dusk, warm lamp light on her face, city lights behind",
      "aspect_ratio": "3:4",
      "seed": 42,
      "loras": [
        { "path": "https://huggingface.co/OWNER/REPO/resolve/main/FILE.safetensors", "scale": 0.9 }
      ]
    }
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A few details that save time:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;?wait=60&lt;/code&gt; holds the connection for up to 60 seconds. If the task finishes in time you get the final record (HTTP 200) straight away; otherwise you get the usual 202 with a &lt;code&gt;taskId&lt;/code&gt; and poll.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;Idempotency-Key&lt;/code&gt; makes a retry of the same request return the same task instead of creating and charging a second one.&lt;/li&gt;
&lt;li&gt;Fix the &lt;code&gt;seed&lt;/code&gt; while you are tuning. It is the only way to tell whether a change came from the LoRA or from luck.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you got a 202, poll until &lt;code&gt;state&lt;/code&gt; is terminal:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="s2"&gt;"https://api.spicyapi.ai/api/v1/jobs/recordInfo?taskId=TASK_ID"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$SPICY_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"code"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"msg"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"success"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"data"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"taskId"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"job_..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"state"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"succeeded"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"cost"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"0.012"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"settled"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"output"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"assets"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"url"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"mime"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"image/jpeg"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The asset URL is a signed link valid for about 20 minutes; poll again for a fresh one, or download the file. Results are kept for 14 days.&lt;/p&gt;

&lt;p&gt;Want the price first? Send the same body to &lt;code&gt;POST /api/v1/jobs/quote&lt;/code&gt;. It returns the estimated charge and the maximum you can be billed, and creates nothing. Image LoRA endpoints are mostly a flat price per image (Z-Image Turbo LoRA is $0.012); video endpoints bill per second of output, so quoting first is worth it there.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Check that the LoRA is actually doing something
&lt;/h2&gt;

&lt;p&gt;Because a mismatched or broken LoRA often fails silently, do one A/B run before you scale up:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Run the request once with the LoRA.&lt;/li&gt;
&lt;li&gt;Run it again with the same seed and &lt;code&gt;"loras": []&lt;/code&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If the two images look alike, something is off: wrong base, missing trigger word, or a scale too low to matter. Fix that before generating a batch.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Tuning strength and stacking
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Start at the card's recommended scale&lt;/strong&gt;, often 0.7 to 1.0. Far above 1 tends to burn in artefacts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stack carefully.&lt;/strong&gt; Two style LoRAs at full strength usually fight. Keep the main one near 1 and bring the second in around 0.5, moving one scale by about 0.2 at a time with the seed fixed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Trigger word first&lt;/strong&gt; in the prompt, then the subject, the setting and the light. Concrete lighting ("a warm lamp on the face, neon behind") moves a realism LoRA further than adjectives like "cinematic".&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Video models can have more than one slot.&lt;/strong&gt; Wan 2.2 uses a two-stage sampler, so its LoRAs come as separate HIGH-noise and LOW-noise files; that endpoint takes them in &lt;code&gt;high_noise_loras&lt;/code&gt; and &lt;code&gt;low_noise_loras&lt;/code&gt;, plus a plain &lt;code&gt;loras&lt;/code&gt; field. Put each file in the slot it was trained for.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Limits differ by model (how many LoRAs, what scale range), so read the &lt;code&gt;loras&lt;/code&gt; part of the model's &lt;code&gt;inputSchema&lt;/code&gt; from the catalog before you hard-code anything.&lt;/p&gt;

&lt;h2&gt;
  
  
  When this route makes sense
&lt;/h2&gt;

&lt;p&gt;Hosted LoRA calls are a good fit when you want results inside your own app without running GPUs, when you switch LoRAs often, or when you are testing many community LoRAs before committing to one. If you generate around the clock on a single fixed LoRA, your own GPU may work out cheaper; at a few cents per image, most side projects and prototypes never get there.&lt;/p&gt;

&lt;p&gt;The full field reference, SDKs (TypeScript, Python, Go, PHP, Java) and a CLI are in the &lt;a href="https://docs.spicyapi.ai" rel="noopener noreferrer"&gt;docs&lt;/a&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



</description>
      <category>ai</category>
      <category>api</category>
      <category>tutorial</category>
      <category>machinelearning</category>
    </item>
  </channel>
</rss>
