<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Victor Tarasov</title>
    <description>The latest articles on DEV Community by Victor Tarasov (@victor_tarasov_057fe5583a).</description>
    <link>https://dev.to/victor_tarasov_057fe5583a</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3430294%2Fa963e743-b51b-4833-aa9c-56ae9e5860a9.jpg</url>
      <title>DEV Community: Victor Tarasov</title>
      <link>https://dev.to/victor_tarasov_057fe5583a</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/victor_tarasov_057fe5583a"/>
    <language>en</language>
    <item>
      <title>Four OpenAI-compatible endpoints are giving away free tokens right now</title>
      <dc:creator>Victor Tarasov</dc:creator>
      <pubDate>Fri, 25 Sep 2026 14:47:07 +0000</pubDate>
      <link>https://dev.to/victor_tarasov_057fe5583a/four-openai-compatible-endpoints-are-giving-away-free-tokens-right-now-3l2l</link>
      <guid>https://dev.to/victor_tarasov_057fe5583a/four-openai-compatible-endpoints-are-giving-away-free-tokens-right-now-3l2l</guid>
      <description>&lt;p&gt;&lt;em&gt;Disclosure: I work on Gonka. Four independent brokers are running the free-token offer described below, and I've tried to keep the technical details accurate enough to be useful even if you skip the offer entirely.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;If you build with LLMs, you've probably hit the same wall I have: you want to try a long-context model on a real workload, and the only honest way to find out how it behaves is to throw a few million tokens at it. That's expensive to do on a whim.&lt;/p&gt;

&lt;p&gt;Right now there are four independent endpoints handing out free tokens, all OpenAI-compatible, so trying them costs you a base URL change and nothing else. Here's what's actually on offer and how to wire it up.&lt;/p&gt;

&lt;h2&gt;
  
  
  The four endpoints
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Broker&lt;/th&gt;
&lt;th&gt;Welcome bonus&lt;/th&gt;
&lt;th&gt;Base host&lt;/th&gt;
&lt;th&gt;Docs&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;DAHL&lt;/td&gt;
&lt;td&gt;100M tokens&lt;/td&gt;
&lt;td&gt;&lt;code&gt;inference.dahl.global&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;/docs/&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gonka GG&lt;/td&gt;
&lt;td&gt;1M tokens&lt;/td&gt;
&lt;td&gt;&lt;code&gt;proxy.gonka.gg&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;/docs&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gonka Router&lt;/td&gt;
&lt;td&gt;$20 in credits&lt;/td&gt;
&lt;td&gt;&lt;code&gt;gonkarouter.io&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;/docs&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gonka API&lt;/td&gt;
&lt;td&gt;10M tokens&lt;/td&gt;
&lt;td&gt;&lt;code&gt;gonka-api.org&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;/dashboard&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These are model tokens, not a cryptocurrency. There's no wallet anywhere in the flow — you sign up, you get a key.&lt;/p&gt;

&lt;p&gt;Each broker is a separate business. They set their own rate limits, data policy and pricing once the bonus runs out, so read the one you pick rather than assuming they match.&lt;/p&gt;

&lt;h2&gt;
  
  
  Models
&lt;/h2&gt;

&lt;p&gt;All four serve the same three models:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;deepseek-ai/DeepSeek-V4-Flash-0731&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;zai-org/GLM-5.3-Flash&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;MiniMaxAI/MiniMax-M2.7&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Context limits are worth checking per broker rather than trusting a spec sheet. On Gonka GG, the published limits are 400K context / 16K output for the two Flash models and 180K / 16K for MiniMax. The DeepSeek and GLM models are served in the 380–400K range depending on who you're pointed at — if your workload actually depends on the top of that window, measure it on your key before you design around it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Wiring it up
&lt;/h2&gt;

&lt;p&gt;Every one of these speaks the OpenAI chat-completions API, so it's a base URL and a key:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl https://BROKER_HOST/v1/chat/completions &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "model": "deepseek-ai/DeepSeek-V4-Flash-0731",
    "messages": [{"role": "user", "content": "Say hi"}]
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Python, with the official SDK:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://BROKER_HOST/v1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;YOUR_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;resp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;zai-org/GLM-5.3-Flash&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Summarise this repo&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;resp&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The same two values drop into anything that takes a custom OpenAI-compatible provider — LibreChat, Cline, Roo, opencode, LiteLLM, Continue. No code changes, just config.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to spend it on
&lt;/h2&gt;

&lt;p&gt;A bonus is only useful if you learn something from it. Some things worth more than a chat demo:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Long-context retrieval.&lt;/strong&gt; Put 300K tokens of your own codebase or docs in and ask questions whose answers sit in the middle. Models degrade very differently across that window.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Agent loops.&lt;/strong&gt; These burn tokens fast, which makes them the honest test. A 1M bonus disappears in one enthusiastic afternoon; a 100M one supports a real evaluation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Structured output under load.&lt;/strong&gt; Ask for JSON a few thousand times and count the malformed responses.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you've got 100M tokens to spend, that's enough to run a proper eval rather than a vibe check, which is the thing free tiers usually can't buy you.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the compute comes from
&lt;/h2&gt;

&lt;p&gt;Worth knowing what you're pointed at: these brokers front &lt;a href="https://gonka.ai" rel="noopener noreferrer"&gt;Gonka&lt;/a&gt;, a decentralized inference network. Instead of one company's GPUs, independent operators serve requests, and correctness is kept honest by re-running a random sample of tasks — somewhere between 1% and 10% — and scoring operators on reputation over time. It's a Go codebase built on a fork of the Cosmos SDK.&lt;/p&gt;

&lt;p&gt;For your purposes as a caller none of that matters: it's an HTTPS endpoint that speaks OpenAI. But it's the reason the token grants can be this large.&lt;/p&gt;

&lt;h2&gt;
  
  
  Getting a key
&lt;/h2&gt;

&lt;p&gt;All four are listed at &lt;strong&gt;&lt;a href="https://aidrop.gnk.space" rel="noopener noreferrer"&gt;aidrop.gnk.space&lt;/a&gt;&lt;/strong&gt; with sign-up links. Pick whichever bonus fits what you want to measure — 1M for a quick look, 10M or 100M for an actual evaluation.&lt;/p&gt;

&lt;p&gt;If you hit something odd with keys, limits, or a model behaving differently than documented, each broker has its own channel on the &lt;a href="https://discord.gg/REcpeYc7P7" rel="noopener noreferrer"&gt;Gonka Discord&lt;/a&gt; and the people running them read it.&lt;/p&gt;

&lt;p&gt;Happy to answer questions in the comments — including sceptical ones about the verification claims above.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>api</category>
      <category>opensource</category>
    </item>
    <item>
      <title>AMA session. Topic: AI Training Beyond the Data Center</title>
      <dc:creator>Victor Tarasov</dc:creator>
      <pubDate>Wed, 15 Oct 2025 09:54:42 +0000</pubDate>
      <link>https://dev.to/victor_tarasov_057fe5583a/ama-session-topic-ai-training-beyond-the-data-center-23e0</link>
      <guid>https://dev.to/victor_tarasov_057fe5583a/ama-session-topic-ai-training-beyond-the-data-center-23e0</guid>
      <description>&lt;p&gt;Join us for an AMA session on Tuesday, October 21, at 9 AM PST / 6 PM CET with special guest - Egor Shulgin, co-creator of Gonka.&lt;/p&gt;

&lt;p&gt;Topic: AI Training Beyond the Data Center: Breaking the Communication Barrier&lt;/p&gt;

&lt;p&gt;Discover how algorithms that "communicate less" are making it possible to train massive AI models over the internet, overcoming the bottleneck of slow networks.&lt;/p&gt;

&lt;p&gt;We will explore:&lt;br&gt;
🔹 The move from centralized data centers to globally distributed training.&lt;br&gt;
🔹 How low-communication frameworks use federated optimization to train billion-parameter models on standard internet connections.&lt;br&gt;
🔹 The breakthrough results: matching data-center performance while reducing communication by up to 500x.&lt;/p&gt;

&lt;p&gt;Click the event link below to set a reminder!&lt;br&gt;
&lt;a href="https://discord.gg/DyDxDsP3Pd?event=1427265849223544863" rel="noopener noreferrer"&gt;https://discord.gg/DyDxDsP3Pd?event=1427265849223544863&lt;/a&gt; &lt;/p&gt;

</description>
      <category>ai</category>
      <category>blockchain</category>
      <category>web3</category>
      <category>computerscience</category>
    </item>
  </channel>
</rss>
