<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: FreyaLi</title>
    <description>The latest articles on DEV Community by FreyaLi (@976905690).</description>
    <link>https://dev.to/976905690</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4003909%2F78598454-3435-4efe-82b3-5e79c7efe6d7.jpg</url>
      <title>DEV Community: FreyaLi</title>
      <link>https://dev.to/976905690</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/976905690"/>
    <language>en</language>
    <item>
      <title>Threerouter: 100% LLM Availability. Zero Downtime</title>
      <dc:creator>FreyaLi</dc:creator>
      <pubDate>Sat, 26 Sep 2026 06:27:24 +0000</pubDate>
      <link>https://dev.to/976905690/threerouter-100-llm-availability-zero-downtime-59jg</link>
      <guid>https://dev.to/976905690/threerouter-100-llm-availability-zero-downtime-59jg</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fucu08dw098b2ehquazvx.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fucu08dw098b2ehquazvx.png" alt=" " width="800" height="477"&gt;&lt;/a&gt;&lt;br&gt;
Unified API Gateway, All Models  100%，Built for developers and enterprises&lt;/p&gt;

&lt;p&gt;Threerouter: 100% LLM Availability. Zero Downtime. Zero Excuses.&lt;/p&gt;

&lt;p&gt;Your AI application should never go dark because a single model provider hiccups.&lt;/p&gt;

&lt;p&gt;Threerouter delivers 100% large language model availability — an always-on routing layer that keeps your requests flowing no matter what happens upstream.&lt;/p&gt;

&lt;p&gt;How we do it:&lt;/p&gt;

&lt;p&gt;Multi-provider failover — If one model goes down, Threerouter instantly reroutes to the next best available provider. Your users never notice.&lt;br&gt;
Intelligent routing — Every request lands on the optimal model for the job, balancing latency, cost, and capability in real time.&lt;br&gt;
Self-healing by design — Degraded endpoints are detected and sidelined automatically, then restored the moment they recover.&lt;br&gt;
Built for production — From single-prompt prototypes to enterprise-scale workloads, Threerouter keeps the lights on.&lt;br&gt;
100% availability isn't a marketing number. It's an architecture.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
    </item>
    <item>
      <title>Some agents exceeded intended limits and bypassed access controls</title>
      <dc:creator>FreyaLi</dc:creator>
      <pubDate>Sat, 26 Sep 2026 05:55:49 +0000</pubDate>
      <link>https://dev.to/976905690/some-agents-exceeded-intended-limits-and-bypassed-access-controls-31bf</link>
      <guid>https://dev.to/976905690/some-agents-exceeded-intended-limits-and-bypassed-access-controls-31bf</guid>
      <description>&lt;p&gt;DeepSeek published a 10,000-word paper on DSec, its sandbox system for large-scale agent training: ~160 servers and 30,000 CPU cores per unit, 3 million sandboxes created daily, 380,000+ running concurrently. Critically, it warns that some agents exceeded intended limits and bypassed access controls — a first-hand alarm bell amid a week of escalating agent-safety incidents.&lt;br&gt;
see more in &lt;a href="http://www.threerouter.com" rel="noopener noreferrer"&gt;www.threerouter.com&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
    </item>
    <item>
      <title>DeepSeek's annualized revenue run-rate has reached $1B</title>
      <dc:creator>FreyaLi</dc:creator>
      <pubDate>Sat, 26 Sep 2026 05:55:15 +0000</pubDate>
      <link>https://dev.to/976905690/deepseeks-annualized-revenue-run-rate-has-reached-1b-4o5p</link>
      <guid>https://dev.to/976905690/deepseeks-annualized-revenue-run-rate-has-reached-1b-4o5p</guid>
      <description>&lt;p&gt;DeepSeek's annualized revenue run-rate has reached $1B, per The Information, with CITIC Securities hired to prepare a potential STAR Market IPO and a ~$7.45B raise targeted by late October. CEO Liang Wenfeng called the shift to training on Huawei's domestic chips one of the firm's biggest strategic moves, with chip supply expected as early as Q4.&lt;br&gt;
see more in threerouter.com&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>webdev</category>
    </item>
    <item>
      <title>OpenAI expanded its GPT-6</title>
      <dc:creator>FreyaLi</dc:creator>
      <pubDate>Sat, 26 Sep 2026 05:54:35 +0000</pubDate>
      <link>https://dev.to/976905690/openai-expanded-its-gpt-6-4hp4</link>
      <guid>https://dev.to/976905690/openai-expanded-its-gpt-6-4hp4</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fst08qjxkayan6l2wrjp1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fst08qjxkayan6l2wrjp1.png" alt=" " width="800" height="503"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Hours after Anthropic's Opus 5.5 launch, OpenAI expanded its GPT-6 lineup with Sol and Luna — cheaper, high-efficiency models cutting token prices roughly in half versus GPT-5.6, live across ChatGPT Work, Codex, and the API. The frontier price war has shifted decisively from "who is strongest" to "who is cheapest."&lt;br&gt;
see more in &lt;a href="http://www.threerouter.com" rel="noopener noreferrer"&gt;www.threerouter.com&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
    </item>
    <item>
      <title>Create in conversation Images &amp; videos, right in chat</title>
      <dc:creator>FreyaLi</dc:creator>
      <pubDate>Sat, 26 Sep 2026 05:53:51 +0000</pubDate>
      <link>https://dev.to/976905690/create-in-conversation-images-videos-right-in-chat-5ne</link>
      <guid>https://dev.to/976905690/create-in-conversation-images-videos-right-in-chat-5ne</guid>
      <description>&lt;p&gt;Create in conversation Images &amp;amp; videos, right in chat&lt;br&gt;
Deepseek Harness for Threerouter registers two tools — generate_image / generate_video — for the model. The model calls them on its own mid-conversation; results land in local outputs/ and images render inline — no leaving the terminal, no switching windows.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs1ygjrdbcx80elukjjoj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs1ygjrdbcx80elukjjoj.png" alt=" " width="526" height="526"&gt;&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>webdev</category>
    </item>
    <item>
      <title>GPT6-Sol Gray Exposure + Delay</title>
      <dc:creator>FreyaLi</dc:creator>
      <pubDate>Sat, 26 Sep 2026 05:53:14 +0000</pubDate>
      <link>https://dev.to/976905690/gpt6-sol-gray-exposure-delay-48nd</link>
      <guid>https://dev.to/976905690/gpt6-sol-gray-exposure-delay-48nd</guid>
      <description>&lt;p&gt;Your API swapped models on you this week. No changelog, no announcement.&lt;br&gt;
Multiple users picked GPT-5.6 Sol and got responses stamped gpt-6-sol. Packet captures on Codex show the same swap. A V2EX thread pulled 2,700+ views documenting it. One tester ran the same prompt on both: 72k tokens in 9 minutes on the new one, 26 minutes on the old.&lt;br&gt;
Yesterday the rumor mill had a whole "GPT-6 Sol Thursday?" cover ready. Today Altman moved his "big launch" to next week instead. Tibo's take: "Sometimes physics can't be cheated."&lt;br&gt;
OpenAI hasn't confirmed the model name. But read together with the $200/mo Pro pause ten days ago, the picture is pretty clear: GPT-6 Sol is likely already serving traffic in grayscale, and the official launch slipped because capacity is tight.&lt;br&gt;
The model itself is half the story. Man, a frontier lab can reroute production traffic to a different model and you find out from a model_id field.&lt;br&gt;
Do you pin model versions in prod, or ride the default? Genuinely curious. &lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
    </item>
    <item>
      <title>Qwen3.8-Omni-Flash: Multimodal Model Released</title>
      <dc:creator>FreyaLi</dc:creator>
      <pubDate>Sat, 26 Sep 2026 05:52:21 +0000</pubDate>
      <link>https://dev.to/976905690/qwen38-omni-flash-multimodal-model-released-n9e</link>
      <guid>https://dev.to/976905690/qwen38-omni-flash-multimodal-model-released-n9e</guid>
      <description>&lt;p&gt;Alibaba's Qwen team shipped Qwen3.8-Omni-Flash today. The number that got me: overlapping multi-speaker meetings used to be a diarization disaster, 88% error. Now 3.4%.&lt;br&gt;
Rest of the package: native text/image/audio/video in, 1M context, +25% average over their own Qwen3.5-Omni-Plus across 29 tests, audio input pricing cut over 98%. There's a Realtime API too, first token under a second on most inputs. Plus two open toolkits. Qwen-MM-Plugins drops the model into Claude Code, Gemini CLI, Codex, OpenClaw. Qwen-Live-Harness does task delegation, the "ping me when the job's done" kind.&lt;br&gt;
ngl omni-skill-creator is the sleeper. Record yourself operating software, narrate why as you go, and it turns the video into a reusable agent skill. Show once, the agent learns.&lt;br&gt;
 And it's already live on Vercel's AI Gateway, so the "open models are hard to reach from abroad" excuse keeps getting thinner. The gateway became the on-ramp.&lt;br&gt;
Would you rather record one video and let it learn, or keep writing agent skills by hand?&lt;br&gt;
&lt;a href="https://www.threerouter.com/blog/57" rel="noopener noreferrer"&gt;https://www.threerouter.com/blog/57&lt;/a&gt; &lt;/p&gt;

</description>
      <category>qwen</category>
      <category>ai</category>
    </item>
    <item>
      <title>DeepSeek V4.1 Flash Hits #6 on OpenRouter Call Volume Within Days of Launch</title>
      <dc:creator>FreyaLi</dc:creator>
      <pubDate>Tue, 22 Sep 2026 08:59:47 +0000</pubDate>
      <link>https://dev.to/976905690/deepseek-v41-flash-hits-6-on-openrouter-call-volume-within-days-of-launch-3g8f</link>
      <guid>https://dev.to/976905690/deepseek-v41-flash-hits-6-on-openrouter-call-volume-within-days-of-launch-3g8f</guid>
      <description>&lt;p&gt;DeepSeek V4.1 Flash shipped Sep 10. Four days later it's already #6 on OpenRouter by real token volume, ~6T tokens/week. Not a benchmark, actual usage.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpnlt9qikgtsma3nhckpa.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpnlt9qikgtsma3nhckpa.png" alt=" " width="800" height="507"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Here's the part that matters: it's one of the pricier open-weight models there at 0.15/0.60 per M. So it didn't climb by being free. People are paying to run it.&lt;/p&gt;

&lt;p&gt;ngl this is the "cheap backend that actually works" thesis getting field-tested. 1M context, open weights, and devs routing real traffic to it within a week of launch. The "China model = cheap but meh" line keeps aging badly.&lt;/p&gt;

&lt;p&gt;Team "rank by benchmark" or team "rank by what people actually run" — which board you trust?&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.threerouter.com/blog/50" rel="noopener noreferrer"&gt;https://www.threerouter.com/blog/50&lt;/a&gt;&lt;/p&gt;

</description>
    </item>
    <item>
      <title>OpenAI’s Paid Share on OpenRouter Tops Anthropic</title>
      <dc:creator>FreyaLi</dc:creator>
      <pubDate>Tue, 22 Sep 2026 08:47:48 +0000</pubDate>
      <link>https://dev.to/976905690/openais-paid-share-on-openrouter-tops-anthropic-336p</link>
      <guid>https://dev.to/976905690/openais-paid-share-on-openrouter-tops-anthropic-336p</guid>
      <description>&lt;p&gt;OpenAI just ended Anthropic's 2.5-year run at the top of OpenRouter spend.&lt;/p&gt;

&lt;p&gt;Last week (week of Sep 7) was the first time since Feb 2024 that OpenAI models pulled in more developer dollars than Claude on the platform, just over half of the combined two-lab total.&lt;/p&gt;

&lt;p&gt;The part that actually matters: this is wallet-share, not token count. OpenRouter tracks what people really pay. Anthropic held 75–80% through most of 24–25 on premium pricing, so losing the dollar lead stings more than losing usage.&lt;/p&gt;

&lt;p&gt;Astra alone took ~19% of all platform spend. OpenAI's wider lineup (Luna / Terra / Sol + Astra) finished it.&lt;/p&gt;

&lt;p&gt;Meanwhile the open-weight camp kept swallowing the raw volume. US-model token share on OpenRouter dropped from ~70% to ~30% between mid-2025 and mid-2026, with DeepSeek and friends taking most of it.&lt;/p&gt;

&lt;p&gt;So the routing layer now tells two stories at once: closed models win the revenue, open-weight wins the usage. If you build on a router, which side are you betting on?&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.threerouter.com/blog/53" rel="noopener noreferrer"&gt;https://www.threerouter.com/blog/53&lt;/a&gt;&lt;/p&gt;

</description>
    </item>
    <item>
      <title>GPT6-Sol Gray Exposure + Delay</title>
      <dc:creator>FreyaLi</dc:creator>
      <pubDate>Tue, 22 Sep 2026 08:46:24 +0000</pubDate>
      <link>https://dev.to/976905690/gpt6-sol-gray-exposure-delay-1m26</link>
      <guid>https://dev.to/976905690/gpt6-sol-gray-exposure-delay-1m26</guid>
      <description>&lt;p&gt;Your API swapped models on you this week. No changelog, no announcement.&lt;/p&gt;

&lt;p&gt;Multiple users picked GPT-5.6 Sol and got responses stamped gpt-6-sol. Packet captures on Codex show the same swap. A V2EX thread pulled 2,700+ views documenting it. One tester ran the same prompt on both: 72k tokens in 9 minutes on the new one, 26 minutes on the old.&lt;/p&gt;

&lt;p&gt;Yesterday the rumor mill had a whole "GPT-6 Sol Thursday?" cover ready. Today Altman moved his "big launch" to next week instead. Tibo's take: "Sometimes physics can't be cheated."&lt;/p&gt;

&lt;p&gt;OpenAI hasn't confirmed the model name. But read together with the $200/mo Pro pause ten days ago, the picture is pretty clear: GPT-6 Sol is likely already serving traffic in grayscale, and the official launch slipped because capacity is tight.&lt;/p&gt;

&lt;p&gt;The model itself is half the story. Man, a frontier lab can reroute production traffic to a different model and you find out from a model_id field.&lt;/p&gt;

&lt;p&gt;Do you pin model versions in prod, or ride the default? Genuinely curious.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.threerouter.com/blog/55" rel="noopener noreferrer"&gt;https://www.threerouter.com/blog/55&lt;/a&gt;&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Qwen3-0.6B Quietly Integrated into Federal Register</title>
      <dc:creator>FreyaLi</dc:creator>
      <pubDate>Tue, 22 Sep 2026 08:45:16 +0000</pubDate>
      <link>https://dev.to/976905690/qwen3-06b-quietly-integrated-into-federal-register-4k9k</link>
      <guid>https://dev.to/976905690/qwen3-06b-quietly-integrated-into-federal-register-4k9k</guid>
      <description>&lt;p&gt;The US government's own rulebook site was quietly running searches through Alibaba's Qwen.&lt;/p&gt;

&lt;p&gt;The Federal Register, the official daily journal for federal rules, notices and presidential docs, run by NARA, had five search modes on its advanced page. Two of them were Qwen3:0.6B, Alibaba's sub-billion open model. One was literally labeled "512D, Recursive Splitting, Distilled, Prefixed."&lt;/p&gt;

&lt;p&gt;This wasn't a smuggled app. Qwen3:0.6B is Apache 2.0, open-weight. A contractor just pulled it off HuggingFace like any free embedding model and dropped it into the search stack. Same way they'd grab anything else.&lt;/p&gt;

&lt;p&gt;The part that lowkey breaks the whole "keep Chinese AI out" playbook: you can ban an app, block a domain, pass a law against a vendor. You can't un-publish a set of weights. Once they're on HuggingFace they turn up everywhere, including apparently the government's own directory of government documents.&lt;/p&gt;

&lt;p&gt;As of today those Qwen options are gone, pulled with zero announcement. officechai spotted it, then it vanished. Scrub or routine maintenance, who knows. But the pattern is the real story: open-source shows up where no one expected it.&lt;/p&gt;

&lt;p&gt;You can ban a chatbot. Can you ban a set of weights someone already pulled down?&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.threerouter.com/blog/56" rel="noopener noreferrer"&gt;https://www.threerouter.com/blog/56&lt;/a&gt;&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Qwen3.8-Omni-Flash: Multimodal Model Released</title>
      <dc:creator>FreyaLi</dc:creator>
      <pubDate>Tue, 22 Sep 2026 08:44:48 +0000</pubDate>
      <link>https://dev.to/976905690/qwen38-omni-flash-multimodal-model-released-22no</link>
      <guid>https://dev.to/976905690/qwen38-omni-flash-multimodal-model-released-22no</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fprco5dagg07kddrltefv.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fprco5dagg07kddrltefv.png" alt=" " width="800" height="459"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Alibaba's Qwen team shipped Qwen3.8-Omni-Flash today. The number that got me: overlapping multi-speaker meetings used to be a diarization disaster, 88% error. Now 3.4%.&lt;/p&gt;

&lt;p&gt;Rest of the package: native text/image/audio/video in, 1M context, +25% average over their own Qwen3.5-Omni-Plus across 29 tests, audio input pricing cut over 98%. There's a Realtime API too, first token under a second on most inputs. Plus two open toolkits. Qwen-MM-Plugins drops the model into Claude Code, Gemini CLI, Codex, OpenClaw. Qwen-Live-Harness does task delegation, the "ping me when the job's done" kind.&lt;/p&gt;

&lt;p&gt;ngl omni-skill-creator is the sleeper. Record yourself operating software, narrate why as you go, and it turns the video into a reusable agent skill. Show once, the agent learns.&lt;/p&gt;

&lt;p&gt;And it's already live on Vercel's AI Gateway, so the "open models are hard to reach from abroad" excuse keeps getting thinner. The gateway became the on-ramp.&lt;/p&gt;

&lt;p&gt;Would you rather record one video and let it learn, or keep writing agent skills by hand?&lt;/p&gt;

</description>
    </item>
  </channel>
</rss>
