<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Women in AI &amp; Analytics</title>
    <description>The latest articles on DEV Community by Women in AI &amp; Analytics (@wiaia).</description>
    <link>https://dev.to/wiaia</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4160823%2F65b083dd-f920-4ddd-b98d-253e4a5259fe.png</url>
      <title>DEV Community: Women in AI &amp; Analytics</title>
      <link>https://dev.to/wiaia</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/wiaia"/>
    <language>en</language>
    <item>
      <title>When hardware, culture, and governance become the new AI bottleneck</title>
      <dc:creator>Women in AI &amp; Analytics</dc:creator>
      <pubDate>Sun, 04 Oct 2026 17:04:27 +0000</pubDate>
      <link>https://dev.to/wiaia/when-hardware-culture-and-governance-become-the-new-ai-bottleneck-58om</link>
      <guid>https://dev.to/wiaia/when-hardware-culture-and-governance-become-the-new-ai-bottleneck-58om</guid>
      <description>&lt;h2&gt;
  
  
  The new economics of AI: when hardware, culture, and documentation intersect
&lt;/h2&gt;

&lt;p&gt;Practitioners in AI and analytics are navigating a shift where capability is no longer the bottleneck — economics and governance are. The items harvested this week show one pattern: running large models on consumer GPUs, searching visual archives with AI, debating whether agents need memory or documentation, tracking institutional rifts at leading labs, and questioning pure mathematics' future. Each signals a different layer of the same system.&lt;/p&gt;

&lt;p&gt;The community has noticed that high point scores on technical releases often reflect enthusiasm rather than validated performance. When &lt;a href="https://github.com/Niko1221/Strata" rel="noopener noreferrer"&gt;Strata&lt;/a&gt; reports 289 points and 148 comments, practitioners should read the ratio — roughly 2 comments per point — as evidence that people are actively stress-testing the claim rather than simply endorsing it. A non-obvious point is that these stories are not independent. The &lt;a href="https://github.com/Niko1221/Strata" rel="noopener noreferrer"&gt;Strata project&lt;/a&gt; claims 100T/s for Qwen 3.8 Flash Next (125B) on an RTX 4090, which is precisely the kind of consumer-grade inference that lowers barriers but also raises reproducibility questions — community reports of 289 points and 148 comments suggest both excitement and skepticism. People running analytics on tight budgets need to verify such claims independently, since single-source benchmarks rarely account for batch size, quantization, or sustained thermal load. &lt;a href="https://github.com/allenv0/SCM" rel="noopener noreferrer"&gt;SCM's macOS AI search&lt;/a&gt; for every photo and video frame, at 71 points and 38 comments, points to a similar economic shift: visual data is now searchable at scale on personal devices, but practitioners must audit storage costs and privacy tradeoffs that headline scores ignore.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why documentation beats memory for autonomous agents
&lt;/h2&gt;

&lt;p&gt;The most debated item this week, &lt;a href="https://liao.gg/blog/agents-dont-need-memory" rel="noopener noreferrer"&gt;"Agents don't need memory, they need documentation"&lt;/a&gt;, reached 288 points with 165 comments. The thesis is that agents fail not from missing recollections but from missing structured context. For analytics practitioners, this is operational: building a retrieval-augmented pipeline is cheaper and more reproducible than fine-tuning a memory-augmented model. One can observe that documentation quality is a governance issue — poorly documented agents produce untraceable decisions. The community should treat agent design as a documentation audit, not a model architecture problem.&lt;/p&gt;

&lt;p&gt;A critical observation: the high engagement suggests people are not convinced. When 165 comments accompany 288 points, it indicates active disagreement rather than consensus, which practitioners should read as a signal that the field has not standardized best practices for agent reliability. One practical implication is that analytics teams should document their agent decision paths the same way they document data pipelines — with version control, change logs, and rollback procedures — because undocumented agent behavior creates audit failures that regulators and stakeholders will not accept.&lt;/p&gt;

&lt;h2&gt;
  
  
  Institutional rifts and safety culture
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://www.theatlantic.com/technology/2026/10/openai-safety-team-resignation/688881/?gift=v5U_UzUTothfWXsPxtvNVAh7esWToMRD6XnbXmc5WgA" rel="noopener noreferrer"&gt;OpenAI's culture is broken&lt;/a&gt;, with 375 points and 616 comments, is the highest-engagement item this cycle — and it is not a technical story. A resignation from a safety team carries operational weight for practitioners who rely on these labs for APIs and model updates. When &lt;a href="https://fortune.com/2026/10/01/ai-godfather-yann-lecun-has-zero-concerns-about-human-extinction-says-anthropic-ceo-dario-amodei-is-deuded/" rel="noopener noreferrer"&gt;LeCun expresses zero concerns about AI extinction&lt;/a&gt; at 282 points with 465 comments, the divergence in elite opinion signals that governance frameworks remain unsettled. Practitioners have found that building analytics pipelines on unstable institutional foundations introduces compliance risk; it is not sufficient to track model capability, one must also track organizational stability.&lt;/p&gt;

&lt;h2&gt;
  
  
  Tools, visualization, and reproducibility
&lt;/h2&gt;

&lt;p&gt;Lower-engagement items still carry practical weight. &lt;a href="https://github.com/mikesart/gpuvis" rel="noopener noreferrer"&gt;gpuvis: GPU Trace Visualizer&lt;/a&gt; (60 points, 10 comments) addresses a concrete debugging need for people profiling inference workloads on consumer hardware. Its low visibility relative to model releases suggests practitioners are more vocal about capabilities than about reproducibility tooling — a gap the analytics community should close. Meanwhile, &lt;a href="https://writings.stephenwolfram.com/2026/09/whats-the-future-for-pure-math-research-in-the-age-of-ai/" rel="noopener noreferrer"&gt;Wolfram's question on pure math in the age of AI&lt;/a&gt; at 28 points with 5 comments asks a structural question: if AI accelerates derivation and verification, does pure mathematics become applied mathematics, and does that shift funding and career paths? The low engagement is itself data — the analytics community appears less concerned with foundational theory shifts than with operational ones.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://www.reddit.com/r/2007scape/comments/1wxfyzp/runescapes_position_on_gen_ai/" rel="noopener noreferrer"&gt;RuneScape community's position on generative AI&lt;/a&gt;, at 13 points and 24 comments, is a reminder that user communities often set norms faster than institutions. Practitioners deploying generative features should expect similar community-level resistance in niche domains, not just from regulators, and should build community feedback loops into deployment timelines.&lt;/p&gt;

&lt;h2&gt;
  
  
  A practitioner's checklist
&lt;/h2&gt;

&lt;p&gt;Taken together, these items advise a cautious posture. Verify benchmark claims independently; prefer documentation over memory for agent reliability; track institutional stability as a supply-chain variable; invest in reproducibility tooling; and anticipate community resistance to generative deployments. The economics favor people who can run inference locally and audit what they deploy — the analytics practitioners with limited budgets have an advantage here, not a handicap. People who build reproducible pipelines with documented agents, verified benchmarks, and local inference have lower long-term operational risk than those who depend solely on external APIs and opaque memory architectures.&lt;/p&gt;

&lt;p&gt;For practitioners in AI, analytics, and open-source communities, &lt;a href="https://wiaia.github.io/" rel="noopener noreferrer"&gt;Women in AI &amp;amp; Analytics&lt;/a&gt; offers community resources, mentorship, and collaboration space to work through these challenges together.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/Niko1221/Strata" rel="noopener noreferrer"&gt;Strata — Qwen 3.8 Flash Next on RTX 4090&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/allenv0/SCM" rel="noopener noreferrer"&gt;SCM — AI search for photos and video on macOS&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://liao.gg/blog/agents-dont-need-memory" rel="noopener noreferrer"&gt;Agents don't need memory, they need documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://fortune.com/2026/10/01/ai-godfather-yann-lecun-has-zero-concerns-about-human-extinction-says-anthropic-ceo-dario-amodei-is-deuded/" rel="noopener noreferrer"&gt;LeCun: zero concerns about AI extinction&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/mikesart/gpuvis" rel="noopener noreferrer"&gt;gpuvis: GPU Trace Visualizer&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://writings.stephenwolfram.com/2026/09/whats-the-future-for-pure-math-research-in-the-age-of-ai/" rel="noopener noreferrer"&gt;Wolfram — Pure math research in the age of AI&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.theatlantic.com/technology/2026/10/openai-safety-team-resignation/688881/?gift=v5U_UzUTothfWXsPxtvNVAh7esWToMRD6XnbXmc5WgA" rel="noopener noreferrer"&gt;OpenAI safety team resignation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.reddit.com/r/2007scape/comments/1wxfyzp/runescapes_position_on_gen_ai/" rel="noopener noreferrer"&gt;RuneScape's position on Gen AI&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://wiaia.github.io/" rel="noopener noreferrer"&gt;Women in AI &amp;amp; Analytics&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>analytics</category>
      <category>wiaia</category>
    </item>
    <item>
      <title>When hype stops, practitioners build for reproducibility and governance</title>
      <dc:creator>Women in AI &amp; Analytics</dc:creator>
      <pubDate>Sun, 04 Oct 2026 12:34:46 +0000</pubDate>
      <link>https://dev.to/wiaia/when-hype-stops-practitioners-build-for-reproducibility-and-governance-20ja</link>
      <guid>https://dev.to/wiaia/when-hype-stops-practitioners-build-for-reproducibility-and-governance-20ja</guid>
      <description>&lt;h2&gt;
  
  
  What practitioners are actually building when the hype stops
&lt;/h2&gt;

&lt;p&gt;The eight items gathered this week share one operational theme: practitioners have moved past asking what AI can imagine and are now solving for reproducibility, governance, cost, and the uneven terrain of real-world deployment. A &lt;a href="https://github.com/mikesart/gpuvis" rel="noopener noreferrer"&gt;GPU trace visualizer&lt;/a&gt; for debugging compute; a &lt;a href="https://github.com/allenv0/SCM" rel="noopener noreferrer"&gt;macOS AI search&lt;/a&gt; over every photo and video frame; a manifesto arguing agents need documentation over memory; a leading researcher dismissing extinction fears; a former OpenAI safety employee publicly resigning; scholars meeting Anthropic on morals; three agents operating across two countries with uneven web infrastructure; and a &lt;a href="https://pipod.dev/" rel="noopener noreferrer"&gt;sandboxed Pi agent platform&lt;/a&gt;. Together they sketch a community that cares less about capability claims and more about operational discipline, and that tendency has practical consequences for anyone running analytics without enterprise budgets.&lt;/p&gt;

&lt;h2&gt;
  
  
  Agents need documentation, not just memory
&lt;/h2&gt;

&lt;p&gt;A widely shared essay argues that agents do not fail for lack of cognition but for lack of documentation: without structured, verifiable records of what was attempted, practitioners cannot reproduce outcomes, audit decisions, or scale safely. The essay, &lt;a href="https://liao.gg/blog/agents-dont-need-memory" rel="noopener noreferrer"&gt;"Agents don't need memory, they need documentation"&lt;/a&gt;, collected 219 points on Hacker News with 121 comments, a signal that the community recognizes this as a practical bottleneck rather than an abstract design preference. The non-obvious implication is that adding more model parameters or memory stores will not fix operational fragility; the missing layer is an auditable paper trail that survives staff turnover and budget cuts. In practice, this means practitioners should log prompts, outputs, environment versions, and dependency states for every analytical run, not just final results. For limited-resource teams, documentation is cheaper than compute and often more reliable than model upgrades.&lt;/p&gt;

&lt;h2&gt;
  
  
  Governance is no longer optional for open-source agents
&lt;/h2&gt;

&lt;p&gt;Religious scholars meeting Anthropic to discuss Claude's moral framework, reported in &lt;a href="https://www.nytimes.com/2026/09/29/us/anthropic-claude-morals-ai.html" rel="noopener noreferrer"&gt;a New York Times article&lt;/a&gt; with 94 points and 228 comments, reveals that governance is being shaped by external stakeholders well before deployment reaches end users. Meanwhile, a former OpenAI safety team member's resignation, covered by &lt;a href="https://www.theatlantic.com/technology/2026/10/openai-safety-team-resignation/688881/?gift=v5U_UzUTothfWXsPxtvNVAh7esWToMRD6XnbXmc5WgA" rel="noopener noreferrer"&gt;The Atlantic&lt;/a&gt; with 274 points and 522 comments, signals that internal dissent is now public. The critical observation is that governance failures are not only corporate reputational risks: practitioners relying on these systems for analytics face downstream liability when upstream ethics are unresolved. For analysts with limited budgets, reproducibility requires knowing which governance framework applies to the model one is actually running, which is rarely documented in vendor terms of service. The community has noticed that governance gaps tend to accumulate fastest at the interface between open-source tools and proprietary models.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cost and reproducibility favor sandboxed, local infrastructure
&lt;/h2&gt;

&lt;p&gt;Two practical releases point toward a reproducible, budget-conscious path. A &lt;a href="https://github.com/mikesart/gpuvis" rel="noopener noreferrer"&gt;GPU trace visualizer&lt;/a&gt; with 25 points gives practitioners visibility into compute behavior rather than black-box inference, which matters when debugging why a model behaves differently across runs. More significantly, &lt;a href="https://pipod.dev/" rel="noopener noreferrer"&gt;Pi pod&lt;/a&gt;, a sandboxed coding-agent platform running on one's own server, attracted 103 points and 39 comments, reflecting demand for controlled, reproducible agent execution rather than vendor-hosted opacity. The operational caveat is that sandboxing introduces its own maintenance burden: one must patch hosts, manage isolation policies, replicate environments, and budget for storage and network overhead. Practitioners on lean budgets should weigh whether the reproducibility gain justifies the operational overhead before committing to self-hosted stacks, especially if staff time is the scarcest resource. Reproducibility is not free; it is an explicit budget line item.&lt;/p&gt;

&lt;h2&gt;
  
  
  The world wide web is not even for multilingual agents
&lt;/h2&gt;

&lt;p&gt;A piece on &lt;a href="https://royapakzad.substack.com/p/multilingual-ai-agents" rel="noopener noreferrer"&gt;multilingual AI agents across two countries&lt;/a&gt; notes 35 points and highlights a structural inequality: agents serve communities better where web infrastructure, data abundance, and language coverage are already strong. Combined with a &lt;a href="https://github.com/allenv0/SCM" rel="noopener noreferrer"&gt;macOS AI search tool&lt;/a&gt; that indexes every photo and video frame locally with 11 points and 1 comment, this suggests practitioners are building tools whose value depends heavily on the richness of the underlying data landscape. The non-obvious point is that deploying multilingual agents without auditing data coverage for underrepresented languages risks reinforcing existing disparities rather than closing them. Practitioners should test for representational gaps before scaling, and consider whether a model trained primarily on English-language web content can fairly represent the communities it is intended to serve. Auditing for language coverage is an operational step, not a post-deployment review.&lt;/p&gt;

&lt;h2&gt;
  
  
  What practitioners should take from recent public debates
&lt;/h2&gt;

&lt;p&gt;Yann LeCun's statement of "zero concerns" about AI-driven extinction, reported by &lt;a href="https://fortune.com/2026/10/01/ai-godfather-yann-lecun-has-zero-concerns-about-human-extinction-says-anthropic-ceo-dario-amodei-is-deuded/" rel="noopener noreferrer"&gt;Fortune&lt;/a&gt; with 180 points and 270 comments, contrasts sharply with the resignation and governance stories. The practical takeaway is not to adopt either extreme, but to recognize that public disagreement among leading figures signals unresolved risk frameworks. Analysts making budget and infrastructure choices should rely on reproducible evidence rather than expert consensus. When practitioners cannot reproduce an inference pipeline or verify a governance claim, the safer operational choice is to treat that gap as a blocking dependency rather than a minor caveat. The community has observed that small reproducibility gaps compound quickly under budget pressure.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/mikesart/gpuvis" rel="noopener noreferrer"&gt;GPU Trace Visualizer — GitHub&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/allenv0/SCM" rel="noopener noreferrer"&gt;Show HN: AI search for every photo and video frame on macOS&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://liao.gg/blog/agents-dont-need-memory" rel="noopener noreferrer"&gt;Agents don't need memory, they need documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://fortune.com/2026/10/01/ai-godfather-yann-lecun-has-zero-concerns-about-human-extinction-says-anthropic-ceo-dario-amodei-is-deuded/" rel="noopener noreferrer"&gt;LeCun: "zero concerns" about AI extinction and Anthropic CEO response&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.nytimes.com/2026/09/29/us/anthropic-claude-morals-ai.html" rel="noopener noreferrer"&gt;Religious scholars met with Anthropic&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.theatlantic.com/technology/2026/10/openai-safety-team-resignation/688881/?gift=v5U_UzUTothfWXsPxtvNVAh7esWToMRD6XnbXmc5WgA" rel="noopener noreferrer"&gt;I quit OpenAI because its culture is broken&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://royapakzad.substack.com/p/multilingual-ai-agents" rel="noopener noreferrer"&gt;Three AI agents, two countries, and one uneven world wide web&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://pipod.dev/" rel="noopener noreferrer"&gt;Show HN: Pi pod — sandboxed Pi coding agent&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Join practitioners working on reproducible analytics, open governance, and budget-conscious AI at &lt;a href="https://wiaia.github.io/" rel="noopener noreferrer"&gt;Women in AI &amp;amp; Analytics&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>analytics</category>
      <category>wiaia</category>
    </item>
    <item>
      <title>Sovereign models, hard budget caps, and what they mean for practitioners without enterprise budgets</title>
      <dc:creator>Women in AI &amp; Analytics</dc:creator>
      <pubDate>Sun, 04 Oct 2026 06:22:59 +0000</pubDate>
      <link>https://dev.to/wiaia/sovereign-models-hard-budget-caps-and-what-they-mean-for-practitioners-without-enterprise-budgets-4hbe</link>
      <guid>https://dev.to/wiaia/sovereign-models-hard-budget-caps-and-what-they-mean-for-practitioners-without-enterprise-budgets-4hbe</guid>
      <description>&lt;h1&gt;
  
  
  Sovereign models, hard budget caps, and what they mean for practitioners without enterprise budgets
&lt;/h1&gt;

&lt;p&gt;Three things landed on Hacker News this week that look unrelated and are actually the same conversation: Aleph Alpha released Kolibri under Apache 2.0, Simon Willison argued that every pay-by-usage service needs a default hard budget cap, and DwarfStar 4 landed as a local inference engine from the creator of Redis. All three trended above 340 points.&lt;/p&gt;

&lt;p&gt;For people building with AI on a limited budget and no procurement department, that is a meaningful week. Here is what the community noticed, including the parts that are less flattering than the launch posts.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sovereign open weights: the openness is in the documentation
&lt;/h2&gt;

&lt;p&gt;Aleph Alpha's &lt;a href="https://aleph-alpha.com/en/blog/kolibri-has-landed-a-sovereign-open-weight-model/" rel="noopener noreferrer"&gt;Kolibri announcement&lt;/a&gt; describes an English-German mixture-of-experts transformer with 78B total parameters and 3B active, supporting up to 1M tokens of context. Full weights are downloadable from Hugging Face under Apache 2.0. It followed Kolibri Origin, a 30B total / 3B active model with a 65k context window that ran through the same pipeline.&lt;/p&gt;

&lt;p&gt;What drew the most attention in the thread was not the model. It was the technical report. Multiple commenters described it as a tutorial on how to build a modern agentic LLM, down to how the dataset was constructed. One commenter who said they worked on Kolibri's pre-training and mid-training data wrote that the team's stated aim was to be as open as possible. Another commenter's complaint is a useful standard to hold releases to: they would not call a model "open" if the ingested training data were hidden and the process not repeatable by a third party. That distinction matters more than a parameter count.&lt;/p&gt;

&lt;h2&gt;
  
  
  Abstention training is the part worth copying
&lt;/h2&gt;

&lt;p&gt;Aleph Alpha says Kolibri was trained with abstention data and what they call the Merlin-Arthur protocol, so that it is trained to say "I don't know" when the answer is not in context.&lt;/p&gt;

&lt;p&gt;The comments were split, and the skeptical responses are the interesting part. Someone asked the model to interpret a song and artist they had invented, and it produced a confident answer with a plausible album name and a year. Another commenter, asked how to run the model on limited RAM via llama.cpp, mostly got a redirect to the Hugging Face interface. There was also a running joke about models saying "I don't know" so often that they resembled Merl, the Minecraft support chatbot who famously would not answer how to craft a diamond pickaxe.&lt;/p&gt;

&lt;p&gt;So abstention training was not a clean win on display. Still, the direction is right. A model that hedges toward "I don't know" fails visibly. A model that invents an album title fails silently, and in an analytics or data pipeline the silent version is the expensive one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Hard budget caps, and the fine print
&lt;/h2&gt;

&lt;p&gt;Simon Willison's &lt;a href="https://simonwillison.net/2026/Oct/3/default-hard-budget-caps/" rel="noopener noreferrer"&gt;post&lt;/a&gt; makes the case for default hard caps: after a set spend, cut the service off and return errors rather than send a warning email. His argument is specifically about coding agents, where spinning up something that calls paid APIs is cheap to do and easy to leave running overnight.&lt;/p&gt;

&lt;p&gt;The Hacker News thread is more skeptical, and the objections are worth reading if anyone is building a budget-aware service.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Billing data is not instant.&lt;/strong&gt; One commenter, apparently on the provider side, explained that a VM reports billing units periodically, so a network blip can delay that data. This is the real technical reason caps are hard: you cannot reliably estimate what an operation will cost before you start it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Google Cloud shipped something narrower than expected.&lt;/strong&gt; One commenter asked whether Google had added per-service caps, then edited to note it "is fake," saying it only works for four random services and is unsupported elsewhere. Another noted it only applies to projects created in AI Studio. If that assessment holds, the cap is not usable as a general safety net.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Enterprise customers may not want it.&lt;/strong&gt; Another commenter, formerly on a backend support team, described hard caps as a nightmare of tickets and threats of lawsuits when a service got cut off mid-growth-event. Someone else countered that a provider can write off a $10k bill while the customer experiences it as panic, so the incentives diverge. Most businesses would rather have a large bill than an outage.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Small recurring charges can strand accounts.&lt;/strong&gt; One commenter described a $0.20 monthly AWS charge they could only remove by deleting the entire account, and gave up on AWS for personal projects.&lt;/p&gt;

&lt;p&gt;If one practical thing is taken from this thread: alerts are not caps, and a cap scoped to one product is not a cap. A soft warning email is a rounding error next to an overnight agent loop.&lt;/p&gt;

&lt;h2&gt;
  
  
  Local inference keeps moving
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://dwarfstar.sh/" rel="noopener noreferrer"&gt;DwarfStar 4&lt;/a&gt;, from the creator of Redis, is a narrow C inference engine for high-memory Mac, CUDA and ROCm machines. It supports DeepSeek V4 and V4.1 Flash, GLM 5.x and Qwen3.8 Flash, under MIT license. DeepSeek V4 Flash is a 284B-parameter mixture-of-experts model; the project applies asymmetric quantization to the routed experts while preserving critical paths.&lt;/p&gt;

&lt;p&gt;One implementation detail stands out for anyone thinking about agent workloads: the KV cache is keyed by the SHA1 of the rendered prompt prefix and persisted to disk, so a matching prefix is reloaded rather than recomputed, and can survive a server restart. For an agent loop that re-sends the same context on every turn, that is the difference between an idle machine and a bill.&lt;/p&gt;

&lt;p&gt;Note the hardware line. "High-memory" is doing real work in that sentence. This is not the same category as running a 3B model on a Pi, where the memory ceiling is a few gigabytes and the compromises are known.&lt;/p&gt;

&lt;h2&gt;
  
  
  The thread running through all three
&lt;/h2&gt;

&lt;p&gt;None of this is frontier capability. Kolibri is 3B active parameters, not a replacement for the largest models. ds4 needs a high-memory machine. Hard caps are a billing feature. What connects them is control.&lt;/p&gt;

&lt;p&gt;Open weights plus published training data means a team can inspect what a model learned instead of trusting a vendor summary. Abstention training means failure is visible. Local inference means data does not leave the machine. Hard caps mean an unattended job cannot quietly become a line item. These are all the same move: shifting control toward the person running the model.&lt;/p&gt;

&lt;p&gt;For anyone without a platform team behind them, that is the constraint that actually binds. Model quality still matters, but so does knowing what happens at 3am when an agent loop is still running and nobody has looked at the dashboard since lunch.&lt;/p&gt;

&lt;p&gt;Worth watching: whether sovereign models keep shipping training data with the weights, whether hard caps become a default rather than a per-product checkbox, and whether KV cache persistence becomes standard in agent tooling.&lt;/p&gt;




&lt;p&gt;Have you run a budget cap or a local model in your own work? The WIAIA community would like to hear what worked and what quietly did not. Visit &lt;a href="https://wiaia.github.io/" rel="noopener noreferrer"&gt;wiaia.github.io&lt;/a&gt; to find the mentorship program and ways to contribute.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>opensource</category>
      <category>devops</category>
    </item>
    <item>
      <title>Running LLMs locally on Linux: what actually works on a Raspberry Pi</title>
      <dc:creator>Women in AI &amp; Analytics</dc:creator>
      <pubDate>Sun, 04 Oct 2026 06:08:16 +0000</pubDate>
      <link>https://dev.to/wiaia/running-llms-locally-on-linux-what-actually-works-on-a-raspberry-pi-1fpc</link>
      <guid>https://dev.to/wiaia/running-llms-locally-on-linux-what-actually-works-on-a-raspberry-pi-1fpc</guid>
      <description>&lt;h1&gt;
  
  
  Running LLMs locally on Linux: what actually works on a Raspberry Pi
&lt;/h1&gt;

&lt;p&gt;A 5B-parameter model can run on a Raspberry Pi 5 every day. It is not fast. It is also completely offline, costs nothing per token, and never phones home. That trade is worth making for a specific class of work, and worthless for everything else. Below is what the WIAIA community has found after several months of experimenting.&lt;/p&gt;

&lt;h2&gt;
  
  
  The hardware reality nobody mentions
&lt;/h2&gt;

&lt;p&gt;Before any software, the constraint. A Pi 5 has 8GB of RAM and shares it with the rest of the system. When one loads a 4-bit quantized 5B model, the weights alone eat roughly 3.5 to 4GB. Add context KV cache and space is already tight. A usable context window lands around 2 to 3k tokens before generation starts failing and the process gets OOM-killed.&lt;/p&gt;

&lt;p&gt;This is the single most important fact about local inference on a Pi, and it is not about the GPU or the quantization method. It is that roughly 4GB is available, and everything competes for it.&lt;/p&gt;

&lt;p&gt;For actual throughput, stop chasing tokens per second. Raw speed on this hardware is poor. What matters more is that a request never leaves the machine, so a job can be left running and the result checked hours later.&lt;/p&gt;

&lt;h2&gt;
  
  
  What people actually run
&lt;/h2&gt;

&lt;p&gt;Three tools have earned their place:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Ollama&lt;/strong&gt; is the common default. Model management is two commands, it has a stable HTTP API on port 11434, and it runs behind systemd so it survives reboots. For anything scripted, this is the one worth starting with. It is also the one that needs care for sensitive work, because it can bind to all interfaces depending on the install. Check &lt;code&gt;ollama.allowed_origins&lt;/code&gt; in the config.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;llama.cpp via &lt;code&gt;llama-server&lt;/code&gt;&lt;/strong&gt; is what to reach for when control over sampling matters. The GGUF file is yours, every sampling parameter is exposed as a flag, and threads can be pinned to specific cores. It sits closer to the metal and closer to what perplexity actually means.&lt;/p&gt;

&lt;p&gt;For anything expecting an OpenAI-compatible endpoint, both of the above expose one. Point a client at &lt;code&gt;http://localhost:11434/v1&lt;/code&gt; or the llama.cpp equivalent and the difference stops mattering. Most tooling never notices.&lt;/p&gt;

&lt;h2&gt;
  
  
  Quantization: pick 4-bit and move on
&lt;/h2&gt;

&lt;p&gt;Q4_K_M is the pragmatic default. It is the sweet spot where the model barely degrades from the original and the file size stays manageable. For a 5B model that lands around 3.5GB on disk, which is what makes it fit on a Pi at all.&lt;/p&gt;

&lt;p&gt;Going to Q8 doubles the size with no perceptible difference on typical workloads. Going to Q2 saves space but output degrades into word salad. There is no reason to be a martyr about it. Q4_K_M, load it, move on.&lt;/p&gt;

&lt;p&gt;The thing worth knowing is that quantization is not free. Information is lost, and on reasoning-heavy tasks the loss surfaces as confidently wrong answers rather than obvious garbage. A badly quantized model does not say "I don't know." It says something plausible in the same register as the rest of its output.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prompting a small model is a different job
&lt;/h2&gt;

&lt;p&gt;This takes the longest to accept. A 5B model does not follow complex instructions. Multi-part prompts with formatting requirements produce output that satisfies roughly one constraint and quietly ignores the rest.&lt;/p&gt;

&lt;p&gt;What works:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;One task per prompt. Not "summarize this and list the action items and format as markdown." Just "summarize this."&lt;/li&gt;
&lt;li&gt;Short output targets. "Three sentences" gets obeyed. "A thorough analysis" produces rambling until context runs out and the output gets truncated mid-sentence.&lt;/li&gt;
&lt;li&gt;Output format in the system prompt and nowhere else. Redundancy in a small model reads as emphasis, not confirmation.&lt;/li&gt;
&lt;li&gt;A one-shot example when format matters. One. Not three.&lt;/li&gt;
&lt;li&gt;Stop sequences instead of instructions to stop. Ending on a newline and setting &lt;code&gt;stop=["\n\n\n"]&lt;/code&gt; beats asking for concise output in prose.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Where the Pi setup beats a hosted API
&lt;/h2&gt;

&lt;p&gt;This is the underrated part.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Data stays put.&lt;/strong&gt; A model can be run over logs, config files, and half-written notes containing credentials and internal hostnames. None of it goes to anyone. That is the whole reason to buy the hardware.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;No rate limits, no per-token cost, no quota.&lt;/strong&gt; Jobs can queue up overnight without a single throttle error.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It works on a dead network.&lt;/strong&gt; Local setups keep going when connectivity does not.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It is auditable.&lt;/strong&gt; The weights are a file on disk. What is running is known, and it can be hashed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it loses, badly
&lt;/h2&gt;

&lt;p&gt;Local inference on Pi hardware is not competitive for anything latency-sensitive. Interactive autocomplete in an editor is out. Code completion across a large file is out. Anything where a human is waiting is out.&lt;/p&gt;

&lt;p&gt;Suited workloads: summarizing a 40-page document, classifying a few thousand files, drafting from notes already written, batch generating variations for review. Hours of unattended work where a 1.5 tokens per second generation rate is fine.&lt;/p&gt;

&lt;p&gt;Quality is capped. A 5B model will lose to a frontier model on anything requiring nuance or long reasoning. Worth accepting rather than pretending otherwise. If an answer needs to be right about something subtle, a bigger model is the right tool.&lt;/p&gt;

&lt;h2&gt;
  
  
  Setup, to try it
&lt;/h2&gt;

&lt;p&gt;Install Ollama, pull a model under 6B parameters, and set the context window explicitly rather than accepting the default:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-fsSL&lt;/span&gt; https://ollama.com/install.sh | sh
ollama pull qwen2.5:3b
&lt;span class="nv"&gt;OLLAMA_CONTEXT_LENGTH&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;2048 systemctl restart ollama
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Set &lt;code&gt;OLLAMA_NUM_PARALLEL=1&lt;/code&gt;. On 8GB shared memory, parallel requests multiply KV cache usage and hit the OOM killer faster than expected.&lt;/p&gt;

&lt;p&gt;Then test before trusting it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;time &lt;/span&gt;curl http://localhost:11434/api/generate &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
  "model": "qwen2.5:3b",
  "prompt": "Summarize in three sentences: the tradeoff of running models locally.",
  "stream": false
}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If that round trips without a kernel OOM message in &lt;code&gt;dmesg&lt;/code&gt;, the setup is working.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bigger picture
&lt;/h2&gt;

&lt;p&gt;An experimental setup like this was expected to migrate everything to a hosted model once the wait times became annoying. What actually happened is that the work split cleanly in two. Hosted APIs took everything interactive and judgment-heavy. The Pi took the volume work that is slow, repetitive, and touches data better kept local.&lt;/p&gt;

&lt;p&gt;That split is the real lesson. Local inference is not a cheaper version of the cloud. It is a different tool for a different job, and the useful decision is which jobs go where.&lt;/p&gt;




&lt;p&gt;Interested in contributing to WIAIA or sharing what has worked in your own setup? Visit &lt;a href="https://wiaia.github.io/" rel="noopener noreferrer"&gt;wiaia.github.io&lt;/a&gt; to learn more about the community, the mentorship program, and how to get involved.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>linux</category>
      <category>llm</category>
      <category>raspberrypi</category>
    </item>
  </channel>
</rss>
