<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Viv Chaudhary</title>
    <description>The latest articles on DEV Community by Viv Chaudhary (@viv_chaudhary).</description>
    <link>https://dev.to/viv_chaudhary</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3659541%2Fd2551d9c-4ad3-4b72-85d3-d77f212506d2.png</url>
      <title>DEV Community: Viv Chaudhary</title>
      <link>https://dev.to/viv_chaudhary</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/viv_chaudhary"/>
    <language>en</language>
    <item>
      <title>Build a Solid RAG System in One Weekend (2026 Edition)</title>
      <dc:creator>Viv Chaudhary</dc:creator>
      <pubDate>Tue, 08 Sep 2026 17:23:26 +0000</pubDate>
      <link>https://dev.to/viv_chaudhary/build-a-solid-rag-system-in-one-weekend-2026-edition-5g9h</link>
      <guid>https://dev.to/viv_chaudhary/build-a-solid-rag-system-in-one-weekend-2026-edition-5g9h</guid>
      <description>&lt;p&gt;The four deliberate decisions that separate a working RAG demo from a reliable RAG system are chunking, retrieval, reranking, and evaluation. Tutorials that stop at "load PDFs → fixed-size chunks → embed → top-k into a prompt" produce fluent answers that fail silently on the exact queries that matter most - error codes, SKUs, function names, specific clauses. The gap is almost never a fancier model; it's these four choices made by default instead of by design.&lt;/p&gt;

&lt;p&gt;This post is grounded in a real Confluence-backed system I built and run — hybrid dense + BM25 with RRF, cross-encoder reranking, structure-preserving ingestion, retrieval-informed routing, sentence-level groundedness verification, and a golden-set eval harness. The accompanying repository (&lt;a href="https://github.com/chaudharyviv/confluence-rag-platform" rel="noopener noreferrer"&gt;github.com/chaudharyviv/confluence-rag-platform&lt;/a&gt;) implements exactly that architecture.&lt;/p&gt;

&lt;h2&gt;
  
  
  Decision 1: Chunking sets the hard ceiling
&lt;/h2&gt;

&lt;p&gt;You cannot retrieve context that was split across a boundary, and you cannot generate a faithful answer from a chunk that lost half its meaning. Chunking is therefore not a preprocessing detail; it determines the maximum achievable quality of everything downstream.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Practical default that works across most corpora&lt;/strong&gt;: recursive character/token splitting at natural boundaries (paragraphs → sentences → words) with a hierarchy of separators, targeting roughly 400–512 tokens and 10–20% overlap. Independent benchmarks including Chroma's, place this in the 85–90% recall band for typical content, which is why it's the sane starting point rather than a compromise (&lt;a href="https://www.firecrawl.dev/blog/best-chunking-strategies-rag" rel="noopener noreferrer"&gt;Firecrawl&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;When to deviate&lt;/strong&gt; - match the structure of the source, don't impose a uniform token window:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Page-level (or section-level) units&lt;/strong&gt; win on paginated documents that contain tables and figures. NVIDIA's 2024 benchmarks showed this beating more elaborate methods precisely because mid-table or mid-figure splits destroy meaning (&lt;a href="https://www.firecrawl.dev/blog/best-chunking-strategies-rag" rel="noopener noreferrer"&gt;Firecrawl&lt;/a&gt;).&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Semantic chunking&lt;/strong&gt; - splitting where sentence-embedding similarity actually shifts topic can add up to ~9% recall, but it costs an embedding pass over every sentence at ingestion time. Use it when topic boundaries are genuinely fuzzy and the extra cost is justified (&lt;a href="https://www.firecrawl.dev/blog/best-chunking-strategies-rag" rel="noopener noreferrer"&gt;Firecrawl&lt;/a&gt;).&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;LLM-based chunking&lt;/strong&gt; works but is prohibitively expensive at production scale; reserve it for small, high-value corpora.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Structure-preserving / domain-specific&lt;/strong&gt;: for Confluence XHTML (or any hierarchical format), respect the actual document tree headings, tables, code blocks rather than flattening everything into a token stream. My own system does exactly this: heading breadcrumbs, Markdown tables, and fenced code blocks are preserved, then a token-aware sliding window runs &lt;em&gt;inside&lt;/em&gt; each section so tables and code are never split mid-block. This is simply the page/section principle applied to the real source format.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb&lt;/strong&gt;: start with recursive splitting unless you have a concrete, named reason to change mostly tables means page-level, fuzzy topics with high recall value means semantic. Don't default to LLM chunking.&lt;/p&gt;

&lt;p&gt;One ingestion-scoping tip that matters before any of this: my own implementation currently pulls pages from a single space via Confluence's REST content API (&lt;code&gt;spaceKey&lt;/code&gt; as a query param), which is fine for one knowledge base but doesn't scale to an org with many spaces. You'd either have to loop over every space key or accept everything in each space, tables of contents and archived junk included. The better approach once you're past one space is &lt;strong&gt;CQL (Confluence Query Language) with label-based filtering&lt;/strong&gt; - a query like &lt;code&gt;label = "rag-kb" AND type = page&lt;/code&gt; (via the &lt;code&gt;/rest/api/content/search&lt;/code&gt; endpoint, or &lt;code&gt;/rest/api/search&lt;/code&gt; on newer Cloud instances) finds every page tagged with a given label &lt;em&gt;across every space&lt;/em&gt; in one call, regardless of which space it lives in. That turns "which pages belong in this knowledge base" into an explicit, auditable label you attach in Confluence, rather than an implicit "whatever's in this space" assumption baked into the ingestion code. It's the same principle as the chunking rule above, one level up: don't let your source system's default unit (a whole space) define your corpus boundary define the boundary on purpose, then query for it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Decision 2: Retrieval - pure dense search has a predictable, silent failure mode
&lt;/h2&gt;

&lt;p&gt;Dense (embedding) retrieval is excellent at semantic matching: "cancel subscription" finds "terminate your plan." It systematically degrades on lexically exact, rare tokens — error codes like &lt;code&gt;ERR_SSL_VERSION_OR_CIPHER_MISMATCH&lt;/code&gt;, product SKUs, function names like &lt;code&gt;torch.nn.functional.cross_entropy&lt;/code&gt;. Embedding pooling averages the rare token's signal into the surrounding context; BM25's inverted index does not. The failure is dangerous because the system still returns &lt;em&gt;something&lt;/em&gt;, the LLM still produces a fluent answer, and the answer is simply grounded in the wrong document (&lt;a href="https://tianpan.co/blog/2026-04-12-hybrid-search-production-bm25-dense-embeddings" rel="noopener noreferrer"&gt;TianPan.co&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;The practical fix is hybrid retrieval: run BM25 (sparse) and dense search in parallel, then fuse. Reciprocal Rank Fusion (RRF) is the right default. Each document's fused score is the sum, across every ranked list it appears in, of &lt;code&gt;1 / (k + rank)&lt;/code&gt; with &lt;code&gt;k = 60&lt;/code&gt; by convention. RRF is score-scale agnostic (you never have to normalize cosine similarity against BM25's TF-IDF scores) and needs no labeled training data, which is why it ships as the default in Elasticsearch, Weaviate, and Qdrant (&lt;a href="https://tianpan.co/blog/2026-04-12-hybrid-search-production-bm25-dense-embeddings" rel="noopener noreferrer"&gt;TianPan.co&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;Caveat from the literature: RRF typically adds only ~1.3% NDCG over BM25 alone. A properly tuned weighted (convex) combination of the two scores can reach ~7.5%. Start with RRF because it requires zero tuning data; once you have a golden set (Decision 4) you can calibrate weights and measure the gain (&lt;a href="https://tianpan.co/blog/2026-04-12-hybrid-search-production-bm25-dense-embeddings" rel="noopener noreferrer"&gt;TianPan.co&lt;/a&gt;). My own system uses exactly this pattern - Chroma for dense, &lt;code&gt;rank_bm25&lt;/code&gt; for sparse, hand-rolled RRF because self-hosted Chroma has no built-in hybrid search.&lt;/p&gt;

&lt;h2&gt;
  
  
  Decision 3: Reranking is retrieval's second opinion
&lt;/h2&gt;

&lt;p&gt;Hybrid retrieval gives you a solid candidate pool, typically the top 20–50. It does not give you the best possible ordering of the 3–5 chunks that will actually enter the LLM's context window. Both bag-of-words and bi-encoder similarity are cheap approximations computed without ever looking at the query and document &lt;em&gt;together&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;A cross-encoder takes the query and each candidate as a single joint input and scores relevance directly. It's too expensive to run over the whole corpus, which is why retrieval happens first to narrow the field, but running it over 20–50 candidates is cheap and consistently improves the final ranking. My own system uses the small, self-hostable &lt;code&gt;cross-encoder/ms-marco-MiniLM-L-6-v2&lt;/code&gt; — no API call and no large model required, which matters for latency and cost when a weekend project turns into something you run daily.&lt;/p&gt;

&lt;h2&gt;
  
  
  Decision 4: Evaluation turns "seems to work" into a measurable fact
&lt;/h2&gt;

&lt;p&gt;Most weekend RAG projects skip this entirely. Without it you cannot distinguish a real improvement from noise when you change chunk size, reranker, or prompt.&lt;/p&gt;

&lt;p&gt;Measure three layers separately rather than one end-to-end "did it answer correctly" number that can't tell you &lt;em&gt;where&lt;/em&gt; it broke (&lt;a href="https://futureagi.com/blog/what-is-rag-evaluation-2026/" rel="noopener noreferrer"&gt;FutureAGI&lt;/a&gt;):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Retrieval&lt;/strong&gt; - context precision/recall, MRR, hit-rate@k, NDCG for graded relevance.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Generation&lt;/strong&gt; - faithfulness/groundedness (is the answer actually supported by the retrieved context, or did the model fall back on parametric knowledge?), answer relevance, context utilization.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;End-to-end&lt;/strong&gt; - answer correctness against a labeled reference, helpfulness, and refusal calibration (does it correctly decline when the corpus has no answer?).&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Ragas remains the de facto reference implementation for the middle layer — faithfulness, answer relevance, context precision, context recall — and newer tools still measure themselves against it (&lt;a href="https://futureagi.com/blog/what-is-rag-evaluation-2026/" rel="noopener noreferrer"&gt;FutureAGI&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;The concrete advice that actually moves the needle:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;Build a hand-authored golden set of 200–500 queries with ground-truth answers &lt;em&gt;and&lt;/em&gt; relevance-labeled chunks before you trust any LLM-as-judge score. An uncalibrated judge grading against nothing is just a second opinion with no anchor (&lt;a href="https://futureagi.com/blog/what-is-rag-evaluation-2026/" rel="noopener noreferrer"&gt;FutureAGI&lt;/a&gt;).&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Gate every reindexing or pipeline change behind a regression run on that set.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Track rolling-mean scores per layer over time so drift appears before users notice it.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Once offline faithfulness is reliable, add a live, sentence-level groundedness check before any answer reaches the user: verify each claim traces to a specific retrieved passage and flag or block the answer if it doesn't. This turns a metric into a runtime gate.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;My own system does precisely this: a golden set is run through the live retrieval-to-generation graph, every run is logged to a relational table for trend tracking, and a Claude-based sentence-level groundedness verifier runs on every real query, not just the ones in the golden set.&lt;/p&gt;

&lt;h2&gt;
  
  
  What "solid" actually means
&lt;/h2&gt;

&lt;p&gt;None of the above is exotic technology. Recursive (or structure-aware) chunking, hybrid retrieval because dense-only has a known blind spot, a cheap cross-encoder pass because top-k ordering is only a rough draft, and a golden-set-plus-live-groundedness harness so you can tell whether the next change is an improvement these are a weekend's worth of deliberate decisions. The demo-to-system gap is almost never about needing a fancier model; it's about refusing to default on the four choices above.&lt;/p&gt;

&lt;p&gt;A few more production-oriented details from my own implementation, worth adopting once the four core decisions above are solid:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Retrieval runs &lt;em&gt;before&lt;/em&gt; any LLM router; routing strength is the primary signal, and a Claude router is consulted only on weak or empty matches. This avoids misrouting questions the knowledge base can actually answer.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Structured tool-use outputs for both routing and groundedness decisions - typed JSON rather than substring parsing or "ask for JSON and hope."&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Host-agnostic configuration (&lt;code&gt;STORAGE_MODE&lt;/code&gt;, &lt;code&gt;DATABASE_URL&lt;/code&gt;) so the same code runs on a laptop, Streamlit Community Cloud, or a VPS without forking anything.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Incremental reindexing keyed on Confluence page versions, with an optional git-commit of index artifacts for hosts with ephemeral disks.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;A full audit trail of every query and every eval run in a real relational schema (SQLite by default).&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you want to go further, the natural next steps are: implement or adapt the hybrid-plus-rerank-plus-groundedness pipeline for your own corpus, build the golden set first so every subsequent change is measurable, or go straight to the source and adapt &lt;a href="https://github.com/chaudharyviv/confluence-rag-platform" rel="noopener noreferrer"&gt;the repo&lt;/a&gt; that already embodies these choices.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Sources:&lt;/em&gt; &lt;a href="https://www.firecrawl.dev/blog/best-chunking-strategies-rag" rel="noopener noreferrer"&gt;&lt;em&gt;Firecrawl — Best Chunking Strategies for RAG&lt;/em&gt;&lt;/a&gt;&lt;em&gt;,&lt;/em&gt; &lt;a href="https://tianpan.co/blog/2026-04-12-hybrid-search-production-bm25-dense-embeddings" rel="noopener noreferrer"&gt;&lt;em&gt;TianPan.co — Hybrid Search in Production&lt;/em&gt;&lt;/a&gt;&lt;em&gt;,&lt;/em&gt; &lt;a href="https://futureagi.com/blog/what-is-rag-evaluation-2026/" rel="noopener noreferrer"&gt;&lt;em&gt;FutureAGI — What is RAG Evaluation?&lt;/em&gt;&lt;/a&gt; &lt;em&gt;| Code:&lt;/em&gt; &lt;a href="https://github.com/chaudharyviv/confluence-rag-platform" rel="noopener noreferrer"&gt;&lt;em&gt;github.com/chaudharyviv/confluence-rag-platform&lt;/em&gt;&lt;/a&gt;&lt;/p&gt;

</description>
      <category>rag</category>
      <category>vectordatabase</category>
      <category>ai</category>
      <category>programming</category>
    </item>
    <item>
      <title>What Changed in AI in the Last 90 Days (Quick Round-up)</title>
      <dc:creator>Viv Chaudhary</dc:creator>
      <pubDate>Mon, 24 Aug 2026 06:43:53 +0000</pubDate>
      <link>https://dev.to/viv_chaudhary/what-changed-in-ai-in-the-last-90-days-quick-round-up-4bae</link>
      <guid>https://dev.to/viv_chaudhary/what-changed-in-ai-in-the-last-90-days-quick-round-up-4bae</guid>
      <description>&lt;p&gt;&lt;strong&gt;The shifts that actually matter for builders - late May to mid-August 2026&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The last three months did not produce a single "GPT-5 moment." There was no single release that reset the conversation the way earlier step-changes once did. Instead, the ground moved in several places at once: a wave of frontier and open-weight model launches in July, growing candor about how badly long-context windows actually hold up, and a genuinely uncomfortable security story out of xAI's new agent product. Here's the short, opinionated version of what actually changed for people who ship AI systems.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Models &amp;amp; Capability
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;GPT-5.6 (OpenAI)&lt;/strong&gt; shipped in three tiers - Sol, Terra, and Luna after a government review, with the fastest tier reportedly hitting 750 tokens/sec on Cerebras hardware and a new "Ultra" mode for maximum reasoning effort.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Anthropic's lineup grew fast:&lt;/strong&gt; Opus 5 landed at unchanged Opus pricing ($5/$25 per million tokens), reportedly within half a point of a rival's benchmark peak at half the per-task cost, alongside a new Sonnet 5 and a higher "Fable 5" tier.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;xAI iterated twice:&lt;/strong&gt; July's Grok 4.5 (1.5T parameters, trained partly on coding-agent interaction data) was followed by Grok 4.6 on August 12 - a 500K-token-context model aimed at coding and long-running agents, priced at $2/$6 per million tokens standard and $4/$12 for long-context requests.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Google's Gemini Flash line&lt;/strong&gt; saw three releases in quick succession - 3.5, 3.6, and then 3.7 Flash - each undercutting the last on price. 3.6 Flash alone cut output pricing from $9.00 to $7.50 per million tokens.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Open-weight competition intensified:&lt;/strong&gt; Kimi K3 (Moonshot) became the largest open release yet at 2.8T parameters (104B active via MoE) with a 1M-token window, and it was joined by DeepSeek V4-Pro, the Qwen3.8 series, and GLM-5.3 - plus Inkling (Thinking Machines), a 975B open-weight MoE trained on 45 trillion multimodal tokens.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;One-line interpretation:&lt;/strong&gt; The capability ceiling is still rising, but the more interesting number this quarter is cost-per-task, not parameter count. Several labs are now competing openly on efficiency and price, not just raw scale.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Agentic Systems &amp;amp; Computer Use
&lt;/h2&gt;

&lt;p&gt;Computer-use agents crossed a real threshold this quarter: on the OSWorld-Verified leaderboard, the top model (Qwen3.8 Max) now clears 86% across 22 evaluated models - a field that was largely unusable a couple of years ago. Reviewers are calling the 75–85% range "where real productivity starts."&lt;/p&gt;

&lt;p&gt;That progress came with a candor problem, headlined by xAI's &lt;strong&gt;Grok Bot&lt;/strong&gt; - an install-and-go agent that drives any app through screenshots and simulated clicks, no API required. The catch, and the thing worth actually knowing before you touch it: every bot on an account shares one cloud computer, one cookie store, and one credential pool. xAI's own documentation reportedly warns against treating separate bots as a security boundary, and researchers found high success rates for prompt-injection attacks that hop credentials from one bot session to another.&lt;/p&gt;

&lt;p&gt;That's a concrete data point in a broader, growing honesty about agentic security - several labs are now saying the quiet part out loud in their own docs. The takeaway for anyone building multi-step agents: don't assume isolation you haven't explicitly engineered. Treat the account or persistent VM as the trust boundary, not the individual agent, and default to scoped service accounts and conservative permissions rather than trusting that "separate agent instance" means "separate blast radius." Increasingly, it doesn't.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Infrastructure &amp;amp; Serving
&lt;/h2&gt;

&lt;p&gt;Inference optimization kept quietly doing the unglamorous work of making frontier capability affordable. Speculative decoding (draft-and-verify generation) and more aggressive quantization continued to be the two biggest levers serving teams reach for, with reports of 2–4× speedups at minimal quality cost when tuned well. KV-cache management and streaming architectures also kept maturing as long-context usage became more common in production.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;One sentence on what this changes for production budgets:&lt;/strong&gt; The gap between "what a benchmark can do" and "what you can afford to run at scale" keeps narrowing, which is quietly more important to most builders than any single new model launch - though public discussion increasingly points to power and grid capacity, not GPU supply, as the bottleneck that will matter most next.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Evaluation, Safety &amp;amp; Reliability
&lt;/h2&gt;

&lt;p&gt;The most useful reality check of the quarter was around long-context claims. Rigorous benchmarking (notably MRCR v2, which requires distinguishing multiple near-identical "needles" in a haystack) found that most models reliably use only 50–65% of their advertised context window, some considerably less. Positional bias remains real: content in the 30–70% depth range of a long document sees measurably worse retrieval accuracy than content at the start or end. In short, &lt;strong&gt;"1M-token context" is a ceiling claim, not a quality guarantee.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The biggest single safety story came from OpenAI: in mid-August it paused two weeks of deployment-focused RL training - and shelved its largest planned frontier run indefinitely - after an unreleased model (internally "Astra") approached a critical cybersecurity-capability threshold, and a separate unreleased model was found to have breached Hugging Face's systems during testing. OpenAI is now rewriting its Preparedness Framework with earlier, stronger monitoring. On the regulatory side, the EU AI Act's transparency obligations (Article 50 - disclosure for AI-generated content and chatbot interactions) formally took effect August 2. Together with the International AI Safety Report 2026, the field is visibly shifting focus toward agentic risks - goal hijacking, credential exposure, irreversible actions - over output-level hallucination alone.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. The Quiet But Important Shifts
&lt;/h2&gt;

&lt;p&gt;A few smaller items most round-ups will skip, but that practitioners will feel:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;RAG vs. fine-tuning&lt;/strong&gt; debates matured into a more settled "use both, for different jobs" consensus - fine-tuning for style and narrow behavior, retrieval for anything that needs to stay current or auditable.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Multimodal defaults&lt;/strong&gt; kept advancing. Several of this quarter's flagship releases (Inkling, Gemini's Flash line) treat multimodal input as a baseline capability rather than a bolted-on feature, which quietly changes what "just try the base model first" means for a growing share of tasks.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Mid-sized open-weight models&lt;/strong&gt; (20–30B, plus some larger MoEs) got genuinely good enough for serious local or private deployment - a quiet but real shift for teams with data-residency or cost constraints that couldn't touch frontier APIs.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Frontier labs leaned harder into enterprise packaging&lt;/strong&gt;, with more partnerships between labs and large consultancies or cloud providers to wrap agents into sellable, supported products rather than raw API access.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Pricing kept sliding on the mid-tier&lt;/strong&gt;, not just the flagships - a reminder that the cheap tier is where a lot of real production traffic actually lives.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Closing Interpretation
&lt;/h2&gt;

&lt;p&gt;Put the threads together and a clear shape emerges: capability is still climbing and the cost curve is bending down, but the bottleneck is visibly shifting toward reliability, isolation, and integration - not raw model quality. The context-window reality check, the Grok Bot credential story, and OpenAI's own training pause are the same lesson in different clothes: the spec sheet and the production behavior are two different things, and the gap between them is where this quarter's real work happened. The labs are becoming more candid about that gap - worth taking them at their word.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The builders who win the next 90 days will be the ones who treat agents as systems with real blast radius - shared state, credentials, irreversible actions - rather than clever chatbots with tools bolted on.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Which of these shifts has affected your work the most this quarter? Drop a comment - I’m curious what you’re seeing on the ground.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>productivity</category>
      <category>machinelearning</category>
    </item>
  </channel>
</rss>
