<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Yashwanth Gadagani</title>
    <description>The latest articles on DEV Community by Yashwanth Gadagani (@yashwanthg).</description>
    <link>https://dev.to/yashwanthg</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4155067%2F86949a6b-3fbd-4705-a811-dbee926af310.png</url>
      <title>DEV Community: Yashwanth Gadagani</title>
      <link>https://dev.to/yashwanthg</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/yashwanthg"/>
    <language>en</language>
    <item>
      <title>Stop Trusting a Single Model's Answer</title>
      <dc:creator>Yashwanth Gadagani</dc:creator>
      <pubDate>Thu, 01 Oct 2026 14:41:20 +0000</pubDate>
      <link>https://dev.to/yashwanthg/stop-trusting-a-single-models-answer-4knn</link>
      <guid>https://dev.to/yashwanthg/stop-trusting-a-single-models-answer-4knn</guid>
      <description>&lt;p&gt;Here's an uncomfortable truth about every AI app shipping today: when your app calls one model and prints what comes back, &lt;strong&gt;you've built an opinion dispenser, not an answer engine.&lt;/strong&gt; The model is confident. The UI is clean. The answer might be wrong. And nothing in the pipeline would ever tell you.&lt;/p&gt;

&lt;p&gt;I learned this building &lt;a href="https://gpai-jade.vercel.app" rel="noopener noreferrer"&gt;Forge&lt;/a&gt;, an AI STEM solver. In education, a confident wrong answer isn't a minor glitch — it's a student copying a false derivation into their exam prep. So I stopped treating the model's output as the answer and started treating it as a &lt;em&gt;claim to be checked&lt;/em&gt;. That one reframing changed the architecture more than any prompt tweak ever did.&lt;/p&gt;

&lt;h2&gt;
  
  
  The core argument: verification is a separate job from generation
&lt;/h2&gt;

&lt;p&gt;We obsess over which model generates the best response — Nemotron vs. DeepSeek vs. Llama, temperature settings, system prompts. That's the generation side, and the industry has gotten very good at it. But generation quality has a ceiling that prompting can't break: &lt;strong&gt;a model cannot reliably grade its own homework.&lt;/strong&gt; Ask it "are you sure?" and it will confidently say yes, because confidence is what it was trained to perform.&lt;/p&gt;

&lt;p&gt;The fix isn't a better prompt. It's a second, independent opinion. In Forge, every solver result goes through a &lt;strong&gt;cross-check&lt;/strong&gt;: a separate model independently solves the same problem, and the UI reports agree, minor difference, or disagree. Two models converging independently is evidence. One model's confidence is theater.&lt;/p&gt;

&lt;p&gt;And for genuinely hard questions, Forge has &lt;strong&gt;debate mode&lt;/strong&gt;: three models answer the same prompt simultaneously, side by side, and then a judge model picks the winner and explains why. Watching three frontier models disagree with each other is the fastest cure for AI sycophancy — you see exactly where the confident consensus breaks down.&lt;/p&gt;

&lt;h2&gt;
  
  
  "Just use a bigger model"
&lt;/h2&gt;

&lt;p&gt;The most common objection: a bigger, smarter model makes verification unnecessary. It doesn't, for three reasons.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;First, errors don't scale away linearly.&lt;/strong&gt; A stronger model makes &lt;em&gt;fewer&lt;/em&gt; mistakes, but its mistakes get &lt;em&gt;harder to spot&lt;/em&gt; — they're fluent, well-structured, and wrong in the middle step of a derivation. Verification catches what capability can't prevent.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Second, the failure mode is silent.&lt;/strong&gt; When a single model is wrong, there's no signal. When two models disagree, the disagreement &lt;em&gt;is&lt;/em&gt; the signal. That signal is worth more than a marginal capability bump, because it tells the user when to be careful — which is precisely when it matters.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Third, diversity beats size.&lt;/strong&gt; A 70B Llama and a DeepSeek model failing &lt;em&gt;differently&lt;/em&gt; gives you more information than a single larger model failing &lt;em&gt;confidently&lt;/em&gt;. Independent errors cancel; correlated confidence compounds.&lt;/p&gt;

&lt;h2&gt;
  
  
  The honest tradeoffs
&lt;/h2&gt;

&lt;p&gt;I'll steelman the other side, because verification isn't free:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Latency.&lt;/strong&gt; Two model calls take longer than one. Forge mitigates this by streaming the primary solution first — the user reads steps while the verifier works in the background. Perceived latency matters more than actual latency.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cost.&lt;/strong&gt; You're paying for 2–3x the tokens. On NVIDIA NIM's pricing this is real money at scale. My answer: verify where errors are expensive (education, code, medical-adjacent) and skip it where they're cheap (brainstorming, drafts).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The judge can be wrong too.&lt;/strong&gt; In debate mode, the judge model is itself a single model with opinions. It's judges all the way down — but each layer converts silent failure into &lt;em&gt;visible disagreement&lt;/em&gt;, and visible disagreement is something a human can act on.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What it looks like in practice
&lt;/h2&gt;

&lt;p&gt;A concrete example of what debate mode is for: take a tricky physics question — say, a rotational dynamics problem with a subtle sign convention. Run three models on it and you might see two agreeing on the setup while one confidently flips a sign in step three. A good judge verdict doesn't just say "Model B wins" — it points at &lt;em&gt;where&lt;/em&gt; the derivations diverge, which is the exact step a student needs to scrutinize.&lt;/p&gt;

&lt;p&gt;That divergence map is the real product. Without debate mode, the student gets one derivation and no reason to doubt step three. With it, they get a highlighted disagreement and an explanation of the contention. The feature isn't "three answers" — it's &lt;em&gt;one answer with its weak points labeled&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;The same principle works outside education. Code review agents that run the linter &lt;em&gt;and&lt;/em&gt; a second model produce fewer escaped bugs than either alone. Support bots that check a generated reply against the knowledge base before sending hallucinate less. The pattern is always the same: generate, then verify with something that fails differently.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this means for your AI app
&lt;/h2&gt;

&lt;p&gt;You don't need debate mode. Start smaller:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Never show a high-stakes answer without a confidence signal.&lt;/strong&gt; A cross-check verdict, a citation, &lt;em&gt;something&lt;/em&gt; beyond the model's own assurance.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Make disagreement visible, not hidden.&lt;/strong&gt; Retry loops that silently re-roll until something looks plausible are worse than showing the conflict — they manufacture false confidence.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Verify with diversity, not repetition.&lt;/strong&gt; The same model twice is a ritual. A different model — or better, a different &lt;em&gt;method&lt;/em&gt; (a calculator for arithmetic, a linter for code) — is a check.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The broader point: we've spent two years making models sound more certain. The next leap in trustworthy AI apps won't come from more capable generators — it'll come from architectures that assume the generator is wrong and check anyway. Build the courtroom, not just the witness.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Forge is open source at &lt;a href="https://github.com/Asdfyash1/GPAI" rel="noopener noreferrer"&gt;github.com/Asdfyash1/GPAI&lt;/a&gt; — the cross-check pipeline, debate mode, and the whole verification architecture are in there.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>discuss</category>
      <category>llm</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>Automating Marketing Reports with n8n and Looker Studio</title>
      <dc:creator>Yashwanth Gadagani</dc:creator>
      <pubDate>Thu, 01 Oct 2026 14:34:32 +0000</pubDate>
      <link>https://dev.to/yashwanthg/automating-marketing-reports-with-n8n-and-looker-studio-d</link>
      <guid>https://dev.to/yashwanthg/automating-marketing-reports-with-n8n-and-looker-studio-d</guid>
      <description>&lt;p&gt;Here's a ritual every marketing team knows: Monday morning, someone opens Meta Ads Manager, Google Ads, and GA4, exports three CSVs, pastes them into a spreadsheet, fixes the date formats, and rebuilds the same charts they rebuilt last week. It takes an hour, it's error-prone, and nobody enjoys it.&lt;/p&gt;

&lt;p&gt;This guide walks through automating that entire pipeline with &lt;strong&gt;n8n&lt;/strong&gt; (self-hosted workflow automation) and &lt;strong&gt;Looker Studio&lt;/strong&gt; (free dashboards). The pattern: n8n pulls spend and performance data from each ad platform on a schedule, normalizes it into one Google Sheet, and Looker Studio visualizes it — refreshed automatically, forever.&lt;/p&gt;

&lt;p&gt;A note on honesty up front: this is a build walkthrough based on my real n8n + Looker Studio automation work — I've wired n8n workflows into Google Sheets and visualized the results in Looker Studio. The Meta Ads and GA4 legs below are walked through from the API docs and working n8n patterns, but I haven't run those two legs end-to-end against live ad accounts myself. The Google Ads API leg I have &lt;em&gt;not&lt;/em&gt; run all the way through either — Google's developer-token approval is the known hard part there, and I'll flag exactly where that bites.&lt;/p&gt;

&lt;h2&gt;
  
  
  The architecture
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Meta Marketing API ─┐
                    ├─▶ n8n (scheduled) ─▶ normalize ─▶ Google Sheets ─▶ Looker Studio
GA4 Data API ───────┘         ▲
                              │  (Google Ads API — see caveat below)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Why Google Sheets in the middle instead of writing straight to Looker Studio? Because Looker Studio's native ad-platform connectors refresh on &lt;em&gt;their&lt;/em&gt; schedule and give you limited control over blending. A Sheet you own means one schema, one timezone, and blended cross-platform metrics (total spend, blended ROAS) computed before visualization. For a weekly marketing report, Sheets is plenty — upgrade to BigQuery when the row counts get silly.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1: Pull Meta Ads data with the HTTP Request node
&lt;/h2&gt;

&lt;p&gt;Meta's Marketing API has an &lt;code&gt;insights&lt;/code&gt; edge that returns exactly what a report needs. In n8n, use the &lt;strong&gt;HTTP Request&lt;/strong&gt; node (not the Meta Ads node — the generic HTTP node gives you full control over fields and breakdowns):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Method:&lt;/strong&gt; GET&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;URL:&lt;/strong&gt; &lt;code&gt;https://graph.facebook.com/v21.0/act_&amp;lt;AD_ACCOUNT_ID&amp;gt;/insights&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Query parameters:&lt;/strong&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;fields&lt;/code&gt;: &lt;code&gt;campaign_name,spend,impressions,clicks,actions&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;date_preset&lt;/code&gt;: &lt;code&gt;last_7d&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;level&lt;/code&gt;: &lt;code&gt;campaign&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;access_token&lt;/code&gt;: your token (store it in n8n &lt;strong&gt;Credentials&lt;/strong&gt;, never hardcoded)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Request &lt;code&gt;actions&lt;/code&gt; and you'll get conversions broken down by type — filter for &lt;code&gt;purchase&lt;/code&gt; or &lt;code&gt;lead&lt;/code&gt; in the next step. One gotcha: Meta returns &lt;code&gt;actions&lt;/code&gt; as an array of &lt;code&gt;{action_type, value}&lt;/code&gt; objects, not flat columns. You'll flatten that in the normalize step.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Token setup:&lt;/strong&gt; create a Meta app, add the Marketing API product, generate a token with &lt;code&gt;ads_read&lt;/code&gt; scope. For a scheduled job, you want a &lt;strong&gt;system user token&lt;/strong&gt; (doesn't expire when someone changes their Facebook password) — this is the detail that breaks most first attempts when the report silently stops updating two months in.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2: Pull GA4 data with the Data API
&lt;/h2&gt;

&lt;p&gt;GA4's Data API uses OAuth2 service accounts, which n8n supports natively:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;In Google Cloud Console, enable the &lt;strong&gt;Google Analytics Data API&lt;/strong&gt;, create a service account, download the JSON key.&lt;/li&gt;
&lt;li&gt;In GA4 Admin → Property Access Management, add the service account email as a &lt;strong&gt;Viewer&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;In n8n, create a &lt;strong&gt;Google Analytics&lt;/strong&gt; credential (OAuth2 / service account) and use the &lt;strong&gt;Google Analytics&lt;/strong&gt; node with the &lt;code&gt;Run Report&lt;/code&gt; operation:

&lt;ul&gt;
&lt;li&gt;Dimensions: &lt;code&gt;date&lt;/code&gt;, &lt;code&gt;sessionDefaultChannelGroup&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Metrics: &lt;code&gt;sessions&lt;/code&gt;, &lt;code&gt;conversions&lt;/code&gt;, &lt;code&gt;totalRevenue&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Date range: last 7 days&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The failure mode to expect here is the property-access step. If n8n returns a 403, it's almost always the service account missing the Viewer role in GA4, not your n8n config.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3: Google Ads — the hard one (partially tested)
&lt;/h2&gt;

&lt;p&gt;Google Ads data comes from the &lt;strong&gt;Google Ads API&lt;/strong&gt;, and this is where you should budget real time. The API requires a &lt;strong&gt;developer token&lt;/strong&gt;, and Google approves those only for accounts in good standing — the application asks about your use case and intended API usage. Test accounts get an unapproved token with heavy limits.&lt;/p&gt;

&lt;p&gt;The honest status: the n8n side is straightforward (HTTP Request node against &lt;code&gt;googleads.googleapis.com&lt;/code&gt;, GAQL query like &lt;code&gt;SELECT campaign.name, metrics.cost_micros, metrics.impressions FROM campaign WHERE segments.date DURING LAST_7_DAYS&lt;/code&gt;), but I have &lt;strong&gt;not&lt;/strong&gt; completed a live pull against a production account because of the developer-token gate. If you plan to run this in production, start the developer-token application on day one — it's the longest pole in the tent.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Pragmatic fallback:&lt;/strong&gt; until the API is approved, the Google Ads &lt;em&gt;scheduled email reports&lt;/em&gt; (CSV to inbox) plus n8n's &lt;strong&gt;IMAP Email&lt;/strong&gt; trigger can land the same data in your Sheet. Less elegant, works Monday morning.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 4: Normalize everything in a Code node
&lt;/h2&gt;

&lt;p&gt;The three APIs return three different shapes, three different date formats, and two different currencies if you're not careful. One n8n &lt;strong&gt;Code&lt;/strong&gt; node (JavaScript) standardizes them into a single schema:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Normalize one platform's rows into the canonical schema&lt;/span&gt;
&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;normalize&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;rows&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;platform&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;rows&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;r&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;({&lt;/span&gt;
    &lt;span class="na"&gt;date&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;date&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;                       &lt;span class="c1"&gt;// YYYY-MM-DD, coerced per platform&lt;/span&gt;
    &lt;span class="nx"&gt;platform&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;                            &lt;span class="c1"&gt;// 'meta' | 'google_ads' | 'ga4'&lt;/span&gt;
    &lt;span class="na"&gt;campaign&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;campaign_name&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nx"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;campaign&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;(organic)&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;spend&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nc"&gt;Number&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;spend&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;         &lt;span class="c1"&gt;// micros → units for Google Ads!&lt;/span&gt;
    &lt;span class="na"&gt;impressions&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nc"&gt;Number&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;impressions&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="na"&gt;clicks&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nc"&gt;Number&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;clicks&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="na"&gt;conversions&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nc"&gt;Number&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;conversions&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="p"&gt;}));&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;normalize&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;$input&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;all&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;json&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;meta&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The Google Ads &lt;code&gt;cost_micros&lt;/code&gt; trap deserves its own warning: costs come back in &lt;strong&gt;micro-units&lt;/strong&gt;, so divide by 1,000,000 or your spend column will claim you spent $4 billion.&lt;/p&gt;

&lt;p&gt;Also normalize timezones here — pull everything in the ad account's timezone and stamp the Sheet with it, or your Monday report will attribute Sunday night spend to the wrong week.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 5: Write to Google Sheets
&lt;/h2&gt;

&lt;p&gt;n8n's &lt;strong&gt;Google Sheets&lt;/strong&gt; node with the &lt;em&gt;Append&lt;/em&gt; operation (service-account credential again) writes the normalized rows to a tab like &lt;code&gt;raw_spend&lt;/code&gt;. Keep one tab per platform plus one blended tab, or just one tab with the &lt;code&gt;platform&lt;/code&gt; column — the latter blends more easily in Looker Studio.&lt;/p&gt;

&lt;p&gt;Idempotency matters for scheduled runs: before appending, use &lt;strong&gt;IF&lt;/strong&gt; + &lt;strong&gt;Google Sheets: Read&lt;/strong&gt; to check whether this week's rows already exist (match on date + platform + campaign). Otherwise a retried run double-counts spend, and nobody notices until the invoice meeting.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 6: Build the Looker Studio dashboard
&lt;/h2&gt;

&lt;p&gt;Connect Looker Studio to the Sheet (native connector, 15-minute minimum refresh — fine for weekly reporting). The charts that actually get used:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Scorecards:&lt;/strong&gt; total spend, total conversions, blended ROAS, week-over-week deltas&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Time series:&lt;/strong&gt; spend vs. conversions by day, with platform breakdown&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Table:&lt;/strong&gt; campaign-level spend / ROAS, sortable, with conditional formatting on ROAS &amp;lt; target&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Blend the platforms with a calculated field for blended ROAS: &lt;code&gt;SUM(conversions * value) / SUM(spend)&lt;/code&gt; — or simpler, a blended &lt;em&gt;cost per conversion&lt;/em&gt; if conversion values aren't tracked consistently across platforms (they usually aren't at first).&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 7: Schedule it and handle failures
&lt;/h2&gt;

&lt;p&gt;n8n's &lt;strong&gt;Schedule Trigger&lt;/strong&gt; (cron: Mondays 6 AM in the account timezone) kicks off the workflow. Then add the unglamorous parts:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Error branch:&lt;/strong&gt; n8n's Error Trigger workflow → send yourself a Telegram/Slack message with the failed node name. A report pipeline that fails silently is worse than no pipeline.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stale-data check:&lt;/strong&gt; a final node that reads the Sheet and alerts if the newest row is older than 8 days — catches an expired Meta token before anyone relying on the report notices.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Credential rotation reminder:&lt;/strong&gt; Meta system-user tokens and Google service-account keys should be rotated; put a quarterly reminder in the workflow's notes.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What this actually buys you
&lt;/h2&gt;

&lt;p&gt;About an hour a week back, per report — but the real win is trust. When the numbers are pulled by a machine on a schedule instead of copy-pasted by a human on a Monday, the "are these numbers right?" conversation disappears. The dashboard becomes the source of truth instead of the spreadsheet someone edited.&lt;/p&gt;

&lt;p&gt;Start with Meta + GA4 (both testable today with free-tier accounts), add Google Ads once the developer token lands, and keep the Sheet schema stable from day one — every downstream chart depends on it. The n8n workflow itself is maybe 10 nodes. The data hygiene around it is the actual project.&lt;/p&gt;

</description>
      <category>n8n</category>
      <category>automation</category>
      <category>tutorial</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Building an AI STEM Solver That Shows Its Work — and Double-Checks It</title>
      <dc:creator>Yashwanth Gadagani</dc:creator>
      <pubDate>Thu, 01 Oct 2026 14:34:04 +0000</pubDate>
      <link>https://dev.to/yashwanthg/building-an-ai-stem-solver-that-shows-its-work-and-double-checks-it-5dj5</link>
      <guid>https://dev.to/yashwanthg/building-an-ai-stem-solver-that-shows-its-work-and-double-checks-it-5dj5</guid>
      <description>&lt;p&gt;Ask any AI chatbot a calculus question and you'll get an answer. Maybe even the right one. But a student doesn't need &lt;em&gt;an&lt;/em&gt; answer — they need to understand &lt;em&gt;how&lt;/em&gt; to get there, and they need to know the answer is actually correct. Those are two different engineering problems, and I built &lt;a href="https://gpai-jade.vercel.app" rel="noopener noreferrer"&gt;Forge&lt;/a&gt; (repo: &lt;a href="https://github.com/Asdfyash1/GPAI" rel="noopener noreferrer"&gt;Asdfyash1/GPAI&lt;/a&gt;) to solve both: a full-stack STEM copilot that turns any problem — typed, photographed, or pasted from a URL — into a step-by-step derivation, then has a &lt;em&gt;second&lt;/em&gt; model independently verify the result.&lt;/p&gt;

&lt;p&gt;The stack: &lt;strong&gt;Next.js 16&lt;/strong&gt; (App Router, Turbopack), &lt;strong&gt;TypeScript&lt;/strong&gt; (strict), &lt;strong&gt;NVIDIA NIM&lt;/strong&gt; models (Nemotron, DeepSeek Flash, Llama 3.3) via the Vercel AI SDK, deployed on &lt;strong&gt;Vercel's Hobby tier&lt;/strong&gt;. Here's how the solver works, step by step, with the real architecture.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1: Stream the solution, don't dump it
&lt;/h2&gt;

&lt;p&gt;Nobody reads a wall of math. Forge's solver streams the derivation step-by-step over &lt;strong&gt;Server-Sent Events&lt;/strong&gt;, so steps appear as they're generated and the UI can reveal them one at a time (collapsible step reveal — "show me the next step" instead of the whole answer).&lt;/p&gt;

&lt;p&gt;The streaming endpoint lives at &lt;code&gt;/api/educate/stream&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// src/app/api/educate/stream/route.ts (simplified)&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;requireAuth&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;@/lib/api-guard&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;streamSolverResponse&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;@/lib/orchestrator&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;POST&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;request&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Request&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;guard&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;requireAuth&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;request&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;guard&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;ok&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;guard&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;body&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;handle&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;streamSolverResponse&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;problem&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;problem&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Response&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;handle&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;textStream&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Content-Type&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;text/plain; charset=utf-8&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On the client, a small &lt;code&gt;useStream&lt;/code&gt; hook consumes the SSE stream and appends steps to state as they arrive. The response parser (&lt;code&gt;src/lib/response-parser.ts&lt;/code&gt;) decomposes the LLM's markdown into structured steps using regex — each step becomes a collapsible card in &lt;code&gt;SolverView.tsx&lt;/code&gt;. That decomposition matters: if the model returns one blob, you can't do progressive reveal, follow-up chips ("Why does this step work?"), or per-step quizzes later.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2: Verify with a second model (cross-check)
&lt;/h2&gt;

&lt;p&gt;This is the part most AI solvers skip. Forge runs a &lt;strong&gt;cross-check&lt;/strong&gt;: after the primary model produces a solution, a second model independently solves the same problem, and the two are compared. The UI shows one of three verdicts: &lt;strong&gt;agree&lt;/strong&gt;, &lt;strong&gt;minor difference&lt;/strong&gt;, or &lt;strong&gt;disagree&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Why it matters: a single model's confident wrong answer is the most dangerous output in education. Two models arriving at the same answer independently is dramatically stronger evidence than one model's confidence score. The cross-check model is configurable via env var:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;NVIDIA_SOLVER_MODEL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;meta/llama-3.3-70b-instruct
&lt;span class="nv"&gt;NVIDIA_VERIFIER_MODEL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;meta/llama-3.3-70b-instruct
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You can point the verifier at a completely different model family (or even a different provider via the &lt;code&gt;ADDITIONAL_OPENAI_COMPATIBLE_*&lt;/code&gt; vars) so the check isn't just the same weights agreeing with themselves. When the verdict is "disagree," the student sees that upfront instead of copying a wrong derivation into their homework.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3: Render math like math
&lt;/h2&gt;

&lt;p&gt;A derivation full of &lt;code&gt;x^2 + 2x + 1 = 0&lt;/code&gt; in monospace is unreadable. Forge renders formulas with &lt;strong&gt;KaTeX&lt;/strong&gt; via &lt;code&gt;rehype-katex&lt;/code&gt; + &lt;code&gt;remark-math&lt;/code&gt;, layered on &lt;code&gt;react-markdown&lt;/code&gt; + &lt;code&gt;remark-gfm&lt;/code&gt;. The &lt;code&gt;MathMarkdown.tsx&lt;/code&gt; component handles the whole pipeline, so model output containing &lt;code&gt;$...$&lt;/code&gt; and &lt;code&gt;$$...$$&lt;/code&gt; becomes properly typeset math inline.&lt;/p&gt;

&lt;p&gt;Small detail, big UX difference: students trust a solution that &lt;em&gt;looks&lt;/em&gt; like their textbook. Plain-text math looks like a chatbot guessing; typeset math looks like a worked example.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 4: Accept photos, not just text
&lt;/h2&gt;

&lt;p&gt;Half of real homework exists as a photo of a notebook or a textbook page. Forge's OCR pipeline:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;User uploads a photo of handwritten work or a printed problem.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Client-side compression&lt;/strong&gt; shrinks it to 1600px at quality 0.85 — this keeps the payload under the &lt;strong&gt;1.7 MB&lt;/strong&gt; Vercel API gateway limit (&lt;code&gt;MAX_INLINE_IMAGE_BYTES&lt;/code&gt;). Do this in the browser, not the server; it saves bandwidth and avoids gateway rejections.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;NVIDIA Nemotron Omni 30B&lt;/strong&gt; reads the image and extracts the text.&lt;/li&gt;
&lt;li&gt;The extracted text feeds straight into the solver pipeline from Step 1.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;PDFs get the same treatment via &lt;code&gt;unpdf&lt;/code&gt; — a pure-JS parser with no native binaries, which matters because Vercel serverless functions don't love native modules.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 5: Make it work with zero API keys (demo mode)
&lt;/h2&gt;

&lt;p&gt;Here's a trick more AI apps should steal: Forge runs &lt;strong&gt;without any API keys at all&lt;/strong&gt;. If &lt;code&gt;NVIDIA_API_KEY&lt;/code&gt; isn't set, the app falls back to &lt;code&gt;demo-solver.ts&lt;/code&gt; — deterministic, hand-written demo output that exercises the entire UI (step reveal, KaTeX, quiz, cross-check display) with no model calls.&lt;/p&gt;

&lt;p&gt;Why bother? Three reasons: (1) anyone can clone and run it in 30 seconds, which is gold for a portfolio project; (2) frontend development doesn't burn API credits; (3) the demo doubles as a UI test fixture. The env table is honest about what's needed for what:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Variable&lt;/th&gt;
&lt;th&gt;Needed for&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;NVIDIA_API_KEY&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;All real AI features&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;JWT_SECRET&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Auth (self-generated)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;RESEND_API_KEY&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;OTP emails (lazy-initialized — build succeeds without it)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;TELEGRAM_BOT_TOKENS&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Cloud storage backend&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Step 6: Secure the endpoints by default
&lt;/h2&gt;

&lt;p&gt;Every AI endpoint sits behind a &lt;code&gt;requireAuth()&lt;/code&gt; guard that runs four checks in order: payload size (→ 413), origin validation for CSRF (→ 403), JWT cookie verification (→ 401), and per-user rate limiting (→ 429). The API key lives &lt;strong&gt;only&lt;/strong&gt; in Vercel env vars — the frontend calls your own &lt;code&gt;/api/*&lt;/code&gt; routes with an HttpOnly &lt;code&gt;SameSite=Lax&lt;/code&gt; cookie and never sees a secret.&lt;/p&gt;

&lt;p&gt;The rate limits are tuned per endpoint: 30 req/min for most AI routes, 10/min for debate mode (which fires 3 model calls per request). In-memory sliding-window counters per Vercel isolate, pruned periodically.&lt;/p&gt;

&lt;h2&gt;
  
  
  Honest gotchas
&lt;/h2&gt;

&lt;p&gt;A few things I'd do differently or that bit me:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The OTP store is in-memory.&lt;/strong&gt; Serverless cold starts wipe pending OTPs. Acceptable for an MVP, not for production — move it to Redis or Postgres before real users.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cloud storage has a ceiling.&lt;/strong&gt; The Telegram-bot storage backend hits a ~4096-char pinned-message cap around ~50 users. Fine for a demo; a document-based registry is needed to scale.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Next.js 16 broke things&lt;/strong&gt; vs. older training data. Read the framework docs in &lt;code&gt;node_modules&lt;/code&gt;, not your memory, when something behaves oddly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One model to avoid:&lt;/strong&gt; &lt;code&gt;mistralai/mistral-large-3-675b-instruct-2512&lt;/code&gt; isn't in the NVIDIA NIM catalog — every call 404s. The README warns about it because I learned the hard way.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The takeaway
&lt;/h2&gt;

&lt;p&gt;The pattern generalizes beyond homework: &lt;strong&gt;stream structured output, verify with an independent model, render it like the domain expects, and degrade gracefully without keys.&lt;/strong&gt; Most AI wrappers stop at "call the API and print the text." The difference between a demo and a tool is everything around the model call — and that's where the interesting engineering lives.&lt;/p&gt;

&lt;p&gt;The repo (&lt;a href="https://github.com/Asdfyash1/GPAI" rel="noopener noreferrer"&gt;Asdfyash1/GPAI&lt;/a&gt;) has the full source, and the live app is at &lt;a href="https://gpai-jade.vercel.app" rel="noopener noreferrer"&gt;gpai-jade.vercel.app&lt;/a&gt; — try photographing a math problem and watch the cross-check verdict.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>nextjs</category>
      <category>tutorial</category>
      <category>typescript</category>
    </item>
  </channel>
</rss>
