<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Rovugui</title>
    <description>The latest articles on DEV Community by Rovugui (@rovugui).</description>
    <link>https://dev.to/rovugui</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3824706%2F1aa25acb-5497-49c8-bf41-63bf25deae08.jpg</url>
      <title>DEV Community: Rovugui</title>
      <link>https://dev.to/rovugui</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/rovugui"/>
    <language>en</language>
    <item>
      <title>Polar rejected my SaaS over two words in my landing page. Here's the appeal process nobody documents.</title>
      <dc:creator>Rovugui</dc:creator>
      <pubDate>Sat, 18 Jul 2026 19:39:21 +0000</pubDate>
      <link>https://dev.to/rovugui/polar-rejected-my-saas-over-two-words-in-my-landing-page-heres-the-appeal-process-nobody-2k0g</link>
      <guid>https://dev.to/rovugui/polar-rejected-my-saas-over-two-words-in-my-landing-page-heres-the-appeal-process-nobody-2k0g</guid>
      <description>&lt;p&gt;Last week my merchant-of-record application got denied.&lt;br&gt;
Automated review. Then my appeal got denied too — same day.&lt;/p&gt;

&lt;p&gt;The reason: my landing page described my product as helping with&lt;br&gt;
"lead generation". Two words. That phrase sits on the prohibited list of basically every merchant of record (Polar, Paddle, Lemon Squeezy), next to gambling and CBD. I didn't know. It's not really in their docs.&lt;/p&gt;

&lt;p&gt;What the product actually does: monitors public conversations and drafts helpful replies a human reviews. Social listening, in industry terms. Same product, different two words — one framing is bannable, the other is fine.&lt;/p&gt;

&lt;p&gt;Here's what got me approved on human review, in case you hit the same wall:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Audited every processor-visible word.&lt;/strong&gt; "Leads" → "mentions and&lt;br&gt;
conversations". "Outreach" → gone. If a word implies contacting people who didn't ask, it's radioactive to underwriters.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Added real legal pages.&lt;/strong&gt; Terms, Privacy, Refunds — linked in the footer. Approval reviewers check for them; most indie landing pages don't have them.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Made the "human in the loop" explicit.&lt;/strong&gt; "No auto-send. You review every reply." Underwriters care about spam potential, not your feature list.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Escalated past the bot.&lt;/strong&gt; Automated denial + denied appeal ≠ dead. I requested human review and got approved the same day.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Since then I've been reading every MoR rejection story I can find&lt;br&gt;
(Paddle's undocumented checks, Lemon Squeezy's Trustpilot meltdown) and the pattern is the same: most rejections are fixable wording/legal-page problems, but founders don't know which words tripped the wire, and the processors won't tell you.&lt;/p&gt;

&lt;p&gt;I ended up packaging what I learned: I put the copy scanner up for free at &lt;a href="https://audeza.com" rel="noopener noreferrer"&gt;https://audeza.com&lt;/a&gt; — paste your landing copy, it flags the trigger vocabulary. The full kit (legal-page templates, per-processor escalation playbook) is waitlist-only for now. Honest framing: it maximizes your odds, nothing can guarantee underwriting.&lt;/p&gt;

&lt;p&gt;If you're facing (or fearing) an MoR rejection: the scanner is free,&lt;br&gt;
no signup. If the full kit would be useful, join the waitlist there —&lt;br&gt;
if enough people want it, I'll finish it; if not, this post is the free version — steal the 4 steps above.&lt;/p&gt;

</description>
      <category>saas</category>
      <category>startup</category>
      <category>payments</category>
      <category>buildinpublic</category>
    </item>
    <item>
      <title>How I Built Audeza — A Real-Time AI Pitch Coach with Gemini Live API and Google Cloud</title>
      <dc:creator>Rovugui</dc:creator>
      <pubDate>Sun, 15 Mar 2026 01:07:45 +0000</pubDate>
      <link>https://dev.to/rovugui/how-i-built-audeza-a-real-time-ai-pitch-coach-with-gemini-live-api-and-google-cloud-27o6</link>
      <guid>https://dev.to/rovugui/how-i-built-audeza-a-real-time-ai-pitch-coach-with-gemini-live-api-and-google-cloud-27o6</guid>
      <description>&lt;p&gt;&lt;em&gt;I created this piece of content for the purposes of entering the &lt;a href="https://googleai.devpost.com/" rel="noopener noreferrer"&gt;Gemini Live Agent Challenge&lt;/a&gt; hackathon.&lt;/em&gt; #GeminiLiveAgentChallenge&lt;/p&gt;




&lt;h2&gt;
  
  
  What is Audeza?
&lt;/h2&gt;

&lt;p&gt;Audeza is the first AI pitch coach that can &lt;strong&gt;see you, hear you, and talk back — all in real time&lt;/strong&gt;. You upload your pitch deck, turn on your camera, and start practicing. The AI watches your body language, listens to your delivery, and coaches you mid-pitch like a mentor sitting across the table.&lt;/p&gt;

&lt;p&gt;It has two modes:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Coaching Mode&lt;/strong&gt; — Real-time feedback on posture, eye contact, filler words, pacing, and content coverage as you present&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Investor Simulator&lt;/strong&gt; — The AI role-plays as a tough VC, listens to your full pitch, then grills you with hard questions based on what you actually said&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;After each session, you get a scorecard with delivery, content, and presence scores, per-slide timing analysis, and actionable improvements.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Core: Gemini Live API
&lt;/h2&gt;

&lt;p&gt;The entire experience is powered by the &lt;strong&gt;Gemini Live API&lt;/strong&gt; through the &lt;code&gt;@google/genai&lt;/code&gt; SDK. Here's why this API was the right choice:&lt;/p&gt;

&lt;h3&gt;
  
  
  Multimodal real-time streaming
&lt;/h3&gt;

&lt;p&gt;Gemini Live accepts &lt;strong&gt;audio and video simultaneously&lt;/strong&gt; over a single WebSocket connection. During a pitch session, Audeza sends:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Microphone audio&lt;/strong&gt; — 16kHz PCM captured via AudioWorklet&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Camera frames&lt;/strong&gt; — 1 FPS JPEG snapshots from the webcam&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The model processes both streams and responds with natural spoken audio. This is what makes Audeza different from tools that only analyze recordings after the fact — the coaching happens &lt;em&gt;while you're presenting&lt;/em&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Connecting to Gemini Live API&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;GoogleGenAI&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;@google/genai&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;ai&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;GoogleGenAI&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;apiKey&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;session&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;ai&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;live&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;connect&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;gemini-2.5-flash-native-audio-preview-12-2025&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;config&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;responseModalities&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;AUDIO&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="na"&gt;systemInstruction&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;parts&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt; &lt;span class="na"&gt;text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;systemPrompt&lt;/span&gt; &lt;span class="p"&gt;}]&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="na"&gt;contextWindowCompression&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;slidingWindow&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{}&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="na"&gt;inputAudioTranscription&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{},&lt;/span&gt;
    &lt;span class="na"&gt;outputAudioTranscription&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{},&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="na"&gt;callbacks&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;onmessage&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;msg&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="c1"&gt;// Handle audio responses, transcriptions, turn completion&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;})&lt;/span&gt;

&lt;span class="c1"&gt;// Send real-time audio&lt;/span&gt;
&lt;span class="nx"&gt;session&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sendRealtimeInput&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;audio&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;data&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;base64PCM&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;mimeType&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;audio/pcm;rate=16000&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;})&lt;/span&gt;

&lt;span class="c1"&gt;// Send camera frames at 1 FPS&lt;/span&gt;
&lt;span class="nx"&gt;session&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sendRealtimeInput&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;video&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;data&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;base64Jpeg&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;mimeType&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;image/jpeg&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;})&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Context window compression
&lt;/h3&gt;

&lt;p&gt;Pitch sessions can run 5-10 minutes. With continuous audio and video, that's a lot of context. I used Gemini's built-in &lt;strong&gt;sliding window compression&lt;/strong&gt; to keep sessions running without hitting token limits:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nx"&gt;contextWindowCompression&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nl"&gt;slidingWindow&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{},&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Input/output transcription
&lt;/h3&gt;

&lt;p&gt;Gemini Live provides real-time transcription of both the user's speech and the model's responses. I capture these to build a full session transcript, which is then used for scorecard generation.&lt;/p&gt;

&lt;h3&gt;
  
  
  Native audio output
&lt;/h3&gt;

&lt;p&gt;Using &lt;code&gt;gemini-2.5-flash-native-audio-preview&lt;/code&gt;, the model speaks with natural vocal variety — it can be warm and encouraging when you're nervous, or push harder when you're confident. The voice style adapts to the coaching context.&lt;/p&gt;

&lt;h2&gt;
  
  
  Scorecard Generation with Gemini 2.5 Flash
&lt;/h2&gt;

&lt;p&gt;After a session ends, I send the full transcript plus slide context to &lt;strong&gt;Gemini 2.5 Flash&lt;/strong&gt; (text model) to generate a structured scorecard:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Overall score (1-100)&lt;/li&gt;
&lt;li&gt;Delivery, Content, and Presence sub-scores&lt;/li&gt;
&lt;li&gt;Per-slide breakdown with timing analysis&lt;/li&gt;
&lt;li&gt;Specific improvements to work on&lt;/li&gt;
&lt;li&gt;Investor verdict (in simulator mode)&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Google Cloud Stack
&lt;/h2&gt;

&lt;p&gt;The full infrastructure runs on Google Cloud:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Service&lt;/th&gt;
&lt;th&gt;Purpose&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Cloud Run&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Hosts the app (containerized with Bun runtime)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Firebase Auth&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Google sign-in for users&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Cloud Firestore&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Session history, scorecards, user preferences, deck metadata&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Artifact Registry&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Docker image storage&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;GitHub Actions&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;CI/CD pipeline → build → push → deploy to Cloud Run&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The app is a &lt;strong&gt;TanStack Start&lt;/strong&gt; (React) application built with Bun, containerized via Docker, and deployed to Cloud Run with automatic deploys on every push to main.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Learned
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Gemini Live API is surprisingly natural&lt;/strong&gt; — The model's ability to process video and audio simultaneously and respond conversationally makes it feel like you're talking to a real person, not an AI.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;AudioWorklet is essential&lt;/strong&gt; — For reliable real-time audio capture, the Web Audio API's AudioWorklet is the way to go. It runs in a separate thread and doesn't drop frames.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;1 FPS is enough for body language&lt;/strong&gt; — You don't need high frame rates for the model to pick up on posture, eye contact, and gestures. One frame per second keeps the context window manageable while still giving useful visual coaching.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;System prompts matter enormously&lt;/strong&gt; — The difference between a generic AI response and a great coaching experience came down to prompt engineering. Defining specific triggers (slide transitions, feedback requests, real-time interjections) made the model proactive instead of passive.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Wanna try it?
&lt;/h2&gt;

&lt;p&gt;Stay tuned for Audeza live soon.&lt;/p&gt;

</description>
      <category>geminiliveagentchallenge</category>
      <category>hackathon</category>
      <category>googlecloud</category>
      <category>ai</category>
    </item>
  </channel>
</rss>
