<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: ADK BBX</title>
    <description>The latest articles on DEV Community by ADK BBX (@adk_bbx).</description>
    <link>https://dev.to/adk_bbx</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4158217%2F579edfae-f460-4c4b-8838-05a5f09ff6bb.png</url>
      <title>DEV Community: ADK BBX</title>
      <link>https://dev.to/adk_bbx</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/adk_bbx"/>
    <language>en</language>
    <item>
      <title>Phone calls are the boss fight of living in Japan. 🗼</title>
      <dc:creator>ADK BBX</dc:creator>
      <pubDate>Sun, 04 Oct 2026 17:42:34 +0000</pubDate>
      <link>https://dev.to/adk_bbx/phone-calls-are-the-boss-fight-of-living-in-japan-1h6h</link>
      <guid>https://dev.to/adk_bbx/phone-calls-are-the-boss-fight-of-living-in-japan-1h6h</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/hacktoberfest-weekend-2026-10-01"&gt;Hacktoberfest Weekend Challenge: Build for a Friend&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjtp63ghvcqumhy7x8rr3.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjtp63ghvcqumhy7x8rr3.gif" alt="A practice call in Before I Call: a Japanese question with furigana, a word's meaning on hover, then a plain-English explanation of the question" width="720" height="405"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I've lived alone in Tokyo for four years now, but when my washing machine started leaking, I stared at the building manager's number for twenty minutes before I pressed call.&lt;/p&gt;

&lt;p&gt;My friend ran into the same thing when she had to cancel her internet contract by phone. The provider's explanation sounded alien to her. The Japanese businesses use on the phone is not the Japanese you use every day, and on a call there's no face to read and no time to look anything up.&lt;/p&gt;

&lt;p&gt;So I built Before I Call, an app where you practice the call with an AI before you make the real one. I made it for myself and for friends in the same situation, and sent it to one of them to see if it actually helps. Their reply is further down.&lt;/p&gt;

&lt;p&gt;Here it is in action:&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/Sw112GoVYqo" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;You describe the call you need to make, or pick one of the examples (home repairs, the clinic, a missed delivery, the city office, dietary requests, lost property, bills). Then you talk to an AI that plays the person on the other end, like a receptionist or your building manager. It speaks polite Japanese, asks one question at a time and waits for your answer. There's no score.&lt;/p&gt;

&lt;p&gt;If you get stuck, you don't have to hang up. &lt;strong&gt;Explain question&lt;/strong&gt; opens a helper next to the call with what the question means and a reply you could give. The AI on the call doesn't see it, so the role-play carries on. &lt;strong&gt;Slow replay&lt;/strong&gt; plays the question again at 0.7x speed, and &lt;strong&gt;Repeat question&lt;/strong&gt; asks the AI to say it again. Kanji have furigana, every line has romaji, and known words show their meaning when you hover or tap.&lt;/p&gt;

&lt;p&gt;When you say goodbye, the AI hangs up and you get a call card: the whole conversation with readings and meanings, and a list of useful words from your call. You can download it as a PDF and keep it next to you when you make the real call.&lt;/p&gt;

&lt;p&gt;It won't make up dates, prices or availability, and it can't book anything for you. It's only for practice.&lt;/p&gt;

&lt;p&gt;Setting up a call. You can type or dictate it in English or Japanese, or start from an example. &lt;strong&gt;Enhance prompt&lt;/strong&gt; tidies up rough notes without adding facts you didn't give it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fytxcbot3v0454u2tctqn.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fytxcbot3v0454u2tctqn.png" alt="The setup form: a washing-machine leak typed in English, with Start from an example, Enhance prompt, Dictate situation and Start voice practice" width="800" height="876"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Asking for help in the middle of the call:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9rkr0mhmh6tyk6aii5ed.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9rkr0mhmh6tyk6aii5ed.png" alt="The explanation dialog: the question's meaning, a note on どうされましたか, and a suggested reply with furigana" width="800" height="822"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The word list on the call card:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwh3dmo649m9jukkb83yo.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwh3dmo649m9jukkb83yo.png" alt="The call card's useful words: 管理会社, 田中 and 洗濯機 with furigana, romaji and English meanings" width="800" height="1011"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  What my friend said
&lt;/h3&gt;

&lt;p&gt;I sent the link to my friend on WhatsApp and asked them to let me know if it helps. This is what came back:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frkznedsexrxwtt7iz67l.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frkznedsexrxwtt7iz67l.png" alt="WhatsApp chat: I share the Before I Call link with my friend. They reply " width="800" height="681"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I didn't expect the PDF to be the part they liked. I made it as a cheat sheet for the real call, but for someone studying for the JLPT it's also a word list from a conversation they actually had.&lt;/p&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;Try it at &lt;a href="https://before-i-call.onrender.com" rel="noopener noreferrer"&gt;before-i-call.onrender.com&lt;/a&gt;. There's no sign-up. Voice practice needs a microphone. Without one, tap &lt;strong&gt;Play guided example&lt;/strong&gt; to watch a recorded call in Japanese or English.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2ky1gf298ag9sk66wstt.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2ky1gf298ag9sk66wstt.png" alt="The PDF call card: the situation, each turn as a speech bubble with furigana, romaji and meaning" width="800" height="1132"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;


&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/adkbbx" rel="noopener noreferrer"&gt;
        adkbbx
      &lt;/a&gt; / &lt;a href="https://github.com/adkbbx/before_I_call" rel="noopener noreferrer"&gt;
        before_I_call
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;div&gt;
&lt;a rel="noopener noreferrer" href="https://github.com/adkbbx/before_I_call/docs/images/logo.svg"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2Fadkbbx%2Fbefore_I_call%2FHEAD%2Fdocs%2Fimages%2Flogo.svg" width="76" alt="Before I Call logo"&gt;&lt;/a&gt;
&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;Before I Call&lt;/h1&gt;
&lt;/div&gt;
&lt;div class="markdown-heading"&gt;
&lt;h3 class="heading-element"&gt;Rehearse the phone call you've been putting off.&lt;/h3&gt;
&lt;/div&gt;
&lt;p&gt;A patient AI voice partner for everyday Japanese and English calls, with furigana on every word,&lt;br&gt;
help that never interrupts the conversation, and a call card to take with you.&lt;br&gt;
Use it in the cloud, or &lt;b&gt;free and offline on your own computer&lt;/b&gt; with Gemma 4, Whisper and Kokoro.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://before-i-call.onrender.com" rel="nofollow noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/73e3fad786c3d941e2b2cf53d224bdb71b0c835aca00f8c398a0a5d0d3b3b520/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f5472795f69745f6c6976652d6265666f72652d2d692d2d63616c6c2e6f6e72656e6465722e636f6d2d3331356234393f7374796c653d666f722d7468652d6261646765" alt="Try it live"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://deepmind.google/models/gemma/" rel="nofollow noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/078f8dc5422b0d1ea9b38b5a6b6c2cf2fd6c5479f3843b538607706d6de97423/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f47656d6d615f342d6f70656e5f776569676874732d3161373365383f7374796c653d666c61742d737175617265" alt="Gemma 4 open weights"&gt;&lt;/a&gt;
&lt;a href="https://elevenlabs.io/agents" rel="nofollow noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/6fe212177af7800b3b9bba35df7a8e02e43a2a7ca1acd47de3bb249896e5aed9/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f456c6576656e4c6162732d4167656e74732d3030303030303f7374796c653d666c61742d737175617265" alt="ElevenLabs Agents"&gt;&lt;/a&gt;
&lt;a href="https://docs.digitalocean.com/products/gradient-ai-platform/how-to/use-serverless-inference/" rel="nofollow noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/662624bae698400e417d417d9ea67a611aa3e82d45b324652098e6a227f8cf6e/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f4469676974616c4f6365616e2d7365727665726c6573735f696e666572656e63652d3030383066663f7374796c653d666c61742d737175617265" alt="DigitalOcean serverless inference"&gt;&lt;/a&gt;
&lt;a href="https://docs.sentry.io/product/insights/ai/agents/" rel="nofollow noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/e005abc142c60935dff2440bcb899265550ac0978fe2eb52f1cbc991bf6d63d2/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f53656e7472792d6167656e745f74726163696e672d3336326435393f7374796c653d666c61742d737175617265" alt="Sentry agent tracing"&gt;&lt;/a&gt;
&lt;a href="https://render.com" rel="nofollow noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/dfc69bd58faad7e0a384a6da7a417d3f61354a98794e6533b3f2517dfa704578/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f52656e6465722d6465706c6f7965642d3436653362373f7374796c653d666c61742d737175617265" alt="Deployed on Render"&gt;&lt;/a&gt;
&lt;a href="https://github.com/adkbbx/before_I_call/LICENSE" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/e04ea5c0d2a58a0233a89eda6eb5b38d982abc818aa2bdb30142f21a1eff0f1b/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f6c6963656e73652d4d49542d3331356234393f7374796c653d666c61742d737175617265" alt="MIT license"&gt;&lt;/a&gt;
&lt;a href="https://github.com/adkbbx/before_I_call#free-local-mode-on-your-own-computer" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/da7581c0e6ea9ed2949c5267fb431ed06f53f9f67871eb3304b6e7e8e548537c/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f667265655f6c6f63616c5f6d6f64652d576869737065725fc2b75f47656d6d615fc2b75f4b6f6b6f726f2d6530613532363f7374796c653d666c61742d737175617265" alt="Free local mode with Whisper, Gemma and Kokoro"&gt;&lt;/a&gt;
&lt;a rel="noopener noreferrer nofollow" href="https://camo.githubusercontent.com/e4f202733aede36bad72b88e8ebe6482bc52a311a8cdebcd913cd87b3dbb969c/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f74657374732d38375f70617373696e672d3331356234393f7374796c653d666c61742d737175617265"&gt;&lt;img src="https://camo.githubusercontent.com/e4f202733aede36bad72b88e8ebe6482bc52a311a8cdebcd913cd87b3dbb969c/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f74657374732d38375f70617373696e672d3331356234393f7374796c653d666c61742d737175617265" alt="87 tests passing"&gt;&lt;/a&gt;
&lt;a href="https://dev.to/challenges/hacktoberfest-weekend-2026-10-01" rel="nofollow"&gt;&lt;img src="https://camo.githubusercontent.com/a5340b509445a5480c2fab7253ba6e314df05f1d4dd45982775e981a0b92cc36/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f4861636b746f626572666573742d323032362d3364356635383f7374796c653d666c61742d737175617265" alt="Hacktoberfest 2026"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://before-i-call.onrender.com" rel="nofollow noopener noreferrer"&gt;&lt;strong&gt;Live app&lt;/strong&gt;&lt;/a&gt; · &lt;a href="https://github.com/adkbbx/before_I_call#how-a-practice-call-works" rel="noopener noreferrer"&gt;How it works&lt;/a&gt; · &lt;a href="https://github.com/adkbbx/before_I_call#free-local-mode-on-your-own-computer" rel="noopener noreferrer"&gt;Free local mode&lt;/a&gt; · &lt;a href="https://github.com/adkbbx/before_I_call/docs/ARCHITECTURE.md" rel="noopener noreferrer"&gt;Architecture&lt;/a&gt; · &lt;a href="https://github.com/adkbbx/before_I_call#run-it-locally" rel="noopener noreferrer"&gt;Run it locally&lt;/a&gt;&lt;/p&gt;
&lt;br&gt;

  
  &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2Fadkbbx%2Fbefore_I_call%2FHEAD%2Fdocs%2Fimages%2Fhome-light.png" width="900" alt="Before I Call home page: practice a phone call in Japanese or English, or play a guided example"&gt;

&lt;/div&gt;
&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;The problem&lt;/h2&gt;
&lt;/div&gt;
&lt;p&gt;A phone call is the hardest everyday conversation in a second language. There is no face to read and no time to look anything up, and the other person speaks at native speed. People who live in Japan put off calling the building manager, the clinic or the delivery company, not because they can't manage the conversation, but because the call itself feels risky.&lt;/p&gt;
&lt;p&gt;Phrasebooks help with the first sentence. They don't…&lt;/p&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/adkbbx/before_I_call" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;


&lt;p&gt;MIT licensed and built during the challenge. It has 87 tests (80 Python, 7 Node), and the README explains how to run it locally.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I Built It
&lt;/h2&gt;

&lt;h3&gt;
  
  
  The call runs on ElevenLabs Agents
&lt;/h3&gt;

&lt;p&gt;The live call is an ElevenLabs agent. It handles all the audio: the WebRTC connection, speech recognition, working out when you've finished speaking, interruptions, and the voice. I use one agent for every scenario. When a call starts, the app sends that call's prompt, first line, language and voice as overrides, so the same agent can be a Japanese building manager or an English clinic receptionist.&lt;/p&gt;

&lt;p&gt;I slowed the agent's speech to 0.85x for learners and capped calls at two minutes, with a daily limit on the server so the credits can't all go in one day. Hanging up uses the built-in &lt;code&gt;end_call&lt;/code&gt; tool. Its description tells the model to use it only after your question is dealt with and you've said you don't need anything else, not when you just say thanks. &lt;strong&gt;Explain question&lt;/strong&gt; opens a second, text-only session with a helper prompt, so the explanation never ends up in the call transcript. &lt;strong&gt;Slow replay&lt;/strong&gt; uses ElevenLabs text to speech (Flash v2.5 at 0.7x).&lt;/p&gt;

&lt;p&gt;The agent has had 113 conversations so far. ElevenLabs groups them by topic on its own, and the top three match the examples in the app: booking an appointment (27), the washing machine (19) and a missed delivery (17).&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fa5lc0pv9s443vs3nf7aw.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fa5lc0pv9s443vs3nf7aw.png" alt="ElevenLabs agent dashboard: 113 conversations, with topics grouped as Appointment Booking and Scheduling (27), Washing Machine Issue Resolution (19) and Missed Delivery Redelivery Request (17)" width="800" height="305"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Why I switched to a custom LLM
&lt;/h3&gt;

&lt;p&gt;When I started, the agent used one of the LLMs hosted by ElevenLabs, Qwen3.5-397B-A17B. Its usage came out of the same ElevenLabs credits I got for Hacktoberfest, and it was going through them fast. ElevenLabs lists it at about $0.016 a minute and $0.003 a message, on top of the voice. A custom LLM costs nothing on the ElevenLabs side:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft9semi240lwnyjkulteu.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft9semi240lwnyjkulteu.png" alt="The ElevenLabs LLM picker: hosted models with their prices, including Qwen3.5-397B-A17B at about $0.0159 per minute and $0.0030 per message, and Custom LLM at $0.0000" width="437" height="363"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I wanted those credits for the voice, so I switched the agent to a custom LLM and pointed it at Gemma 4 31B on DigitalOcean, which bills my DigitalOcean account instead. With the per-token prices my tracing code uses for Gemma ($0.18 per million input tokens and $0.50 per million output tokens), the turn in the trace further down cost about $0.0002. The ElevenLabs credits now only pay for the audio.&lt;/p&gt;

&lt;p&gt;At first the custom LLM pointed straight at DigitalOcean. Later I put my own server in between so I could trace and check every reply. This is the agent's LLM setting now:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzqy1dx7uofaix46lqauc.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzqy1dx7uofaix46lqauc.png" alt="The agent's Custom LLM setting in ElevenLabs: Chat Completions, server URL https://before-i-call.onrender.com/llm/v1/ and model ID gemma-4-31B-it" width="434" height="491"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;On each turn, the agent sends the conversation to my FastAPI server on Render, the server forwards it to Gemma 4 31B on DigitalOcean's serverless inference, and the response streams straight back, so the voice can start as soon as the first words arrive.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fihxfyvrlids4vsau6p4s.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fihxfyvrlids4vsau6p4s.png" alt="Architecture: the browser talks to ElevenLabs over WebRTC; ElevenLabs calls the Gemma proxy on Render, which streams Gemma 4 31B from DigitalOcean and sends traces to Sentry" width="800" height="365"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  The proxy on Render
&lt;/h3&gt;

&lt;p&gt;The custom LLM URL points at my own server: a FastAPI app that Render runs as one Docker web service. The same service serves the React app, the API and the proxy. The whole setup is a &lt;code&gt;render.yaml&lt;/code&gt; Blueprint in the repo (trimmed here):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;services&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;web&lt;/span&gt;
    &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;before-i-call&lt;/span&gt;
    &lt;span class="na"&gt;runtime&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;docker&lt;/span&gt;
    &lt;span class="na"&gt;plan&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;starter&lt;/span&gt;
    &lt;span class="na"&gt;healthCheckPath&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;/api/health&lt;/span&gt;
    &lt;span class="na"&gt;disk&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;usage-data&lt;/span&gt;
      &lt;span class="na"&gt;mountPath&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;/var/data&lt;/span&gt;
      &lt;span class="na"&gt;sizeGB&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;1&lt;/span&gt;
    &lt;span class="na"&gt;envVars&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;LLM_PROXY_KEY&lt;/span&gt;
        &lt;span class="na"&gt;sync&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;GRADIENT_MODEL&lt;/span&gt;
        &lt;span class="na"&gt;value&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;gemma-4-31B-it&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;MAX_CONCURRENT_CALLS&lt;/span&gt;
        &lt;span class="na"&gt;value&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;2'&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;MAX_CALL_SECONDS&lt;/span&gt;
        &lt;span class="na"&gt;value&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;120'&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;MAX_CALLS_PER_VISITOR_DAY&lt;/span&gt;
        &lt;span class="na"&gt;value&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;3'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Keys are &lt;code&gt;sync: false&lt;/code&gt;, so they're set in the Render dashboard and never committed. The limits that protect my credits are plain environment variables: two calls at once, two minutes each, three calls per visitor a day. The 1 GB disk holds the SQLite files for those limits and the analytics, so they survive redeploys.&lt;/p&gt;

&lt;p&gt;The proxy is a single route. It accepts only the agent's key, refuses every model except Gemma 4 31B, and passes DigitalOcean's stream through unchanged while it reads each chunk for tracing and the rule checks (trimmed):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nd"&gt;@router.post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;/llm/v1/chat/completions&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;chat_completions&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Request&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="bp"&gt;...&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="nf"&gt;isinstance&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;body&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="n"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="c1"&gt;# The key is never usable for other, more expensive models.
&lt;/span&gt;        &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;HTTPException&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;400&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Only &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; is available.&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="bp"&gt;...&lt;/span&gt;
    &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;relay&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
        &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;chunk&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;aiter_bytes&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
            &lt;span class="n"&gt;stream&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;feed&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="k"&gt;yield&lt;/span&gt; &lt;span class="n"&gt;chunk&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nc"&gt;StreamingResponse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;relay&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="n"&gt;media_type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;text/event-stream&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                             &lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Cache-Control&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;no-cache&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;X-Accel-Buffering&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;no&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;})&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;An always-on web service suits this. Each reply is a streamed response that stays open while Gemma talks. The server reuses one HTTPS client for DigitalOcean, so a turn doesn't pay for a new TLS handshake. And when you interrupt the AI, ElevenLabs drops the request, so the proxy closes the stream and records the turn as cancelled in Sentry.&lt;/p&gt;

&lt;p&gt;Here's the proxy at work in Render's logs, one line per Gemma request from the agent:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fup9w3t6i5dpvgdhpwmzy.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fup9w3t6i5dpvgdhpwmzy.png" alt="Render application logs filtered to POST /llm/v1/chat/completions: a steady list of 200 OK responses on October 4 and 5" width="800" height="544"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The service is managed by the Blueprint, which stays synced to the repo, and every push to main deploys on its own:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fc7j667vtciwcaf9ajyih.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fc7j667vtciwcaf9ajyih.png" alt="Render Blueprints page: before-i-call, synced with adkbbx/before_I_call on main" width="799" height="295"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5qpdffdna3conozbvnee.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5qpdffdna3conozbvnee.png" alt="Render events for the before-i-call web service (Docker, Starter, Blueprint managed, live): deploys started by " width="799" height="586"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Fixing how numbers are read
&lt;/h3&gt;

&lt;p&gt;The first version read 7時 as "nana-ji". Nobody says that. Japanese numbers change their reading depending on the counter after them: 7時 is しちじ, 10分 is じゅっぷん, 4月1日 is しがつついたち and 2人 is ふたり. Dates, times and the number of people are exactly what you say on a phone call, so this had to be right.&lt;/p&gt;

&lt;p&gt;I fixed it with a pronunciation dictionary: 118 entries I wrote by hand plus 5,989 generated ones for dates, clock times, durations and counters, 6,107 in total. A script builds the rules and uploads them to the ElevenLabs agent as a pronunciation dictionary, and the app uses the same list for slow replay and the recorded demos. The text on screen doesn't change. This is the dictionary on the agent:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhqhetmcb99e2dyr9jv40.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhqhetmcb99e2dyr9jv40.png" alt="ElevenLabs agent settings: the pronunciation dictionary " width="444" height="232"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs2tp1dwigc5iwkzv7luj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs2tp1dwigc5iwkzv7luj.png" alt="Japanese voices trip over dates and counters, so I wrote 6,107 fixes: 7時 is shichiji, not nana-ji" width="800" height="615"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Tuning the prompt for Gemma
&lt;/h3&gt;

&lt;p&gt;Switching to Gemma meant rewriting the prompt. I tested it with text-only runs of the agent. The old prompt never called &lt;code&gt;end_call&lt;/code&gt; (0 of 2 scenarios), and the tuned one hung up in all 4. It also keeps ordinary words in kanji and writes numbers in kana so the voice reads them correctly, and it asks for practice details instead of your real ones. It still missed the hang-up sometimes in real calls, and Sentry caught that. The start of the tuned prompt:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqz6ak1gyrnawv8cenotg.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqz6ak1gyrnawv8cenotg.png" alt="The start of the agent's system prompt in ElevenLabs: polite spoken Japanese, one question per turn, kanji for ordinary words, numbers with counters in hiragana such as 7時→しちじ, and never inventing dates, prices or availability" width="762" height="304"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Checking every reply with Sentry
&lt;/h3&gt;

&lt;p&gt;Since every turn goes through my server, I can check each reply against the rules of the practice call. If you said goodbye and the AI didn't hang up, if it "confirmed" a booking, or if it put digits or English letters into Japanese speech, that becomes a Sentry issue tagged with the prompt version and turn number. Each request is also a &lt;code&gt;gen_ai&lt;/code&gt; trace with time to first token, token counts and estimated cost. After a call you can tap &lt;strong&gt;Report this reply&lt;/strong&gt;, and the report shows up next to that turn's trace. Conversation text only goes to Sentry if you choose to attach it.&lt;/p&gt;

&lt;p&gt;Here's a real one. On October 3, the practice partner didn't hang up after the learner said goodbye, and Sentry caught it 8 times in four minutes. Each rule gets its own issue:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5wphabsqbkzn8v5rtg3c.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5wphabsqbkzn8v5rtg3c.png" alt="Sentry issue feed showing one unresolved issue, Practice partner: missed hang up, on /llm/v1/chat/completions with 8 events" width="800" height="301"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The issue lists every time it happened and which release it came from:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3qvsbv911uq67nkpqwih.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3qvsbv911uq67nkpqwih.png" alt="The missed hang up issue in Sentry: 8 events between 5:57 and 6:01 PM UTC on October 3, all on the same release in production" width="800" height="632"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;For comparison, here's a turn where the partner did hang up. The &lt;code&gt;execute_tool end_call&lt;/code&gt; span is the hang-up. The Gemma call took 1.27 s (Sentry's average for it is 1.49 s), used about 1,100 input tokens and 43 output tokens, and cost less than a cent. Traces don't include the conversation text, so the Input tab is empty.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr216xseyxbgjo70k5uaz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr216xseyxbgjo70k5uaz.png" alt="A Sentry trace of one practice turn: invoke_agent, chat gemma-4-31B-it taking 1.27 s, the end_call tool and the request to DigitalOcean, with 1.1K input tokens, 43 output tokens and a cost under $0.01" width="799" height="333"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;All of those Gemma calls run on DigitalOcean's serverless inference, so I never had to run a GPU server. In two days, October 3 and 4, the app sent about 1.28 million input tokens to Gemma 4 31B and got about 75,000 back. Most of it is input because every turn sends the instructions and the conversation so far, while the replies are short, like the 43 tokens in the trace above.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh4c1w707sh96pv9yp65t.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh4c1w707sh96pv9yp65t.png" alt="DigitalOcean Serverless Inference, Analyze tab: 1,278,483 input tokens and 74,719 output tokens, all used on October 3 and 4" width="800" height="728"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Bugs I hit along the way
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Gemma sometimes wrote voice directions like &lt;code&gt;[calm]&lt;/code&gt; or &lt;code&gt;[pause]&lt;/code&gt; into its replies. The app strips them from the text on screen.&lt;/li&gt;
&lt;li&gt;Small models sometimes say a tool call out loud, like &lt;code&gt;end_call(reason="done")&lt;/code&gt;, instead of calling the tool. The app removes that text before it's spoken and treats it as a hang-up.&lt;/li&gt;
&lt;li&gt;Running Gemma 4 locally, the first replies came back empty because the model spent all its tokens on hidden reasoning. Setting &lt;code&gt;reasoning_effort: "none"&lt;/code&gt; fixed it, and replies now take about half a second on my laptop's GPU.&lt;/li&gt;
&lt;li&gt;Kokoro's Japanese voice wouldn't install on Windows, because pyopenjtalk needs CMake to build. pyopenjtalk-plus has prebuilt wheels, so I switched to that.&lt;/li&gt;
&lt;li&gt;Romaji showed 10月4日 as "10tsuki4hi". Romaji and furigana now use the same reading list as the voice, so it reads &lt;em&gt;juugatsu yokka&lt;/em&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Why Does Open Innovation Matter?
&lt;/h2&gt;

&lt;p&gt;The hosted version runs on paid APIs. I also wanted a version that's free and private, so the same code can run entirely on a laptop. With &lt;code&gt;LOCAL_VOICE=1&lt;/code&gt;, Whisper handles speech recognition, Gemma 4 E2B (4.3 GB, through Ollama) plays the other person, and Kokoro-82M speaks. Explanations, slow replay, Enhance prompt, the word list, furigana and the PDF all still work. Nothing is sent to a provider, it works offline, and you can switch models with one setting (&lt;code&gt;LOCAL_LLM_MODEL&lt;/code&gt;).&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx5cw1xbzbmhd035ce2sw.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx5cw1xbzbmhd035ce2sw.png" alt="Same app, two brains: Gemma 4 31B on DigitalOcean with ElevenLabs, or Gemma 4 E2B on a laptop with Whisper and Kokoro" width="800" height="414"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;On my laptop (Ryzen 9 5900HS, RTX 3060 with 6 GB), the AI greets you in 2.4 s, answers what you say in about 6 s, and an explanation takes 1.4 s.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb012bqobkf76hoyugtik.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb012bqobkf76hoyugtik.png" alt="A free local practice call: Gemma's reply with furigana and romaji, a tap-to-speak microphone and the call controls" width="800" height="500"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The local version isn't as good. The ElevenLabs voice sounds more natural, responds faster and lets you interrupt it. Locally you tap to talk, and the small model doesn't follow the role-play rules as closely as the 31B. It's possible at all because Gemma's weights are open: the same model family runs on DigitalOcean and on my laptop, with the same prompt, rules and pronunciation list. Furigana and romaji don't need a model. The open-source pykakasi and Janome libraries handle them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prize Categories
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Best Use of ElevenLabs:&lt;/strong&gt; one agent over WebRTC for every call, with per-call prompt, first line, language and voice overrides, the built-in &lt;code&gt;end_call&lt;/code&gt; tool, a custom LLM pointed at Gemma, a separate text-only session for explanations, text to speech for slow replay, and a 6,107-entry pronunciation dictionary on the agent.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Best Use of Sentry Agent Tracing:&lt;/strong&gt; every Gemma request is a &lt;code&gt;gen_ai&lt;/code&gt; trace with time to first token, tokens and cost. Replies are checked against the role-play rules, and user reports link to the exact turn.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Best Use of DigitalOcean:&lt;/strong&gt; serverless inference runs Gemma 4 31B for every call turn, explanation, prompt enhancement and word list, with no GPU to manage.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Best Use of Gemma:&lt;/strong&gt; Gemma does four jobs (the other person on the call, the explanation helper, the prompt enhancer and the word picker), as 31B in the cloud and E2B on a laptop.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Best Use of Render:&lt;/strong&gt; one Docker web service from a &lt;code&gt;render.yaml&lt;/code&gt; Blueprint serves the app, the API and the custom LLM proxy that streams every Gemma reply to ElevenLabs, with a persistent disk for usage limits and analytics, keys kept out of the repo, and the credit limits as environment variables.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Built solo.&lt;/p&gt;

&lt;p&gt;If you live in Japan and put off phone calls too, try it and tell me in the comments what's missing.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>weekendchallenge</category>
      <category>hf26challenge</category>
      <category>ai</category>
    </item>
    <item>
      <title>Hello world!</title>
      <dc:creator>ADK BBX</dc:creator>
      <pubDate>Fri, 02 Oct 2026 18:16:26 +0000</pubDate>
      <link>https://dev.to/adk_bbx/hello-world-58d3</link>
      <guid>https://dev.to/adk_bbx/hello-world-58d3</guid>
      <description></description>
    </item>
  </channel>
</rss>
