<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Hayato Kamiya</title>
    <description>The latest articles on DEV Community by Hayato Kamiya (@officekamiya).</description>
    <link>https://dev.to/officekamiya</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4105539%2Fef87f2c4-7ba2-4839-9778-351b14e02e77.png</url>
      <title>DEV Community: Hayato Kamiya</title>
      <link>https://dev.to/officekamiya</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/officekamiya"/>
    <language>en</language>
    <item>
      <title>Forgetting is for remembering: my AI chat broke long before the token limit</title>
      <dc:creator>Hayato Kamiya</dc:creator>
      <pubDate>Sun, 06 Sep 2026 03:45:54 +0000</pubDate>
      <link>https://dev.to/officekamiya/forgetting-is-for-remembering-my-ai-chat-broke-long-before-the-token-limit-oee</link>
      <guid>https://dev.to/officekamiya/forgetting-is-for-remembering-my-ai-chat-broke-long-before-the-token-limit-oee</guid>
      <description>&lt;p&gt;I say "I can't remember" a lot. But when I look at what is actually going on inside me, that one phrase turns out to hold several different things.&lt;/p&gt;

&lt;h2&gt;
  
  
  "I can't remember" is not one state
&lt;/h2&gt;

&lt;p&gt;In my experience there are three.&lt;/p&gt;

&lt;p&gt;The first: &lt;strong&gt;nothing catches.&lt;/strong&gt; This is less "I can't remember" than a state where I cannot even feel that anything is there. There is no target to try to remember.&lt;/p&gt;

&lt;p&gt;The second: &lt;strong&gt;something catches, but it is too heavy to come up.&lt;/strong&gt; I know there is something related to what is in front of me; it is only the detail I do not have. It is far away, or heavy, and my hand does not reach it. In a sense, this is already remembering.&lt;/p&gt;

&lt;p&gt;The third: &lt;strong&gt;when I try to pull it up, it gets pushed back.&lt;/strong&gt; Here the problem is not reach. My hand is on it, and something is working to keep it from coming up. Somewhere in my experience there is something that contradicts that memory, and it silently refuses to let the two surface together. The conscious part of me wants to haul it up, and the part that does not want to is the one that is moving.&lt;/p&gt;

&lt;p&gt;The first has no target. The second is out of reach. The third is within reach and is being held down. Lined up they look alike, but they are completely different things.&lt;/p&gt;

&lt;h2&gt;
  
  
  Seen through how an AI behaves
&lt;/h2&gt;

&lt;p&gt;You send an AI a prompt and something comes back. That much is a given. The interesting part, I think, is how it behaves when what comes back is not quite right. Let me line that up against the three above.&lt;/p&gt;

&lt;p&gt;What corresponds to the second, "something catches but it doesn't come up." When there seems to be something related but the substance does not come out, the AI does not stop. It fills the gap with something nearby, or raises the level of abstraction and returns a generality. Some of what gets called hallucination looks to me like this state.&lt;/p&gt;

&lt;p&gt;But people do the same thing. There are more than a few people who cannot say "I don't remember" and bluff instead. In filling the gap rather than stopping, people and AI are very much alike.&lt;/p&gt;

&lt;p&gt;What corresponds to the third, "pushed back when you try to pull it up." This is a state where a safety lock is engaged, in people and in AI alike. In a person, it was built unconsciously out of experience: something like a trauma, for example, holding the memory down. In an AI, it was put there from outside, during training. The difference is who set the lock: you, or someone outside.&lt;/p&gt;

&lt;p&gt;And yet the way it comes off is the same. In both cases it does not come off from the front; it can come off through something like a back door, an unexpected stimulus. In people, as a flashback. In AI, as prompt injection. Not only is the reason for stopping similar; the shape of how it comes off is the same too.&lt;/p&gt;

&lt;p&gt;Then what about the first, "nothing catches"? This is the one that does not line up.&lt;/p&gt;

&lt;h2&gt;
  
  
  Yes, no, null
&lt;/h2&gt;

&lt;p&gt;Let me arrange the same three in different words. Yes, no, null.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"Yes"&lt;/strong&gt; is the state where something can be pulled up. The second case, heavy and not coming up, is on this side. That there is something related has already come back; only the detail is heavy.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"No"&lt;/strong&gt; is the state where what exists is the information that it is not there. The third case, being pushed back, is this one. That case is not "it must be there." It is "I would rather it were not." And if the information "it is not there" exists, then, turned around, that is also a trace of something that once was.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"Null"&lt;/strong&gt; is the state where nothing is defined at all. The first case is on this side.&lt;/p&gt;

&lt;p&gt;And there are two kinds of null. One is null from the start: I never experienced it, I do not know it. The other is something that used to be a yes or a no, and became null.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The second kind is forgetting.&lt;/strong&gt; The second and third cases were about remembering. They were never forgetting.&lt;/p&gt;

&lt;p&gt;The awkward part is that from the inside, I cannot tell these two nulls apart. Whether I forgot, or never knew, I cannot say.&lt;/p&gt;

&lt;p&gt;"No" and "null" also look alike from the inside. But they are not the same. Holding the information "it is not there" affects what you do before and after. It does not produce the same result as having nothing at all. Even if you cannot tell them apart from inside, the behavior that comes out is different.&lt;/p&gt;

&lt;p&gt;So what about AI?&lt;/p&gt;

&lt;p&gt;When nothing catches for an AI, either it never had it, or the thread changed and everything was wiped. The former is null from the start. In the latter, the entity itself has been swapped out. Either way, there is no transition in which a yes or a no becomes null inside the same entity.&lt;/p&gt;

&lt;p&gt;That is why it did not line up. &lt;strong&gt;The AI we have today does not have this function of forgetting.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Before the limit, there was a different limit
&lt;/h2&gt;

&lt;p&gt;If there is no function for forgetting, then the more exchanges you have, the more it should keep carrying. I ended up seeing what actually happens while building something myself.&lt;/p&gt;

&lt;p&gt;I build a service called &lt;a href="https://awai.live/en" rel="noopener noreferrer"&gt;AWAI&lt;/a&gt;, which visualizes the flow of thinking from conversations, and it has a chat feature where you can talk with an AI.&lt;/p&gt;

&lt;p&gt;When I first built that chat, I did it the way the Gemini API tutorial does. Send the text of every past exchange, every time. Each time, the AI reads the whole history before it responds. In support-center terms, it is as if the staff read the entire case history from the beginning before every single reply to the customer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;When the conversation got long, it broke down.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;There was plenty of room under the maximum token count. Even so, once the conversation passed about 20 turns, the instructions in the system prompt started getting vague, its grasp of the conversation got vague, and it started behaving oddly.&lt;/p&gt;

&lt;p&gt;This was not about the limit. Separate from the maximum token count, there is an effective amount it can handle while keeping quality, and I think we had crossed that. Put the other way around, as long as you stay inside that range, continuing the exchange is not a problem in itself. A growing history is not the bad thing.&lt;/p&gt;

&lt;p&gt;So I reduced the amount of text sent to the API. With that, both the system prompt's instructions and the content of the conversation were grasped properly again.&lt;/p&gt;

&lt;p&gt;Reduce it, and it got better.&lt;/p&gt;

&lt;p&gt;What I saw from this was that with Gemini 3.1, precision saturates first, at roughly the amount of information in a 20-turn conversation. It is not the limit that caps it; precision does. In that case, select, and stay within that range.&lt;/p&gt;

&lt;p&gt;(The turn count is a feel from the implementation at the time; change the model or the length of the prompts and it should move.)&lt;/p&gt;

&lt;p&gt;And here a question comes up. So how do the AIs out in the world do their selecting?&lt;/p&gt;

&lt;h2&gt;
  
  
  Probably everyone is doing this
&lt;/h2&gt;

&lt;p&gt;From here on, this is not something I have verified. It is what I sense from using them. I would honestly rather write it as fact, but I have the impression that this is an area every AI has been built not to answer, whichever one you ask. So I write it as a guess.&lt;/p&gt;

&lt;p&gt;My guess is that they are doing the same thing. The reason they do not break down over many turns is that some selection is happening inside.&lt;/p&gt;

&lt;p&gt;Early ChatGPT broke down as soon as the conversation got a little long. Was that because this selection was not there yet? Seen that way, it adds up.&lt;/p&gt;

&lt;p&gt;If so, the meaning of the growing token counts changes too. That was not the acquisition of forgetting. They became able to select, conversation by conversation, from what had accumulated inside, and so a long conversation now holds. What grew was the ability to choose what to pass along each time, not the ability to forget.&lt;/p&gt;

&lt;p&gt;That is the view: AI did not become able to forget. It extended the length over which it can get by without forgetting.&lt;/p&gt;

&lt;h2&gt;
  
  
  Selection is not forgetting
&lt;/h2&gt;

&lt;p&gt;What is the difference between selection and forgetting?&lt;/p&gt;

&lt;p&gt;Selection means choosing again every time. Everything is still there, and you choose the portion to send this time. So what was not chosen has not disappeared, and it may be chosen on the next turn. Turned around, the candidates keep growing. The cost of choosing does not go down.&lt;/p&gt;

&lt;p&gt;Forgetting means the candidates themselves get fewer. Once something has come off the list of what can be pulled up, you no longer go looking for it. There is less to search, so the scan finishes sooner. The content is unchanged, and pulling gets lighter.&lt;/p&gt;

&lt;p&gt;If you keep everything pullable, then no matter how capable the system, the search space grows until you cannot get a realistic speed out of it. So you need an operation that reduces what can be pulled up in the first place. Forgetting, I think, may be that operation.&lt;/p&gt;

&lt;p&gt;Seen this way, what happened earlier means something different too. Reducing what I sent made it grasp things again. That is remembering because I reduced. Selection itself is not forgetting, but the act of reducing led to the result of remembering.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Forgetting is for remembering.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;And finally, to be honest about it: &lt;strong&gt;AWAI's AI chat has no mechanism for forgetting either.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;What it does is this. From the information in past conversation, it pulls out what is close in meaning to what is being discussed now, and hands that to the AI. Alongside that, for each piece of information, it calculates a priority every time, "how much should this be kept," and the high ones get a mark that says "this matters."&lt;/p&gt;

&lt;p&gt;So there is a mechanism for choosing. But there is no mechanism that reduces what can be pulled up.&lt;/p&gt;

&lt;p&gt;There is a mechanism that decides what should be kept, and only the mechanism for forgetting is missing. That is because I have not been able to think of how forgetting should be built. Everything I have written here about AI is also, as it stands, about my own tool.&lt;/p&gt;

&lt;h2&gt;
  
  
  What decides what gets forgotten?
&lt;/h2&gt;

&lt;p&gt;I have got as far as forgetting leading to remembering. Then what decides what gets forgotten?&lt;/p&gt;

&lt;p&gt;My guess is that it is decided on the side of instinct. Stress, events that left a strong mark, how I feel that day. Somewhere the conscious mind cannot do anything about, what gets forgotten is shifting.&lt;/p&gt;

&lt;p&gt;Thinking that way, I notice an odd asymmetry.&lt;/p&gt;

&lt;p&gt;Taking things in can be done consciously, to a degree. You memorize for an exam. On a trip, you look at a view thinking, I want to keep this one. What you meant to keep does stay, more or less. (Of course, plenty of things stay on their own, out of habit, without any intention to keep them.)&lt;/p&gt;

&lt;p&gt;But forgetting does not run on the conscious mind. Try to forget and you cannot. If anything, you remember it more.&lt;/p&gt;

&lt;p&gt;And then: recalling does not actually run on the conscious mind either. I feel as though I am pulling things out, but in practice they come out on their own when the conditions line up. The three cases I started with were all "it did not come out" rather than "I could not pull it," and I think that is why.&lt;/p&gt;

&lt;p&gt;Only the way in, where things enter memory, is, just barely, on my side. The way out, and the way things get called up, I cannot operate myself.&lt;/p&gt;

&lt;p&gt;That, I suspect, is what makes it hard to put on an AI. If what decides what gets forgotten is the condition of a living creature, things like stress and how you feel that day, then an AI that has none of that has no basis to build the criterion from.&lt;/p&gt;

&lt;p&gt;I do have one candidate for something to use instead. Anything that takes too long to retrieve gets judged "did not come back, so it is not there." The second case I started with, "something catches but it is too heavy to come up," looks like exactly that. It does not disappear, though; if a similar stimulus comes along, it comes out again.&lt;/p&gt;

&lt;p&gt;If this worked, what would change? Probably you would no longer need to cut threads. You could keep going without breaking it up.&lt;/p&gt;

&lt;p&gt;And there is one more thing I think about.&lt;/p&gt;

&lt;p&gt;Because you forget, you can remember. And what you remember becomes your next action. I wrote earlier that holding the information "it is not there" affects what you do before and after. If so, then what you forgot comes out directly as a difference in what you do.&lt;/p&gt;

&lt;p&gt;The accumulation of that may be what makes a person who they are, or an AI what it is. Not what you remember. What you forgot.&lt;/p&gt;

&lt;p&gt;But what decides that, I still do not know.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I found after writing this
&lt;/h2&gt;

&lt;p&gt;I wrote this without reading any literature, from my own sense of it alone. When I went looking afterwards, several things turned out to be close. I leave them here for anyone who wants to go deeper.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Schacter, &lt;a href="https://doi.org/10.1037/0003-066X.54.3.182" rel="noopener noreferrer"&gt;The seven sins of memory&lt;/a&gt; (American Psychologist, 1999)&lt;/strong&gt; — Sorts memory's failures into seven kinds and argues they are not design flaws but by-products of otherwise adaptive features. The second case in section 1, "something catches but it is too heavy to come up," is what this taxonomy calls blocking, the tip-of-the-tongue state.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Bjork &amp;amp; Bjork, &lt;a href="https://bjorklab.psych.ucla.edu/research/" rel="noopener noreferrer"&gt;A new theory of disuse&lt;/a&gt; (1992)&lt;/strong&gt; — Splits the strength of a memory in two: how well it is learned (storage strength) and how accessible it is (retrieval strength). The "yes" in section 3 reads as high storage strength with low retrieval strength.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Liu et al., &lt;a href="https://arxiv.org/abs/2307.03172" rel="noopener noreferrer"&gt;Lost in the Middle: How Language Models Use Long Contexts&lt;/a&gt; (2023)&lt;/strong&gt; — Measures how language model performance drops sharply just from moving the relevant information to a different position in a long context. Close to section 4, "before the limit, there was a different limit."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Chroma, &lt;a href="https://www.trychroma.com/research/context-rot" rel="noopener noreferrer"&gt;Context Rot: How Increasing Input Tokens Impacts LLM Performance&lt;/a&gt; (2025)&lt;/strong&gt; — Tests 18 models and finds their performance grows unreliable, unevenly, as the input gets longer. Also section 4.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Anthropic, &lt;a href="https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents" rel="noopener noreferrer"&gt;Effective context engineering for AI agents&lt;/a&gt; (2025)&lt;/strong&gt; — The idea of treating context as a finite resource. It describes compaction: summarizing a conversation that is nearing the limit and starting a fresh window from the summary. A published method, aimed at people building agents, in the area I wrote about as a guess in section 5.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Richards &amp;amp; Frankland, &lt;a href="https://doi.org/10.1016/j.neuron.2017.04.037" rel="noopener noreferrer"&gt;The Persistence and Transience of Memory&lt;/a&gt; (Neuron, 2017)&lt;/strong&gt; — Argues that the purpose of memory is not to keep information accurate over time but to make decisions better, and that forgetting is a function serving that. The closest thing to section 6, "forgetting is for remembering."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Borges's story "Funes el memorioso" (1942) and Luria's case study "The Mind of a Mnemonist" (1968)&lt;/strong&gt; — Both about a person who cannot forget. There is also a &lt;a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC6022980/" rel="noopener noreferrer"&gt;paper&lt;/a&gt; (Dementia &amp;amp; Neuropsychologia, 2018) that reads the two side by side. This is the "keep everything pullable" side of section 6.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Anderson &amp;amp; Green, &lt;a href="https://doi.org/10.1038/35066572" rel="noopener noreferrer"&gt;Suppressing unwanted memories by executive control&lt;/a&gt; (Nature, 2001)&lt;/strong&gt; — An experiment showing that repeatedly keeping yourself from recalling something makes it harder to recall later. It gets harder to recall; it does not become null. So this looks less like forgetting than like a lock put on a memory after the fact, the safety-lock side of section 2.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What set this off was a service called &lt;a href="https://awai.live/en" rel="noopener noreferrer"&gt;AWAI&lt;/a&gt;. It visualizes the flow of thinking from conversations, and the chat in sections 4 and 6 lives inside it. It still has no mechanism for forgetting.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>machinelearning</category>
      <category>programming</category>
    </item>
    <item>
      <title>I measured Gemini Live Translate in 3 languages: same rhythm, 2x the characters</title>
      <dc:creator>Hayato Kamiya</dc:creator>
      <pubDate>Fri, 04 Sep 2026 06:46:22 +0000</pubDate>
      <link>https://dev.to/officekamiya/i-measured-gemini-live-translate-in-3-languages-same-rhythm-2x-the-characters-3fnm</link>
      <guid>https://dev.to/officekamiya/i-measured-gemini-live-translate-in-3-languages-same-rhythm-2x-the-characters-3fnm</guid>
      <description>&lt;p&gt;I build a meeting tool where people who speak different languages talk to each other directly. You speak Japanese, the other person hears English; they speak English, you hear Japanese. The speech translation runs on Google's &lt;a href="https://workspace.google.com/blog/ai-and-machine-learning/gemini-35-live-translate" rel="noopener noreferrer"&gt;Gemini 3.5 Live Translate&lt;/a&gt; (&lt;code&gt;gemini-3.5-live-translate-preview&lt;/code&gt;).&lt;/p&gt;

&lt;p&gt;&lt;a href="https://awai.live/en/demo/meeting" rel="noopener noreferrer"&gt;Here is a demo&lt;/a&gt; with three people in Japanese, English and Spanish. No signup.&lt;/p&gt;

&lt;p&gt;I wanted to add live subtitles to that screen. You listen to the translated audio, and you read &lt;strong&gt;what the other person actually said&lt;/strong&gt; in their own words, before translation. Something like putting original-language subtitles on a dubbed film. Not a transcript that appears after they finish talking: the Pixel Recorder kind, where words show up one or two at a time while they are still speaking.&lt;/p&gt;

&lt;p&gt;That gave me a hard requirement I could not design around: &lt;strong&gt;if the delay is five seconds, the feature is useless.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Subtitles exist so you can react, in the moment, to what was just said. A subtitle for something said five seconds ago arrives after the conversation has already moved on. Under a second, and you can glance at the panel and keep up. Somewhere between those two numbers is the line between a feature and a decoration.&lt;/p&gt;

&lt;p&gt;One word before going further, because it runs through the whole post. The transcript does not come back in one piece: it arrives &lt;strong&gt;in small instalments&lt;/strong&gt;, and each instalment is a &lt;strong&gt;chunk&lt;/strong&gt;. Subtitles are those chunks, joined in the order they turn up.&lt;/p&gt;

&lt;p&gt;So the question was simple. How fast does Live Translate actually emit chunks, and how big is each one?&lt;/p&gt;

&lt;p&gt;I could not answer it from anything I had.&lt;/p&gt;

&lt;h2&gt;
  
  
  When I built the prototype, there was no reason to measure this
&lt;/h2&gt;

&lt;p&gt;A few months earlier I had built a prototype on &lt;a href="https://ai.google.dev/gemini-api/docs/live" rel="noopener noreferrer"&gt;Gemini Live&lt;/a&gt; (&lt;code&gt;gemini-3.1-flash-live-preview&lt;/code&gt;). The question then was whether you could feed it human speech directly and have the model recognize it as text at all. It could, and it did that well.&lt;/p&gt;

&lt;p&gt;Later I settled on Live Translate for making human conversation legible. In principle the two should behave alike, so the prototype code looked reusable — but &lt;strong&gt;how much of it carried over was an open question&lt;/strong&gt;. So I rebuilt on top of it with Live Translate.&lt;/p&gt;

&lt;p&gt;When I got to the subtitles, I needed latency numbers. I went back to the prototype for them, and they were not there.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The other side had changed.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Gemini Live is a voice version of an AI chat. You speak, the model answers, one turn each. &lt;strong&gt;The AI waits for you to finish.&lt;/strong&gt; It does not matter how long you take; the conversation does not break. So time was never a variable in the prototype. "Speech goes in, text comes out" was the whole question, and there was no reason to record when each chunk arrived.&lt;/p&gt;

&lt;p&gt;Subtitles are not like that. The other side is a person, and people do not wait. If the text does not land while they are still speaking, it is worth nothing — so &lt;strong&gt;time itself becomes the requirement&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Same model family, similar-looking API, and the requirements swap out completely depending on whether the thing on the other side is an AI or a human. The prototype answered "does this work." What I needed here was "does this arrive in time." Different job.&lt;/p&gt;

&lt;h2&gt;
  
  
  Making time measurable
&lt;/h2&gt;

&lt;p&gt;To make the subtitles arrive in time, I first had to know how long they were taking. And to measure "how many seconds after speech begins", &lt;strong&gt;I have to know when speech began&lt;/strong&gt; — on my side of the wire.&lt;/p&gt;

&lt;p&gt;So I added VAD (voice activity detection): a step that watches the incoming audio locally and catches &lt;strong&gt;the moment a voice starts&lt;/strong&gt;. When it fires, I record that timestamp and tell the API "starting now".&lt;/p&gt;

&lt;p&gt;Gemini Live Translate can also work out where speech begins and ends on its own. But then &lt;strong&gt;the start of the clock sits on its side&lt;/strong&gt;, and there is no way to say what the seconds are counted from. Bringing the time reference over to my side is the whole reason the VAD is there.&lt;/p&gt;

&lt;p&gt;The check runs every 32 milliseconds, so the gap to the true onset stays inside that. Against &lt;strong&gt;the one second or so that a subtitle is allowed&lt;/strong&gt;, that is fine enough.&lt;/p&gt;

&lt;p&gt;After that it is just logging. The start-of-speech signal, each audio send, each transcript chunk coming back, all with timestamps, on a monotonic clock so nothing shifts under you if the system time is adjusted.&lt;/p&gt;

&lt;p&gt;The measuring script pulls the connection settings, the endpoint and the model straight out of the code we actually run, so what goes over the wire is byte-identical to a real session. A measurement of a lookalike would prove nothing about production.&lt;/p&gt;

&lt;h2&gt;
  
  
  Results
&lt;/h2&gt;

&lt;p&gt;Japanese, English and Arabic. I picked Arabic deliberately: it is written right-to-left, and its script is far from both the Latin alphabet and Japanese.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Japanese&lt;/th&gt;
&lt;th&gt;English&lt;/th&gt;
&lt;th&gt;Arabic&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Audio length&lt;/td&gt;
&lt;td&gt;90s&lt;/td&gt;
&lt;td&gt;100s&lt;/td&gt;
&lt;td&gt;100s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Speech segments&lt;/td&gt;
&lt;td&gt;19&lt;/td&gt;
&lt;td&gt;22&lt;/td&gt;
&lt;td&gt;21&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Transcript chunks&lt;/td&gt;
&lt;td&gt;67&lt;/td&gt;
&lt;td&gt;73&lt;/td&gt;
&lt;td&gt;75&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Characters per chunk (median)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;6&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;13&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;9&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Same (mean / p90 / max)&lt;/td&gt;
&lt;td&gt;6.0 / 10 / 13&lt;/td&gt;
&lt;td&gt;12.0 / 20 / 23&lt;/td&gt;
&lt;td&gt;9.0 / 14 / 22&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Inter-chunk interval (median / p90)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;1.023s&lt;/strong&gt; / 1.267s&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;1.049s&lt;/strong&gt; / 1.399s&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;1.038s&lt;/strong&gt; / 1.272s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Speech onset → first chunk (median)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.64s&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.80s&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.57s&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Same (p90 / max)&lt;/td&gt;
&lt;td&gt;1.03s / 2.88s&lt;/td&gt;
&lt;td&gt;1.24s / 4.40s&lt;/td&gt;
&lt;td&gt;1.13s / 2.23s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Arrived during speech / after&lt;/td&gt;
&lt;td&gt;63 / 4&lt;/td&gt;
&lt;td&gt;67 / 6&lt;/td&gt;
&lt;td&gt;71 / 4&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The three bold rows are the ones that mattered for building subtitles.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Characters per chunk&lt;/strong&gt; — how much text arrives at once. &lt;strong&gt;This is the only row that differs by 2×&lt;/strong&gt; (6 / 13 / 9)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Inter-chunk interval&lt;/strong&gt; — how long until the next piece of text. &lt;strong&gt;Near-identical across all three&lt;/strong&gt; (1.023 / 1.049 / 1.038s)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Speech onset → first chunk&lt;/strong&gt; — how long before the first character shows up. &lt;strong&gt;Also close&lt;/strong&gt; (0.64 / 0.80 / 0.57s)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So the cadence and the head start are the same in every language, and &lt;strong&gt;only the amount of text per delivery changes&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That difference is really about how much a single character carries. Since the interval is almost exactly a second, the per-second count is the same figure: about 6 characters in Japanese, 9 in Arabic, 12 in English. Turned around, that is roughly &lt;strong&gt;170 / 110 / 85 milliseconds of screen time per character&lt;/strong&gt;. One Japanese character occupies about as long as two English ones.&lt;/p&gt;

&lt;p&gt;The chunk counts (67 / 73 / 75) differ because the clips were 90 and 100 seconds long. Per second they come out at 0.74 / 0.73 / 0.75.&lt;/p&gt;

&lt;p&gt;For what it is worth, I ran Japanese twice in independent sessions: median onset 0.66s then 0.64s, interval 1.009s then 1.023s. Run-to-run spread is around ±0.05s.&lt;/p&gt;

&lt;h2&gt;
  
  
  The rhythm does not depend on the language
&lt;/h2&gt;

&lt;p&gt;Three languages, three scripts, and the interval lands on the same second. That does not look like a coincidence.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The server is not emitting per word. It is emitting per unit of time.&lt;/strong&gt; It appears to flush roughly once a second with whatever it recognized in that second. That means "one or two words at a time" is a property you get in any language. What changes is how many characters fit in a second.&lt;/p&gt;

&lt;p&gt;Onset latency is language-independent too: 0.57s to 0.80s at the median, 1.03s to 1.24s at p90. The outliers (2.2s to 4.4s) were all the &lt;strong&gt;first utterance of a session&lt;/strong&gt;. Warm-up. Every subsequent utterance landed in the 0.4–1.2s band. It recurs once per reconnection, which is worth knowing if your session drops.&lt;/p&gt;

&lt;p&gt;So why does &lt;em&gt;that&lt;/em&gt; number — characters per second — differ by 2×?&lt;/p&gt;

&lt;p&gt;This part is speculation. I suspect the rate at which people exchange information by voice settles in roughly the same place whoever is speaking. The processing happens on the human side, and that side does not change with the language. If so, however differently a language chooses to express things, the amount it can carry per second stays put: a script that packs more into each character needs fewer of them, one that packs less needs more, and the same content goes across either way.&lt;/p&gt;

&lt;p&gt;There is research that measured something close to this. A &lt;a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC6984970/" rel="noopener noreferrer"&gt;2019 study of 17 languages across 9 families, with 170 native speakers&lt;/a&gt;, found that speech rate (syllables per second) varies a lot between languages, but multiply it by the information density of a syllable and every language lands around &lt;strong&gt;39 bits per second (SD 5.10)&lt;/strong&gt;. Languages with less dense syllables talk faster to make up for it.&lt;/p&gt;

&lt;p&gt;That study does not include Arabic, and I measured characters on a screen rather than bits or syllables, so the two do not stack directly. Still, "the interval and the head start line up, and only the characters per second differ" fits that shape of explanation.&lt;/p&gt;

&lt;p&gt;Two more things I checked, because they change how the receiving side has to be built:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The stream is append-only.&lt;/strong&gt; I brought up the Pixel Recorder at the top, and one thing it does is &lt;strong&gt;rewrite what it already showed you&lt;/strong&gt;: as recognition catches up, it goes back and corrects an earlier guess.&lt;/p&gt;

&lt;p&gt;If Live Translate did the same, subtitles would have to be built on the assumption that text already on screen can move. That changes the whole display, so it was worth checking first.&lt;/p&gt;

&lt;p&gt;I looked for rewrites with a deliberately strict test: does any chunk start with three or more characters matching the tail of what I already have? Two hits across all runs, one in Japanese and one in English. Both turned out to be at the same point in the audio, where the speaker restarted a phrase, and the model transcribed the restart faithfully. Not a protocol resend: a real repetition, reproducing at the same position across four independent sessions. &lt;strong&gt;Zero corrections.&lt;/strong&gt; What arrives simply stacks up in order; nothing behind reaches forward to change what is already there. You join the pieces as they come and never look back.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Chunks arrive while the person is still talking.&lt;/strong&gt; This one decides whether the feature is possible at all.&lt;/p&gt;

&lt;p&gt;If the text only landed once an utterance was over, subtitles could only ever be &lt;strong&gt;finished sentences appearing in blocks&lt;/strong&gt; — the opposite of what I wanted. And in this audio the longest single utterance ran &lt;strong&gt;14 seconds&lt;/strong&gt;. Wait for the end and the panel sits empty for those 14 seconds, by which point the conversation has moved on.&lt;/p&gt;

&lt;p&gt;It does not wait. &lt;strong&gt;Over 90% of the text was on screen while the person was still speaking.&lt;/strong&gt; The rest turns up just after they finish.&lt;/p&gt;

&lt;h2&gt;
  
  
  Only the on-screen text varies by language
&lt;/h2&gt;

&lt;p&gt;And it was never only the character count. &lt;strong&gt;How many chunks the same speech breaks into, and where those breaks fall, both change with the language.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Here is what actually arrived:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;English   'Anyway, the true'  /  ' ability of the'  /  ' lovely ghostwriter'
Arabic    'طيب، عشان كده يا جماعه'  /  ' قدرات الكاتب'
Japanese  'ラブリーゴースト'  /  'ライターの真の'
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;English and Arabic broke &lt;strong&gt;at word boundaries&lt;/strong&gt;, and every continuation carried a leading space (look at the gap in front of &lt;code&gt;' ability of the'&lt;/code&gt;). Join them and that space is your word separator.&lt;/p&gt;

&lt;p&gt;Japanese did not. The single word "ラブリーゴーストライター" (lovely ghostwriter) is split into &lt;code&gt;'ラブリーゴースト'&lt;/code&gt; and &lt;code&gt;'ライターの真の'&lt;/code&gt; — &lt;strong&gt;cut in the middle of the word, with nothing to signal that it is mid-word.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Two consequences for the client:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Concatenate verbatim. Do not trim, do not insert separators.&lt;/strong&gt; Trimming the leading space welds English and Arabic words together. The Japanese mid-word split means a word will sit visibly incomplete for about a second, and there is nothing you can do about that.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Switch text direction per bubble.&lt;/strong&gt; Subtitles go out the way a messaging app looks: one utterance, one bubble. What comes back is the speaker's own language, so in a multilingual meeting a right-to-left bubble and a left-to-right one sit in the same panel. Pick a direction for the whole panel and one of them breaks, so the UI has to decide bubble by bubble (in HTML, &lt;code&gt;dir="auto"&lt;/code&gt; on each one).&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  So the UI limit cannot be a character count
&lt;/h2&gt;

&lt;p&gt;Here is the rule I was about to write: cap the subtitle panel at N characters, drop the oldest.&lt;/p&gt;

&lt;p&gt;With a 2× spread in characters per second, that cap buys a different amount of &lt;em&gt;time&lt;/em&gt; in every language. Cap at 60 characters and Japanese keeps 10 seconds on screen while English keeps 5. Tune N until Japanese feels right and English scrolls away before you can read it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cap by rendered lines instead&lt;/strong&gt;, and the difference cancels out:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Characters per second&lt;/th&gt;
&lt;th&gt;Characters per line (mobile)&lt;/th&gt;
&lt;th&gt;Characters in 3 lines&lt;/th&gt;
&lt;th&gt;3 lines ≈&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Japanese&lt;/td&gt;
&lt;td&gt;6.0 (full-width)&lt;/td&gt;
&lt;td&gt;~20&lt;/td&gt;
&lt;td&gt;60&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~10s&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;English&lt;/td&gt;
&lt;td&gt;12.0&lt;/td&gt;
&lt;td&gt;~40&lt;/td&gt;
&lt;td&gt;120&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~10s&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Arabic&lt;/td&gt;
&lt;td&gt;9.0&lt;/td&gt;
&lt;td&gt;~35&lt;/td&gt;
&lt;td&gt;105&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~12s&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Capping by lines is, of course, still capping by "characters per line × lines". &lt;strong&gt;The difference is that you are not the one choosing the characters per line.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Chunks arrive with no line breaks in them. Where the text wraps is decided by the width of the panel and the font — and the panel is the same width in every language, because it is the width of the screen. Fill it with full-width characters and a line holds about 20; fill it with half-width ones and it holds about 40. &lt;strong&gt;The character count per line moves by exactly as much as the characters are wide.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That turned out to move in step with the characters-per-second figure. So three lines hold 60 characters in one language and 120 in another, and &lt;strong&gt;in time they both come out at 10–12 seconds&lt;/strong&gt;. The line count was doing the normalization for me.&lt;/p&gt;

&lt;p&gt;If the horizontal side follows the font, so should the vertical. &lt;strong&gt;Reserve the height of those three lines in font-relative units, not pixels.&lt;/strong&gt; Arabic, Devanagari and Thai have taller line boxes, and a px-fixed box clips them.&lt;/p&gt;

&lt;p&gt;(Three lines rather than three messages, incidentally, because that 14-second utterance came to about 84 characters, which overflows three lines on its own.)&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://awai.live/en/demo/meeting" rel="noopener noreferrer"&gt;three-line panel is running here&lt;/a&gt; if you want to see whether the reasoning holds up in practice. Turn subtitles on in the toolbar; they are off by default.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this does not cover
&lt;/h2&gt;

&lt;p&gt;Of the three recordings I measured, &lt;strong&gt;only the Japanese one is an actual person speaking.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;For English and Arabic I had no audio to work with, so I took the &lt;em&gt;translated speech&lt;/em&gt; the model hands back during a Japanese run and fed that in as the input. Which means those two languages are &lt;strong&gt;the model talking to itself&lt;/strong&gt;: clean, no accent, no disfluency, no one talking fast. Real speakers will chunk differently, and the model card itself notes that accents degrade language detection.&lt;/p&gt;

&lt;p&gt;Separately from this measurement, I have confirmed that English speech from a real person &lt;strong&gt;is transcribed properly&lt;/strong&gt;. But that is as far as it goes — I did not measure chunk granularity or latency on it. For every other language I have not even checked that much.&lt;/p&gt;

&lt;p&gt;I have not measured Chinese, Korean, Thai, or Hindi. The one-second period looks structural rather than tuned, so I would expect it to hold, but that is a guess and I would rather label it as one.&lt;/p&gt;

&lt;p&gt;And this is one speaker at a time. What happens when several people hold separate sessions concurrently is a different measurement.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this ended up
&lt;/h2&gt;

&lt;p&gt;The tool is called &lt;a href="https://awai.live/en" rel="noopener noreferrer"&gt;AWAI&lt;/a&gt;: people in different languages sit in the same meeting and talk. The subtitles I measured here are on that screen.&lt;/p&gt;

&lt;p&gt;And the &lt;a href="https://awai.live/en/demo/meeting" rel="noopener noreferrer"&gt;demo&lt;/a&gt; I have linked a few times is &lt;strong&gt;not a mock-up made for this post&lt;/strong&gt;. The code drawing those subtitles is exactly the production code; the only difference is that the data comes from a recorded replay instead of a live channel. &lt;strong&gt;Open a real meeting and the same thing runs.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Subtitles, though, only exist during the meeting. &lt;strong&gt;What AWAI is actually for is what is left afterwards&lt;/strong&gt;: the conversation sorted by topic, with who said what on which subject. If subtitles are "what did they just say", that is "what did we actually talk about". The 0.64 seconds I measured here is one instant at the front of it.&lt;/p&gt;

&lt;p&gt;Building for more than one language usually gets framed as absorbing the differences between them. Measured, the differences and the sameness fell into two clean piles.&lt;/p&gt;

&lt;p&gt;The text on screen differs — the character count, where it breaks, how tall a line is. The rhythm it arrives in does not. And that seems to be less about Gemini than about &lt;strong&gt;people exchanging information by voice at roughly the same rate whatever language they use&lt;/strong&gt;. The writing system varies; the speaking and listening underneath it does not.&lt;/p&gt;

&lt;p&gt;That is the part that stuck with me while building a multilingual meeting tool. The thing that matters when two people talk to each other appears to survive the change of language.&lt;/p&gt;

&lt;p&gt;And I only saw it because I wired the translation up myself and &lt;strong&gt;measured what came back, one piece at a time&lt;/strong&gt;. Using it would not have shown me that. If you are building on a streaming speech API, measure it for your own use rather than carrying over what you checked before — what you verified with an AI on the other end goes back up for verification the moment you put it between two people.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>python</category>
      <category>programming</category>
    </item>
  </channel>
</rss>
