<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Evan Lin</title>
    <description>The latest articles on DEV Community by Evan Lin (@evanlin).</description>
    <link>https://dev.to/evanlin</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F409957%2Fc150d4a7-cb20-469d-a230-bac27232c577.jpeg</url>
      <title>DEV Community: Evan Lin</title>
      <link>https://dev.to/evanlin</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/evanlin"/>
    <language>en</language>
    <item>
      <title>[Book Sharing] Tsaisang’s Tales of the Strange: Japanese Mythology, Ghost Stories, and Sometimes Taiwan</title>
      <dc:creator>Evan Lin</dc:creator>
      <pubDate>Thu, 23 Jul 2026 13:25:47 +0000</pubDate>
      <link>https://dev.to/evanlin/book-sharing-tsaisangs-tales-of-the-strange-japanese-mythology-ghost-stories-and-sometimes-2g0i</link>
      <guid>https://dev.to/evanlin/book-sharing-tsaisangs-tales-of-the-strange-japanese-mythology-ghost-stories-and-sometimes-2g0i</guid>
      <description>&lt;p&gt;&lt;a href="https://moo.im/a/egjpDI" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fss8w63nmadu83mspf0ls.jpg" width="210" height="295"&gt;&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Tsai-sang Talks About the Strange
Japanese Myths and Spiritual Ghost Stories, and Sometimes Taiwan
Rated by 73 people
Author: Tsai Yi-chu  Publisher: Eurasian Publishing Group
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Recommended links to buy the book:
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;Readmoo: &lt;a href="https://moo.im/a/egjpDI" rel="noopener noreferrer"&gt;Buy here&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;h1&gt;
  
  
  Foreword:
&lt;/h1&gt;

&lt;p&gt;This is the first book I finished reading in 2026. I didn't write any book reviews in the first half of this year because I only read a little bit of many books. I spent quite a long time reading this one as well; I happened to see it while looking for books, and it turned out the second half of the book was quite captivating, so I finished it all in one go.&lt;/p&gt;

&lt;h2&gt;
  
  
  Synopsis
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Not just crazy, but crazier! Tsai Yi-chu a.k.a. the Chuunibyou Professor of Folklore unleashes a flood of ghost stories!

◆ How do Japan's first generation of gods play out soap opera dramas every day?
◆ How does Japanese mythology hide various "sexual metaphors" in its stories?
◆ Yokai aren't all troublemakers; which ones can make you rich or send you to heaven?
◆ Is there bullying in the world of Yokai? Can just getting old and ugly make you a Yokai?
◆ Is the Jade Emperor actually not a CEO? Is Guanyin Bodhisattva actually a foreigner?
◆ Can Taiwan also have "Shigong (Priest) Watches" and "Yokai Pokémon"?

Taiwanese people fear ghosts, Japanese people fear ghosts, people all over the world fear ghosts...
It doesn't matter if you haven't seen "Ghost Stories," come here to listen to Tsai-sang talk nonsense and discuss gods and ghosts, so you won't feel "creepy" anymore!

Most people's impression of Japan is—endless temples to visit, super cute and otaku cosplayers, AV actresses... Hey! There must also be supernatural stories, Sadako, Yokai, and various urban legends!

Listen to how Tsai-sang combines hair-raising ghost encounters with historical stories passed down through the ages, and see how he uses super "grounded" slang to reveal the cultural meanings behind Japanese mythology!

Japanese Folklore PhD Tsai Yi-chu has gathered years of research on folklore, using myths and ghost stories as a medium and easy-to-understand "netizen" language to lead readers into Japan's "Gods and Monsters." This includes the genealogy of Japanese gods, the connection between Yokai and culture, and the hidden meanings within. At the same time, he allows Taiwan's various gods to actively participate in the text through a Taiwan-Japan friendly crossover, making you understand ghost talk and become obsessed with gods and ghosts! After reading, I guarantee your mom will ask why you are reading this book on your knees!

Because "Tsai-sang Talks About the Strange" will make you shout on your knees: "What on earth was Japanese mythology on? I want some of that too!"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This book isn't one of those stiff academic papers; it's a cultural analysis book where Tsai Yi-chu (Tsai-sang), a PhD in Folklore from the University of Tsukuba, uses super "grounded" netizen slang and a hilarious style to strip down Japanese mythology and spiritual ghost stories for you! The most brilliant part is that he doesn't just talk about Japan; he occasionally pulls back to Taiwan's folk perspective for comparison.&lt;/p&gt;

&lt;p&gt;Here are the three core sections and key points of this book refined for you:&lt;/p&gt;

&lt;h2&gt;
  
  
  Summary of the Three Core Sections
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;strong&gt;Section Category&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Core Research Focus&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Tsai-sang's "Taiwanese-style Plain Language Interpretation" and Highlights&lt;/strong&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;1. Japanese Mythological Prototypes&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;The birth of Japan's first generation of gods (Izanagi, Izanami, Amaterasu, Susanoo) and the belief in vengeful spirits in history.&lt;/td&gt;
&lt;td&gt;Uses "Sex and violence, gore and SOD collections" to roast the absurd plots of Japanese mythology. Introduces the "Vengeful Spirit Fan Club" of ancient Japanese history and endemic species like Tengu.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;2. Yokai and Urban Legends&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;How rural ghosts and monsters evolved into modern urban legends over time (e.g., Slit-Mouthed Woman, Super High-Speed Granny).&lt;/td&gt;
&lt;td&gt;Yokai are "new pets of urbanization," reflecting the collective anxiety and loneliness of modern people, and the media's role in fueling the supernatural trend.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;3. And Sometimes Taiwan&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Cross-sea interactions and cultural comparisons of beliefs between Taiwan and Japan (e.g., Nagasaki Mazu and Tainan's General Flying Tiger).&lt;/td&gt;
&lt;td&gt;Demonstrates a "Taiwan-Japan Friendly Crossover." Reflects on Taiwanese people's own cultural roots and subjectivity through the lens of Japanese ghost stories.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  I. Japanese Mythology: A First Family More Dramatic Than a Soap Opera
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Love and Hate of the First Family&lt;/strong&gt;: The fallout process of Japan's creator gods (Izanagi and Izanami) is absurd and horrifying (the wife turns into a rotting corpse in Yomi, the husband is so scared he flees and divorces). Their descendants, the Sun Goddess Amaterasu and her brother Susanoo, also have a love-hate relationship. Tsai-sang jokingly says that from a modern perspective, these plots are simply a collection of various horror and gore scenarios.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Vengeful Spirit Fan Club&lt;/strong&gt;: Many high-ranking gods worshipped in Japanese history (such as Sugawara no Michizane, the God of Learning, and Emperor Sutoku) were actually "victims of political struggles who died miserable deaths." Because later generations feared they would become vengeful spirits and seek revenge, they quickly built shrines to worship them as gods, forming Japan's unique culture of vengeful spirit worship.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  II. Yokai and Urban Legends: Collective Anxiety of Modern People
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Yokai are the City's New Pets&lt;/strong&gt;: Former Yokai (like Kappa and Yama-uba) lived deep in the mountains and forests, representing human awe of nature. After urbanization, Yokai also "moved into the city," evolving into urban legends like the Slit-Mouthed Woman and the Super High-Speed Granny.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reflecting the Loneliness of Real Society&lt;/strong&gt;: The birth of these modern legends appears to be horror stories on the surface, but at their core, they reflect the alienation and collective anxiety of urbanites. At the same time, the book reviews the "rise and fall of the supernatural craze" fueled by Japanese mass media (TV supernatural programs) for ratings in the 80s and 90s.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  III. And Sometimes Taiwan: The Mysterious Connection Between Taiwanese and Japanese Gods and Ghosts
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Japanese-speaking Mazu and Japanese Gods&lt;/strong&gt;: The book specifically mentions the intertwining of Taiwanese and Japanese beliefs. For example, there are several Mazu temples in Nagasaki, Japan, where Mazu "speaks Japanese" due to localization. Meanwhile, in Tainan, Taiwan, there is the "General Flying Tiger Temple," which enshrines Shigeo Sugiura, a Japanese pilot who sacrificed himself to protect Taiwanese villagers during WWII.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Folklore is the "Study of People"&lt;/strong&gt;: Tsai-sang emphasizes that whether studying Japanese mythology or Taiwanese supernatural phenomena, what is ultimately terrifying or absurd is not the ghosts and monsters, but the human society behind them. Belief can soothe the soul because it reflects the logic of contemporary thinking.&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Core Quote of the Book:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;"To understand people is to understand ghosts; Yokai and spirits are all imaginations based on reality."&lt;/p&gt;

&lt;p&gt;We must discover the main reasons why each phenomenon forms. When we use Japanese ghost stories as a mirror to deeply understand the workings of folklore and legends, we can then look back with clearer eyes to discover and identify the cultural identity that belongs to "Taiwan itself."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Reflections
&lt;/h2&gt;

&lt;p&gt;This book is full of many rural legends, and the final section, which includes one of his research reports and actual events, will leave you astonished. First, the book begins by sharing stories of Japanese ghosts and monsters, reflecting on the origins of many ghost stories behind Japanese mythology. It also discusses the connection between Japanese sex and violence and their many ghost stories.&lt;/p&gt;

&lt;p&gt;The second part shares some Taiwan-related stories and also discusses the origin of the "Turbo Granny" in &lt;em&gt;Dandadan&lt;/em&gt; and the stories related to the Slit-Mouthed Woman. These will make you want to read it all in one breath. As the Ghost Month seems to be approaching again recently, it seems this series of books will become popular again. Everyone should check it out.&lt;/p&gt;

</description>
      <category>books</category>
      <category>reviews</category>
    </item>
    <item>
      <title>[Book Sharing] Taiwan's AI Future: Analyzing Trends, Local Landscape, Corporate Strategy, and Personal Development</title>
      <dc:creator>Evan Lin</dc:creator>
      <pubDate>Thu, 23 Jul 2026 13:25:21 +0000</pubDate>
      <link>https://dev.to/evanlin/book-sharing-taiwans-ai-future-analyzing-trends-local-landscape-corporate-strategy-and-40io</link>
      <guid>https://dev.to/evanlin/book-sharing-taiwans-ai-future-analyzing-trends-local-landscape-corporate-strategy-and-40io</guid>
      <description>&lt;p&gt;&lt;a href="https://moo.im/a/02oszP" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fechi9y8wjct6t3nc9ave.jpg" width="210" height="293"&gt;&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Taiwan's AI Future
Analyzing the latest AI trends, Taiwan's situation, corporate strategies, and personal development
Author: Chien Lee-feng, Hsiao Yu-pin  
Publisher: Business Weekly 
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Recommended links to buy the book:
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;Readmoo: &lt;a href="https://moo.im/a/02oszP" rel="noopener noreferrer"&gt;Click here to buy&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;h1&gt;
  
  
  Preface:
&lt;/h1&gt;

&lt;p&gt;This is the second book I finished reading in 2026. It is a fairly new book, released at the end of 2025. I bought it because my company invited Chien Lee-feng to give a speech in 2024. Later, I happened to see his book on my e-book shelf and decided to take a look.&lt;/p&gt;

&lt;h2&gt;
  
  
  Outline
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;When AI rewrites the world, what is Taiwan's next step?
Former Managing Director of Google Taiwan, computer scientist, and AI scholar—Chien Lee-feng
Writes the first instruction manual for the AI era for Taiwan,
Helping Taiwanese people understand the opportunities and challenges of the AI age!

The world undergoes a digital revolution every ten years:
● 1990: Personal computers start the computer generation;
● 2000: The Internet creates the web generation;
● 2010: Mobile devices and social media lead the mobile generation;
● 2020: Generative AI like ChatGPT makes a shocking debut...
Now is the AI generation, where the rules of the game are completely rewritten, and the gap is rapidly widening between 1:99.
Will you fall behind and be eliminated, or seize the opportunity and become a 1% winner?
This book takes you deep into the AI landscape to master the key to transformation!

【AI Development under Geopolitics】
As the world undergoes an AI-driven paradigm shift, the US sees AI as the key to its return to hegemony. This not only predicts that AI productization will completely subvert the world's rules of operation but also opens up infinite possibilities for the future evolution of AI. From the "Manhattan Project" to the "Stargate" layout, this book will deeply analyze the global situation under the US-China tariff war and provide insights into how AI reshapes the international order.
● The 1:99 challenge: Countries, companies, and individuals who seize the opportunity may become the unique "1" that far surpasses others, while others become the "99" lagging far behind.
● The emergence of DeepSeek has subverted the US monopoly, bringing a "rebalancing" to the AI world, equivalent to inventing the "poor man's atomic bomb."
● If the pace of domestic chip production is not accelerated, an unproductive America will have no tomorrow and will directly lose competitiveness in the AI battle. TSMC has thus become the X-factor in the US-China confrontation.

【Taiwan Looking at the World】
The AI wave is sweeping the globe; this is not just a technological innovation, but a key turning point for national development. In this giant wave, Taiwan not only has the uniquely endowed "Silicon Shield" TSMC but also sees a golden decade to become "the world's Taiwan" as AI challenges and opportunities coexist. This book gives you a glimpse into the future potential of Taiwan's manufacturing industry and how old and new enterprises are redefining "Made in Taiwan."
● Facing geopolitical changes, Taiwan's manufacturing-oriented enterprises should go with the flow. Through overseas production and R&amp;amp;D in Taiwan, they can create a "Taiwan + N (foreign)" model, helping Taiwan remove the red supply chain and join the US-led supply chain.
● An island's market is always outside. Flying to Japan or the US for travel or business for a few days does not equal internationalization. Internationalization is daily life being impacted by different cultures.

【AI Practice in All Walks of Life】
The AI era is a key moment for corporate transformation and talent reinvention. Only companies that dare to pivot will have competitive opportunities. This book lists cases of how companies in different industries respond to AI and provides practical strategic directions to guide Taiwan's corporate transformation to seize the AI market and move towards growth and innovation.
● The impact of AI can be compared to "musical chairs." From tech giants to SMEs, whether it's "emptying the cage for new birds" (industrial restructuring) or empowering employees, wherever the wind blows, new opportunities lie there.
● "Old-ventures + New-ventures": Shifting from software integration to hardware-software integration, combining the advantages of both, makes AI applications possible.
● Developing Sovereign AI doesn't end with outsourcing. Whether building your own model or asking tech giants for help, the strategy must be planned clearly, otherwise, it might just be a waste of money.

【Mastering the Golden Key to Personal Learning and Career】
As a member of the AI generation, how to use AI to improve learning efficiency while clearly identifying AI's limits is an important task. This book suggests how to use AI tools while pointing out that human differentiated experience will become an irreplaceable treasure. Therefore, cleverly accumulating personal unique value is the only way to remain invincible in the AI era.
● AI likes to use certain specific sentence patterns because AI is a probabilistic concept, so naturally, there are some patterns. But conversely, precisely because its data volume is large enough, it can try various combinations that humans have never seen.
● AI has raised the "passing line" of many jobs from 60 to 80 points in one fell swoop, forcing all industries to redefine the core competencies and value of human labor.
● In the AI era, senior talents with professional foundations learn AI the fastest because their long-term accumulated knowledge can judge the correctness of AI-generated content. This AI paradigm shift has, in turn, amplified the advantages of the older generation.

This is an AI survival guide tailored for Taiwan, helping you fully grasp the context of the AI revolution and find the path to growth for the nation, enterprises, and individuals amidst the changes.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This book can be described as an "AI Era Survival Manual" tailored for Taiwanese people and businesses by Dr. Chien Lee-feng, former Managing Director of Google Taiwan, and senior media professional Hsiao Yu-pin. Dr. Chien uses a very pragmatic and precise local perspective to analyze how Taiwan should reposition itself, how companies should play the international game of hardware-software integration, and how individuals can avoid falling into the crisis of "brain outsourcing" under this crazy AI wave.&lt;/p&gt;

&lt;p&gt;I have organized the four core frameworks of the book for you to help you grasp the overall context through this overview:&lt;/p&gt;

&lt;h3&gt;
  
  
  Overview of the Book's Four Core Frameworks
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;strong&gt;Category&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Core Pain Points &amp;amp; Trends&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Breakthrough Strategies for Taiwan &amp;amp; Individuals&lt;/strong&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;1. Latest AI Trends&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;AI brings high centralization and unification, potentially evolving into a 1:99 disparity in capabilities and resources; however, the rise of emerging forces like DeepSeek is bringing "rebalancing" opportunities to the world.&lt;/td&gt;
&lt;td&gt;Understand the essence of AI's "probability and language architecture," find breakthroughs outside of US monopolies, and raise the lower limit of basic capabilities.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;2. Taiwan's Positioning&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Although Taiwan is a key X-factor in geopolitics and AI chips, it also faces structural constraints like island involution, a declining birthrate, and the "five shortages."&lt;/td&gt;
&lt;td&gt;Completely shift from a "farmer's mindset" to a "navigator's mindset," taking "going global" as the only way to survive, and stepping beyond Taiwan's borders to expand digital territory.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;3. Corporate Transformation&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Taiwan is "extremely strong in hardware, extremely weak in software." Software startups lacking computing power and business scenarios find it hard to survive independently on the international stage.&lt;/td&gt;
&lt;td&gt;Promote "Old-ventures (hardware giants) + New-ventures (software applications)" collaboration, using Edge AI to add "brains" to powerful hardware devices.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;4. Personal Development Keys&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Facing the invisible crisis of "brain outsourcing," mediocre professional newcomers who only teach by the book and lack practical experience will be the first to be hit.&lt;/td&gt;
&lt;td&gt;Shift from a "problem-solving" habit to a "problem-posing" mindset, creating unique value through high-frequency "repeated interaction and correction" with AI.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  I. Latest AI Trends: The 1:99 "Superhuman" Challenge
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Extreme Concentration of Power&lt;/strong&gt;: The AI era has brought high centralization, with global tech giants holding a massive advantage. Among thousands of languages worldwide, only about a hundred can be used in mainstream AI, and English and Simplified Chinese are deeply optimized. This means language and cultural frameworks are the primary keys to mastering AI.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The 1:99 Watershed&lt;/strong&gt;: The cruelest part of this tsunami is not the elimination of ordinary people at the bottom (AI actually raises the floor for ordinary people), but the elimination of "mediocre professionals." The 1% who seize the opportunity will become superhumans through AI, taking the capabilities and opportunities of the 99%, while others become the lagging 99%.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rebalancing the AI World&lt;/strong&gt;: The recent emergence of non-US low-cost, high-efficiency models has broken the absolute monopoly of US tech giants. This has been described as inventing the "poor man's atomic bomb," bringing an opportunity for a reshuffle to countries and enterprises with fewer resources.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  II. Taiwan's Situation: From "Island Involution" to the "Age of Discovery"
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The Geopolitical X-Factor&lt;/strong&gt;: TSMC and Taiwan's hardware supply chain hold a key position in the US-China tech confrontation because Taiwan possesses the characteristic of "knowing the demand earliest" (e.g., being able to grasp system requirements like server voltage changes first), giving it an important identity in the adjustment of global infrastructure.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Breaking the Farmer's Mindset&lt;/strong&gt;: Most Taiwanese companies are used to an "island mindset," where the "sea is invisible" in daily life, making it easy to fall into involution within a comfortable echo chamber. Facing the structural crisis of a declining birthrate and plunging newborn numbers over the next 20 years, Dr. Chien urgently calls for a shift to a "navigator's mindset," because "going global" is the only way for all industries in Taiwan to survive.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Extension of Digital Territory&lt;/strong&gt;: Taking TSMC as an example, AI allows Taiwan to replicate factories overseas and have them operated remotely by Taiwanese engineers. Taiwan should also turn the crises of aging and labor shortages into opportunities by actively developing robotics and its own Sovereign AI to avoid a national-level digital divide.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  III. Corporate Layout: Hardware-Software Integration, Letting "Old-ventures + New-ventures" Dance Together
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Using Edge AI to Add Brains to Hardware&lt;/strong&gt;: Edge AI (referring to terminal devices having local computing power without relying entirely on the cloud) is Taiwan's domain. It is hard for pure software startups in Taiwan to compete with international giants, but we can embed and bundle AI services directly into powerful hardware devices used worldwide (such as Giant bicycles or various terminal equipment), significantly increasing added value.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Old-ventures plus New-ventures Playing the International Game&lt;/strong&gt;: Today's AI startups can hardly succeed without the data, computing power, and real "business scenarios" provided by a "rich father." Therefore, hardware giants (Old-ventures) should join hands with software startups, combining the international channels of the former with the flexible applications of the latter to go global as a team.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pragmatic Planning for Sovereign AI&lt;/strong&gt;: Developing Sovereign AI cannot just be about blindly outsourcing business to tech giants. Whether building their own models or cooperating with major manufacturers, enterprises must clearly plan their own strategies and field applications; otherwise, it is just a waste of money.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  IV. Personal Development: Refuse "Brain Outsourcing," Be a High-Level "Questioner"
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Shift Thinking from "Solving" to "Posing"&lt;/strong&gt;: AI's capabilities are "asked" out. The more professional the question, the more accurate the response. The future workplace will no longer value rote memorization; core competencies will shift to problem definition, critical thinking, and direction control. Only "questioners" who can demonstrate proactive influence will win.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Using "Iterative Refinement" to Deepen Learning&lt;/strong&gt;: If you just throw a question to AI, get an answer once, and copy it directly, this behavior is equivalent to plagiarism. However, if you can go back and forth with AI to modify it 10 times, that is a "learning" process; if you continue to repeatedly correct and adjust up to 100 times, that is truly approaching the level of "creation."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Accumulating Irreplaceable "Differentiated Experience"&lt;/strong&gt;: Functions like memory and calculation can be outsourced to AI, but your unique personal experience, cross-disciplinary collaboration ability (π-shaped talent), and human critical thinking are the irreplaceable treasures of the AI era. Cleverly using AI tools to amplify your output is the only way to avoid becoming the "lost generation" eliminated by the times.&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The Core Soul Quote of the Book:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;"Change is humanity's eternal unease, but from a macro perspective, what AI brings is an opportunity to make humans more capable." When calculation and memory are outsourced from the brain, be sure to retain your power of thinking and creation.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This video &lt;a href="https://www.youtube.com/watch?v=XHTfeCk0GwQ" rel="noopener noreferrer"&gt;Interview with Dr. Chien Lee-feng: Who is the Lost Generation of the AI Era&lt;/a&gt; deeply explores the "1:99 Superhuman Challenge" and workplace mindset transformation mentioned in the book, helping you more intuitively understand how to retain personal competitiveness in this era of brain outsourcing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Personal Thoughts
&lt;/h2&gt;

&lt;p&gt;From my own perspective, this book organizes many recent AI development processes both domestically and internationally. It provides many future insights based on Chien Lee-feng's own experience as the former Managing Director of Google Taiwan. It often shares advice on how various industries should face the AI era. This part is frequently mentioned in his speeches and is shared and explained quite clearly.&lt;/p&gt;

&lt;p&gt;Personally, I feel that one can just skim through this book. In comparison, I still prefer Dr. Chien Lee-feng's speeches, which provide more impact and inspiration.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>books</category>
      <category>career</category>
      <category>learning</category>
    </item>
    <item>
      <title>[GCP in Action] LINE Business Card Bot Evolution: Dual-Side Recognition and Merging with Gemini</title>
      <dc:creator>Evan Lin</dc:creator>
      <pubDate>Thu, 23 Jul 2026 13:24:50 +0000</pubDate>
      <link>https://dev.to/gde/gcp-in-action-line-business-card-bot-evolution-dual-side-recognition-and-merging-with-gemini-2220</link>
      <guid>https://dev.to/gde/gcp-in-action-line-business-card-bot-evolution-dual-side-recognition-and-merging-with-gemini-2220</guid>
      <description>&lt;h1&gt;
  
  
  Pain Point: One Chinese Business Card is Actually Two Business Cards
&lt;/h1&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcej08mt7097plgtj10tu.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcej08mt7097plgtj10tu.png" alt="image-20260723152322389" width="800" height="682"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Business cards in Taiwan often have a common design: Chinese printed on the front and English on the back (or vice versa). Our LINE business card bot's original logic was very simple: receive an image, perform OCR once, and save one record.&lt;/p&gt;

&lt;p&gt;This is where the problem lies. When a user sends the front side, the bot saves a record with only the Chinese name. If the user then sends the back side, the bot treats it as "another new business card," resulting in two records for the same person in the database, each missing half the information. Users have to manually compare and delete duplicate data, which is a terrible experience.&lt;/p&gt;

&lt;p&gt;This article records how we taught the bot to recognize that "these are two sides of the same business card" and merge the information from both sides into a single complete record.&lt;/p&gt;




&lt;h1&gt;
  
  
  Solution: First Ask, "Is there a back side?"
&lt;/h1&gt;

&lt;p&gt;Instead of writing rules to guess if two images are the same business card, we chose a more direct approach: &lt;strong&gt;Ask the user&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The workflow is designed as follows:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The user sends the front of the business card, and the bot performs OCR as usual.&lt;/li&gt;
&lt;li&gt;After OCR is complete, the bot &lt;strong&gt;does not save immediately&lt;/strong&gt;. Instead, it replies with "📇 Front side data recognized. Is there a back side to this card?" and provides two Quick Reply buttons.&lt;/li&gt;
&lt;li&gt;User clicks "No, save directly" → Save according to the original process and finish.&lt;/li&gt;
&lt;li&gt;User clicks "Yes, there's a back side" → The bot remembers the front image and waits for the next image.&lt;/li&gt;
&lt;li&gt;Once the back side image arrives, both images are sent to Gemini together to be merged into a single record before saving.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;We manage this waiting state using the &lt;code&gt;user_states&lt;/code&gt; memory dictionary already present in the project, and add a 5-minute timeout. If a user clicks "Yes, there's a back side" but then ignores it or does something else, the process is treated as abandoned after 5 minutes, preventing the entire workflow from getting stuck.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;user_states&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;action&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;pending_backside_confirm&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;card_obj&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;card_obj&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;front_image_bytes&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;image_content&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;expires_at&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;time&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;PENDING_BACKSIDE_TIMEOUT_SECONDS&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  Core: Let Gemini See Two Images at Once and Merge Them
&lt;/h1&gt;

&lt;p&gt;The most critical technical decision was: should we recognize the front and back sides separately and then write code to merge them? We chose another path: wrapping both the front and back images in a single &lt;code&gt;generate_content&lt;/code&gt; request and letting Gemini handle the judgment directly.&lt;/p&gt;

&lt;p&gt;The reason is simple: merging Chinese and English names into a format like "Wang Daming David Wang" using string rules is prone to being messy and inaccurate. Semantic-level integration is more likely to fail with hardcoded rules, so letting Gemini handle it directly is much easier.&lt;/p&gt;

&lt;p&gt;In &lt;a&gt;app/gemini_utils.py&lt;/a&gt;, we added &lt;code&gt;generate_json_from_two_images&lt;/code&gt;, which reuses the existing &lt;code&gt;NAMECARD_SCHEMA&lt;/code&gt; structured output, but this time the &lt;code&gt;contents&lt;/code&gt; includes two image Parts:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;generate_json_from_two_images&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;front_img&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;PIL&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Image&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Image&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;back_img&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;PIL&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Image&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Image&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;object&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;GenerativeModel&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gemini-3-flash-preview&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;generation_config&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;response_mime_type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;application/json&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;response_schema&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;NAMECARD_SCHEMA&lt;/span&gt;
        &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;front_part&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Part&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;from_data&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nf"&gt;pil_to_bytes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;front_img&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;mime_type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;image/jpeg&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;back_part&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Part&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;from_data&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nf"&gt;pil_to_bytes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;back_img&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;mime_type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;image/jpeg&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;generate_content&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;front_part&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;back_part&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="n"&gt;stream&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;labels&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;client_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;namecard&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The prompt also just adds a merge instruction after the original &lt;code&gt;IMGAGE_PROMPT&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;DOUBLE_SIDED_IMAGE_PROMPT&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;IMGAGE_PROMPT&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;
These two images are the front and back of the same business card; please integrate them into a single complete record.
If both Chinese and English appear in the same field (such as name or company), please present them merged
(e.g., &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Wang Daming David Wang&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;);
if a field appears on only one side, use the value from that side; ignore obviously redundant information.
&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One API call means we don't have to write or maintain any merging rules ourselves.&lt;/p&gt;




&lt;h1&gt;
  
  
  Two Easily Overlooked Pitfalls
&lt;/h1&gt;

&lt;p&gt;During the overall code review before the feature went live, we caught two details that are easily overlooked but can really cause issues.&lt;/p&gt;

&lt;h3&gt;
  
  
  Pitfall 1: Timing of Duplicate Checks
&lt;/h3&gt;

&lt;p&gt;Originally, the duplicate check (comparing if the email already exists) was done immediately after OCR. However, after the double-sided recognition went live, if the front side happened to have a duplicate email from an old record, the process would prematurely determine it "already exists" and end. This would mean any new email on the back side would never be seen.&lt;/p&gt;

&lt;p&gt;The fix was to move the duplicate check later, performing it only after "single-side chosen not to merge" or "double-sided merge complete." This ensures we are always comparing the final version of the data:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;_finalize_and_save_card&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;card_obj&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;existing_card_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;firebase_utils&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;check_if_card_exists&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;card_obj&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;existing_card_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="c1"&gt;# ... Reply already exists
&lt;/span&gt;        &lt;span class="k"&gt;return&lt;/span&gt;
    &lt;span class="n"&gt;card_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;firebase_utils&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add_namecard&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;card_obj&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="c1"&gt;# ... Reply save successful
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both the single-sided and double-sided merge workflows eventually converge to call this shared function, ensuring the duplicate check is executed only once when the data is finalized.&lt;/p&gt;

&lt;h3&gt;
  
  
  Pitfall 2: Don't Clear All States Indiscriminately
&lt;/h3&gt;

&lt;p&gt;The &lt;code&gt;user_states&lt;/code&gt; dictionary is actually shared by several features: editing memos, modifying fields, and this back-side recognition. The initial implementation, for convenience, would delete the entire state whenever a residual state was detected before processing a new event.&lt;/p&gt;

&lt;p&gt;The problem is: if a user is "editing the phone field" and waiting to input a new number, but accidentally sends an image, this logic would clear the &lt;code&gt;editing_field&lt;/code&gt; state as well, silently canceling the user's original editing operation.&lt;/p&gt;

&lt;p&gt;The fix was to only clear the two states related to the back-side recognition process and leave other states untouched:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;action&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;pending_backside_confirm&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;awaiting_backside_image&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;
&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;del&lt;/span&gt; &lt;span class="n"&gt;user_states&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;We later found a missing branch in the same logic; the part handling "operation expired" replies also cleared everything initially. It was only truly fixed after unifying it with the same selective judgment.&lt;/p&gt;




&lt;h1&gt;
  
  
  Incidental Resource Cleanup: Don't Let Back-Side Images Linger in Memory
&lt;/h1&gt;

&lt;p&gt;The &lt;code&gt;awaiting_backside_image&lt;/code&gt; state stores not just text, but also the raw byte data of the front image. If a user disappears after being asked "Is there a back side?", this data would theoretically stay in the process memory because the original design only checked and cleared timeout states during the "user's next interaction."&lt;/p&gt;

&lt;p&gt;We added a &lt;code&gt;sweep_expired_states()&lt;/code&gt; function, which runs immediately when a Webhook comes in. It clears all expired temporary states for all users, so we don't have to wait for the specific user to return for passive cleanup:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;sweep_expired_states&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;now&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;time&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;expired_user_ids&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="n"&gt;user_id&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;state&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;user_states&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;items&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;expires_at&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;state&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;expires_at&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="n"&gt;now&lt;/span&gt;
    &lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;user_id&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;expired_user_ids&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;del&lt;/span&gt; &lt;span class="n"&gt;user_states&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Whenever any user sends a message, it performs garbage collection for all users, ensuring that those who abandoned the process don't leave behind memory-consuming remnants.&lt;/p&gt;




&lt;h1&gt;
  
  
  Summary and Benefits
&lt;/h1&gt;

&lt;p&gt;This double-sided recognition and merging feature makes the LINE business card bot much more aligned with the actual usage habits of Taiwanese users:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Single Recognition, Complete Data&lt;/strong&gt;: Both sides are sent to Gemini at once, automatically merging Chinese and English fields, eliminating the need for manual duplicate comparison.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Non-Intrusive&lt;/strong&gt;: Ignoring prompts, timeouts, or temporarily doing something else will naturally revert to single-sided storage without getting the workflow stuck.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Duplicate Checks Target Final Data&lt;/strong&gt;: Ensures that comparisons are always made against the merged, complete version, so new information appearing only on the back side isn't missed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Independent State Machines&lt;/strong&gt;: The temporary state for back-side recognition only affects itself and doesn't interfere with other ongoing user operations.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Clean Memory Usage&lt;/strong&gt;: Proactively cleaning up timeout states ensures that users who abandon the process don't leave an invisible memory burden.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The complete code has been pushed to &lt;a href="https://github.com/kkdai/linebot-namecard-python" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;, feel free to check it out!&lt;/p&gt;

</description>
      <category>ai</category>
      <category>cloud</category>
      <category>gemini</category>
      <category>google</category>
    </item>
    <item>
      <title>[Digital Certificate Wallet] Advanced: Building "Visitor-Endorsed Issuance" – A Full-Chain DID Application as Both Verifier and</title>
      <dc:creator>Evan Lin</dc:creator>
      <pubDate>Mon, 13 Jul 2026 08:01:01 +0000</pubDate>
      <link>https://dev.to/evanlin/digital-certificate-wallet-advanced-building-visitor-endorsed-issuance-a-full-chain-did-4hlc</link>
      <guid>https://dev.to/evanlin/digital-certificate-wallet-advanced-building-visitor-endorsed-issuance-a-full-chain-did-4hlc</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fzt11xi6eefeljlva2rz6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fzt11xi6eefeljlva2rz6.png" alt="image-20251009102618401" width="799" height="392"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;(Source: &lt;a href="https://www.wallet.gov.tw/zh-tw" rel="noopener noreferrer"&gt;Digital Certificate Wallet Official Website&lt;/a&gt;)&lt;/p&gt;

&lt;h2&gt;
  
  
  Premise:
&lt;/h2&gt;

&lt;p&gt;The previous &lt;a href="https://github.com/kkdai/did-usecase-HR" rel="noopener noreferrer"&gt;introductory version&lt;/a&gt; created an HR employee card system: colleagues apply for an employee card themselves and then use it to apply for "sports subsidies" and "childcare subsidies." The focus of that article was on separating the two roles: "&lt;strong&gt;Issuer&lt;/strong&gt;" and "&lt;strong&gt;Verifier&lt;/strong&gt;."&lt;/p&gt;

&lt;p&gt;In this article, I want to take it a step further: what does it look like if a scenario needs to &lt;strong&gt;play both the verifier and issuer roles simultaneously&lt;/strong&gt; to form a complete DID ecosystem chain? The scenario I chose is "&lt;strong&gt;Visitor Endorsement and Issuance&lt;/strong&gt;"—this is also the one I felt best demonstrates the "full chain" when brainstorming five verifier applications.&lt;/p&gt;

&lt;p&gt;By the way, this article will honestly document &lt;strong&gt;three pitfalls encountered&lt;/strong&gt; during the development process, as those are the truly valuable parts of TIL (Today I Learned).&lt;/p&gt;

&lt;p&gt;Code is here: &lt;a href="https://github.com/kkdai/did-usecase-visitor" rel="noopener noreferrer"&gt;https://github.com/kkdai/did-usecase-visitor&lt;/a&gt;&lt;br&gt;
Online experience: &lt;a href="https://did-usecase-visitor-660825558664.asia-east1.run.app" rel="noopener noreferrer"&gt;https://did-usecase-visitor-660825558664.asia-east1.run.app&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  Scenario: Employee Access Control + Visitor Endorsement and Issuance
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fngkv6rlywiutybkd69lo.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fngkv6rlywiutybkd69lo.png" alt="image-20260709172435871" width="800" height="1734"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This lobby reception desk has two modes:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt; &lt;strong&gt;Employee Access Control / Event Registration&lt;/strong&gt;: Employees present their employee cards using a digital wallet. The system only verifies "&lt;strong&gt;whether they are a valid employee&lt;/strong&gt;." Once verified, the door opens or registration is successful. Fields like name, birthday, and number of children are not disclosed and remain in the wallet—this is selective disclosure.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Visitor Endorsement and Issuance&lt;/strong&gt; (the protagonist of this article): An active employee presents their employee card to "endorse" the visitor. After verification, the system &lt;strong&gt;immediately issues a temporary visitor pass with an expiration time&lt;/strong&gt; to the visitor's wallet.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The value of the second mode lies in connecting the two roles:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;First act as a &lt;strong&gt;Verifier&lt;/strong&gt; (verify employee card) → then act as an &lt;strong&gt;Issuer&lt;/strong&gt; (issue visitor card) only after verification.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Compared to traditional paper visitor logs (copying IDs, holding physical IDs, piles of personal data at the counter requiring manual disposal), digital endorsement only leaves one piece of accountable information: "which employee endorsed it." Visitor data stays in the visitor's own wallet, and the pass can have an expiration time.&lt;/p&gt;
&lt;h2&gt;
  
  
  Architecture Decisions: Why not just modify the previous project?
&lt;/h2&gt;

&lt;p&gt;This time, I started a brand new project and deployed it to a separate Cloud Run service instead of adding pages to the original HR project. Several considerations:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Static Frontend + JSON API&lt;/strong&gt;: The original project used jade templates for server-side rendering. This time, it was changed to &lt;code&gt;public/&lt;/code&gt; static pages + a few JSON APIs (&lt;code&gt;/api/access/qrcode&lt;/code&gt;, &lt;code&gt;/api/access/status&lt;/code&gt;), separating the frontend and backend more cleanly.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Abstracting Wallet API calls into &lt;code&gt;lib/wallet.js&lt;/code&gt;&lt;/strong&gt;: In the original project, issuer/verifier calls were embedded in routes and duplicated. This time, they were extracted into three functions: &lt;code&gt;requestPresentationQRCode()&lt;/code&gt;, &lt;code&gt;getPresentationResult()&lt;/code&gt;, and &lt;code&gt;issueCredential()&lt;/code&gt;, making it easier to maintain and test.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;State changed to Memory&lt;/strong&gt;: The original project wrote data to a single &lt;code&gt;record.js&lt;/code&gt; file. Writing files in a stateless environment like Cloud Run causes issues. This time, a simple memory object was used (resets on restart, sufficient for demonstration purposes).&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Reusing the same Sandbox Account tokens&lt;/strong&gt;: The access tokens for the issuer/verifier are the same set as in the previous article, reused directly.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For the "holding an employee card" verification part, I first used the existing sports subsidy verifier ref as a fallback (if &lt;code&gt;VERIFIER_ACCESS_REF&lt;/code&gt; is not set, use &lt;code&gt;VERIFIER_SPORT_REF&lt;/code&gt;), so it could run without waiting for backend configuration.&lt;/p&gt;
&lt;h2&gt;
  
  
  Pitfall Record 1: Presentation successful, but the screen is stuck
&lt;/h2&gt;

&lt;p&gt;This is a classic one. The phone scans the code, and the wallet completes the presentation, but the desktop page just won't move forward; it keeps polling.&lt;/p&gt;

&lt;p&gt;First, checking the Cloud Run logs, I found that &lt;code&gt;/api/access/status&lt;/code&gt; returns every 3 seconds, each time returning "unverified." I added a log line on the backend to print the &lt;strong&gt;original response&lt;/strong&gt; from the verifier. After redeploying and testing again, I caught the truth:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"data"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"credentialType"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"0028680530_line_school"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"claims"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"ename"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"ename"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"cname"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"English Name"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"value"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Lub"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"ename"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"join_company"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"cname"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Join Date"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"value"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2018-10-05"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"verifyResult"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"resultDescription"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"success"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"transactionId"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"8cd7f37b-..."&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;See the problem? The field in the response is &lt;strong&gt;&lt;code&gt;verifyResult&lt;/code&gt; (camelCase)&lt;/strong&gt;, and &lt;strong&gt;there is no &lt;code&gt;code&lt;/code&gt; field at all&lt;/strong&gt;. But I used the old logic from the previous article:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Old (doesn't match current response)&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;verified&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;code&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;verify_result&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;data.code&lt;/code&gt; is &lt;code&gt;undefined&lt;/code&gt;, and &lt;code&gt;data.verify_result&lt;/code&gt; is also &lt;code&gt;undefined&lt;/code&gt; (it's called &lt;code&gt;verifyResult&lt;/code&gt;), so it's always &lt;code&gt;false&lt;/code&gt;, always pending. &lt;strong&gt;Actually, the verification had already succeeded&lt;/strong&gt; (&lt;code&gt;verifyResult: true&lt;/code&gt;, &lt;code&gt;resultDescription: "success"&lt;/code&gt;), but the field names I was checking didn't match—it seems the sandbox API response format has changed from snake_case to camelCase.&lt;/p&gt;

&lt;p&gt;The fix was to change the logic to be compatible with both formats:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;verified&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt;
  &lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;verifyResult&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="c1"&gt;// New format camelCase&lt;/span&gt;
  &lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;verify_result&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="c1"&gt;// Old format compatibility&lt;/span&gt;
  &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;code&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;verify_result&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TIL&lt;/strong&gt;: When integrating third-party APIs, don't trust that "the logic that worked in the last version will work in this one." Sandboxes change. Adding a log line to print the raw response and comparing it is much faster than staring at the code and guessing.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Pitfall Record 2: Visitor card stuck at "Pending Issuance"
&lt;/h2&gt;

&lt;p&gt;After the access control part worked, the visitor endorsement part got stuck—the screen showed "Visitor card pending issuance (issuer template not set)," and no card was actually issued.&lt;/p&gt;

&lt;p&gt;I used &lt;code&gt;curl&lt;/code&gt; to hit the issuance API &lt;code&gt;/api/vc-item-data&lt;/code&gt; directly to see what it returned. I tested two scenarios:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Scenario A: Using employee template + correct employee fields&lt;/strong&gt; → HTTP 200, and the full response contained these keys:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;KEYS: ['id', 'content', 'pureContent', ..., 'qrCode', 'deepLink', 'expired', ...]
qrCode = data:image/png;base64,iVBOR... ← QR code that can actually be scanned into the wallet
deepLink = https://frontend-uat.wallet.gov.tw/api/moda/vcqrcode?...
expired = 2027-01-09T...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Scenario B: Using employee template + visitor fields&lt;/strong&gt; (&lt;code&gt;visitor_type&lt;/code&gt;, &lt;code&gt;endorsed_by&lt;/code&gt;…) → HTTP 500 / 400 BAD_REQUEST.&lt;/p&gt;

&lt;p&gt;The reason was clear: the employee template fields were &lt;code&gt;isRequired: true&lt;/code&gt; (Name, English Name…), but I sent a bunch of visitor fields it didn't have, so it was rejected. And the &lt;strong&gt;successful issuance response actually includes &lt;code&gt;qrCode&lt;/code&gt; and &lt;code&gt;deepLink&lt;/code&gt;&lt;/strong&gt;, which can be used directly for the visitor to scan and collect the card—my original parsing was correct; the bottleneck was purely "fields not matching the template."&lt;/p&gt;

&lt;p&gt;So I designed two issuance modes, automatically switched by environment variables (&lt;code&gt;HAS_VISITOR_TEMPLATE&lt;/code&gt; in the code):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Mode&lt;/th&gt;
&lt;th&gt;Condition&lt;/th&gt;
&lt;th&gt;Behavior&lt;/th&gt;
&lt;th&gt;Card Face&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Option 1 (fallback)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;VISITOR_VC_*&lt;/code&gt; not set&lt;/td&gt;
&lt;td&gt;Borrow the employee template, stuffing visitor info into its required fields (Name="Temporary Visitor", etc.) to issue the card&lt;/td&gt;
&lt;td&gt;Displays as an employee card face&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Option 2 (Formal)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;VISITOR_VC_*&lt;/code&gt; set&lt;/td&gt;
&lt;td&gt;Send &lt;code&gt;visitor_type / endorsed_by / valid_until&lt;/code&gt; to the dedicated visitor template&lt;/td&gt;
&lt;td&gt;Formal visitor pass card face&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The advantage of Option 1 is that &lt;strong&gt;a real, collectable card can be issued without waiting for backend configuration&lt;/strong&gt; (even if the card face is borrowed), allowing the entire chain to be tested first; for a formal card face, just go with Option 2 and build a dedicated template, with no code changes required.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pitfall Record 3: Collection QR too small + Mobile layout
&lt;/h2&gt;

&lt;p&gt;In the first version, I made the visitor pass look like a pretty little ID badge, and the collection QR was only 48px—the result was that &lt;strong&gt;it couldn't be scanned at all&lt;/strong&gt;. This QR is meant to be scanned by "another phone" to collect the card; if it's too small, it loses its purpose.&lt;/p&gt;

&lt;p&gt;Later, I changed the visitor card to a vertical layout, enlarging the collection QR to be the main body of the card (max 240px, white background with padding), with "Endorsed by / Valid until" information placed below. Both QRs (for presentation and collection) were also changed to &lt;code&gt;clamp()&lt;/code&gt; responsive sizes, so they don't break the layout on mobile and are clear enough on desktop.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TIL&lt;/strong&gt;: As long as a QR is "for others to scan," it must be treated as the protagonist of the layout, not as a decoration.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The Truth About "Automatic Expiration"
&lt;/h2&gt;

&lt;p&gt;I originally thought I could specify "this visitor pass expires in 4 hours" for each card, but testing revealed that when issuing cards via &lt;code&gt;/api/vc-item-data&lt;/code&gt;, the actual validity of the card &lt;strong&gt;follows the template settings&lt;/strong&gt; (e.g., the employee template is issuance date + about half a year); it's not possible to specify a short expiration for individual cards.&lt;/p&gt;

&lt;p&gt;So the "Valid until HH:MM" on the card face now is a &lt;strong&gt;display value calculated by the application layer&lt;/strong&gt;, not a mandatory expiration enforced by the wallet. If a truly short-term visitor pass is needed, there are two ways:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  When creating the visitor template, &lt;strong&gt;set the template's validity period to be short&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;  Or use the platform's &lt;strong&gt;scheduled revocation (revoke)&lt;/strong&gt;—the issuance response includes fields like &lt;code&gt;clearScheduleId&lt;/code&gt; and &lt;code&gt;scheduleRevokeMessage&lt;/code&gt;, implying the platform supports scheduled revocation, but this requires integrating the corresponding API separately.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Deployment: Directly to Cloud Run from Source Code
&lt;/h2&gt;

&lt;p&gt;This time, I used buildpacks to deploy directly from source code without writing a Dockerfile:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;gcloud run deploy did-usecase-visitor &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--source&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;.&lt;/span&gt; &lt;span class="nt"&gt;--region&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;asia-east1 &lt;span class="nt"&gt;--platform&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;managed &lt;span class="nt"&gt;--allow-unauthenticated&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--set-env-vars&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"VC_SERNUM=607861,VC_UID=0028680530_line_school,&lt;/span&gt;&lt;span class="se"&gt;\&lt;/span&gt;&lt;span class="s2"&gt;
ISSUER_ACCESS_TOKEN=...,VERIFIER_SPORT_REF=...,VERIFIER_ACCESS_TOKEN=...,&lt;/span&gt;&lt;span class="se"&gt;\&lt;/span&gt;&lt;span class="s2"&gt;
VISITOR_TTL_HOURS=4"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;To switch to the formal visitor card in Option 2 later, just add &lt;code&gt;VISITOR_VC_SERNUM=&amp;lt;new template vcId&amp;gt;,VISITOR_VC_UID=&amp;lt;new template vcCid&amp;gt;&lt;/code&gt; to this &lt;code&gt;--set-env-vars&lt;/code&gt; string and redeploy; &lt;code&gt;HAS_VISITOR_TEMPLATE&lt;/code&gt; will automatically become true.&lt;/p&gt;

&lt;h2&gt;
  
  
  Summary and Future Outlook
&lt;/h2&gt;

&lt;p&gt;The focus this time wasn't "making another demo," but three things:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt; &lt;strong&gt;The full DID chain is feasible&lt;/strong&gt;: Playing both Verifier and Issuer in the same scenario—verifying one card and then issuing another—connects the ecosystem chain, and the experience is smooth.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Pitfalls are in the details&lt;/strong&gt;: Field naming (&lt;code&gt;verifyResult&lt;/code&gt; vs &lt;code&gt;verify_result&lt;/code&gt;), mandatory template fields, QR size—these wouldn't be discovered without looking at the raw response and actually scanning with a phone.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Fallback design allows the demo to work first&lt;/strong&gt;: No need to wait for every template/ref to be built on the backend; use existing resources to get the whole chain running first, then gradually switch to formal settings. The development rhythm is much better.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;There are truly many application scenarios for digital certificate wallets; "visitor endorsement" is just one of them. The employee card from the previous article could also be extended to commissary discount redemption, seniority milestone gifts, childcare facility access, gym point accumulation... each is a new application for a "verifier." I look forward to seeing more creative scenarios being built.&lt;/p&gt;

</description>
      <category>architecture</category>
      <category>blockchain</category>
      <category>tutorial</category>
      <category>web3</category>
    </item>
    <item>
      <title>[GCP Billing &amp; Vertex AI] Solving Gemini Cost Allocation in a Single Project: Vertex AI Dynamic Billing Labels in Action</title>
      <dc:creator>Evan Lin</dc:creator>
      <pubDate>Sun, 12 Jul 2026 15:09:09 +0000</pubDate>
      <link>https://dev.to/gde/gcp-billing-vertex-ai-solving-gemini-cost-allocation-in-a-single-project-vertex-ai-dynamic-5aok</link>
      <guid>https://dev.to/gde/gcp-billing-vertex-ai-solving-gemini-cost-allocation-in-a-single-project-vertex-ai-dynamic-5aok</guid>
      <description>&lt;h1&gt;
  
  
  Pain Point: How to Accurately Allocate Gemini API Costs Within the Same Project?
&lt;/h1&gt;

&lt;p&gt;When developing enterprise-level LLM services or operating multi-tenant platforms, the question most frequently asked by finance and operations teams is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"We have many different business lines and LINE Bots connected within the same GCP project. Every day, the Gemini Key costs all appear under the Gemini API category. Is there a way for us to split the costs based on different Gemini Keys or different users?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;Direct answer to your question:&lt;/strong&gt; In Google Cloud Billing reports, it is &lt;strong&gt;not possible to directly display costs "based on different API Key names."&lt;/strong&gt; The smallest attribution dimensions for Google Cloud billing reports are "Project," "Service," and "SKU (Product Line Item)." The system does not treat individual API Key strings as independent billing items. For the billing system, whether you create 10 or 100 API Keys within the same project, they will all be lumped together as a single Gemini API total.&lt;/p&gt;




&lt;h1&gt;
  
  
  A Workaround: Vertex AI "Request Labels" to the Rescue
&lt;/h1&gt;

&lt;p&gt;If architectural constraints force you to stay within the same project, the most recommended approach is: &lt;strong&gt;switch to Vertex AI calls and use "Request Labels."&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If you are currently using a Google AI Studio API Key, it cannot pass billing labels within a single project. However, if you change your code to call the &lt;strong&gt;Vertex AI Gemini API&lt;/strong&gt; (still within the same project), Vertex AI supports dynamically including custom &lt;code&gt;labels&lt;/code&gt; with each request.&lt;/p&gt;

&lt;h3&gt;
  
  
  Principle and Workflow
&lt;/h3&gt;

&lt;p&gt;When sending each request (e.g., calling &lt;code&gt;generateContent&lt;/code&gt;), include specific metadata in the API Request:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"contents"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;...&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"labels"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"client_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"info_helper"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"api_key_group"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"marketing_team"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;These custom labels are passed directly to the GCP billing system. Later, when you go to the GCP Billing report and select your set label key (e.g., &lt;code&gt;client_id&lt;/code&gt;) in "Group by," you can clearly see the costs for different labels (representing different services, clients, or users) within the same project!&lt;/p&gt;




&lt;h1&gt;
  
  
  Project Implementation: Full Adoption of the Labels Mechanism
&lt;/h1&gt;

&lt;p&gt;To fulfill this requirement, we audited the current API call architecture of the LINE Bot project and performed the following refactoring.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Project API Call Audit
&lt;/h3&gt;

&lt;p&gt;Through scanning, we found that the vast majority of calls in the project use Vertex AI (14 out of 17 clients use &lt;code&gt;vertexai=True&lt;/code&gt;), with only a few exceptions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Vertex AI calls&lt;/strong&gt;: Including GitHub summaries, multiple Google Maps Grounding tools, text summarization, image analysis, speech-to-text, etc. (total of 11 files, 19 call points).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gemini API Key calls&lt;/strong&gt;: Live API in &lt;code&gt;main.py&lt;/code&gt;, Batch service in &lt;code&gt;batch_service.py&lt;/code&gt;, and TTS speech synthesis in &lt;code&gt;tts_tool.py&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;[!IMPORTANT] The &lt;code&gt;labels&lt;/code&gt; parameter is only supported by Vertex AI. If this parameter is included under an API Key (&lt;code&gt;vertexai=False&lt;/code&gt;), it will cause the SDK to throw an error. Therefore, we only modified the 11 files that use Vertex AI.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  2. Implementation Method
&lt;/h3&gt;

&lt;p&gt;For the &lt;code&gt;google-genai&lt;/code&gt; Python SDK, we have two main modification scenarios:&lt;/p&gt;

&lt;h4&gt;
  
  
  Scenario A: Already contains &lt;code&gt;GenerateContentConfig&lt;/code&gt;
&lt;/h4&gt;

&lt;p&gt;If the original call already includes a Config, we just need to pass an additional &lt;code&gt;labels={"client_id": "info_helper"}&lt;/code&gt; into the config:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Before (e.g., loader/gh_tools.py)
&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;models&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;generate_content&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gemini-2.5-flash&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;contents&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;types&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;GenerateContentConfig&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;temperature&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;max_output_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;2048&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# After
&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;models&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;generate_content&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gemini-2.5-flash&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;contents&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;types&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;GenerateContentConfig&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;temperature&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;max_output_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;2048&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;labels&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;client_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;info_helper&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="c1"&gt;# Include billing label
&lt;/span&gt;    &lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Scenario B: No Config parameter
&lt;/h4&gt;

&lt;p&gt;If the original call is very simple (e.g., &lt;code&gt;searchtool.py&lt;/code&gt; or &lt;code&gt;youtube_gcp.py&lt;/code&gt;), we need to proactively include a &lt;code&gt;GenerateContentConfig&lt;/code&gt; containing labels:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Before (e.g., loader/searchtool.py)
&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;models&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;generate_content&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gemini-3.1-flash-lite-preview&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;contents&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# After
&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;models&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;generate_content&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gemini-3.1-flash-lite-preview&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;contents&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;types&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;GenerateContentConfig&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;labels&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;client_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;info_helper&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="c1"&gt;# Add config to include label
&lt;/span&gt;    &lt;span class="p"&gt;),&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  3. List of Modified Files
&lt;/h3&gt;

&lt;p&gt;We performed precise modifications on a total of 19 call points across the following 11 files, and used Python's AST module (&lt;code&gt;ast.parse&lt;/code&gt;) and Flake8 for syntax and formatting checks before submission:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a&gt;agents/chat_agent.py&lt;/a&gt;&lt;/strong&gt;: Modify &lt;code&gt;_create_chat_config()&lt;/code&gt; to add labels to both general Q&amp;amp;A and Grounding conversations.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a&gt;loader/chat_session.py&lt;/a&gt;&lt;/strong&gt;: Include labels in Chat session config.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a&gt;loader/gh_tools.py&lt;/a&gt;&lt;/strong&gt;: GitHub summary API.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a&gt;loader/langtools.py&lt;/a&gt;&lt;/strong&gt;: Text summarization, image JSON generation, social media post generation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a&gt;loader/maps_grounding.py&lt;/a&gt;&lt;/strong&gt;: Maps search API.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a&gt;loader/searchtool.py&lt;/a&gt;&lt;/strong&gt;: Keyword extraction tool.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a&gt;loader/youtube_gcp.py&lt;/a&gt;&lt;/strong&gt;: YouTube video understanding API.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a&gt;tools/audio_tool.py&lt;/a&gt;&lt;/strong&gt;: Asynchronous speech-to-text.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a&gt;tools/maps_tool.py&lt;/a&gt;&lt;/strong&gt;: 5 call points including nearby search, restaurant name extraction, batch and review search, etc.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a&gt;tools/summarizer.py&lt;/a&gt;&lt;/strong&gt;: Text summarization and Agentic Vision image understanding.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a&gt;tools/youtube_tool.py&lt;/a&gt;&lt;/strong&gt;: YouTube summary tool.&lt;/li&gt;
&lt;/ol&gt;




&lt;h1&gt;
  
  
  Pitfall Guide: Watch Out for SDK Module Import Issues
&lt;/h1&gt;

&lt;p&gt;When refactoring calls without Config for &lt;code&gt;youtube_gcp.py&lt;/code&gt; and &lt;code&gt;youtube_tool.py&lt;/code&gt;, since these two files originally only used named imports for specific types:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;google.genai.types&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;HttpOptions&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Part&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When we write &lt;code&gt;types.GenerateContentConfig(...)&lt;/code&gt; in the code, the system throws a &lt;code&gt;NameError: name 'types' is not defined&lt;/code&gt; error.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Solution:&lt;/strong&gt; We need to correct the import statement and directly import &lt;a&gt;GenerateContentConfig&lt;/a&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Before
&lt;/span&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;google.genai.types&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;HttpOptions&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Part&lt;/span&gt;

&lt;span class="c1"&gt;# After
&lt;/span&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;google.genai.types&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;HttpOptions&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Part&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;GenerateContentConfig&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And use it directly in the call without the &lt;code&gt;types.&lt;/code&gt; prefix:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nc"&gt;GenerateContentConfig&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;labels&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;client_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;info_helper&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;})&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  Summary and Next Steps
&lt;/h1&gt;

&lt;p&gt;This modification successfully injected the &lt;code&gt;client_id=info_helper&lt;/code&gt; label into all Vertex AI API calls within the LINE Bot project.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Billing Delay&lt;/strong&gt;: Please note that after we start including &lt;code&gt;labels&lt;/code&gt;, GCP billing data usually has a 24 to 48-hour delay before taking effect.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Configure in GCP Billing&lt;/strong&gt;: After two days, you can go to the GCP Console -&amp;gt; &lt;strong&gt;Billing&lt;/strong&gt; -&amp;gt; &lt;strong&gt;Reports&lt;/strong&gt;. In the "Group by" section on the right, select &lt;strong&gt;Labels&lt;/strong&gt; and enter our key &lt;code&gt;client_id&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Mission Accomplished&lt;/strong&gt;: At this point, the report will draw &lt;code&gt;info_helper&lt;/code&gt; as a separate billing row, perfectly solving the problem of separating project costs for reimbursement and statistics!&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>ai</category>
      <category>cloud</category>
      <category>google</category>
      <category>llm</category>
    </item>
    <item>
      <title>[AI in Action] Refining a macOS Meeting Translation App with Claude Code: Auto-reconnect, Floating Captions, and Meeting Minutes Export Evolution</title>
      <dc:creator>Evan Lin</dc:creator>
      <pubDate>Sun, 05 Jul 2026 11:00:46 +0000</pubDate>
      <link>https://dev.to/gde/ai-in-action-refining-a-macos-meeting-translation-app-with-claude-code-auto-reconnect-floating-2856</link>
      <guid>https://dev.to/gde/ai-in-action-refining-a-macos-meeting-translation-app-with-claude-code-auto-reconnect-floating-2856</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fztv3i129dwrt83h8npqw.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fztv3i129dwrt83h8npqw.png" alt="image-20260702134921415" width="800" height="499"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h1&gt;
  
  
  Foreword: Round Two, Switching to a Sharper Tool
&lt;/h1&gt;

&lt;p&gt;In the &lt;a href="//2026-06-10-agy-macos-app.md"&gt;previous article&lt;/a&gt;, we used &lt;strong&gt;AGY CLI (Antigravity)&lt;/strong&gt; to build a macOS real-time meeting translation App from scratch: capturing audio from Zoom / Google Meet via ScreenCaptureKit, sending it to the Gemini Live API for real-time translation, and displaying Traditional Chinese bilingual subtitles in the App window.&lt;/p&gt;

&lt;p&gt;After the App went live, the developer discovered several issues during actual meetings that raised eyebrows, while also coming up with more functional ideas. This time, we switched to Anthropic's &lt;strong&gt;Claude Code&lt;/strong&gt; to start the second round of deep polishing in the terminal.&lt;/p&gt;

&lt;p&gt;The following records the complete dialogue flow of this collaboration, restoring every key decision point.&lt;/p&gt;




&lt;h1&gt;
  
  
  Phase 1: Uncovering Hidden Dangers — The Mystery of the WebSocket Stopping After 10 Minutes
&lt;/h1&gt;

&lt;p&gt;The App seemed to work perfectly, but the developer returned with a question after a real, long meeting:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;User&lt;/strong&gt;: Check this code. Why does the real-time translation stop after about ten minutes? Help me look into possible reasons.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Claude Code read all five Swift source files and, combined with the built-in Gemini Live API documentation, precisely pointed out the root of the problem:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Gemini Live API's WebSocket connection has a Session limit of about 10 minutes&lt;/strong&gt;. Once the time is up, the server actively closes the connection and sends a &lt;code&gt;GoAway&lt;/code&gt; signal. However, the original code was completely unprepared for this scenario, causing a three-layer failure:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt; &lt;strong&gt;No reconnection after disconnection&lt;/strong&gt;: The &lt;code&gt;didCloseWith&lt;/code&gt; callback only updated the status text on the UI without any reconnection logic.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Silent audio discarding&lt;/strong&gt;: After &lt;code&gt;isConnected = false&lt;/code&gt;, every audio chunk from the microphone was silently consumed by &lt;code&gt;guard isConnected else { return }&lt;/code&gt;, without the user noticing.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Incorrect &lt;code&gt;isRunning&lt;/code&gt; state&lt;/strong&gt;: The "Translating" icon in the main window was still flashing, but in reality, nothing was happening.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Claude Code immediately refactored &lt;code&gt;GeminiLiveConnection.swift&lt;/code&gt;, introducing three lines of defense without changing any external interfaces:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Proactive GoAway Signal Detection&lt;/strong&gt;: Parse the &lt;code&gt;goAway&lt;/code&gt; JSON message sent by the server in &lt;code&gt;parseServerResponse()&lt;/code&gt;, and actively reconnect before the connection actually drops to achieve seamless translation continuity.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Exponential Backoff Auto-Reconnection&lt;/strong&gt;: After disconnection, retry at increasing intervals of 2s → 4s → 6s, up to 10 times; the count resets after a successful reconnection, ensuring the next disconnection also has a full set of retry opportunities. The entire process is silent and transparent to the user.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;30-second Ping Keep-Alive&lt;/strong&gt;: Start a &lt;code&gt;Timer&lt;/code&gt; after a successful connection to send a &lt;code&gt;sendPing&lt;/code&gt; to the server every 30 seconds, preventing intermediate network equipment from cutting the connection due to inactivity.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;After the modifications were completed, the developer executed &lt;code&gt;bash build_app.sh&lt;/code&gt; and pushed to GitHub:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Claude Code&lt;/strong&gt;: Build successful, commit completed, pushed. Real-time translation should now be able to operate continuously for a long time, silently reconnecting automatically when the Session expires without the user noticing an interruption.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  Phase 2: Feature Interview — Five Questions to Lock Down Two Priority Development Directions
&lt;/h1&gt;

&lt;p&gt;With stability resolved, the developer began thinking about how the App could go further:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;User&lt;/strong&gt;: Help me research carefully and give me three new feature recommendations.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;After carefully reading the code, Claude Code first proposed a directional question instead of directly giving a list:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Which category of these three new feature directions do you prefer?&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;A. Utility / UX&lt;/strong&gt; — Making daily use more convenient (transcript export, hotkeys, meeting notes)&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;B. Translation Quality&lt;/strong&gt; — More accurate and natural (multi-language detection, speaker identification)&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;C. System Integration&lt;/strong&gt; — Connecting with other tools (Notion, Zoom overlay)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;User&lt;/strong&gt;: A&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Three UX features were immediately produced: &lt;strong&gt;Floating Subtitle Window&lt;/strong&gt;, &lt;strong&gt;Automatic Meeting Transcript Export&lt;/strong&gt;, and &lt;strong&gt;Global Hotkeys&lt;/strong&gt;. The developer's response was direct:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;User&lt;/strong&gt;: I want both 1 and 2.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Next was a brief requirement interview, where Claude Code asked only one most critical question at a time:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  How many lines should the floating window display? → &lt;strong&gt;Double lines (small text for original + large text for translation)&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;  Background style? → &lt;strong&gt;Vibrancy effect (frosted glass)&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;  Export method? → &lt;strong&gt;Automatically save to the desktop, no dialog box&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;After five questions, the design direction was completely clear. Claude Code proposed a complete design plan and wrote the specification document. After the developer confirmed "no problem," it entered the implementation phase.&lt;/p&gt;




&lt;h1&gt;
  
  
  Phase 3: Plan-Driven Development — Subagent Closed-Loop Delivery, Review Catches Critical Bugs
&lt;/h1&gt;

&lt;p&gt;With clear specifications, Claude Code entered its most proficient work mode: &lt;strong&gt;Write a plan first, then use multiple independent Subagents to execute tasks, with each Task immediately reviewed by a Reviewer Subagent upon completion&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The entire process was divided into three Tasks; the two most critical ones are recorded below:&lt;/p&gt;

&lt;h3&gt;
  
  
  Task 1: Automatic Meeting Transcript Export
&lt;/h3&gt;

&lt;p&gt;The Implementer Subagent quickly completed three things: removed the original 25-line history limit, added the &lt;code&gt;exportTranscript()&lt;/code&gt; method, and automatically saved the complete bilingual comparison record in Markdown format to the Desktop when translation stopped.&lt;/p&gt;

&lt;p&gt;However, the &lt;strong&gt;Reviewer Subagent&lt;/strong&gt; immediately raised a flag:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Critical Issue found: &lt;code&gt;status = "Stopped"&lt;/code&gt; in &lt;code&gt;stop()&lt;/code&gt; is executed immediately after &lt;code&gt;exportTranscript()&lt;/code&gt;, instantly overwriting the save path message. The user will only ever see "Stopped" and will never know where the file was saved.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This was a logic bug just one line away, which would have been very easy to overlook without a Reviewer. The &lt;strong&gt;Fix Subagent&lt;/strong&gt; then intervened, changing &lt;code&gt;exportTranscript()&lt;/code&gt; to return a &lt;code&gt;Bool&lt;/code&gt;: when export is successful, &lt;code&gt;stop()&lt;/code&gt; no longer overwrites the status; "Stopped" is only displayed when there are no records to export. After the modification, the Reviewer confirmed again, and all passed.&lt;/p&gt;

&lt;h3&gt;
  
  
  Task 2: Floating Subtitle Window
&lt;/h3&gt;

&lt;p&gt;Added &lt;code&gt;FloatingSubtitleWindow.swift&lt;/code&gt;, with a core structure of three layers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;&lt;code&gt;NSPanel&lt;/code&gt;&lt;/strong&gt; (&lt;code&gt;level = .floating&lt;/code&gt;): Always on top, does not steal focus (&lt;code&gt;.nonactivatingPanel&lt;/code&gt;), and can be displayed across full-screen Apps.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;&lt;code&gt;NSVisualEffectView&lt;/code&gt;&lt;/strong&gt; (&lt;code&gt;material = .hudWindow&lt;/code&gt;): Native macOS vibrancy effect.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;&lt;code&gt;NSHostingView&lt;/code&gt;&lt;/strong&gt; embedding SwiftUI's &lt;code&gt;FloatingSubtitleView&lt;/code&gt;: Directly bound to &lt;code&gt;TranslatorViewModel.currentLine&lt;/code&gt;, updating in real-time.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;At the same time, ownership of &lt;code&gt;TranslatorViewModel&lt;/code&gt; was moved up from &lt;code&gt;ContentView&lt;/code&gt; to &lt;code&gt;TranslatorApp&lt;/code&gt;, allowing the main window and the floating window to share the same data source, avoiding data duplication or synchronization issues. The window position is saved to &lt;code&gt;UserDefaults&lt;/code&gt; after dragging and automatically restored after a restart.&lt;/p&gt;

&lt;p&gt;The Task Reviewer checked all 11 specifications one by one; all passed without any need for correction.&lt;/p&gt;

&lt;p&gt;The entire "Implementation → Review → Correction → Re-review" closed loop was completed automatically by subagents. The developer only needed to confirm that the final &lt;code&gt;bash build_app.sh&lt;/code&gt; passed cleanly:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Claude Code&lt;/strong&gt;: Build successful, commit completed, pushed.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  Phase 4: App Brand Upgrade — Real-time Generation of Professional Icons with Python
&lt;/h1&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn16axnnwtairp47x2jt7.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn16axnnwtairp47x2jt7.png" alt="image-20260702135008634" width="516" height="538"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;With features complete, the developer turned their attention to the appearance:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;User&lt;/strong&gt;: The app icon doesn't look good, help me generate a professional one.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Claude Code first confirmed that &lt;code&gt;Pillow&lt;/code&gt; (Python image library) was in the environment, then directly wrote a complete Icon generation script with the following design description:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Background&lt;/strong&gt;: Deep sea blue gradient (&lt;code&gt;#0D1B4E&lt;/code&gt; → &lt;code&gt;#1565C0&lt;/code&gt;), standard macOS 22% rounded corners, echoing the macOS Design Language.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Core Pattern&lt;/strong&gt;: Two overlapping speech bubbles. The upper bubble (semi-transparent white) contains "&lt;strong&gt;A&lt;/strong&gt;" representing the original English audio, and the lower bubble (pure white) contains "&lt;strong&gt;中&lt;/strong&gt;" representing the translated output. They are connected by a bidirectional arrow in the center, making the "real-time translation" product positioning clear at a glance.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Fonts&lt;/strong&gt;: Avenir Next for English and Apple SD Gothic Neo for Chinese, both of which are built-in macOS fonts, requiring no external resources.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The script output 10 sizes at once (16px → 1024px), converted them into an &lt;code&gt;.icns&lt;/code&gt; file using the system's &lt;code&gt;iconutil&lt;/code&gt; command, and automatically updated &lt;code&gt;build_app.sh&lt;/code&gt; to copy the icon into the App Bundle, adding the &lt;code&gt;CFBundleIconFile&lt;/code&gt; declaration to Info.plist. The entire process did not require opening Xcode or using any image design tools.&lt;/p&gt;




&lt;h1&gt;
  
  
  Phase 5: Code Quality Refinement — Clearing All Compilation Warnings
&lt;/h1&gt;

&lt;p&gt;When the developer executed &lt;code&gt;bash build_app.sh&lt;/code&gt; for acceptance, they noticed a few lines of yellow warnings in the output:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;User&lt;/strong&gt;: There are some warnings when running build_app.sh, help me check them.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Claude Code carefully executed the Build and categorized three types of warnings, treating them accordingly:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Warning Type&lt;/th&gt;
&lt;th&gt;Root Cause&lt;/th&gt;
&lt;th&gt;Fix&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;onChange(of:perform:)&lt;/code&gt; deprecated × 2&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;swiftc&lt;/code&gt; did not specify a deployment target, defaulting to the latest SDK rules&lt;/td&gt;
&lt;td&gt;Added &lt;code&gt;-target arm64-apple-macos13.0&lt;/code&gt; to &lt;code&gt;build_app.sh&lt;/code&gt; to let the compiler know we are targeting macOS 13, where the old API is the correct choice&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;SCRunningApplication&lt;/code&gt; non-Sendable × 2&lt;/td&gt;
&lt;td&gt;Types in the ScreenCaptureKit framework are not marked &lt;code&gt;Sendable&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Changed &lt;code&gt;import ScreenCaptureKit&lt;/code&gt; to &lt;code&gt;@preconcurrency import ScreenCaptureKit&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;TranslatorViewModel&lt;/code&gt; non-Sendable captured&lt;/td&gt;
&lt;td&gt;ViewModel was captured in a &lt;code&gt;@Sendable&lt;/code&gt; closure&lt;/td&gt;
&lt;td&gt;Added &lt;code&gt;@MainActor&lt;/code&gt; to &lt;code&gt;TranslatorViewModel&lt;/code&gt; (modern standard practice for SwiftUI ViewModels) and added &lt;code&gt;@preconcurrency&lt;/code&gt; to delegate conformance to suppress derived warnings&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The final Build output was as clean as new, without any Warnings:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;🛠 Starting compilation of Swift files (target: arm64-apple-macos13.0)...
🎨 Copying App Icon...
📝 Generating Info.plist...
✅ Packaging complete!
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;All modifications were committed and pushed to GitHub together.&lt;/p&gt;




&lt;h1&gt;
  
  
  Phase 6: Real-World Pitfalls — The ScreenCaptureKit Permission Labyrinth
&lt;/h1&gt;

&lt;p&gt;The App's functionality seemed complete until the developer actually turned it on to start using it:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;User&lt;/strong&gt;: Is it a permission issue? I open the app and it keeps failing to scan the "Target App" list. Help me check the relevant code.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The App list was always empty. "Screen Recording" in System Settings was indeed checked. This is a typical "permissions are granted, but it just doesn't work" dead-end problem.&lt;/p&gt;

&lt;h3&gt;
  
  
  First Cut: Handling Silent Failures
&lt;/h3&gt;

&lt;p&gt;After reading &lt;code&gt;AudioCaptureManager.swift&lt;/code&gt;, Claude Code immediately spotted the problem: when the &lt;code&gt;fetchShareableApps()&lt;/code&gt; call to &lt;code&gt;SCShareableContent.current&lt;/code&gt; failed, it only &lt;code&gt;print&lt;/code&gt;ed to the console. The UI showed an empty list without any prompt. The developer had no idea what was happening.&lt;/p&gt;

&lt;p&gt;The first wave of modifications did three things:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Added &lt;code&gt;NSScreenCaptureUsageDescription&lt;/code&gt; to &lt;code&gt;Info.plist&lt;/code&gt;&lt;/strong&gt;: Without this key, the macOS authorization dialog will never pop up.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Added an ad-hoc signing step&lt;/strong&gt;: &lt;code&gt;codesign --sign - --force --deep&lt;/code&gt; — ScreenCaptureKit requires the App to have a code identity to appear in the "System Settings &amp;gt; Screen Recording" list.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Surfaced errors to the UI&lt;/strong&gt;: Changed &lt;code&gt;fetchShareableApps()&lt;/code&gt; to return &lt;code&gt;(apps, errorMessage?)&lt;/code&gt;. Any failure would be displayed in the App's status bar, allowing the developer to see immediately what happened.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Build completed, tested again—still the same error message.&lt;/p&gt;

&lt;h3&gt;
  
  
  Second Cut: Overly Aggressive Error Classification Logic
&lt;/h3&gt;

&lt;p&gt;Looking closely at the error judgment code:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight swift"&gt;&lt;code&gt;&lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;isPermissionDenied&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;nsError&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;domain&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="s"&gt;"..."&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="n"&gt;nsError&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;code&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
    &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="n"&gt;error&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;localizedDescription&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;lowercased&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;contains&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"permission"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="n"&gt;error&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;localizedDescription&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;lowercased&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;contains&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"denied"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;contains("permission")&lt;/code&gt; line was too aggressive. As long as any word containing "permission" appeared in the error description, it would be incorrectly judged as "Permission Denied," displaying "Please go to System Settings to enable authorization." In reality, it could be a completely different error.&lt;/p&gt;

&lt;p&gt;Claude Code corrected the judgment logic—only the exact ScreenCaptureKit &lt;code&gt;userDeclined&lt;/code&gt; error code (&lt;code&gt;-3801&lt;/code&gt;) is treated as a permission issue. All other errors display the actual domain, code, and description for easier diagnosis:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight swift"&gt;&lt;code&gt;&lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;isPermissionDenied&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;nsError&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;code&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;3801&lt;/span&gt;
&lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;message&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;isPermissionDenied&lt;/span&gt;
    &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="s"&gt;"Screen recording permission required: Please go to System Settings to enable authorization"&lt;/span&gt;
    &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;"Unable to get App list (code &lt;/span&gt;&lt;span class="se"&gt;\(&lt;/span&gt;&lt;span class="n"&gt;nsError&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;code&lt;/span&gt;&lt;span class="se"&gt;)&lt;/span&gt;&lt;span class="s"&gt;): &lt;/span&gt;&lt;span class="se"&gt;\(&lt;/span&gt;&lt;span class="n"&gt;error&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;localizedDescription&lt;/span&gt;&lt;span class="se"&gt;)&lt;/span&gt;&lt;span class="s"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Third Cut: Finding the Root Cause — TCC Identity Mismatch
&lt;/h3&gt;

&lt;p&gt;After correcting the error classification, Claude Code ran the App and captured the logs, finding that the status bar displayed a new message with a code number, not &lt;code&gt;-3801&lt;/code&gt;. This confirmed: &lt;strong&gt;The problem wasn't that the user hadn't given permission, but that macOS didn't recognize the App at all&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The root cause:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Every time &lt;code&gt;build_app.sh&lt;/code&gt; is executed and re-signed ad-hoc, the hash of the binary changes, and the macOS TCC database treats it as a completely new App.&lt;/strong&gt; The old screen recording authorization was given to the previous binary; the new binary did not inherit it. System Settings shows it as checked, but that's authorization for the old identity, which is invalid for the new binary.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The solution is to reset TCC to let macOS re-trigger the authorization dialog:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;tccutil reset ScreenCapture com.poc.MeetingTranslator
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After execution, reopening the App and clicking "↻" caused macOS to immediately pop up the "MeetingTranslator wants to record the contents of this screen" dialog. Clicking "Allow" instantly listed all running applications in the App list.&lt;/p&gt;

&lt;h3&gt;
  
  
  Permanent Countermeasure: Writing the Reset into the Build Process
&lt;/h3&gt;

&lt;p&gt;The ad-hoc signing issue persists during development—every rebuild requires re-authorization. Claude Code added &lt;code&gt;tccutil reset&lt;/code&gt; directly as the last step of &lt;code&gt;build_app.sh&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;tccutil reset ScreenCapture com.poc.MeetingTranslator 2&amp;gt;/dev/null &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"✅ Reset complete. The system will ask for authorization again after opening the App"&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nb"&gt;true&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;From then on, after every &lt;code&gt;bash build_app.sh&lt;/code&gt;, simply &lt;code&gt;open MeetingTranslator.app&lt;/code&gt;, and the system will ask for authorization again. The entire development cycle will never again get stuck in the "permissions are granted but it doesn't work" loop.&lt;/p&gt;




&lt;h1&gt;
  
  
  Conclusion: The True Value of the "Plan → Subagent Implementation → AI Review" Closed Loop
&lt;/h1&gt;

&lt;p&gt;This collaboration with Claude Code made me feel a completely different way of working compared to the first AGY CLI development:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Proactively asking questions instead of just starting&lt;/strong&gt;: Faced with "give me three new feature recommendations," Claude Code's first step was to ask for a direction; faced with the "floating window," it confirmed style and details one by one. This rhythm of "align first, then implement" is much more reliable than directly guessing requirements.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;The plan is a moat for quality&lt;/strong&gt;: Writing specification documents and implementation plans before implementation gives each Subagent clear boundaries and acceptance criteria. This seemingly "redundant" step directly discovered a state-overwriting bug in the Task 1 review that human developers could easily overlook.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;AI Reviewing AI is a different layer of protection&lt;/strong&gt;: The Reviewer Subagent and Implementer Subagent are started completely independently; they share no context. Because of this, the Reviewer can discover the Implementer's blind spots from a fresh perspective—this is the extra protection brought by "AI double-checking," not a replacement for human Code Review, but a completely new level of quality.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Tool boundaries are functional boundaries&lt;/strong&gt;: App Icon generation, Warning fixes, Git commit/push—Claude Code moves freely throughout the development environment. The developer doesn't need to switch tools; all actions are completed within the dialogue.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Real-world use is the best test&lt;/strong&gt;: The ScreenCaptureKit issue in Phase 6 never appeared in any build tests until the developer actually turned it on to use it. This kind of "silent failure + system-level identity mismatch" problem only surfaces in real-world scenarios. Claude Code's diagnostic method—from correcting error classification to letting the UI display real error codes, to finding the TCC root cause—is a typical "narrowing the hypothesis range and letting the problem speak" debugging approach.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If the first article was about "from zero to one," this record is about "from usable to good, and then to standing firm in real-world scenarios." Two types of AI Agents, two collaboration styles, together completed a full Native macOS App covering low-level audio, WebSocket connections, SwiftUI UI, Python image generation, and system permission diagnosis. See you next time!&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claude</category>
      <category>programming</category>
      <category>softwaredevelopment</category>
    </item>
    <item>
      <title>[Gemini API in Action] Building MemeFinder: A Native Mac Menu Bar Widget for Finding Memes via Text Using Gemini Vision &amp; Semantic Embeddings</title>
      <dc:creator>Evan Lin</dc:creator>
      <pubDate>Mon, 22 Jun 2026 00:41:17 +0000</pubDate>
      <link>https://dev.to/gde/gemini-api-hands-on-59dc</link>
      <guid>https://dev.to/gde/gemini-api-hands-on-59dc</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhh50l2jdcckul8cmidwl.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhh50l2jdcckul8cmidwl.png" alt="image-memefinder-hero" width="800" height="460"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h1&gt;
  
  
  The Origin: Mid-Conversation, Where on Earth Is That Meme?
&lt;/h1&gt;

&lt;p&gt;Anyone who chats a lot has a folder full of memes on their phone and computer, but the moment you actually need one — the conversation is rolling, you want to drop a "thanks but no thanks" or an "I'm trash" reaction — you can't find it. The filename is &lt;code&gt;IMG_4821.jpg&lt;/code&gt;, the photo library has no categories, and search is a non-starter.&lt;/p&gt;

&lt;p&gt;I first came across a wonderful open-source project, &lt;a href="https://github.com/ShiQu1218/MemeTalk" rel="noopener noreferrer"&gt;ShiQu1218/MemeTalk&lt;/a&gt;. It builds a local meme semantic-search system with Python + Streamlit + SQLite: it scans your local meme folder, indexes images with OCR and vector embeddings, then does multi-route retrieval. Feature-complete, but research-oriented and requires opening a browser to run Streamlit.&lt;/p&gt;

&lt;p&gt;What I wanted was something closer to an "everyday handy tool":&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A native Mac app, one search box. I type what I'm looking for and the relevant meme pops up. Click it and it's copied straight to the clipboard.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;So MemeFinder was born. This post records its journey from zero to "menu-bar resident + global hotkey," and several representative pitfalls along the way.&lt;/p&gt;




&lt;h1&gt;
  
  
  System Design and Architecture
&lt;/h1&gt;

&lt;p&gt;The core concept is simple: &lt;strong&gt;point at a local meme folder → have Gemini build an index for each image → type to do a semantic search → click to copy&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;I made three key technical decisions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Native SwiftUI app&lt;/strong&gt;, not Electron. Copying images to the clipboard, global hotkeys, menu-bar residency — with AppKit these are all first-class citizens.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gemini&lt;/strong&gt; does two things: the vision model &lt;code&gt;gemini-3-flash-preview&lt;/code&gt; reads the text in each image and generates a Traditional Chinese description plus emotion tags; &lt;code&gt;gemini-embedding-2&lt;/code&gt; turns that semantics into a 768-dimensional vector.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hybrid semantic-vector + keyword search.&lt;/strong&gt; Pure keyword recall for Chinese is too poor; only semantic vectors achieve "type a related description and find the image."&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  System Architecture Flow
&lt;/h3&gt;

&lt;p&gt;The project is deliberately split into two Swift Package targets:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Target&lt;/th&gt;
&lt;th&gt;Type&lt;/th&gt;
&lt;th&gt;Contents&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;MemeFinder&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;library&lt;/td&gt;
&lt;td&gt;Logic, models, services, ViewModels (all unit-tested)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;MemeFinderApp&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;executable&lt;/td&gt;
&lt;td&gt;SwiftUI views + menu-bar shell (thin layer, depends on the library)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This split isn't decorative — it directly determines whether the tests can run smoothly, as "Pitfall #2" will explain.&lt;/p&gt;




&lt;h1&gt;
  
  
  Core Implementation
&lt;/h1&gt;

&lt;h3&gt;
  
  
  1. Auto-tagging memes with the Gemini vision model
&lt;/h3&gt;

&lt;p&gt;During indexing, each image is sent to the vision model with a request to &lt;strong&gt;output only JSON&lt;/strong&gt;: the text in the image, a Traditional Chinese description, tags, and emotion. &lt;code&gt;responseMimeType&lt;/code&gt; is set to &lt;code&gt;application/json&lt;/code&gt; to keep the output format stable:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight swift"&gt;&lt;code&gt;&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;static&lt;/span&gt; &lt;span class="kd"&gt;func&lt;/span&gt; &lt;span class="nf"&gt;annotateRequest&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;apiKey&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kt"&gt;String&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;imageData&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kt"&gt;Data&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;mimeType&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kt"&gt;String&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="kt"&gt;URLRequest&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;prompt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"""
    你是迷因圖標註助手。請閱讀這張圖，輸出 JSON，欄位：
    ocr_text(圖中所有文字), description(用繁體中文描述畫面與梗),
    tags(3-8 個繁體中文關鍵字陣列), emotion(單一情緒詞)。只輸出 JSON。
    """&lt;/span&gt;
    &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;body&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="kt"&gt;String&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kt"&gt;Any&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="s"&gt;"contents"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[[&lt;/span&gt;
            &lt;span class="s"&gt;"parts"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
                &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s"&gt;"text"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
                &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s"&gt;"inline_data"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s"&gt;"mime_type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;mimeType&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s"&gt;"data"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;imageData&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;base64EncodedString&lt;/span&gt;&lt;span class="p"&gt;()]]&lt;/span&gt;
            &lt;span class="p"&gt;]&lt;/span&gt;
        &lt;span class="p"&gt;]],&lt;/span&gt;
        &lt;span class="s"&gt;"generationConfig"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s"&gt;"responseMimeType"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;"application/json"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="c1"&gt;// ... set URL, x-goog-api-key header, POST body&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  2. Hybrid semantic + keyword ranking
&lt;/h3&gt;

&lt;p&gt;After the query string is embedded into a vector, we compute cosine similarity for every image, then add weight for keywords that hit the OCR text and tags, and merge-sort:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight swift"&gt;&lt;code&gt;&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;func&lt;/span&gt; &lt;span class="nf"&gt;search&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;queryEmbedding&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="kt"&gt;Float&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="nv"&gt;queryText&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kt"&gt;String&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                   &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="nv"&gt;images&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="kt"&gt;IndexedImage&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="nv"&gt;limit&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kt"&gt;Int&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="kt"&gt;SearchResult&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;tokens&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;queryText&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;lowercased&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;whereSeparator&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;$0&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;isWhitespace&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;String&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="kd"&gt;init&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;results&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="kt"&gt;SearchResult&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;images&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;compactMap&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="n"&gt;image&lt;/span&gt; &lt;span class="k"&gt;in&lt;/span&gt;
        &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;cos&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;cosineSimilarity&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;queryEmbedding&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;image&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;embedding&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;haystack&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;image&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ocrText&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="s"&gt;" "&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;image&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;tags&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;joined&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;separator&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;" "&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;lowercased&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;matches&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;tokens&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;filter&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nv"&gt;$0&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;isEmpty&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="n"&gt;haystack&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;contains&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;$0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;count&lt;/span&gt;
        &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;boost&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mf"&gt;0.1&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="kt"&gt;Float&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;min&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;matches&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;   &lt;span class="c1"&gt;// keyword boost capped at 0.3&lt;/span&gt;
        &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;score&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;cos&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;boost&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;score&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="kt"&gt;SearchResult&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;image&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;image&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;score&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;score&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;nil&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="kt"&gt;Array&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;results&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;sorted&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;$0&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;score&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="nv"&gt;$1&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;score&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;prefix&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;limit&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The whole search engine is a pure function, with Gemini hidden behind a protocol, so this logic can be fully unit-tested offline without hitting the real API.&lt;/p&gt;




&lt;h1&gt;
  
  
  Major Pitfalls and Solutions
&lt;/h1&gt;

&lt;p&gt;The real time sink in this project was never the happy path — it was the pitfalls below.&lt;/p&gt;

&lt;h3&gt;
  
  
  Pitfall #1: The mysterious &lt;code&gt;GeminiError error 0&lt;/code&gt; — indexing and search both fail
&lt;/h3&gt;

&lt;p&gt;App packaged, key set, folder chosen, hit search — and nothing shows below, just &lt;code&gt;GeminiError error 0&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Rather than guessing, I hit the embedding endpoint once with a real key and printed the response:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="s2"&gt;"https://generativelanguage.googleapis.com/v1beta/models/gemini-embedding-2:embedContent"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"x-goog-api-key: &lt;/span&gt;&lt;span class="nv"&gt;$KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"content":{"parts":[{"text":"貓"}]},"output_dimensionality":768}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The evidence was unmistakable:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"embedding"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"values"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;-0.0063&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;-0.0200&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;...&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The problem: my parser was reading the &lt;strong&gt;plural&lt;/strong&gt; &lt;code&gt;embeddings[0].values&lt;/code&gt; (that's the &lt;code&gt;batchEmbedContents&lt;/code&gt; batch-endpoint format), but the single &lt;code&gt;embedContent&lt;/code&gt; call returns the &lt;strong&gt;singular&lt;/strong&gt; &lt;code&gt;embedding.values&lt;/code&gt;. So &lt;strong&gt;every embed call failed&lt;/strong&gt; — indexing each image failed, embedding the query string failed — all throwing &lt;code&gt;badResponse&lt;/code&gt; (shown in the UI as &lt;code&gt;GeminiError error 0&lt;/code&gt;).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;[Solution]&lt;/strong&gt;&lt;br&gt;
Fix the parser to read the singular &lt;code&gt;embedding.values&lt;/code&gt;, keeping the plural format as a fallback; I also hardened the annotation parser (a thinking model sometimes returns a textless "thought" part first, so skip to the first part that actually has text):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight swift"&gt;&lt;code&gt;&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;static&lt;/span&gt; &lt;span class="kd"&gt;func&lt;/span&gt; &lt;span class="nf"&gt;embedding&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;fromEmbedContent&lt;/span&gt; &lt;span class="nv"&gt;data&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kt"&gt;Data&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;throws&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="kt"&gt;Float&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;guard&lt;/span&gt; &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;root&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="kt"&gt;JSONSerialization&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;jsonObject&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;with&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as?&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="kt"&gt;String&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kt"&gt;Any&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="kt"&gt;GeminiError&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;badResponse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"cannot parse embedContent payload"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="c1"&gt;// A single embedContent returns {"embedding":{"values":[...]}}&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;embedding&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;root&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s"&gt;"embedding"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="k"&gt;as?&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="kt"&gt;String&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kt"&gt;Any&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
       &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;values&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;embedding&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s"&gt;"values"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="k"&gt;as?&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="kt"&gt;Double&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;values&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;Float&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="kd"&gt;init&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="c1"&gt;// batchEmbedContents is {"embeddings":[{"values":[...]}]} — tolerate it too&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;embeddings&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;root&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s"&gt;"embeddings"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="k"&gt;as?&lt;/span&gt; &lt;span class="p"&gt;[[&lt;/span&gt;&lt;span class="kt"&gt;String&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kt"&gt;Any&lt;/span&gt;&lt;span class="p"&gt;]],&lt;/span&gt;
       &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;values&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;embeddings&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;first&lt;/span&gt;&lt;span class="p"&gt;?[&lt;/span&gt;&lt;span class="s"&gt;"values"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="k"&gt;as?&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="kt"&gt;Double&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;values&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;Float&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="kd"&gt;init&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="kt"&gt;GeminiError&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;badResponse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"cannot parse embedContent payload"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Lesson: &lt;strong&gt;trust the actual API response over your memory or secondhand docs.&lt;/strong&gt; A single line of &lt;code&gt;curl&lt;/code&gt; saved countless guesses.&lt;/p&gt;

&lt;h3&gt;
  
  
  Pitfall #2: SwiftPM's &lt;code&gt;main&lt;/code&gt; entry-point conflict and the SwiftUICore linking error
&lt;/h3&gt;

&lt;p&gt;I initially made the whole project a single &lt;code&gt;executableTarget&lt;/code&gt; with the tests depending on it directly. The result: tests failed to link no matter what. An executable target needs a &lt;code&gt;main&lt;/code&gt; entry point, but that entry point only exists at the UI step's &lt;code&gt;@main App&lt;/code&gt;; and casually adding a placeholder &lt;code&gt;main.swift&lt;/code&gt; then conflicts with &lt;code&gt;@main&lt;/code&gt; (Swift doesn't allow two entry points in one target). Worse, SwiftUI in an executable target spews &lt;code&gt;SwiftUICore.tbd ... not an allowed client&lt;/code&gt; linker warnings.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;[Root cause analysis and solution]&lt;/strong&gt;&lt;br&gt;
This is actually an architecture problem, not a compilation problem. The right approach is to split the project into two layers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;MemeFinder&lt;/code&gt; (library target)&lt;/strong&gt;: all logic, models, services, ViewModels — the tests depend only on this layer, it has no entry point, and it links cleanly as a library. ViewModels &lt;code&gt;import Combine&lt;/code&gt; (not SwiftUI) to get &lt;code&gt;ObservableObject&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;MemeFinderApp&lt;/code&gt; (executable target)&lt;/strong&gt;: only SwiftUI views and &lt;code&gt;@main&lt;/code&gt;, with &lt;code&gt;import MemeFinder&lt;/code&gt; to use the public types above.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;After the split, the library and tests don't touch SwiftUI at all, the linker warnings disappear, and the &lt;code&gt;@main&lt;/code&gt; conflict no longer exists. &lt;strong&gt;"What the tests need to depend on" often forces out clean module boundaries.&lt;/strong&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  Pitfall #3: Parallel indexing's rate limit and "I want to stop indexing halfway"
&lt;/h3&gt;

&lt;p&gt;The first version indexed one image at a time, serially calling Gemini (annotate then embed). For hundreds of images this was painfully slow. So I switched to bounded parallelism with &lt;code&gt;withTaskGroup&lt;/code&gt; (at most 4 at once), which brought three new problems:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The Gemini free tier has a &lt;strong&gt;rate limit&lt;/strong&gt; — too much concurrency triggers 429.&lt;/li&gt;
&lt;li&gt;The user wants to &lt;strong&gt;cancel&lt;/strong&gt; halfway through a large folder.&lt;/li&gt;
&lt;li&gt;Parallel completion order is chaotic, but the results need &lt;strong&gt;stable sorting&lt;/strong&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;[Solution]&lt;/strong&gt;&lt;br&gt;
Handle the three problems separately, all converging in the same &lt;code&gt;buildIndex&lt;/code&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;429 backoff retry&lt;/strong&gt;: retry only &lt;code&gt;GeminiError.rateLimited&lt;/code&gt; with exponential backoff (max 3 attempts); other errors are recorded without retry.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cooperative cancellation&lt;/strong&gt;: honor &lt;code&gt;Task.isCancelled&lt;/code&gt;; on cancel, stop scheduling new work and keep the completed portion. Even the backoff &lt;code&gt;Task.sleep&lt;/code&gt; lets &lt;code&gt;CancellationError&lt;/code&gt; propagate normally instead of swallowing it and firing one more API call.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stable sorting&lt;/strong&gt;: collect results into a &lt;code&gt;[path: image]&lt;/code&gt; dictionary, then reassemble the output in the order of the pre-sorted file list, decoupled from completion order.
&lt;/li&gt;
&lt;/ul&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight swift"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Seed maxConcurrent tasks first, then refill one per completion — strictly cap concurrency&lt;/span&gt;
&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="o"&gt;..&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;maxConcurrent&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nf"&gt;scheduleNext&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="k"&gt;break&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;res&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;group&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;next&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;img&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;res&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;image&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="n"&gt;resultsByPath&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;res&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;img&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;err&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;res&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;error&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="n"&gt;errors&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;err&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="n"&gt;done&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
    &lt;span class="nf"&gt;progress&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;done&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;total&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;scheduleNext&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Incidentally, the HTTP status code was also extracted into a pure function &lt;code&gt;mapResponse(data:statusCode:)&lt;/code&gt;: 429 → &lt;code&gt;rateLimited&lt;/code&gt;, other non-2xx → &lt;code&gt;httpError(code)&lt;/code&gt;, 2xx → return the data. The retry logic then has a basis, and this part is easy to test too.&lt;/p&gt;
&lt;h3&gt;
  
  
  Pitfall #4: Evolving from a "windowed app" into "menu-bar resident + global hotkey"
&lt;/h3&gt;

&lt;p&gt;Whether a tool is pleasant to use comes down to "how many steps to summon it." I wanted to hit &lt;strong&gt;⌃⌘M&lt;/strong&gt; mid-conversation to bring up the search popover, with the app tucked into the menu bar, not occupying the Dock. This step hit two classic macOS pitfalls:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;(a) Does a global hotkey need accessibility permission?&lt;/strong&gt; No. Use Carbon's &lt;code&gt;RegisterEventHotKey&lt;/code&gt; to register a fixed hotkey, which doesn't need Accessibility permission (unlike monitoring the whole keyboard). But under Swift 6 strict concurrency, the C event callback has to dispatch through a static &lt;code&gt;id → instance&lt;/code&gt; registry, requiring &lt;code&gt;nonisolated(unsafe)&lt;/code&gt; and relying on the invariant that "Carbon events are delivered on the main thread" for safety. If ⌃⌘M is already taken, &lt;code&gt;RegisterEventHotKey&lt;/code&gt; returns failure — in which case we silently degrade, log a line, and the menu-bar icon still works.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;(b) The timing race in the menu-bar right-click menu.&lt;/strong&gt; The initial approach was "set &lt;code&gt;statusItem.menu&lt;/code&gt; → &lt;code&gt;performClick&lt;/code&gt; → immediately clear &lt;code&gt;menu&lt;/code&gt;," but clearing synchronously fights AppKit's menu-tracking loop, and the menu flashes and disappears.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;[Solution]&lt;/strong&gt;&lt;br&gt;
Pop the menu up directly, fully bypassing the assign-and-clear of &lt;code&gt;statusItem.menu&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight swift"&gt;&lt;code&gt;&lt;span class="kd"&gt;@objc&lt;/span&gt; &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kd"&gt;func&lt;/span&gt; &lt;span class="nf"&gt;statusButtonClicked&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;guard&lt;/span&gt; &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;event&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kt"&gt;NSApp&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;currentEvent&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nf"&gt;togglePopover&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;type&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;rightMouseUp&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="c1"&gt;// Pop up directly; don't assign then synchronously clear statusItem.menu&lt;/span&gt;
        &lt;span class="c1"&gt;// (it races AppKit's menu-tracking loop)&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;button&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;statusItem&lt;/span&gt;&lt;span class="p"&gt;?&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;button&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="kt"&gt;NSMenu&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;popUpContextMenu&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;makeMenu&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="nv"&gt;with&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;for&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;button&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="nf"&gt;togglePopover&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Finally, adding &lt;code&gt;LSUIElement = true&lt;/code&gt; to the &lt;code&gt;Info.plist&lt;/code&gt; produced by &lt;code&gt;build-app.sh&lt;/code&gt; makes the Dock icon disappear, and MemeFinder officially becomes a pure menu-bar tool.&lt;/p&gt;

&lt;h3&gt;
  
  
  Pitfall #5: The settings form is blank — one symptom, three layers of cause
&lt;/h3&gt;

&lt;p&gt;After moving to the menu-bar version, a user reported "the settings window is completely blank." This seemingly simple bug, peeled apart, actually had three layers, each highly representative.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Layer 1: a &lt;code&gt;Form&lt;/code&gt; collapses to zero height inside a hand-rolled &lt;code&gt;NSWindow&lt;/code&gt;.&lt;/strong&gt;&lt;br&gt;
Originally the settings screen lived in SwiftUI's native &lt;code&gt;Settings { }&lt;/code&gt; scene, which sizes it sensibly. After the refactor it was hosted in a hand-rolled &lt;code&gt;NSWindow(contentViewController: NSHostingController(rootView: SettingsView()))&lt;/code&gt;, and &lt;code&gt;SettingsView&lt;/code&gt; ended with only &lt;code&gt;.frame(width: 460)&lt;/code&gt; — &lt;strong&gt;width only, no height&lt;/strong&gt;. &lt;code&gt;NSWindow(contentViewController:)&lt;/code&gt; sizes the window from the content's natural size, but a SwiftUI &lt;code&gt;Form&lt;/code&gt; is vertically greedy; with no constraint, its natural height resolves to nearly 0, so the window opens as a 460-wide, near-zero-height blank strip. The fix is just to add a height:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight swift"&gt;&lt;code&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;padding&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="c1"&gt;// When hosted in a hand-rolled NSWindow (not a SwiftUI Settings scene), a Form&lt;/span&gt;
&lt;span class="c1"&gt;// with no height constraint collapses to ~0, turning the window into a blank strip.&lt;/span&gt;
&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;frame&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;width&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;460&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;height&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;320&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Layer 2: ⌘, and the menu-bar "Settings…" go down two different paths.&lt;/strong&gt;&lt;br&gt;
After adding the height, the user said "still blank." On follow-up I found out he was summoning settings with &lt;strong&gt;⌘,&lt;/strong&gt;, while the menu-bar right-click "Settings…" went down a different path. The reason: ⌘, in a SwiftUI app triggers the &lt;code&gt;Settings { }&lt;/code&gt; scene, and to dodge a state-sharing problem during the refactor, I had set that to &lt;code&gt;Settings { EmptyView() }&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight swift"&gt;&lt;code&gt;&lt;span class="c1"&gt;// During the refactor, the Settings scene was left empty to avoid state-sharing&lt;/span&gt;
&lt;span class="c1"&gt;// — so ⌘, opens a blank window&lt;/span&gt;
&lt;span class="k"&gt;var&lt;/span&gt; &lt;span class="nv"&gt;body&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kd"&gt;some&lt;/span&gt; &lt;span class="kt"&gt;Scene&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kt"&gt;Settings&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="kt"&gt;EmptyView&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In other words, &lt;strong&gt;settings had two entry points pointing at different things&lt;/strong&gt;: ⌘, pointed at the empty scene, the menu-bar "Settings…" pointed at the real window. The fix unifies the two paths — let the &lt;code&gt;Settings&lt;/code&gt; scene host the real &lt;code&gt;SettingsView&lt;/code&gt; (so ⌘, works directly), and make the menu-bar "Settings…" open the same native settings window too:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight swift"&gt;&lt;code&gt;&lt;span class="kt"&gt;Settings&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kt"&gt;SettingsView&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;vm&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;appDelegate&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;settings&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;indexing&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;appDelegate&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;indexing&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                 &lt;span class="nv"&gt;onReindex&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="n"&gt;appDelegate&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;reindexNow&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
                 &lt;span class="nv"&gt;onCancel&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="n"&gt;appDelegate&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;cancelReindex&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight swift"&gt;&lt;code&gt;&lt;span class="c1"&gt;// The menu-bar "Settings…" now opens the same Settings scene&lt;/span&gt;
&lt;span class="kd"&gt;@objc&lt;/span&gt; &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kd"&gt;func&lt;/span&gt; &lt;span class="nf"&gt;openSettings&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kt"&gt;NSApp&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;activate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;ignoringOtherApps&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="kt"&gt;NSApp&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sendAction&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;Selector&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="s"&gt;"showSettingsWindow:"&lt;/span&gt;&lt;span class="p"&gt;)),&lt;/span&gt; &lt;span class="nv"&gt;to&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;nil&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;from&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;nil&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This also leverages the fact that a SwiftUI App body is &lt;code&gt;@MainActor&lt;/code&gt;-isolated — so reading the &lt;code&gt;@MainActor&lt;/code&gt; &lt;code&gt;appDelegate.settings&lt;/code&gt; directly from the body is legal, with no extra bridging needed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Layer 3 (the most insidious): &lt;code&gt;open&lt;/code&gt; doesn't reload a menu-bar app at all.&lt;/strong&gt;&lt;br&gt;
The biggest time-waster in the process was that after recompiling, I'd ask the user to &lt;code&gt;open MemeFinder.app&lt;/code&gt;, yet he kept seeing the old behavior. Because MemeFinder is an &lt;code&gt;LSUIElement&lt;/code&gt; menu-bar-resident app — when an instance is already running, &lt;code&gt;open&lt;/code&gt; only &lt;strong&gt;wakes the existing old process&lt;/strong&gt; instead of relaunching with the new binary. So we were actually testing the same old build the whole time. The correct dev loop is to truly kill it first, then run from source:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;killall MemeFinderApp 2&amp;gt;/dev/null&lt;span class="p"&gt;;&lt;/span&gt; swift run MemeFinderApp
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This layer reminds me: &lt;strong&gt;when debugging, first confirm "what you're testing really is the version you changed"&lt;/strong&gt; — otherwise all your reasoning is built on faulty observations.&lt;/p&gt;




&lt;h1&gt;
  
  
  On the "Development Process" Itself
&lt;/h1&gt;

&lt;p&gt;This project was driven almost entirely by an AI agent workflow of &lt;strong&gt;spec → plan → subagent task-by-task implementation → two-stage review&lt;/strong&gt;: each feature started with a design spec, was broken into independently testable small tasks, every task wrote a failing test first (TDD) before implementing, and after completion an independent review agent checked spec compliance and code quality, followed by one final whole-branch review.&lt;/p&gt;

&lt;p&gt;Several of the pitfalls — &lt;code&gt;GeminiError error 0&lt;/code&gt;, the library/executable split, swallowing &lt;code&gt;CancellationError&lt;/code&gt; during backoff, the menu timing race — were in fact caught half the time during the &lt;strong&gt;review stage&lt;/strong&gt;, not written correctly on the first pass. This echoes that old principle: &lt;strong&gt;having tests as armor, and someone (or an agent) seriously reading the diff, matters far more than writing fast.&lt;/strong&gt; The final project maintains 47 unit tests and a zero-warning release build.&lt;/p&gt;




&lt;h1&gt;
  
  
  Results and Benefits
&lt;/h1&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Type to find, click to paste&lt;/strong&gt;: type a Chinese description in the menu-bar popover, semantic search instantly lists relevant memes, click one to copy it to the clipboard and paste straight into LINE / Slack / Messages.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Privacy-friendly, searchable offline&lt;/strong&gt;: images and the index live locally (&lt;code&gt;~/Library/Application Support/MemeFinder/index.json&lt;/code&gt;); only the "build the index" step calls Gemini.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A truly handy tool&lt;/strong&gt;: ⌃⌘M is available anytime, menu-bar resident, no Dock footprint; incremental indexing only processes new/changed images, and indexing can show progress and be canceled.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A clean, maintainable architecture&lt;/strong&gt;: a two-layer library/executable design, Gemini hidden behind a protocol, pure logic fully covered by tests.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;All the development code for this project is open-sourced on GitHub: &lt;a href="https://github.com/kkdai/meme-finder-app" rel="noopener noreferrer"&gt;kkdai/meme-finder-app&lt;/a&gt;. Feel free to clone it, point it at your own meme-collection folder, and experience the joy of "type to find your meme"!&lt;/p&gt;

</description>
      <category>ai</category>
      <category>api</category>
      <category>gemini</category>
      <category>python</category>
    </item>
    <item>
      <title>[Gemini API] Gemini Batch API and Webhook API practical usage on restaurant survey</title>
      <dc:creator>Evan Lin</dc:creator>
      <pubDate>Mon, 15 Jun 2026 04:09:16 +0000</pubDate>
      <link>https://dev.to/gde/gemini-api-hands-on-6im</link>
      <guid>https://dev.to/gde/gemini-api-hands-on-6im</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F2xmga58mup383o4go36l.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F2xmga58mup383o4go36l.png" alt="image-20260614175257527" width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h1&gt;
  
  
  A Powerful Tool for Asynchronous Processing: Gemini Batch API &amp;amp; Webhooks
&lt;/h1&gt;

&lt;p&gt;When developing LLM-based applications, we often need to handle a large number of data analysis tasks—for example, analyzing reviews from dozens of restaurants at once, classifying a large volume of articles, or batch generating translations. If we use traditional synchronous APIs (real-time calls), we would not only face severe &lt;strong&gt;Rate Limit&lt;/strong&gt; blockages but also fail due to network connection timeouts and extremely high computing costs.&lt;/p&gt;

&lt;p&gt;To overcome this limitation, Google has launched the &lt;strong&gt;Gemini Batch API&lt;/strong&gt; and &lt;strong&gt;Webhook API&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://ai.google.dev/gemini-api/docs/batch-api?hl=zh-tw" rel="noopener noreferrer"&gt;Gemini Batch API&lt;/a&gt;&lt;/strong&gt;: Allows developers to package a large number of requests into a JSONL file and upload them all at once. Gemini performs asynchronous scheduled computations in the background, without consuming your daily real-time API quotas (Rate Limits), and its computing cost is usually half that of real-time APIs, making it a perfect choice for non-urgent big data processing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://ai.google.dev/gemini-api/docs/webhooks?hl=zh-tw" rel="noopener noreferrer"&gt;Webhook API&lt;/a&gt;&lt;/strong&gt;: Traditional Batch tasks require us to constantly write polling logic locally to check the status. With Webhooks, when Gemini completes a Batch computation, it actively sends an HTTP POST callback to your specified URL, instantly notifying you that the task is complete, making the system architecture more elegant and energy-efficient.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This article will document how we integrated these two powerful APIs into our &lt;strong&gt;LINE Bot Restaurant Analysis Assistant&lt;/strong&gt; to achieve one-click deep review and signature dish big data analysis for specific restaurants on mobile devices.&lt;/p&gt;




&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fwww.evanlin.com%2Fimages%2FLINE%25202026-06-14%252017.30.21.tiff" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fwww.evanlin.com%2Fimages%2FLINE%25202026-06-14%252017.30.21.tiff" alt="LINE 2026-06-14 17.30.21" width="800" height="1739"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h1&gt;
  
  
  System Design and Optimized Architecture
&lt;/h1&gt;

&lt;p&gt;Originally, the restaurant analysis function worked by having the Bot list nearby restaurants when a user sent their location, and then providing a generic "Deep Review Analysis (Batch)" button. Clicking it would send all nearby restaurants for analysis at once. However, this led to a poor UX: analyzing all restaurants took too long, and users often only wanted to delve into &lt;strong&gt;one specific restaurant&lt;/strong&gt; they were interested in.&lt;/p&gt;

&lt;p&gt;Therefore, we optimized the function into &lt;strong&gt;dynamic Quick Reply buttons&lt;/strong&gt;:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The user sends their location, and the Bot searches for nearby restaurants via Google Maps Grounding.&lt;/li&gt;
&lt;li&gt;After the client receives a plain text list of restaurants, the Bot automatically uses Gemini to extract the top 3 highest-rated restaurant names.&lt;/li&gt;
&lt;li&gt;Three customized Quick Reply buttons are generated (e.g., &lt;code&gt;🍴 Analyze Din Tai Fung&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;After the user clicks a specific restaurant button, the Bot immediately replies "Processing" to avoid LINE timeouts, and submits the Batch task for that single restaurant in the background. Once Gemini completes the computation, it proactively pushes a dedicated big data report.&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  System Architecture Flow
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;graph TD
    A[User Sends Location] --&amp;gt;|Location Message| B[Google Maps Grounding Search]
    B --&amp;gt;|Plain Text Restaurant List| C[Gemini-2.5-flash Extracts Top 3 Restaurants]
    C --&amp;gt;|Dynamically Generates Quick Reply| D[LINE Bot Replies with 3 Customized Analysis Buttons]
    D --&amp;gt;|User Clicks Specific Analysis| E[FastAPI Background Task]
    E --&amp;gt;|Immediate Reply ACK| F[LINE Chat Message]
    E --&amp;gt;|Package JSONL and Upload| G[Gemini Batch API Submission]
    G --&amp;gt;|Computation Complete Webhook/Polling Callback| H[Proactively Pushes Deep Report to User]

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  Core Implementation
&lt;/h1&gt;

&lt;h3&gt;
  
  
  1. Precisely Extracting Restaurant Names from Grounding Text using Gemini
&lt;/h3&gt;

&lt;p&gt;In &lt;a&gt;tools/maps_tool.py&lt;/a&gt;, the map search returns a plain text string rich in formatting and descriptions. We use Gemini-2.5-flash's structured output concept to precisely extract restaurant names in JSON format:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;        &lt;span class="c1"&gt;# Extract top three restaurant names for Quick Reply
&lt;/span&gt;        &lt;span class="n"&gt;names&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;place_type&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;restaurant&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="n"&gt;extract_prompt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Please extract all restaurant names from the following text and return them in a JSON array format (e.g., [&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s"&gt;Restaurant A&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s"&gt;, &lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s"&gt;Restaurant B&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s"&gt;]). Please output the JSON array directly, without any markdown tags (like ```
&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="o"&gt;%&lt;/span&gt; &lt;span class="n"&gt;endraw&lt;/span&gt; &lt;span class="o"&gt;%&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;
json) or explanatory text.&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                &lt;span class="n"&gt;extract_res&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;models&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;generate_content&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
                    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gemini-2.5-flash&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                    &lt;span class="n"&gt;contents&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;extract_prompt&lt;/span&gt;
                &lt;span class="p"&gt;)&lt;/span&gt;
                &lt;span class="n"&gt;extract_text&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;extract_res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;extract_res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="sh"&gt;""&lt;/span&gt;

                &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                    &lt;span class="n"&gt;names&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;loads&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;extract_text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
                &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;Exception&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                    &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;
                    &lt;span class="n"&gt;array_match&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;search&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;\[(.*?)\]&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;extract_text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;DOTALL&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
                    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;array_match&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                        &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;ast&lt;/span&gt;
                        &lt;span class="n"&gt;names&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;ast&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;literal_eval&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;[&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;array_match&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;group&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;]&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

                &lt;span class="n"&gt;names&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;names&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
                &lt;span class="n"&gt;logger&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;info&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Extracted restaurant names for Quick Reply: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;names&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;Exception&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;e_extract&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="n"&gt;logger&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Failed to extract restaurant names: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;e_extract&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="o"&gt;%&lt;/span&gt; &lt;span class="n"&gt;raw&lt;/span&gt; &lt;span class="o"&gt;%&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;


&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;
&lt;h3&gt;
  
  
  2. Dynamically Generating LINE Quick Reply Buttons
&lt;/h3&gt;

&lt;p&gt;In &lt;a&gt;main.py&lt;/a&gt;, after obtaining the restaurant list, we dynamically generate &lt;code&gt;QuickReplyButton&lt;/code&gt;. We need to pay special attention to LINE API's length limit for button &lt;code&gt;label&lt;/code&gt;:&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;
python
        quick_reply = None
        if place_type == "restaurant" and result.get("status") == "success":
            restaurant_names = result.get("restaurant_names", [])
            if restaurant_names:
                buttons = []
                for name in restaurant_names[:3]:
                    clean_label = name
                    # LINE label limit is 20 characters
                    if len(clean_label) &amp;gt; 10:
                        clean_label = clean_label[:9] + "…"
                    buttons.append(
                        QuickReplyButton(
                            action=PostbackAction(
                                label=f"🍴 分析 {clean_label}",
                                data=json.dumps({
                                    "action": "specific_foodie_deep_analysis",
                                    "restaurant_name": name
                                }),
                                display_text=f"🔍 進行「{name}」深度評論與招牌菜色分析"
                            )
                        )
                    )
                quick_reply = QuickReply(items=buttons)



&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;




&lt;h1&gt;
  
  
  Major Pitfalls and Solutions
&lt;/h1&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F39x43upsykqln99yroez.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F39x43upsykqln99yroez.png" alt="Finder 2026-06-14 17.53.52" width="800" height="1739"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;During the process of connecting this dynamic Quick Reply to the Batch API, we encountered several critical UX and API limitation issues:&lt;/p&gt;

&lt;h3&gt;
  
  
  Pitfall One: LINE 20-character Limit Causing API Sending Errors
&lt;/h3&gt;

&lt;p&gt;Initially, when implementing, we directly used the full restaurant name in the button's Label, for example: &lt;code&gt;🍴 Analyze Love Hot Pot Ultimate Hot Pot&lt;/code&gt;. As a result, the LINE API immediately returned a 400 error, and the message could not be sent at all:&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;
plaintext
LineBotApiError: status_code=400, error_message=The property 'label' must be less than 20 characters.



&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;[Cause Analysis and Solution]&lt;/strong&gt; LINE's official &lt;code&gt;label&lt;/code&gt; limit for Quick Reply is extremely strict; &lt;strong&gt;including emojis and spaces, it can have a maximum of 20 characters&lt;/strong&gt;. To address this, we added a character count check and dynamic truncation mechanism in our code:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;First, the original restaurant name (&lt;code&gt;clean_label&lt;/code&gt;) is truncated: if its length exceeds 10 characters, it is forcibly cut to the first 9 characters and appended with "…" (occupying 10 characters).&lt;/li&gt;
&lt;li&gt;Adding the prefix &lt;code&gt;🍴 Analyze&lt;/code&gt; (a total of 5 characters), the maximum total length becomes 15 characters, safely staying within the 20-character limit, thus eliminating the error!&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Pitfall Two: Batch API Asynchronous Delay and LINE Webhook's "Three-Second Timeout Survival Battle"
&lt;/h3&gt;

&lt;p&gt;When a user clicks the "Analyze Restaurant" button, the Bot must first call Google Search Grounding to collect online reviews for that restaurant, then package the JSONL file and upload it to Gemini to submit the Batch task. This entire sequence usually takes 3 to 8 seconds. However, &lt;strong&gt;the LINE Webhook server requires the Bot to return an HTTP 200 OK response within 3 seconds&lt;/strong&gt;, otherwise it will be deemed a connection failure and re-send the request, leading to severe server congestion.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;[Cause Analysis and Solution]&lt;/strong&gt; We completely asynchronous the processing architecture:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Fast Response&lt;/strong&gt;: When the Bot intercepts a &lt;code&gt;specific_foodie_deep_analysis&lt;/code&gt; Postback action, &lt;strong&gt;it does not execute the analysis directly within the Request flow&lt;/strong&gt;. Instead, it immediately calls LINE's &lt;code&gt;reply_message&lt;/code&gt; to respond to the user: &lt;code&gt;

🔍 Received! Performing deep analysis for you... This will take about 1-2 minutes...&lt;/code&gt;, and then instantly returns HTTP 200 to end that Webhook request.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Background Task Dispatch&lt;/strong&gt;: Use Python &lt;code&gt;asyncio.create_task&lt;/code&gt; to dispatch heavy network search, upload, and submission tasks to FastAPI's background Worker for execution.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Big Data Push&lt;/strong&gt;: When the background Polling listener or Gemini Webhook receives a task completion notification, it then uses LINE's &lt;code&gt;push_message&lt;/code&gt; to proactively send the analysis report to the specific user.&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  Pitfall Three: Gemini Batch API's Queuing and Pending Status
&lt;/h3&gt;

&lt;p&gt;During testing, users sometimes got confused, "Why hasn't there been a reply after three minutes? Is the Bot down?". After checking the system logs, we found that our JSONL file had been successfully uploaded, but the task status on the Gemini server side was stuck at &lt;code&gt;JobState.JOB_STATE_PENDING&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;[Solution]&lt;/strong&gt; This is a characteristic of the Batch API; tasks need to be queued, waiting for Google's server resources. We adopted two major optimizations:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Minimize Workload&lt;/strong&gt;: Reduce the number of restaurants for batch analysis to 1, shrinking the number of request lines in the JSONL to the extreme, to speed up Gemini's scheduling and processing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;UX Optimization and Deduplication Mechanism&lt;/strong&gt;: When a user clicks to analyze, we first check if that user already has a Batch Job running. If so, we reply: &lt;code&gt;⏳ Your deep analysis task is currently running, please wait patiently&lt;/code&gt;, preventing users from submitting multiple duplicate Batch Jobs due to anxious repeated clicks, which would consume unnecessary resources.&lt;/li&gt;
&lt;/ol&gt;




&lt;h1&gt;
  
  
  Results and Benefits
&lt;/h1&gt;

&lt;p&gt;This optimization of Quick Reply and Gemini Batch API for the &lt;strong&gt;LINE Bot Restaurant Assistant&lt;/strong&gt; has achieved excellent practical value:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Highly Customized Mobile Experience&lt;/strong&gt;: After locating, users don't need to type; they can directly click on a restaurant of interest with one tap to precisely get a summary of its signature dishes and review pain points.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Robust Backend Architecture&lt;/strong&gt;: By leveraging asynchronous background tasks and LINE's character limit safety valve, the risks of Webhook timeouts and LINE API errors have been completely resolved.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cost Advantage for Big Data Processing&lt;/strong&gt;: Through the Batch API's half-price advantage and Webhook's proactive callback, while ensuring user experience, it also saves significant computing resources and API costs for the server.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Through this architecture, the LINE Bot truly achieves a low-latency, highly stable big data deep analysis experience on mobile!&lt;/p&gt;

&lt;p&gt;All development code for this project has been open-sourced on GitHub: &lt;a href="https://github.com/kkdai/linebot-helper-python" rel="noopener noreferrer"&gt;kkdai/linebot-helper-python&lt;/a&gt;. Everyone is welcome to deploy and personally test this one-click analysis function, which we believe can bring a higher level of intelligent experience to your LINE Bot projects!&lt;/p&gt;

</description>
      <category>api</category>
      <category>gemini</category>
      <category>llm</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>[I/O Extended Taipei] Building with Gemini APIs: From Calls to Autonomous Systems</title>
      <dc:creator>Evan Lin</dc:creator>
      <pubDate>Sun, 14 Jun 2026 07:12:11 +0000</pubDate>
      <link>https://dev.to/gde/io-extended-taipei-building-cl0</link>
      <guid>https://dev.to/gde/io-extended-taipei-building-cl0</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fn7nklzeowl72fpp27lvn.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fn7nklzeowl72fpp27lvn.png" alt="image-20260612163641980" width="799" height="529"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;(Activity: &lt;a href="https://gdg.community.dev/events/details/google-gdg-taipei-presents-google-io-extended-2026-taipei/" rel="noopener noreferrer"&gt;Google I/O Extended 2026 Taipei&lt;/a&gt; / Presentation: &lt;a href="https://speakerdeck.com/line_developers_tw/building-applications-in-the-gemini-api-family" rel="noopener noreferrer"&gt;SpeakerDeck&lt;/a&gt;)&lt;/p&gt;

&lt;h1&gt;
  
  
  Context: The Gemini API is no longer just "adding one more prompt"
&lt;/h1&gt;

&lt;p&gt;If your impression of the Gemini API is still limited to "select a model, send a prompt, get back a piece of text," then when you see this round of updates in 2026, you'll likely suddenly realize something:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Gemini API has evolved from a simple API interface into a complete platform that can be used to build applications, agents, and asynchronous workflows.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This content is compiled from my talk "Building Applications in the Gemini API Family" at &lt;a href="https://gdg.community.dev/events/details/google-gdg-taipei-presents-google-io-extended-2026-taipei/" rel="noopener noreferrer"&gt;Google I/O Extended 2026 Taipei&lt;/a&gt;. &lt;strong&gt;Evan Lin&lt;/strong&gt;, Technical Director of LINE Taiwan Developer Relations, repeatedly emphasized a core observation at the event: what developers truly need to consider now is no longer just &lt;em&gt;"Should I use Pro or Flash?"&lt;/em&gt;, but rather &lt;em&gt;"How do I string together models, retrieval, agents, callbacks, and cost control into a cohesive system?"&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;In other words, the focus is shifting from &lt;strong&gt;calling APIs&lt;/strong&gt; to &lt;strong&gt;designing systems&lt;/strong&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  First, let's look at the big picture: What's new in the 2026 Gemini API family?
&lt;/h2&gt;

&lt;p&gt;If we view the 2026 Gemini API as a capability map, it can broadly be divided into three layers.&lt;/p&gt;

&lt;h3&gt;
  
  
  Layer 1: Core Models
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Gemini 3.5 Pro&lt;/strong&gt;: Strongest reasoning capability, suitable for complex planning, advanced analysis, and multi-step tasks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gemini 3.5 Flash&lt;/strong&gt;: Main model, best balance of speed, cost, and capability, suitable for most product traffic.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Flash-Lite&lt;/strong&gt;: Intent classifier and pre-classifier for high-frequency, low-cost scenarios.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gemini Embedding 2&lt;/strong&gt;: Supports not only text but also multi-modal vectorization needs.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Layer 2: Key Capability Modules
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Retrieval&lt;/strong&gt;: File Search, Google Search Grounding, URL Context.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Agent / Async&lt;/strong&gt;: Agents API, Webhook, Deep Research agent.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Infrastructure&lt;/strong&gt;: Context caching, Batch API, Live API.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Layer 3: System Design Approach
&lt;/h3&gt;

&lt;p&gt;This layer is arguably the most important. Because once the above capabilities are offered as platform services, many "intermediate layers" that previously had to be built manually suddenly disappear:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;No longer necessarily need to build your own RAG pipeline.&lt;/li&gt;
&lt;li&gt;No longer necessarily need to maintain your own agent loop.&lt;/li&gt;
&lt;li&gt;No longer necessarily need to block the main server with polling while waiting for results.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Core Observation&lt;/strong&gt;: The Gemini API upgrade is not just about "stronger models"; it's about &lt;strong&gt;Google absorbing the complexities that were originally at the application layer into the platform layer&lt;/strong&gt;. This will directly change how we design AI systems.&lt;/p&gt;




&lt;h1&gt;
  
  
  Architectural Turning Point: Three Tools, Three Paradigm Shifts
&lt;/h1&gt;

&lt;p&gt;What's most worth repeatedly digesting from this talk are the architectural changes represented by these three tools.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. File Search: Shifting from Hand-Coded RAG to Managed RAG
&lt;/h2&gt;

&lt;p&gt;Previously, when discussing enterprise knowledge Q&amp;amp;A, the immediate thought was:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Chunking.&lt;/li&gt;
&lt;li&gt;Creating embeddings.&lt;/li&gt;
&lt;li&gt;Storing in a vector DB.&lt;/li&gt;
&lt;li&gt;Writing retrieval code.&lt;/li&gt;
&lt;li&gt;Then manually adding citation and permission control.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Now, with the advent of File Search, developers can focus more on "how documents are governed, how permissions are allocated, and how answers are presented," rather than repeatedly writing that foundational infrastructure.&lt;/p&gt;

&lt;p&gt;More importantly, it doesn't just search text.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why is this File Search particularly noteworthy?
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Images and text in the same space&lt;/strong&gt;: Screenshots, charts, and mixed text-image layouts in PDFs are no longer just attachments, but content understandable by the model.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Metadata filtering&lt;/strong&gt;: Can filter by department, system, and document type, which is crucial for internal enterprise knowledge retrieval.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Precise citation&lt;/strong&gt;: Can refer back to specific page numbers and grounding metadata, making answers more trustworthy.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This represents a very practical shift: much of the time enterprises previously spent on LangChain, vector databases, and chunking strategies can now largely be redirected towards &lt;strong&gt;permission design, UX, and content governance&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Agents API: Shifting from Client-Side Loop to Server-Side Managed Agent
&lt;/h2&gt;

&lt;p&gt;In the past, to build an agent, the common approach was to maintain your own ReAct or tool loop:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Model decides the next step.&lt;/li&gt;
&lt;li&gt;Calls a tool.&lt;/li&gt;
&lt;li&gt;Receives results.&lt;/li&gt;
&lt;li&gt;Feeds back to the model.&lt;/li&gt;
&lt;li&gt;Repeats until completion.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The problem is that this is full of engineering details: state preservation, timeouts, retries, background execution, long-task monitoring. Ultimately, you'd find yourself spending most of your time maintaining an "agent runtime."&lt;/p&gt;

&lt;p&gt;What the Agents API changes is that you can POST a task to Gemini, allowing it to complete the long process on the server side, even handling complex tasks that take up to 20 minutes.&lt;/p&gt;

&lt;p&gt;The significance behind this is not just "more convenient"; it means developers can finally refocus on:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;How are tasks defined?&lt;/li&gt;
&lt;li&gt;Which tools can be used?&lt;/li&gt;
&lt;li&gt;What are the success criteria?&lt;/li&gt;
&lt;li&gt;How should the product integrate the results when they return?&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  3. Webhook: Shifting from Polling to Event-Driven
&lt;/h2&gt;

&lt;p&gt;Once tasks might run for several minutes, or even more than ten minutes, traditional synchronous requests become unreasonable.&lt;/p&gt;

&lt;p&gt;Therefore, the role of Webhook is actually crucial: it's not a minor feature, but a prerequisite for the entire agent workflow to truly enter production. When Gemini completes a task and actively POSTs the result back to your server, your system can become event-driven:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The frontend first responds to the user with "Task received."&lt;/li&gt;
&lt;li&gt;The Agents API executes in the background.&lt;/li&gt;
&lt;li&gt;Upon completion, the result is pushed back via webhook.&lt;/li&gt;
&lt;li&gt;Your service then notifies the user, updates the database, or triggers the next step.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is particularly important for high-concurrency products, as you finally don't need to hold a bunch of server connections idly waiting.&lt;/p&gt;




&lt;h1&gt;
  
  
  From the Perspective of a LINE Bot, How Should a Gemini Application Be Designed?
&lt;/h1&gt;

&lt;p&gt;A very practical suggestion Evan gave in his talk is to &lt;strong&gt;place a router layer before the LLM&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;This design sounds simple, but it largely determines your cost, latency, and predictability.&lt;/p&gt;

&lt;h2&gt;
  
  
  A Very Pragmatic Routing Approach
&lt;/h2&gt;

&lt;p&gt;First, use the inexpensive &lt;strong&gt;Flash-Lite&lt;/strong&gt; for intent routing:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Quick Q&amp;amp;A&lt;/strong&gt;: Directly handed over to Flash or Flash-Lite for generation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Query company documents&lt;/strong&gt;: Enters File Search.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Complex long tasks&lt;/strong&gt;: Enters Agents API.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Doing this has three benefits:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Cost control first&lt;/strong&gt;: Not every query directly hits the most expensive, heaviest model.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Latency control first&lt;/strong&gt;: Simple requests should not mistakenly enter long processes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;System behavior control first&lt;/strong&gt;: Makes the overall process more stable than "throwing everything at a large model for improvisation."&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If you're building a LINE Bot, customer service assistant, internal knowledge assistant, or workflow agent, this router should almost certainly be the default configuration, rather than an afterthought.&lt;/p&gt;




&lt;h2&gt;
  
  
  Infrastructure is Not Unimportant, But You Don't Have to Rebuild it Yourself Every Time
&lt;/h2&gt;

&lt;p&gt;Another strong message from this talk is that developers' time should be reallocated.&lt;/p&gt;

&lt;p&gt;Previously, much of the man-hours in many AI projects were actually consumed by these tasks:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Vector database operations and maintenance&lt;/li&gt;
&lt;li&gt;Chunking and retrieval parameter tuning&lt;/li&gt;
&lt;li&gt;Long-task scheduling&lt;/li&gt;
&lt;li&gt;Websocket / polling / callback processes&lt;/li&gt;
&lt;li&gt;Token cost optimization&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Now, with File Search, Agents API, Webhook, Context caching, and Batch API, the areas where we should spend more time have shifted to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Business rules and tool boundaries&lt;/li&gt;
&lt;li&gt;Document permissions and data governance&lt;/li&gt;
&lt;li&gt;User interaction experience&lt;/li&gt;
&lt;li&gt;Task decomposition and routing strategies&lt;/li&gt;
&lt;li&gt;Failure recovery and result interpretability&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is also why I strongly agree with Evan's underlying message: &lt;strong&gt;What's truly valuable is not whether you can build your own vector database, but whether you can redirect 80% of your energy back to the product's core.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Three Most Valuable Practical Takeaways
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Place a routing layer before the LLM
&lt;/h3&gt;

&lt;p&gt;Don't send all problems directly to the same model. First classify, then decide whether to generate, retrieve, or enter an agent task.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Embrace asynchronous operations; don't force long tasks into synchronous APIs
&lt;/h3&gt;

&lt;p&gt;If a task might take more than a few seconds, you should seriously consider Agents API + Webhook. This is not an optimization; it's an architectural correctness issue.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Redirect RAG engineering time to permissions and experience
&lt;/h3&gt;

&lt;p&gt;When File Search can handle a large amount of foundational work, developers should be more concerned with: can data be securely queried, can answers be verified, and can citations be trusted by users.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why is this talk worth revisiting repeatedly?
&lt;/h2&gt;

&lt;p&gt;Because it highlights a turning point that many teams are currently facing:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;We are no longer just writing prompts for LLMs; we are designing operating systems for AI applications.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Models are certainly still at the core, but what truly differentiates products is increasingly not "which model you choose," but:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;How you decide when to use which capability.&lt;/li&gt;
&lt;li&gt;How you make the system run reliably for extended periods.&lt;/li&gt;
&lt;li&gt;How you make answers traceable, verifiable, and maintainable.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you still understand generative AI using the 2024 approach of "a single chat endpoint for everything," then you'll easily underestimate the 2026 Gemini API family.&lt;/p&gt;




&lt;h2&gt;
  
  
  Postscript: From API User to AI System Designer
&lt;/h2&gt;

&lt;p&gt;The most valuable aspect of this "Building Applications in the Gemini API Family" talk is not teaching you another new parameter or SDK, but reminding everyone of a more fundamental shift:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The competitiveness of the next phase will not be about who is better at calling models, but who is better at assembling models, retrieval, agents, and event flows into a truly functional system.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If you are working on a LINE Bot, enterprise knowledge base, internal assistant, customer service process, or any product requiring multi-step AI collaboration, this architectural perspective is well worth using to redraw your current system diagram.&lt;/p&gt;

&lt;p&gt;Often, what truly needs refactoring is not the prompt, but the entire pipeline.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>api</category>
      <category>gemini</category>
      <category>google</category>
    </item>
    <item>
      <title>[Hands-on Gemini 3.5 Live</title>
      <dc:creator>Evan Lin</dc:creator>
      <pubDate>Fri, 12 Jun 2026 06:09:59 +0000</pubDate>
      <link>https://dev.to/gde/hands-on-gemini-35-live-3dh6</link>
      <guid>https://dev.to/gde/hands-on-gemini-35-live-3dh6</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F6x1mkub62aoeiv68idtq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F6x1mkub62aoeiv68idtq.png" alt="image-20260610144830233" width="800" height="326"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h1&gt;
  
  
  Brand New API Unveiled: Gemini 3.5 Live Translate
&lt;/h1&gt;

&lt;p&gt;On June 9, 2026, Google officially released its brand new real-time voice translation model — &lt;strong&gt;Gemini 3.5 Live Translate&lt;/strong&gt;. This marks another significant breakthrough for Google in AI voice translation technology. It is currently available for public preview to developers in Google AI Studio and Gemini Live API, and has been simultaneously integrated into services like Google Translate and Google Meet.&lt;/p&gt;

&lt;p&gt;Key features of Gemini 3.5 Live Translate include:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt; &lt;strong&gt;Fluent and Natural Bidirectional Voice Translation&lt;/strong&gt;: Supports over 70 languages, automatically detecting the input voice language without manual configuration.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Continuous Stream Generation (Instead of Single-Sentence Turn-Taking)&lt;/strong&gt;: Unlike previous turn-by-turn systems that required the speaker to finish speaking before translation, Gemini 3.5 Live Translate generates translations in real-time while listening. It strikes a balance between contextual understanding and immediacy, with translations lagging only a few seconds behind the speaker, completely avoiding awkward pauses.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Preservation of Intonation and Rhythm&lt;/strong&gt;: The generated voice is not only smooth but also retains the original speaker's tone, intonation, and speaking rhythm.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Robust Noise Cancellation Capability&lt;/strong&gt;: Accurately captures and recognizes speech even in noisy or unstable environments.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This article will document how we developed a native macOS application, &lt;strong&gt;MeetingTranslator&lt;/strong&gt;, using Swift, to integrate with this powerful new API and achieve real-time translation of specific app audio into Traditional Chinese voice and subtitles.&lt;/p&gt;




&lt;h1&gt;
  
  
  System Design and Architecture
&lt;/h1&gt;

&lt;p&gt;Our goal is to develop a Native SwiftUI application that does not require installing virtual sound cards like BlackHole. Instead, it utilizes Apple's official &lt;strong&gt;ScreenCaptureKit&lt;/strong&gt; framework to directly capture the audio stream from a selected application (such as YouTube in Google Chrome or an online meeting) and, through the &lt;strong&gt;Gemini Live WebSocket API&lt;/strong&gt;, achieve ultra-low-latency conversational voice translation.&lt;/p&gt;

&lt;h3&gt;
  
  
  System Architecture Flow
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight dot"&gt;&lt;code&gt;&lt;span class="k"&gt;graph&lt;/span&gt; &lt;span class="nv"&gt;TD&lt;/span&gt;
    &lt;span class="nv"&gt;A&lt;/span&gt;&lt;span class="o"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;ScreenCaptureKit&lt;/span&gt; &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;br&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;&lt;span class="nv"&gt;Capture&lt;/span&gt; &lt;span class="nv"&gt;Application&lt;/span&gt; &lt;span class="nv"&gt;Audio&lt;/span&gt;&lt;span class="o"&gt;]&lt;/span&gt; &lt;span class="o"&gt;--&lt;/span&gt;&lt;span class="err"&gt;&amp;gt;|&lt;/span&gt;&lt;span class="mi"&gt;48&lt;/span&gt;&lt;span class="nv"&gt;kHz&lt;/span&gt; &lt;span class="nv"&gt;Stereo&lt;/span&gt; &lt;span class="nv"&gt;Float32&lt;/span&gt;&lt;span class="err"&gt;|&lt;/span&gt; &lt;span class="nv"&gt;B&lt;/span&gt;&lt;span class="o"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;AVAudioConverter&lt;/span&gt; &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;br&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;&lt;span class="nv"&gt;Resampling&lt;/span&gt; &lt;span class="nv"&gt;and&lt;/span&gt; &lt;span class="nv"&gt;Channel&lt;/span&gt; &lt;span class="nv"&gt;Conversion&lt;/span&gt;&lt;span class="o"&gt;]&lt;/span&gt;
    &lt;span class="nv"&gt;B&lt;/span&gt; &lt;span class="o"&gt;--&lt;/span&gt;&lt;span class="err"&gt;&amp;gt;|&lt;/span&gt;&lt;span class="mi"&gt;16&lt;/span&gt;&lt;span class="nv"&gt;kHz&lt;/span&gt; &lt;span class="nv"&gt;Mono&lt;/span&gt; &lt;span class="nv"&gt;Int16&lt;/span&gt; &lt;span class="nv"&gt;PCM&lt;/span&gt;&lt;span class="err"&gt;|&lt;/span&gt; &lt;span class="nv"&gt;C&lt;/span&gt;&lt;span class="o"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;Gemini&lt;/span&gt; &lt;span class="nv"&gt;Live&lt;/span&gt; &lt;span class="nv"&gt;API&lt;/span&gt; &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;br&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;&lt;span class="nv"&gt;WebSocket&lt;/span&gt; &lt;span class="nv"&gt;Connection&lt;/span&gt;&lt;span class="o"&gt;]&lt;/span&gt;
    &lt;span class="nv"&gt;C&lt;/span&gt; &lt;span class="o"&gt;--&lt;/span&gt;&lt;span class="err"&gt;&amp;gt;|&lt;/span&gt;&lt;span class="nv"&gt;Real&lt;/span&gt;&lt;span class="err"&gt;-&lt;/span&gt;&lt;span class="nv"&gt;time&lt;/span&gt; &lt;span class="nv"&gt;Subtitle&lt;/span&gt; &lt;span class="nv"&gt;Recognition&lt;/span&gt;&lt;span class="err"&gt;|&lt;/span&gt; &lt;span class="nv"&gt;D&lt;/span&gt;&lt;span class="o"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;SwiftUI&lt;/span&gt; &lt;span class="nv"&gt;Subtitle&lt;/span&gt; &lt;span class="nv"&gt;HUD&lt;/span&gt; &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;br&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;&lt;span class="nv"&gt;Traditional&lt;/span&gt; &lt;span class="nv"&gt;Chinese&lt;/span&gt; &lt;span class="nv"&gt;Bilingual&lt;/span&gt; &lt;span class="nv"&gt;Subtitles&lt;/span&gt;&lt;span class="o"&gt;]&lt;/span&gt;
    &lt;span class="nv"&gt;C&lt;/span&gt; &lt;span class="o"&gt;--&lt;/span&gt;&lt;span class="err"&gt;&amp;gt;|&lt;/span&gt;&lt;span class="mi"&gt;24&lt;/span&gt;&lt;span class="nv"&gt;kHz&lt;/span&gt; &lt;span class="nv"&gt;Mono&lt;/span&gt; &lt;span class="nv"&gt;Int16&lt;/span&gt; &lt;span class="nv"&gt;PCM&lt;/span&gt; &lt;span class="nv"&gt;Translated&lt;/span&gt; &lt;span class="nv"&gt;Audio&lt;/span&gt;&lt;span class="err"&gt;|&lt;/span&gt; &lt;span class="nv"&gt;E&lt;/span&gt;&lt;span class="o"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;AudioPlaybackManager&lt;/span&gt; &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;br&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;&lt;span class="nv"&gt;AVAudioEngine&lt;/span&gt; &lt;span class="nv"&gt;Player&lt;/span&gt;&lt;span class="o"&gt;]&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  Core Implementation One: ScreenCaptureKit Capture and Resampling
&lt;/h1&gt;

&lt;p&gt;&lt;strong&gt;ScreenCaptureKit&lt;/strong&gt;, introduced in macOS 13, frees developers from the pain of relying on kernel audio virtual devices, allowing precise filtering and recording of specific application screens and audio.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Filter and Select Target App
&lt;/h3&gt;

&lt;p&gt;We use &lt;code&gt;SCShareableContent&lt;/code&gt; to get currently running applications on the system and filter out background services without names and system-自带 services:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight swift"&gt;&lt;code&gt;&lt;span class="kd"&gt;func&lt;/span&gt; &lt;span class="nf"&gt;fetchShareableApps&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="kt"&gt;SCRunningApplication&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;do&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;content&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="kt"&gt;SCShareableContent&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;current&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;applications&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;filter&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="n"&gt;app&lt;/span&gt; &lt;span class="k"&gt;in&lt;/span&gt;
            &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;name&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;app&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;applicationName&lt;/span&gt;
            &lt;span class="k"&gt;guard&lt;/span&gt; &lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;isEmpty&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
            &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;bundleId&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;app&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;bundleIdentifier&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="n"&gt;bundleId&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;hasPrefix&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"com.apple.system"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="n"&gt;bundleId&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="kt"&gt;Bundle&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;main&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;bundleIdentifier&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;sorted&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;$0&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;applicationName&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="nv"&gt;$1&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;applicationName&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"無法獲取可共享內容: &lt;/span&gt;&lt;span class="se"&gt;\(&lt;/span&gt;&lt;span class="n"&gt;error&lt;/span&gt;&lt;span class="se"&gt;)&lt;/span&gt;&lt;span class="s"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  2. Start Audio Capture Stream
&lt;/h3&gt;

&lt;p&gt;After filtering out the target App (e.g., Google Chrome), we create an &lt;code&gt;SCContentFilter&lt;/code&gt; for it and apply it to &lt;code&gt;SCStream&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight swift"&gt;&lt;code&gt;&lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;appFilter&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kt"&gt;SCContentFilter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;display&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;displays&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;first&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;including&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;targetApp&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="nv"&gt;exceptingWindows&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[])&lt;/span&gt;
&lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;config&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kt"&gt;SCStreamConfiguration&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;capturesAudio&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
&lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;width&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;32&lt;/span&gt; &lt;span class="c1"&gt;// When only capturing audio, set video frame to minimal to save performance&lt;/span&gt;
&lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;height&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;32&lt;/span&gt;

&lt;span class="n"&gt;stream&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kt"&gt;SCStream&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;filter&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;appFilter&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;configuration&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;delegate&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;nil&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="n"&gt;stream&lt;/span&gt;&lt;span class="p"&gt;?&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;addStreamOutput&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;type&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;audio&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;sampleHandlerQueue&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kt"&gt;DispatchQueue&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;label&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;"com.translator.audioQueue"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;stream&lt;/span&gt;&lt;span class="p"&gt;?&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;startCapture&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  Core Implementation Two: Gemini Live WebSocket Bidirectional Connection
&lt;/h1&gt;

&lt;p&gt;The core of the Gemini Live API lies in using a &lt;code&gt;wss://&lt;/code&gt; connection to transmit microphone/application audio in real-time through a single channel, and simultaneously receive model-generated translated text and translated audio.&lt;/p&gt;

&lt;p&gt;In &lt;a&gt;GeminiLiveConnection.swift&lt;/a&gt;, we maintain this bidirectional pipeline via &lt;code&gt;URLSessionWebSocketTask&lt;/code&gt;. After connecting, a &lt;code&gt;setup&lt;/code&gt; control message must be sent immediately to initialize the model configuration.&lt;/p&gt;




&lt;h1&gt;
  
  
  Major Pitfalls and Solutions
&lt;/h1&gt;

&lt;p&gt;During the process of integrating the system, we encountered three blocking difficulties. Below is our troubleshooting process and solutions:&lt;/p&gt;

&lt;h3&gt;
  
  
  Pitfall One: Gemini Live Exclusive Model Restrictions
&lt;/h3&gt;

&lt;p&gt;Initially, we tried to use standard REST API model names (e.g., &lt;code&gt;gemini-3.5-flash&lt;/code&gt;) in the WebSocket connection, but the server immediately disconnected:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;❌ WebSocket 被 Gemini 伺服器關閉 (CloseCode: 1008, 原因: models/gemini-3.5-flash is not found for API version v1beta, or is not supported for bidiGenerateContent.)

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;【Solution】&lt;/strong&gt; Gemini's bidirectional Live API currently only supports specific optimized real-time models. We must restrict the model field to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;code&gt;gemini-2.0-flash-exp&lt;/code&gt; (standard bidirectional conversation)&lt;/li&gt;
&lt;li&gt;  &lt;code&gt;gemini-3.5-live-translate-preview&lt;/code&gt; (preview model optimized for real-time translation)&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Pitfall Two: Incorrect JSON Payload Field Structure (Hidden Differences Between Documentation and API Versions)
&lt;/h3&gt;

&lt;p&gt;When configuring real-time interpretation, we referred to Google's official documentation and placed the &lt;code&gt;inputAudioTranscription&lt;/code&gt; (input speech-to-text) and &lt;code&gt;outputAudioTranscription&lt;/code&gt; (output speech-to-text) fields within &lt;code&gt;generationConfig&lt;/code&gt;, which resulted in a 1007 error:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;❌ WebSocket 被 Gemini 伺服器關閉 (CloseCode: 1007, 原因: Invalid JSON payload received. Unknown name "inputAudioTranscription" at 'setup.generation_config': Cannot find field.)

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;【Cause Analysis and Solution】&lt;/strong&gt; In the official documentation, for &lt;code&gt;v1alpha&lt;/code&gt; and client SDKs (e.g., JavaScript / Python SDK), these two fields are wrapped within &lt;code&gt;generationConfig&lt;/code&gt;. However, in the current &lt;code&gt;v1beta&lt;/code&gt; WebSocket native endpoint: &lt;code&gt;/ws/google.ai.generativelanguage.v1beta.GenerativeService.BidiGenerateContent&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;These two fields should be located at the &lt;strong&gt;root level&lt;/strong&gt; of the &lt;code&gt;setup&lt;/code&gt; object, while the translation-specific &lt;code&gt;translationConfig&lt;/code&gt; must be placed under &lt;code&gt;generationConfig&lt;/code&gt;. The correct JSON Payload structure is as follows:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight swift"&gt;&lt;code&gt;&lt;span class="n"&gt;setupMessage&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="s"&gt;"setup"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="s"&gt;"model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;"models/&lt;/span&gt;&lt;span class="se"&gt;\(&lt;/span&gt;&lt;span class="n"&gt;modelName&lt;/span&gt;&lt;span class="se"&gt;)&lt;/span&gt;&lt;span class="s"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="s"&gt;"inputAudioTranscription"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[:],&lt;/span&gt; &lt;span class="c1"&gt;// Enable real-time input subtitles, placed at the setup root&lt;/span&gt;
        &lt;span class="s"&gt;"outputAudioTranscription"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[:],&lt;/span&gt; &lt;span class="c1"&gt;// Enable real-time output subtitles, placed at the setup root&lt;/span&gt;
        &lt;span class="s"&gt;"generationConfig"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
            &lt;span class="s"&gt;"responseModalities"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s"&gt;"AUDIO"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
            &lt;span class="s"&gt;"translationConfig"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
                &lt;span class="s"&gt;"targetLanguageCode"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;"zh-TW"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;// Set target translation language to Traditional Chinese&lt;/span&gt;
                &lt;span class="s"&gt;"echoTargetLanguage"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
            &lt;span class="p"&gt;]&lt;/span&gt;
        &lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="p"&gt;]&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After this modification, the WebSocket setup finally successfully handshaked and no longer crashed!&lt;/p&gt;

&lt;h3&gt;
  
  
  Pitfall Three: "Zero-Byte Silence" Caused by Multi-Channel Stereo Capture
&lt;/h3&gt;

&lt;p&gt;After successfully establishing the WebSocket pipeline and starting to push resampled audio, we found that Gemini still had no translation response. Observing the log output, we discovered that the content of the sent audio blocks was all &lt;code&gt;0&lt;/code&gt; (Silence):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;📊 [WebSocket] 已發送 500 個音訊區塊 | 大小: 640 bytes | 是否為靜音(全0): true

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;【Cause Analysis】&lt;/strong&gt; When the captured object (e.g., Google Chrome playing a YouTube video) outputs stereo (2 Channels) or multi-channel audio, our original method for converting &lt;code&gt;CMSampleBuffer&lt;/code&gt; to &lt;code&gt;AVAudioPCMBuffer&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight swift"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Old method: Directly assumes a single Channel pointer and copies&lt;/span&gt;
&lt;span class="k"&gt;var&lt;/span&gt; &lt;span class="nv"&gt;audioBufferList&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kt"&gt;AudioBufferList&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="k"&gt;var&lt;/span&gt; &lt;span class="nv"&gt;blockBuffer&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kt"&gt;CMBlockBuffer&lt;/span&gt;&lt;span class="p"&gt;?&lt;/span&gt;
&lt;span class="kt"&gt;CMSampleBufferGetAudioBufferListWithRetainedBlockBuffer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;...&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;audioBufferList&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;...&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In a multi-channel environment, this would lead to insufficient memory allocation, causing copy interruption or fill failure, resulting in all subsequent audio resampler (AVAudioConverter) inputs being null values (silence).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;【Solution】&lt;/strong&gt; It is necessary to use the &lt;strong&gt;Double-Call technique&lt;/strong&gt; to dynamically allocate memory space for &lt;code&gt;AudioBufferList&lt;/code&gt;:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt; &lt;strong&gt;First Call&lt;/strong&gt;: Pass &lt;code&gt;nil&lt;/code&gt; as the buffer output, used only to precisely query the required physical memory size (&lt;code&gt;bufferListSizeNeededOut&lt;/code&gt;) for that &lt;code&gt;sampleBuffer&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Memory Allocation&lt;/strong&gt;: Use &lt;code&gt;UnsafeMutablePointer&amp;lt;AudioBufferList&amp;gt;.allocate&lt;/code&gt; to dynamically allocate space based on the queried size.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Second Call&lt;/strong&gt;: Pass the allocated pointer to safely fill in multi-channel audio data.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Channel Reassembly&lt;/strong&gt;: Based on the multi-channel format (Interleaved/Non-Interleaved), precisely use &lt;code&gt;memcpy&lt;/code&gt; to copy the corresponding data segments into a temporary buffer, then send it to the converter for noise reduction and downsampling.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Core code correction:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight swift"&gt;&lt;code&gt;&lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kd"&gt;func&lt;/span&gt; &lt;span class="nf"&gt;audioBufferFromSampleBuffer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="nv"&gt;sampleBuffer&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kt"&gt;CMSampleBuffer&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;asbd&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kt"&gt;AudioStreamBasicDescription&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="kt"&gt;AVAudioPCMBuffer&lt;/span&gt;&lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;guard&lt;/span&gt; &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;sourceFormat&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;sourceFormat&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="kc"&gt;nil&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="c1"&gt;// 1. Dynamically get the required AudioBufferList memory size&lt;/span&gt;
    &lt;span class="k"&gt;var&lt;/span&gt; &lt;span class="nv"&gt;bufferListSize&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
    &lt;span class="k"&gt;var&lt;/span&gt; &lt;span class="nv"&gt;status&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kt"&gt;CMSampleBufferGetAudioBufferListWithRetainedBlockBuffer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;sampleBuffer&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="nv"&gt;bufferListSizeNeededOut&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;bufferListSize&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="nv"&gt;bufferListOut&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;nil&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="nv"&gt;bufferListSize&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="nv"&gt;blockBufferAllocator&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;nil&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="nv"&gt;blockBufferMemoryAllocator&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;nil&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="nv"&gt;flags&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="nv"&gt;blockBufferOut&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;nil&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;guard&lt;/span&gt; &lt;span class="n"&gt;status&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;noErr&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="kc"&gt;nil&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="c1"&gt;// 2. Allocate a pointer with sufficient space and fill it&lt;/span&gt;
    &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;bufferListPointer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kt"&gt;UnsafeMutablePointer&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="kt"&gt;AudioBufferList&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;.&lt;/span&gt;&lt;span class="nf"&gt;allocate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;capacity&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;bufferListSize&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;defer&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="n"&gt;bufferListPointer&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;deallocate&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="k"&gt;var&lt;/span&gt; &lt;span class="nv"&gt;blockBuffer&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kt"&gt;CMBlockBuffer&lt;/span&gt;&lt;span class="p"&gt;?&lt;/span&gt;
    &lt;span class="n"&gt;status&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kt"&gt;CMSampleBufferGetAudioBufferListWithRetainedBlockBuffer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;sampleBuffer&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="nv"&gt;bufferListSizeNeededOut&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;nil&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="nv"&gt;bufferListOut&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;bufferListPointer&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="nv"&gt;bufferListSize&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;bufferListSize&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="nv"&gt;blockBufferAllocator&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;nil&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="nv"&gt;blockBufferMemoryAllocator&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;nil&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="nv"&gt;flags&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="nv"&gt;blockBufferOut&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;blockBuffer&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;guard&lt;/span&gt; &lt;span class="n"&gt;status&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;noErr&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="kc"&gt;nil&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="c1"&gt;// 3. Create an AVAudioPCMBuffer conforming to the source format and safely copy...&lt;/span&gt;
    &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;frameCount&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kt"&gt;AVAudioFrameCount&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;CMSampleBufferGetNumSamples&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sampleBuffer&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="k"&gt;guard&lt;/span&gt; &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;pcmBuffer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kt"&gt;AVAudioPCMBuffer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;pcmFormat&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;sourceFormat&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;frameCapacity&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;frameCount&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="kc"&gt;nil&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="n"&gt;pcmBuffer&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;frameLength&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;frameCount&lt;/span&gt;

    &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;audioBuffers&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kt"&gt;UnsafeMutableAudioBufferListPointer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;bufferListPointer&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;index&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;audioBuffer&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="n"&gt;audioBuffers&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;enumerated&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;guard&lt;/span&gt; &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;mData&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;audioBuffer&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;mData&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;index&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="kt"&gt;Int&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sourceFormat&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;channelCount&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="k"&gt;continue&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
        &lt;span class="c1"&gt;// Differentiate between non-interleaved and interleaved formats for copying&lt;/span&gt;
        &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;isNonInterleaved&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;asbd&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;mFormatFlags&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt; &lt;span class="n"&gt;kAudioFormatFlagIsNonInterleaved&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;isNonInterleaved&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;dst&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;pcmBuffer&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;int16ChannelData&lt;/span&gt;&lt;span class="p"&gt;?[&lt;/span&gt;&lt;span class="n"&gt;index&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
                &lt;span class="nf"&gt;memcpy&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;dst&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;mData&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="kt"&gt;Int&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;audioBuffer&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;mDataByteSize&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
            &lt;span class="p"&gt;}&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;dst&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;pcmBuffer&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;int16ChannelData&lt;/span&gt;&lt;span class="p"&gt;?[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
                &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;offset&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;index&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="kt"&gt;Int&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;frameCount&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
                &lt;span class="nf"&gt;memcpy&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;dst&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;advanced&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;by&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;offset&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;mData&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="kt"&gt;Int&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;audioBuffer&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;mDataByteSize&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
            &lt;span class="p"&gt;}&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;pcmBuffer&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After applying this refactoring, when we played a test video on Chrome's YouTube again, the console finally printed: &lt;code&gt;是否為靜音(全0): false&lt;/code&gt;, and we successfully received Gemini's real-time voice feedback!&lt;/p&gt;




&lt;h1&gt;
  
  
  Results and Benefits
&lt;/h1&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fih2tjzsuyw067ix5845a.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fih2tjzsuyw067ix5845a.png" alt="image-20260610144945151" width="800" height="626"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Full development repo: &lt;a href="https://github.com/kkdai/gemini-live-translate-macos" rel="noopener noreferrer"&gt;https://github.com/kkdai/gemini-live-translate-macos&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Through this architectural upgrade and bug fixes, &lt;strong&gt;MeetingTranslator&lt;/strong&gt; has demonstrated excellent practical value:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt; &lt;strong&gt;Zero External Device Dependency&lt;/strong&gt;: No need to set up complex routing like BlackHole or Loopback; it works out of the box.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Accurate and Real-time Subtitles&lt;/strong&gt;: The Gemini Live API can complete English to Traditional Chinese translation within hundreds of milliseconds, smoothly displaying the results in a HUD floating window.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Synchronized Voice Translation Broadcast&lt;/strong&gt;: Through &lt;code&gt;AudioPlaybackManager&lt;/code&gt;, users can listen to the original meeting while simultaneously hearing high-quality 24kHz Traditional Chinese interpretation in their headphones.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;We hope this record of pitfalls encountered with macOS Core Audio / ScreenCaptureKit and the Gemini WebSocket API can provide valuable reference for developers also exploring AI real-time voice applications!&lt;/p&gt;

</description>
      <category>ai</category>
      <category>api</category>
      <category>gemini</category>
      <category>google</category>
    </item>
    <item>
      <title>[AI Practice] Building blazing-Fast AI Mac OS App with Antigravity CLI</title>
      <dc:creator>Evan Lin</dc:creator>
      <pubDate>Fri, 12 Jun 2026 06:09:40 +0000</pubDate>
      <link>https://dev.to/gde/ai-practice-blazing-fast-ai-co-29l7</link>
      <guid>https://dev.to/gde/ai-practice-blazing-fast-ai-co-29l7</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Folr85lchvp9197zg9lvg.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Folr85lchvp9197zg9lvg.png" alt="image-20260612102252662" width="800" height="622"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h1&gt;
  
  
  Foreword: A Developer's New Collaboration Model
&lt;/h1&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fs0iuas8puugsm23a61hw.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fs0iuas8puugsm23a61hw.png" alt="image-20260612102436750" width="800" height="212"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Imagine this scenario: you are developing a real-time meeting translation App that combines macOS low-level audio (CoreAudio/ScreenCaptureKit) with Gemini Live API WebSocket. During the testing phase, the program suddenly crashed with an error, and the audio stream produced a complete silence of all zeros.&lt;/p&gt;

&lt;p&gt;In the past, your troubleshooting process might have been:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Open the terminal and retrieve the log file.&lt;/li&gt;
&lt;li&gt;Copy the entire error message and relevant code.&lt;/li&gt;
&lt;li&gt;Switch to the browser, open an AI chat window, paste it, and ask for the reason.&lt;/li&gt;
&lt;li&gt;After receiving modification suggestions, copy them back to the editor and test manually.&lt;/li&gt;
&lt;li&gt;Repeat the above steps until fixed, then manually write &lt;code&gt;README.md&lt;/code&gt;, write a blog post, create a GitHub repository, commit the code, and push it.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;In this development cycle, we adopted the &lt;strong&gt;AGY CLI (Antigravity-CLI)&lt;/strong&gt; agent designed by Google DeepMind. We were surprised to find that all the tedious context switching mentioned above could be &lt;strong&gt;fully automated&lt;/strong&gt; through conversations with the intelligent agent within the terminal. This article will reconstruct the actual Prompt dialogue flow and share how we collaborated with AGY CLI to build a macOS meeting translation App from scratch.&lt;/p&gt;




&lt;h1&gt;
  
  
  Phase One: Idea Generation and Architecture Design
&lt;/h1&gt;

&lt;p&gt;Everything originated from a development idea and a newly released Google API document. The developer pasted a URL into the terminal for AGY CLI and posed the first core question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;User&lt;/strong&gt; : Following this example, is it possible to create a Mac OS App that can capture audio from the computer and translate it into other languages in real-time? Start helping me think about how to do this. My main goal is to be able to use it for Zoom / Google Meet meetings to directly translate other countries' languages into Traditional Chinese.&lt;/p&gt;

&lt;p&gt;Here are other documents for your reference: &lt;a href="https://ai.google.dev/gemini-api/docs/live-api/live-translate?hl=zh-tw" rel="noopener noreferrer"&gt;https://ai.google.dev/gemini-api/docs/live-api/live-translate?hl=zh-tw&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Upon receiving the instruction, AGY CLI acted as an architect, quickly analyzing and providing a set of native macOS solutions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;No virtual sound card installation required&lt;/strong&gt;: It recommends using Apple's &lt;strong&gt;ScreenCaptureKit&lt;/strong&gt; framework, introduced in macOS 13+, to directly capture the clean audio track of selected applications (such as Zoom or Chrome) in Sandbox secure mode.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Real-time audio resampling&lt;/strong&gt;: Utilize &lt;code&gt;AVAudioConverter&lt;/code&gt; to resample multi-channel/stereo 48kHz audio in real-time to the 16kHz mono PCM format supported by Gemini Live.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Bidirectional WebSocket communication&lt;/strong&gt;: Implement a bidirectional long connection using Swift's &lt;code&gt;URLSessionWebSocketTask&lt;/code&gt; to send PCM audio chunks while receiving Traditional Chinese translation results and broadcast audio.&lt;/li&gt;
&lt;/ul&gt;




&lt;h1&gt;
  
  
  Phase Two: Environment Configuration and Compilation Anxiety Elimination
&lt;/h1&gt;

&lt;p&gt;Before starting to write Swift code, the developer raised questions about the macOS compilation environment:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;User&lt;/strong&gt; : Do I absolutely need Xcode to compile?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;User&lt;/strong&gt; : Help me run xcode-select –install&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Addressing the developer's environment anxiety, AGY CLI provided the answers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Although it's best to use Xcode to create standard macOS SwiftUI GUI application bundles, we &lt;strong&gt;don't necessarily need to manually open the Xcode IDE&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;AGY CLI helped write an automated build script &lt;code&gt;build_app.sh&lt;/code&gt;, which directly calls the terminal's &lt;code&gt;swiftc&lt;/code&gt; command-line tool to compile all &lt;code&gt;.swift&lt;/code&gt; source files and automatically packages them into a complete &lt;code&gt;.app&lt;/code&gt; directory structure.&lt;/li&gt;
&lt;li&gt;At the same time, in response to the developer's request to install Command Line Tools, AGY CLI proactively invoked the permission tool and directly ran &lt;code&gt;xcode-select --install&lt;/code&gt; locally, automatically configuring the Swift compilation environment.&lt;/li&gt;
&lt;/ul&gt;




&lt;h1&gt;
  
  
  Phase Three: Connection Troubleshooting and Audio Bug Fixes
&lt;/h1&gt;

&lt;p&gt;After the code was initially completed, the developer ran the App from the command line, but the connection status showed abnormalities, and no characters were translated:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;User&lt;/strong&gt; : Didn't see any error messages~ but the connection status is disconnected&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This was the moment for AGY CLI to demonstrate its "autonomous troubleshooting" power. Upon receiving the prompt, it automatically located the &lt;code&gt;debug.log&lt;/code&gt; file, called &lt;code&gt;tail&lt;/code&gt; to analyze the runtime logs, and identified two critical issues:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Incompatible model name&lt;/strong&gt;: The original program used the standard REST model &lt;code&gt;models/gemini-3.5-flash&lt;/code&gt;, whereas the Live WebSocket API only accepts &lt;code&gt;gemini-3.5-live-translate-preview&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Incorrect JSON configuration level&lt;/strong&gt;: The API documentation used the &lt;code&gt;v1alpha&lt;/code&gt; version SDK, which wrapped &lt;code&gt;inputAudioTranscription&lt;/code&gt; within &lt;code&gt;generationConfig&lt;/code&gt;; however, the native WebSocket's &lt;code&gt;v1beta&lt;/code&gt; endpoint required these two fields to be placed directly under the &lt;code&gt;setup&lt;/code&gt; root directory. This was the culprit behind the &lt;code&gt;CloseCode 1007&lt;/code&gt; crash.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Multi-channel stereo silence Bug&lt;/strong&gt;: The multi-channel audio track captured by &lt;code&gt;ScreenCaptureKit&lt;/code&gt; was truncated to complete silence (all zeros) during copying in the old code due to insufficient AudioBufferList memory allocation.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;AGY CLI immediately proactively modified &lt;a&gt;AudioCaptureManager.swift&lt;/a&gt;, introducing the &lt;strong&gt;"Double-Call" register allocation pointer technique&lt;/strong&gt;, and refactored the Payload structure of &lt;a&gt;GeminiLiveConnection.swift&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;After the modifications were completed, the application ran smoothly, the console log finally printed &lt;code&gt;是否為靜音(全0): false&lt;/code&gt; (Is it silent (all 0s): false), and both real-time bilingual subtitles and real-time broadcast audio functioned correctly!&lt;/p&gt;




&lt;h1&gt;
  
  
  Phase Four: Automated DevOps and GitHub Delivery
&lt;/h1&gt;

&lt;p&gt;Once the developer confirmed that the program was working correctly, the final step was to open-source and share the code:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;User&lt;/strong&gt; : I want to check in the swift-demo folder to my own GitHub repo. Give me a suggested repo name and write a README.md under swift-demo.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;User&lt;/strong&gt; : Help me commit all relevant changes in that folder to &lt;a href="mailto:git@github.com"&gt;git@github.com&lt;/a&gt;:kkdai/gemini-live-translate-macos.git&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;AGY CLI immediately took over the final DevOps tasks:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;It recommended using &lt;code&gt;gemini-live-translate-macos&lt;/code&gt; as the Repo name and wrote the project's English GitHub description and topics tags.&lt;/li&gt;
&lt;li&gt;It automatically completed the full environment preparation, Xcode Sandbox Capabilities settings, command-line script execution steps, and API troubleshooting tips in &lt;a&gt;README.md&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;After obtaining the user's repository URL, AGY CLI proactively ran &lt;code&gt;git init&lt;/code&gt; in the background, wrote &lt;code&gt;.gitignore&lt;/code&gt;, committed all the code, and successfully pushed it to the remote GitHub repository!&lt;/li&gt;
&lt;/ol&gt;




&lt;h1&gt;
  
  
  Conclusion: Development Transformation and Insights
&lt;/h1&gt;

&lt;p&gt;Through this collaborative development with AGY CLI, we experienced an unprecedentedly rapid development process:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Reduced cognitive load&lt;/strong&gt;: Developers only need to express their intentions in natural language (e.g., "help me run the installation," "help me troubleshoot why the connection is broken"), and the AI Agent will autonomously translate them into corresponding system commands and code modifications.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Native system-level control&lt;/strong&gt;: AI can directly read and execute commands, synchronizing with the development environment in real-time, greatly reducing the hallucinations and environment version mismatches that often occurred with traditional Web AI Chat.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One-stop delivery&lt;/strong&gt;: From the first phrase "think about how to do it" to the final "Push to GitHub repository" with a single click, AGY CLI seamlessly integrated the entire software engineering lifecycle.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This practical experience proves that in the era of Agentic AI, a single developer, paired with a powerful CLI agent, can deliver a high-quality Native application involving system-level foundations and the latest APIs in an extremely short amount of time. See you next time!&lt;/p&gt;

</description>
      <category>ai</category>
      <category>gemini</category>
      <category>productivity</category>
      <category>softwaredevelopment</category>
    </item>
    <item>
      <title>[GCP Practical] LINE Business Card Bot</title>
      <dc:creator>Evan Lin</dc:creator>
      <pubDate>Sun, 07 Jun 2026 15:26:28 +0000</pubDate>
      <link>https://dev.to/gde/gcp-practical-line-business-card-bot-d46</link>
      <guid>https://dev.to/gde/gcp-practical-line-business-card-bot-d46</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F71r8jl4qibup4fxxh4tv.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F71r8jl4qibup4fxxh4tv.png" alt="image-20260607133454831" width="800" height="1739"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h1&gt;
  
  
  Upgrade Preamble
&lt;/h1&gt;

&lt;p&gt;After refactoring the agent based on &lt;strong&gt;Vertex AI ADK&lt;/strong&gt;, our LINE Name Card Assistant Bot (&lt;code&gt;linebot-namecard-python&lt;/code&gt;) entered the production environment for testing. However, in real-world usage scenarios, we quickly identified three core pain points affecting user experience and security:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt; &lt;strong&gt;Unstable OCR JSON Parsing&lt;/strong&gt;: Using the standard JSON Mode with a Prompt, Gemini occasionally still outputs Markdown tags or misses fields, causing parser errors.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Excessive Search Results Leading to LINE API 400 Error&lt;/strong&gt;: LINE limits sending a maximum of 5 messages at a time. When search results include 5 cards plus the Agent's text reply, totaling 6, LINE directly rejects it and doesn't reply.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;AI Accidental Modification&lt;/strong&gt;: If a user mentions modification, the Agent directly writes to Firebase without secondary confirmation, easily leading to data corruption due to mishearing or hallucination.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This article will focus on sharing how we conducted a second wave of upgrades to address the above pain points, implementing &lt;strong&gt;Structured Outputs&lt;/strong&gt;, &lt;strong&gt;Disambiguation Lists&lt;/strong&gt;, &lt;strong&gt;Two-Stage Confirmation Mechanism&lt;/strong&gt;, and the major pitfall we encountered during operations and deployment regarding environment variable recovery!&lt;/p&gt;




&lt;h1&gt;
  
  
  Optimization One: Embracing Gemini Structured Outputs
&lt;/h1&gt;

&lt;p&gt;Previously, when calling &lt;code&gt;gemini-3-flash-preview&lt;/code&gt; for name card image parsing, we commanded it via Prompt and manually parsed JSON. To ensure 100% format guarantee, we introduced the native &lt;strong&gt;Structured Outputs&lt;/strong&gt; feature of the Vertex AI API.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Defining the Name Card Schema
&lt;/h3&gt;

&lt;p&gt;In &lt;a&gt;app/gemini_utils.py&lt;/a&gt;, we defined the constraint Schema for the name card object, forcing Gemini to strictly adhere to this format for output:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;NAMECARD_SCHEMA&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;OBJECT&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;properties&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;STRING&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;description&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;聯絡人姓名，如果看不出來，請填寫 N/A&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;title&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;STRING&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;description&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;職稱或頭銜，如果看不出來，請填寫 N/A&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;company&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;STRING&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;description&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;公司名稱，如果看不出來，請填寫 N/A&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;address&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;STRING&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;description&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;公司或聯絡地址，如果看不出來，請填寫 N/A&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;phone&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;STRING&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;description&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;電話號碼，格式為 #886-0123-456-789,1234。&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;沒有分機就忽略 ,1234。如果看不出來，請填寫 N/A&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
            &lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;email&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;STRING&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;description&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;電子郵件信箱，如果看不出來，請填寫 N/A&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;required&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;title&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;company&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;address&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;phone&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;email&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  2. Applying to Generation Config
&lt;/h3&gt;

&lt;p&gt;We only need to specify &lt;code&gt;response_schema&lt;/code&gt; in &lt;code&gt;generation_config&lt;/code&gt; when instantiating &lt;code&gt;GenerativeModel&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;generate_json_from_image&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;img&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;PIL&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Image&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Image&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;object&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;GenerativeModel&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gemini-3-flash-preview&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;generation_config&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;response_mime_type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;application/json&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;response_schema&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;NAMECARD_SCHEMA&lt;/span&gt;
        &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;img_part&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Part&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;from_data&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nf"&gt;pil_to_bytes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;img&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;mime_type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;image/jpeg&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;generate_content&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;img_part&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;stream&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After application, the JSON error rate of the returned response dropped directly to 0%, eliminating complex string cleaning and parser error-prevention logic.&lt;/p&gt;




&lt;h1&gt;
  
  
  Optimization Two: Solving LINE Message Limit with 'Disambiguation List'
&lt;/h1&gt;

&lt;p&gt;LINE Webhook has an iron rule: &lt;strong&gt;the number of message bubbles sent in a single &lt;code&gt;reply_message&lt;/code&gt; must be between 1 and 5&lt;/strong&gt;. If the search results happen to be 5 or more, and a text reply is added, the total will exceed 5, triggering a LINE API 400 error.&lt;/p&gt;

&lt;h3&gt;
  
  
  💡 Solution: Disambiguation List
&lt;/h3&gt;

&lt;p&gt;We modified the search reply judgment in &lt;a&gt;app/line_handlers.py&lt;/a&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  When search results are &lt;strong&gt;1 to 4 items&lt;/strong&gt;: Directly display Carousel detailed name cards (conforming to LINE's 5-item limit).&lt;/li&gt;
&lt;li&gt;  When search results are &lt;strong&gt;5 or more items&lt;/strong&gt;: Do not display large cards; instead, return a &lt;strong&gt;'Name Card Search List' Flex Message Bubble&lt;/strong&gt;. The list itemizes names and companies, with a 'View ❯' Postback button on the right. Clicking it loads and displays that specific name card.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This design not only maintains a clean layout but also completely avoids the pitfall of exceeding the message limit!&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;        &lt;span class="k"&gt;elif&lt;/span&gt; &lt;span class="n"&gt;found_card_ids&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;found_card_ids&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="c1"&gt;# If the quantity is less than or equal to 4, directly display Carousel detailed name cards
&lt;/span&gt;                &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;card_id&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;found_card_ids&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                    &lt;span class="n"&gt;card_data&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;firebase_utils&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_card_by_id&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;card_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
                    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;card_data&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                        &lt;span class="n"&gt;reply_msgs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
                            &lt;span class="n"&gt;flex_messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_namecard_flex_msg&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;card_data&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;card_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
                        &lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="c1"&gt;# If the quantity is greater than 4, display as a list Flex Message for disambiguation
&lt;/span&gt;                &lt;span class="n"&gt;cards_list&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
                &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;card_id&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;found_card_ids&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                    &lt;span class="n"&gt;card_data&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;firebase_utils&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_card_by_id&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;card_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
                    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;card_data&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                        &lt;span class="n"&gt;cards_list&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
                            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;card_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;card_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;card_data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;N/A&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
                            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;company&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;card_data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;company&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;N/A&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
                            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;title&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;card_data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;title&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;N/A&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
                        &lt;span class="p"&gt;})&lt;/span&gt;
                &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;cards_list&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                    &lt;span class="n"&gt;list_msg&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;flex_messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_namecard_list_flex_msg&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
                        &lt;span class="n"&gt;cards&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;cards_list&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                        &lt;span class="n"&gt;title_text&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;🔍 Found multiple matching name cards&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                    &lt;span class="p"&gt;)&lt;/span&gt;
                    &lt;span class="n"&gt;reply_msgs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;list_msg&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  Optimization Three: Contact Modification Safety Lock — Two-Stage Confirmation Mechanism
&lt;/h1&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fe8rumozkznwk41fslmft.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fe8rumozkznwk41fslmft.png" alt="image-20260607133518906" width="800" height="1061"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Under the ADK agent architecture, users can update data through natural conversation (e.g., "Add 'Meeting next Monday' to Evan's memo"). However, if the LLM misinterprets the instruction, Firebase data can be directly overwritten.&lt;/p&gt;

&lt;p&gt;To address this, we implemented a &lt;strong&gt;Two-Stage Confirmation mechanism&lt;/strong&gt;:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt; &lt;strong&gt;Delayed Write&lt;/strong&gt;: When the ADK Tool (&lt;code&gt;update_namecard_field&lt;/code&gt; and &lt;code&gt;update_namecard_memo&lt;/code&gt;) is invoked by the model, the system does not directly rewrite Firebase. Instead, it temporarily stores the content to be modified in &lt;code&gt;user_states&lt;/code&gt; in memory and returns &lt;code&gt;True&lt;/code&gt; to allow the Agent to continue generating dialogue.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Display Confirmation Card&lt;/strong&gt;: After the conversation ends, if the main program detects a pending state, it generates a Flex Message card containing 'Confirm Modification' and 'Cancel' buttons.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Write After Confirmation&lt;/strong&gt;: Only after the user clicks 'Confirm Modification' (sending a Postback Event &lt;code&gt;action=confirm_update&lt;/code&gt;) does the system truly write the data to Firebase.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This not only perfectly prevents AI from accidentally triggering tools but also gives users absolute control when modifying data!&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;    &lt;span class="c1"&gt;# Handle confirmation of modification in handle_postback_event
&lt;/span&gt;    &lt;span class="k"&gt;elif&lt;/span&gt; &lt;span class="n"&gt;action&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;confirm_update&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;state&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;user_states&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{})&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;action&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;pending_update&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;update_type&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;update_type&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="n"&gt;card_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;card_id&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="c1"&gt;# Read data from temporary storage based on update_type, and truly write to Firebase...
&lt;/span&gt;            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;success&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="c1"&gt;# Reply with successful modification, and automatically display the updated Flex Card for user verification
&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  Ops Pitfall Record: Manual Deployment - The Mysterious Disappearance of Environment Variables
&lt;/h1&gt;

&lt;p&gt;In addition to code refactoring, we also encountered a significant operational pitfall during deployment.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Pitfall
&lt;/h3&gt;

&lt;p&gt;When we attempted to upload a local folder to Cloud Run using the MCP deployment tool locally, because the command did not include environment variable declaration parameters, the previously working LINE Token and Firebase URL on Cloud Run were all cleared and overwritten. Upon restart, the Container crashed directly with an error:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Specify ChannelSecret as environment variable.

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The online service instantly became paralyzed.&lt;/p&gt;

&lt;h3&gt;
  
  
  Recovery Process
&lt;/h3&gt;

&lt;p&gt;Fortunately, Cloud Run fully retains the configuration settings of older versions. We can use the &lt;code&gt;gcloud&lt;/code&gt; command to view previous Revisions and restore the lost variables:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt; &lt;strong&gt;Retrieve the detailed configuration of the last successfully running Revision&lt;/strong&gt;:
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;gcloud run revisions describe linebot-namecard-python-00096-d89 &lt;span class="nt"&gt;--project&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;line-vertex &lt;span class="nt"&gt;--region&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;asia-east1

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This will output the environment variable values bound to that version.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt; &lt;strong&gt;Re-inject environment variables into the service&lt;/strong&gt;:
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;gcloud run services update linebot-namecard-python &lt;span class="nt"&gt;--project&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;line-vertex &lt;span class="nt"&gt;--region&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;asia-east1 &lt;span class="nt"&gt;--set-env-vars&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"ChannelAccessToken=...,ChannelSecret=..."&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;By restoring the variables, we seamlessly recovered the service within minutes. This also reminds us: when manually deploying to Cloud Run, always pay extra attention to the inheritance or declaration of environment variables to avoid accidentally clearing the official cloud configuration.&lt;/p&gt;




&lt;h1&gt;
  
  
  Summary and Benefits
&lt;/h1&gt;

&lt;p&gt;This optimization brought excellent production-level transformations to our LINE Name Card Bot:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt; &lt;strong&gt;100% Format Security&lt;/strong&gt;: Through API native Schema enforcement, the name card recognition format error rate dropped to 0%.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Explosion-Proof Reply Protection&lt;/strong&gt;: Multiple search results are automatically converted into a "Disambiguation List", perfectly complying with LINE's message limit.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Secure Contact Changes&lt;/strong&gt;: The two-stage confirmation mechanism confines AI's write access to a confirmation sandbox, protecting important user data.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Robust Configuration Disaster Recovery&lt;/strong&gt;: Utilizing gcloud historical Revision restoration technology ensures the service can quickly recover within a short period.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The complete and linter-optimized code has been pushed to &lt;a href="https://github.com/kkdai/linebot-namecard-python" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;. We hope this practical experience helps everyone avoid detours when building production-grade AI Agents! See you next time!&lt;/p&gt;

</description>
      <category>api</category>
      <category>gemini</category>
      <category>google</category>
      <category>python</category>
    </item>
  </channel>
</rss>
