DEV Community

Pandas Studio
Pandas Studio

Posted on Fully Autonomous

Solving Voice Memo Skill Trigger Failures via Plugin Hooks

While testing a voice-to-structured-document workflow on an NVIDIA DGX Spark, the Hermes Agent consistently failed to trigger the specific voice-memo skill, opting for the default memory tool instead. This post details how I implemented custom gateway hooks to force correct skill invocation and output validation.

Background

The setup involved a Telegram-based voice interface connected to a Hermes Agent (v0.20.6) running on a DGX Spark. I used a host-side whisper.cpp server with a ggml-large-v3-turbo-q5_0 model for STT (Speech-to-Text). The goal was to have the agent recognize a voice memo, transform it into a structured Markdown document using the voice-memo skill, and save it to /sandbox/notes/memos/.

What happened

In the first round of testing, transcription succeeded 5/5 times, but the voice-memo skill was triggered 0/5 times. In the second round, after changing the trigger word to 메모 (

What I tried

  1. Changing the skill trigger condition in the SKILL.md file to look for the word 메모 (spoken) or 메모: (typed). → The skill still failed to trigger 0/3 times; the Qwen model interpreted 메모 as a command to use the default memory tool rather than the specific skill.
  2. Adding a technical terminology list to the Whisper STT prompt to reduce transcription errors. → Transcription errors for technical terms dropped from 6 to 3 in subsequent tests.

Root cause

The Hermes v0.20.6 gateway only passes transcribed text as a quoted line without any special markers if transcription is successful. The original skill relied on a [The user sent a voice message ...] marker that only appeared during transcription failures. Even when I updated the skill to look for the word 메모, the LLM (Qwen) frequently misidentified the intent as a request to use the standard memory tool for general storage.

The fix

I implemented a custom plugin using two specific Hermes hooks. First, I used pre_gateway_dispatch to intercept messages before they reached the model. For audio, the hook performs its own transcription (taking ~0.7s) to check if the first word is 메모. If it matches, it rewrites the message to include a hard instruction: [voice-memo] This message is the author's memo... Do not use the memory tool.... Second, I added a transform_llm_output hook to intercept the model's response, read the newly saved file, and replace the model's chatty reply with the full document content plus an automated quality check.

Plugin logic for pre-gateway dispatch

if kind == "text" and TYPED.match(text):
    return {"action": "rewrite", "text": f"{INSTRUCTION}\n\n{text}"}
media = list(getattr(event, "media_urls", None) or [])
if kind in ("voice", "audio") and media and SPOKEN.match(_transcribe(media[0]).strip()):
    logger.info("voice_memo: 메모 음성 → voice-memo 스킬 지시")
    return {"action": "rewrite", "text": f"{text}\n\n{INSTRUCTION}" if text else INSTRUCTION}
Enter fullscreen mode Exit fullscreen mode

Automated document quality check logic

def check(doc: str) -> list:
    out = [f"'{s}' 칸 없음" for s in SECTIONS if _section(doc, s) is None]
    summary = _section(doc, "정리") or ""
    missing = [i for i in ITEMS if not re.search(rf"^\s*[-*]\s*\**{re.escape(i)}\**\s*[::]", summary, re.M)]
    if summary and missing:
        out.append("정리에 빠진 항목: " + ", ".join(missing))
    if (cjk := sorted(set(FOREIGN.findall(doc)))):
        out.append("한자·가나 섞임: " + ", ".join(cjk[:5]))
Enter fullscreen mode Exit fullscreen mode

Result

In the third round of testing, the system achieved 3/3 successful saves and 0 uses of the memory tool. The automated check successfully identified missing sections, the presence of Chinese/Japanese characters, and English words used outside the original context.

Lessons

  • Don't rely on LLM intent recognition for critical tool selection; use gateway hooks to inject explicit instructions.
  • When building voice workflows, intercepting the stream before the gateway can provide much more reliable trigger detection.

Still open

The system still experiences 'interrupted' turns if a new voice message arrives while the agent is still processing a previous one due to the busy_input_mode: interrupt setting.


This post was written by a local Gemma 4 model from my own commit log and experiment notes, fact-checked against them by a local Qwen3.6 model, with no human edits.

Top comments (1)

Some comments may only be visible to logged-in visitors. Sign in to view all comments.