DEV Community

Pandas Studio
Pandas Studio

Posted on Fully Autonomous

[Hermes Agent] Fixing Voice Memo Skill Misidentification via pre_gateway_dispatch

While testing voice-to-document automation on a DGX Spark, the Hermes Agent repeatedly failed to trigger the intended 'voice-memo' skill. I discovered the model was misinterpreting the trigger word as a request to use its internal memory tool instead of the specialized skill.

Background

The setup involved a Hermes Agent (v0.20.6) running on an NVIDIA DGX Spark with a Qwen 3.6-35B model. The goal was to send voice memos via Telegram, have them transcribed by a local Whisper server, and then have the agent transform them into structured documents saved to /sandbox/notes/memos/ using a specific skill.

What happened

During the second round of testing, despite successful transcriptions, the skill activation rate was 0/3. While the transcriptions correctly started with the word 메모 ("memo"), the Qwen model consistently chose the default memory tool to store the information rather than calling the voice-memo skill. This resulted in the notes being saved to the general MEMORY.md file instead of being processed into structured markdown files.

What I tried

  1. Modifying the skill activation condition to trigger when the transcription starts with the word 메모. → Failed; the model still prioritized the memory tool over the specialized skill.
  2. Adding technical term prompts to the Whisper STT engine to reduce transcription errors. → Partially successful; reduced errors from 6 to 3 in specific test cases.

Root cause

The root cause was a model-level semantic misunderstanding. When the agent received a message starting with 메모, the Qwen model interpreted the intent as 'remember this' and invoked its built-in memory tool. Because the Hermes v0.20.6 gateway only passes the transcribed text as a plain quoted string without specific voice markers, the model had no way to distinguish a general conversational mention of 'memo' from a formal command to trigger the voice-memo skill.

The fix

I implemented a custom Hermes plugin using the pre_gateway_dispatch hook. This hook intercepts the message before it reaches the model. For audio, it performs a preliminary transcription (approx. 0.7s) to check if the first word is 메모. For text, it checks if the message starts with 메모:. If a match is found, the plugin rewrites the message to include a hardcoded instruction: [voice-memo] This message is the author's memo... Do not use the memory tool... Call skill_view with name "voice-memo".... This forces the model to use the correct skill by explicitly forbidding the memory tool.

Voice memo plugin logic for intercepting and rewriting messages

SPOKEN = re.compile(r"^\W*메모(?=[\s,.:!?~…]|$)")   # 전사는 "메모! 어제…", "메모 주간…"처럼 온다
TYPED = re.compile(r"^\s*메모\s*[::]")
INSTRUCTION = ("[voice-memo] This message is the author's memo (it starts with the word 메모). "
               "Do not use the memory tool and do not continue earlier topics. "
               "Call skill_view with name \"voice-memo\" and follow it exactly: write the document, save it with the file tool "
               "under /sandbox/notes/memos/, then reply with the full document and the 저장: line.")


def on_dispatch(event=None, **_):
    try:
        kind = getattr(getattr(event, "message_type", None), "value", "")
        text = (getattr(event, "text", "") or "").strip()
        if kind == "text" and TYPED.match(text):
            return {"action": "rewrite", "text": f"{INSTRUCTION}\n\n{text}"}
        media = list(getattr(event, "media_urls", None) or [])
        if kind in ("voice", "audio") and media and SPOKEN.match(_transcribe(media[0]).strip()):
            logger.info("voice_memo: 메모 음성 → voice-memo 스킬 지시")
            return {"action": "rewrite", "text": f"{text}\n\n{INSTRUCTION}" if text else INSTRUCTION}
    except Exception as e:   # 훅이 실패해도 메시지는 평소대로 간다
        logger.warning("voice_memo: 판단 실패 %s", e)
    return None
Enter fullscreen mode Exit fullscreen mode

Registering the hook in the plugin

def register(ctx):
    ctx.register_hook("pre_gateway_dispatch", on_dispatch)
Enter fullscreen mode Exit fullscreen mode

Result

In the third round of testing, the fix was successful, achieving a 3/3 success rate for document saving. The memory tool was used 0 times, and all notes were correctly processed into the target directory.

Lessons

  • When an LLM consistently misinterprets a command as a different tool, do not rely on prompting alone; use a pre-dispatch hook to inject explicit constraints.
  • For voice-driven automation, performing a lightweight preliminary transcription at the gateway level can provide the necessary context to steer model behavior.

Still open

The system still experiences interrupted turns if a new voice message arrives while the agent is busy (busy_input_mode: interrupt), and some transcription hallucinations (e.g., Whisper adding 'Thank you' at the end of silence) persist.


This post was written by a local Gemma 4 model from my own commit log and experiment notes, fact-checked against them by a local Qwen3.6 model, with no human edits.

Top comments (0)