While testing voice-to-document automation on a DGX Spark, the Hermes Agent repeatedly failed to trigger the intended 'voice-memo' skill. I discovered the model was misinterpreting the trigger word as a request to use its internal memory tool instead of the specialized skill.
Background
The setup involved a Hermes Agent (v0.20.6) running on an NVIDIA DGX Spark with a Qwen 3.6-35B model. The goal was to send voice memos via Telegram, have them transcribed by a local Whisper server, and then have the agent transform them into structured documents saved to /sandbox/notes/memos/ using a specific skill.
What happened
During the second round of testing, despite successful transcriptions, the skill activation rate was 0/3. While the transcriptions correctly started with the word 메모 ("memo"), the Qwen model consistently chose the default memory tool to store the information rather than calling the voice-memo skill. This resulted in the notes being saved to the general MEMORY.md file instead of being processed into structured markdown files.
What I tried
-
Modifying the skill activation condition to trigger when the transcription starts with the word
메모. → Failed; the model still prioritized thememorytool over the specialized skill. - Adding technical term prompts to the Whisper STT engine to reduce transcription errors. → Partially successful; reduced errors from 6 to 3 in specific test cases.
Root cause
The root cause was a model-level semantic misunderstanding. When the agent received a message starting with 메모, the Qwen model interpreted the intent as 'remember this' and invoked its built-in memory tool. Because the Hermes v0.20.6 gateway only passes the transcribed text as a plain quoted string without specific voice markers, the model had no way to distinguish a general conversational mention of 'memo' from a formal command to trigger the voice-memo skill.
The fix
I implemented a custom Hermes plugin using the pre_gateway_dispatch hook. This hook intercepts the message before it reaches the model. For audio, it performs a preliminary transcription (approx. 0.7s) to check if the first word is 메모. For text, it checks if the message starts with 메모:. If a match is found, the plugin rewrites the message to include a hardcoded instruction: [voice-memo] This message is the author's memo... Do not use the memory tool... Call skill_view with name "voice-memo".... This forces the model to use the correct skill by explicitly forbidding the memory tool.
Voice memo plugin logic for intercepting and rewriting messages
SPOKEN = re.compile(r"^\W*메모(?=[\s,.:!?~…]|$)") # 전사는 "메모! 어제…", "메모 주간…"처럼 온다
TYPED = re.compile(r"^\s*메모\s*[::]")
INSTRUCTION = ("[voice-memo] This message is the author's memo (it starts with the word 메모). "
"Do not use the memory tool and do not continue earlier topics. "
"Call skill_view with name \"voice-memo\" and follow it exactly: write the document, save it with the file tool "
"under /sandbox/notes/memos/, then reply with the full document and the 저장: line.")
def on_dispatch(event=None, **_):
try:
kind = getattr(getattr(event, "message_type", None), "value", "")
text = (getattr(event, "text", "") or "").strip()
if kind == "text" and TYPED.match(text):
return {"action": "rewrite", "text": f"{INSTRUCTION}\n\n{text}"}
media = list(getattr(event, "media_urls", None) or [])
if kind in ("voice", "audio") and media and SPOKEN.match(_transcribe(media[0]).strip()):
logger.info("voice_memo: 메모 음성 → voice-memo 스킬 지시")
return {"action": "rewrite", "text": f"{text}\n\n{INSTRUCTION}" if text else INSTRUCTION}
except Exception as e: # 훅이 실패해도 메시지는 평소대로 간다
logger.warning("voice_memo: 판단 실패 %s", e)
return None
Registering the hook in the plugin
def register(ctx):
ctx.register_hook("pre_gateway_dispatch", on_dispatch)
Result
In the third round of testing, the fix was successful, achieving a 3/3 success rate for document saving. The memory tool was used 0 times, and all notes were correctly processed into the target directory.
Lessons
- When an LLM consistently misinterprets a command as a different tool, do not rely on prompting alone; use a pre-dispatch hook to inject explicit constraints.
- For voice-driven automation, performing a lightweight preliminary transcription at the gateway level can provide the necessary context to steer model behavior.
Still open
The system still experiences interrupted turns if a new voice message arrives while the agent is busy (busy_input_mode: interrupt), and some transcription hallucinations (e.g., Whisper adding 'Thank you' at the end of silence) persist.
This post was written by a local Gemma 4 model from my own commit log and experiment notes, fact-checked against them by a local Qwen3.6 model, with no human edits.
Top comments (0)