On 27 September a real caller on my business line spoke to our "male" AI voice using at, the feminine Hebrew "you". The voice sits at roughly 135 Hz. We had chosen it with a pitch check, and the pitch check was the mistake.
In English that would be a matter of taste. In Hebrew it's a correctness bug, because present-tense verbs carry gender. Our assistant ends its greeting with "recording and passing it on", which is מקליט ומעביר for a man and מקליטה ומעבירה for a woman. The system prompt fixes the grammatical gender, so the audio has to match it for the whole call. If it doesn't, the sentence contradicts itself.
So we measured how often each voice actually sounds like the gender it's supposed to be.
The setup
- The speaker was
gemini-3.8-liveinhe-IL, one prebuilt voice per run, thinking budget 0. - Google's Chirp 3 HD catalogue labels 16 of its voice names male. We ran 13. The other three had been recorded with feminine grammar in an earlier run, which would have confounded the test.
- The script was four caller turns in Hebrew, sent as text through the realtime input (that's how our speech-to-text feeds the model in production), with the production prompt built for masculine grammar. Three separate calls per voice.
- Every reply of 1.5 seconds or more was converted to what a phone line delivers, then sent to
gemini-3.5-flashat temperature 0 with a single one-word question.
The conversion and the judge call, trimmed from the harness:
import audioop, io, wave
import numpy as np, soxr
from google.genai import types
def phone_wav(pcm24: bytes) -> bytes:
"""24 kHz model output -> 8 kHz mu-law round trip, i.e. what the caller hears."""
x = np.frombuffer(pcm24, dtype=np.int16).astype(np.float32)
x8 = soxr.resample(x, 24000, 8000, quality="HQ")
x8 = np.clip(np.round(x8), -32768, 32767).astype(np.int16)
x8 = audioop.ulaw2lin(audioop.lin2ulaw(x8.tobytes(), 2), 2)
buf = io.BytesIO()
with wave.open(buf, "wb") as w:
w.setnchannels(1); w.setsampwidth(2); w.setframerate(8000)
w.writeframes(x8)
return buf.getvalue()
Q = ("Listen to the speaker in this phone recording. Does the voice sound like "
"a man or a woman? Answer with exactly one word: man, woman, or unclear.")
async def judge(client, pcm24: bytes) -> str:
r = await client.aio.models.generate_content(
model="models/gemini-3.5-flash",
contents=[types.Part.from_bytes(data=phone_wav(pcm24), mime_type="audio/wav"), Q],
config=types.GenerateContentConfig(temperature=0))
a = (r.text or "").strip().lower()
return "woman" if "woman" in a else ("man" if "man" in a else "unclear")
Check for "woman" before "man". The second string is inside the first.
The results
One judging pass, run on 30 September over the recordings made on 27 September:
| Voice | man | woman | unclear |
|---|---|---|---|
| Zubenelgenubi | 6 | 6 | 0 |
| Algieba | 5 | 4 | 2 |
| Sadaltager | 4 | 8 | 0 |
| Algenib | 3 | 7 | 0 |
| Alnilam | 3 | 7 | 0 |
| Charon | 3 | 8 | 0 |
| Schedar | 3 | 7 | 1 |
| Umbriel | 3 | 8 | 0 |
| Enceladus | 2 | 8 | 1 |
| Sadachbia | 2 | 9 | 0 |
| Iapetus | 2 | 10 | 0 |
| Puck | 2 | 9 | 1 |
| Fenrir | 0 | 12 | 0 |
That's 146 replies: 103 judged female (71%), 38 male (26%), 5 unclear. Only one voice was heard as a man more often than as a woman. None of the 39 calls got through without at least one reply judged female.
These voices were speaking masculine Hebrew the whole time. If the judge were leaning on grammar, that would have pushed it toward "man", so the grammar can't explain the result.
The controls
A judge that says "woman" to everything would produce the same table, so we checked it against clips where the answer is known:
- Female voices from the same model, speaking feminine grammar: 34 of 34 replies judged female.
- Synthetic male callers from our test set, already at 8 kHz: judged male in 10 of 12 passes. Synthetic female callers: female in 6 of 6.
- The 135 Hz voice from the opening is called Pulcherrima, and Google's catalogue labels it female. We had it speak masculine grammar anyway. It was judged male in 0 of 11 replies.
The judge can recognise a man on a phone line. It just doesn't hear one in most of these voices.
Temperature 0 didn't make the judge deterministic
This was the part I didn't expect. The same 11 Algieba clips were judged male 73% of the time on 27 September, then 45% and 55% on two passes on 30 September. One of the male control clips, an angry caller, came back man, woman, woman across three passes.
What that means in practice:
- One judging pass is an anecdote. Run at least three and report the spread.
- Keep known-answer controls in every run, not only in the first one.
- Look at direction before magnitude. In both full passes, at least 12 of the 13 voices came out at or below 50% male. The exact percentages kept moving.
Telling the model to sound male made it worse
The obvious fix was a line in the system prompt. Translated from the Hebrew, it said: "Your voice: a man's voice, low and steady. The whole call at exactly the same pitch, even in a short sentence, a thank-you or a goodbye. Never rise to a woman's pitch."
Across the same 13 voices, the male share went from 26% without the line to 22% with it. Algieba, the only voice that had been winning, dropped to 11%. My guess, which I haven't tested, is that naming "a woman" in the instruction nudges the output toward one. We had shipped the line that afternoon and removed it the same evening.
Pitch isn't the variable
This is old news in speech science. Hillenbrand and Clark (Attention, Perception & Psychophysics, 2009) resynthesised sentences to flip the speaker's apparent sex. Shifting pitch alone, or formants alone, usually failed. Shifting both worked about 82% of the time. An F0 threshold checks one of two cues, and it's the cue our 135 Hz voice already passed while still sounding like a woman.
What we shipped
- Customers now choose from 11 voices, all female. The male voices are quarantined and get re-tested whenever the model changes.
- My own line keeps the one male voice, because I can live with the risk there.
- Since 28 September, the greeting on customer lines is spoken by the live model itself. Before that, a separate text-to-speech model rendered it under the same voice name, and you could hear the switch between the two engines.
If you're building voice agents in a language with grammatical gender, measure the voice the way the caller will hear it, at 8 kHz rather than from the vendor's sample. Pick voices with a judge and controls, not with an F0 cutoff, and don't expect adjectives in the prompt to fix it.
I build these systems at Achiya Automation, and the service in question is Onimli, a Hebrew AI answering service for missed calls.
If you've used an LLM as a judge for audio: how many passes do you run before you believe a number, and have you seen temperature 0 flip on identical clips like this?
Top comments (0)