an open research benchmark
Do AI voice agents hear how you feel?
Say "fine, whatever, just book it" cheerfully — book it. Say it in a defeated sigh — a good agent checks first. Same words, different right answer.
Take a dozen calls: can you out-hear the AI?hears
models can often name the caller's tone when asked
acts?
whether hearing it changes what they do, versus reading a transcript
you?
every game session is an anonymous data point in the human baseline
How the benchmark works
- Every test item is one caller line, heard in two deliveries with identical words — so only the voice can change the right action.
- Models (and you) hear a clip and choose from the same action menu: proceed, confirm first, ask a question, escalate to a human…
- Controls keep it honest: every model also gets a transcript-only twin of each call, and a words-only speech-to-text pipeline sets the floor for what the voice adds.
- Everything is open source — items, audio, scoring, and (soon) a harness to test your own voice agent against the same calls.