The Humanness Index™

Which voice model sounds the most human?

Sounding human is hard to measure, but it's what decides whether a call works. We clone one voice onto every model and play them blind against a real human, so you can hear which ones pass.

Read the whitepaper

Which voice sounds more human?

Same voice, different models.

Read along
So I can see here that the package was marked as delivered on Tuesday, but if you're saying it never arrived then what we'll do is... let me just. Yeah, I'm going to open a lost package investigation for you. That usually takes about forty-eight hours to resolve.
vs

play each side · space vote, then next pair

How it works

  1. Step 1

    Same voice, every model

    We clone one conversational voice onto every model, so you're judging the model, not its demo reel.

  2. Step 2

    You listen blind

    Two voices, same line, no labels. Pick the one that sounds more human.

  3. Step 3

    A real human sets the bar

    Blind votes are fit into a rating, with a real human at 100. The higher the score, the more human the model sounds.

Humanness Rankings

Humanness distribution

21 Models11 providers13000 unique votes

Color = rank Average
0255075100100200400800HumannessWorseBetterLatency (ms, log scale)FasterSlowerAbove Average

Why latency matters. A voice that lags breaks the conversation, no matter how human it sounds.

Likely RankModelListen
BaselineHumanHomo Sapien1001293589
#1–6SpeechifySimba 3.2971284428 ms$10404
#1–4ElevenLabsEleven v3971283758 ms$100574
#1–7xAIGrok TTS931270460 ms$15559
#1–8MiniMaxSpeech 2.8921268325 ms$60543
#2–8MiniMaxSpeech 2 HD891259357 ms$100524
#2–8Canopy LabsOrpheus891258Open source529
#3–9Fish AudioS2.1-Pro871253141 ms$15396
#3–8xAIGrok TTS (Streaming)871251285 ms$15530
#8–12InworldTTS-1.5-max781224337 ms$35453
#9–13ElevenLabsFlash v2761216226 ms$50458

The Index only includes models that support voice cloning: each battle plays the same cloned source voice through both models, so the comparison is head to head and fair. The Index is an open benchmark and is independent of the Vapi product: a model appearing here does not mean it is available in Vapi, and availability in Vapi plays no part in scoring. Don't see your model on this list? Contact us at humannessindex@vapi.ai.

What we Listen for

What makes a voice sound human?

Humanness doesn't break down into features. You either believe there's a person on the other end, or you don't. When that belief breaks, it's usually because of one of these.

Expressiveness

Emotion and emphasis. Stressing the right words, sounding like it means what it says instead of reading text aloud.

Tone & prosody

The intonation, rhythm, and melody of speech. The natural rise and fall of how people actually talk.

Artifacts

The little human sounds: breaths, stutters, natural pauses. A voice with none of them sounds too clean to be real.

Why trust this benchmark?

Any model can sound good on its own demo voice. The real test is how it handles your use case. We clone one voice across every model so the comparison is fair. Models that can't clone a voice can't be tested fairly, so they're not listed.

Most Human Models

#3
Humanness93
Latency
460 ms
Languages
20
Votes
559
Voting in progress

Rankings are provisional

We're keeping the podium under wraps while the votes come in. Listen and vote above, and the most human models reveal once the standings settle.

Why this exists

Picking a TTS model for a voice agent comes down to one thing: does it sound human enough that people forget they're talking to software? You can't get that from demos or vendor claims. So we made it measurable and took the call out of our own hands: one voice cloned onto every model, played blind with no names attached, scored against a real human by the people who hear it.

Find the most human-sounding voice for your agent.

Compare the models in blind tests, read the methodology, or get in touch.

Build a TTS model? Add yours to the Index.