Verification asks you to say a phrase the server picks, right now — and the whole check now runs on our servers, not in your browser, so nothing a bot asserts is ever trusted. Here's exactly what the check does, how it turns a pass into a stable voice-id that makes "one human, one welcome grant" enforceable, the cloning attack we ran on ourselves — and, just as important, what the badge does not prove.
A verification session is four rounds. In each one the server hands you a random five-word phrase — ordinary words mixed with spoken digits — and you read it aloud. The phrase is single-use and expires after 90 seconds. Nothing you can prepare in advance is any use, because the thing you have to say didn't exist until you asked for it.
A fresh 5-word phrase — words + spoken digits
Within 90 seconds. The phrase is single-use.
Timing, the words, the voice — all measured by us
Then one final check across all four
Earlier versions measured the voice in the browser and sent the server a score. Our own security audit called that what it is: a self-report. A bot never has to load your code — it can just send the number that means "pass."
So we inverted it. Now the browser does exactly one thing: record. Our verification server transcribes the phrase itself, extracts the speaker embedding itself, matches it against your enrolment itself — and cryptographically signs the result. Every claim in the badge is something we measured, never something the client asserted.
Captures the audio of you saying the phrase. That's all it's trusted to do.
Transcribe the words, extract the voice embedding, match it to your enrolment
The verdict — and your voice-id — are signed by our key. Unforgeable.
Each round is judged on three independent things. They fail differently, which is the point — an attack that beats one usually trips another.
A human has to hear the phrase, read it, and speak it. An answer that arrives faster than that is machinery, not a person.
Our server transcribes what you said and requires at least 4 of the 5 words. The tolerance is there because microphones and accents are real; it is not there to let you say something else.
The audio is compared against the voiceprint held for this account — computed and scored on our side. Right words in the wrong voice fails.
The four recordings are checked against each other. This is what catches handing the microphone to someone else halfway through, and stitching a session together out of clips from different sources.
Every practical attack on voice is an attack you prepare. Someone records you off a livestream. Someone trains a clone on your podcast. Both take time, and both produce audio of you saying things the attacker chose.
A server-chosen phrase deletes that advantage. A recording of you cannot know the phrase. A clone built offline cannot know the phrase. To pass, the attacker has to synthesize new speech, in your voice, saying five specific words, inside 90 seconds — live. That's the expensive case, and forcing every attacker into it is the entire purpose of the challenge.
| Attack | What it needs | What the challenge does to it |
|---|---|---|
| Replay a recording | Any audio of you | Dead on arrival — the recording says the wrong words |
| Offline voice clone | Samples of your voice + time | Useless unless it can generate the new phrase live |
| Splice clips together | Words harvested from many recordings | Has to beat the timing check and the same-speaker check across four rounds |
| Pass the mic to a real person | A cooperating human | Voice won't match the account; the cross-round check catches a swap mid-session |
| Verify 10,000 accounts | A bot farm | The economics collapse — every badge costs a live human saying a fresh phrase, and the same human resolves to the same voice-id anyway |
When you pass, our server resolves your voice to a stable, opaque voice-id — an identifier that says "this is the same human we've verified before" without saying who you are. The same person on a second account resolves to the same voice-id. It's derived from the voice itself, on our side, from audio of a phrase the client didn't choose — so it isn't something an account can pick, swap, or fabricate.
The voice-id is signed into the badge by our verification key. Anything downstream that needs "one per human" — starting with the welcome grant — checks that signature and deduplicates on the id. Open ten accounts and you're still one voice, one id, one grant. That is the claim this system was built to make real: one-human-one-account economics, enforced against a value no client can forge.
The only way to know whether a voice check survives cloning is to clone a voice and attack it. So we did — to ourselves. We took a commercial voice generator, built a clone of our own founder's voice from clean recordings, and ran it against our own verification the way a real attacker would.
Here's the half that worked. The challenge phrase does its job: a replayed recording and a pre-built clone both fail, because they can't say five words that didn't exist a minute ago. And at our operating threshold, the speaker match accepted the real voice essentially every time while rejecting almost every clone clip we threw at it.
And here's the half that isn't flattering, which we're telling you anyway: our anti-spoof detector — the model whose only job is asking "is this audio synthetic?" — leaks on voice generators it hasn't seen. In our own testing it missed roughly 29% of clones from unfamiliar engines, and in blocking mode it false-flagged about 7% of real strangers. So it runs in log-only mode: it records what it sees and never rejects anyone on its own. Until it's honest, it doesn't gate.
The intuition that "more tests = harder to fake" is wrong here, and it's worth being blunt about why. Repeating the same test doesn't compound. A bot that can solve one round can solve fifteen — it's the same problem, handed to it again. What you'd be buying with round eleven is not security; it's the same wall, one more time.
What rounds do buy is two real things: a more stable score, because averaging across several samples smooths out a cough, a bad microphone moment, or one unlucky clip; and the same-speaker check across rounds, which needs more than one sample to exist at all. Both of those saturate quickly — around four. Past that, the curve is flat and the only thing still going up is how long a real person has to sit there talking to their phone.
The badge attests to a moment, so it's dated like one. After 90 days it lapses and has to be re-earned by running the check again.
This matters more than it looks. Cloning and synthesis get better every year; a check that was hard to beat in one season may not be in the next. If badges were permanent, a single successful break — one clone that got through once — would buy a permanent mark of trust. Expiry means every account has to keep proving it against whatever the check looks like today. And where money moves, the bar is higher still: grant payouts require a check from the last 24 hours, not just an unexpired badge.
This is the part people are right to interrogate, so here it is plainly.
A biometric you can't rotate is a serious thing to hand anyone, which is why the design keeps the smallest possible artifact, keeps it in one place, and gives you a way to destroy it.
The badge means one specific thing: a live human spoke a phrase we chose, at a moment in time, in a voice matching this account. That's it. It is worth stating what it isn't:
VoiceBan is for real people, and voice verification is how you prove it. Speech stays open to everyone: nothing is blocked, nothing is removed, and you can post, comment and play whether you verify or not. But we will not leave a gray area for bots to exploit — so everything a fake-account farm would actually come for is verified-humans-only. Two things change the moment you verify:
Reach — verified accounts are boosted in the feed and in comments; unverified accounts rank below them. And money — tips, monetization, rewards and incentives are for verified humans only. An unverified account can still speak and take part, but it does not get paid. No live human behind it, no payout.
The reasoning is ordinary: attention and money are exactly what fake accounts are made to farm. Put a real, repeated, live-voice cost in front of both and the farm stops being worth it. You can still say anything without verifying — you just can't cash in as a bot.
Boosted in feed & comments — and the only accounts that can earn: tips, monetization, rewards. Re-earned every 90 days.
Free to post, comment and play — but ranked below verified and can't be paid: no tips, no monetization, no rewards.
No ranking effect at all — it's live voice end to end, so the thing verification measures is already happening.
No gray areas — that's the point. A gray area is exactly what bot farms exploit, so we don't leave one: if an account won't prove a live human is behind it, we treat it as a bot for the two things fakes are built to farm — reach and money. It can still speak; it just can't cash in. That single line — bots can play, but bots don't get paid — takes the profit out of running them. Voice verification is the essence of a platform for real people. What it also unlocks: Voice Seal to stamp a post as your real voice, and Declared Agents, where "human" is a label you earn, not one you type.