Settings

A pitch shift moves the fundamental frequency of a voice, which is the first thing a listener reads and the thing they are most confident about. Move it far enough and a stranger will not place the voice, and a voice that read as clearly male or clearly female stops doing so. It leaves everything else exactly where it was — the accent, the vocabulary, where the pauses fall, the laugh, the throat-clearing, the room — and somebody who knows you often recognises a shifted voice within a sentence, because they were never recognising the pitch. It is also reversible by one number. VOICEGUARD does the shift in your browser; this page is about what it is and is not worth.

What the shift actually moves

Pitch is the rate the vocal folds open and close, forty to fifty times a second for a very low voice and two hundred and something for a high one. Shifting it the simple way — the way a tape machine does, by playing the samples faster or slower — moves the whole spectrum together, and moves the clip's length with it.

That is measurable rather than a matter of opinion. A one-second tone at 220 Hz, shifted down four semitones, comes out of VOICEGUARD at 174.6 Hz and 1.26 seconds long: the interval is a ratio of two to the power of minus four twelfths, and the duration is the same ratio the other way up. Those numbers are checked on every build, because a page that says what a tool does should be wrong loudly rather than quietly.

The lengthening is usually described as a defect of the cheap method. For disguise it is the opposite, because the tempo and rhythm of somebody's speech is itself identifying, and this moves that too.

Everything it leaves alone

This is the list that decides whether a pitch shift is any use to you:

  • How you speak. Accent, the words you reach for, the phrases you repeat, how you build a sentence, what you say when you are thinking. None of it is in the pitch.
  • The non-speech sounds. A laugh, a cough, an intake of breath before a hard word. People recognise these more reliably than they recognise a voice.
  • The room. Reverberation, the hum of a particular extractor fan, a clock, a road outside, a dog. A recording carries where it was made.
  • Anything else in the file. An audio file has metadata like any other — the device, the software, sometimes a date and a location. AVSCRUB deals with that separately, and a pitch shift does not touch it.

So the question is never "is this unrecognisable" in the abstract. It is "unrecognisable to whom". To a stranger, a few semitones is usually enough. To a colleague, a family member or anybody who has spent an hour on a call with you, it frequently is not.

And it is reversible

A shift is one number. Anybody who suspects a recording has been pitched can shift it back by the same interval and hear the original — in this tool, or in any audio editor, in about ten seconds. There is no key and nothing to break. It is a disguise, in the sense that a hat is a disguise.

Which means it should be read as protection from a listener who is not looking for you, and never as anonymity from one who is.

What actually protects a recording

If the point is that a particular person's voice should not be identifiable, the reliable answers are not audio processing:

What you doWhat it costs, what it gives
Send a transcript instead of the audioThe strongest option by a distance: there is no voice in a text file. SCRIBE transcribes in your browser without uploading the recording. You lose tone, and tone is often the evidence
Have somebody else read it aloudKeeps the words and the delivery, removes the speaker entirely. Standard practice in broadcasting for a reason
Remove the identifying passagesHUSH bleeps, silences or garbles named words. Useful when the problem is what is said rather than who said it
Pitch shiftCheap, fast, reversible, and defeats a stranger only

What this is and is not

VOICEGUARD shifts a recording's pitch in your browser and exports a WAV. Nothing is uploaded, and it works with your connection off.

It is not voice anonymisation in the sense a researcher would mean it, it does not separate a speaker from the way they speak, and it does not touch the file's metadata or the background of the room. If somebody being identified from a recording would do them real harm, treat the audio as unusable and work from a transcript.

Questions people ask about Can you make a voice unrecognisable by changing the pitch?

Does a pitch shift make a voice unrecognisable?

To a stranger, usually. To somebody who knows you, often not. The shift moves the fundamental frequency, which is what a listener reads first and is most confident about, and a few semitones is enough that a stranger will not place the voice. It leaves the accent, the vocabulary, the phrasing, the pauses and the laugh exactly where they were, and those are what a familiar listener is actually recognising.

How much does it move?

Exactly the interval you ask for. A one-second tone at 220 Hz shifted down four semitones comes out at 174.6 Hz and 1.26 seconds long — the interval is a ratio of two to the power of minus four twelfths and the duration is the same ratio the other way up. That is measured on every build rather than asserted.

Why does the recording get longer?

Because the shift is done the way a tape machine does it: the samples are played slower, so pitch and duration move together. It is usually called a defect of the cheap method, and for disguise it is the opposite — the tempo and rhythm of somebody's speech is itself identifying, and this moves that too.

Can somebody shift it back?

Yes, in about ten seconds. A shift is one number, so anybody who suspects a recording has been pitched can shift it back by the same interval and hear the original, in this tool or any audio editor. There is no key and nothing to break. It is a disguise in the sense that a hat is a disguise.

What does it not hide?

How you speak — accent, vocabulary, sentence shape, what you say while thinking. The non-speech sounds, which people recognise more reliably than voices: a laugh, a cough, a breath before a hard word. The room, including reverberation, a particular fan, a clock, traffic. And the file's own metadata, which is a separate job for AVSCRUB.

What actually protects somebody on a recording?

Not audio processing. Sending a transcript instead is the strongest answer by a distance, because there is no voice in a text file — SCRIBE transcribes in your browser without uploading anything, and what you lose is tone, which is sometimes the evidence. Having somebody else read it aloud keeps the words and removes the speaker, which is why broadcasters do it. HUSH helps when the problem is what was said rather than who said it.

When should I not rely on this at all?

When somebody being identified from the recording would do them real harm. Treat the audio as unusable in that case and work from a transcript. A pitch shift is protection from a listener who is not looking for you, never anonymity from one who is.

Related tools