Settings

1. Add the audio or video

Most common formats work: mp3, m4a, wav, ogg, webm, mp4, mov. It is read straight from your device into this tab. Nothing is uploaded.

Drop the recording here, or choose it Read in memory only. For now, up to about 30 minutes works best.

The first time you transcribe, SCRIBE downloads the speech model once (about 40 MB of program code, not your audio) and caches it. After that it works with your internet off.

2. Your transcript

What this is, and what it is not

It runs on your device. The speech model is downloaded once as program code and then runs inside this tab on your own processor. Your recording is decoded and transcribed here. It is never sent anywhere. Open the privacy badge and watch: transcribing a file adds nothing to what the page uploads.

The model download is code, not your audio. The first run fetches the model from a public code CDN, the same way a website loads a font or a library. That is the model's weights, not your file. Once it is cached, you can turn your internet off and SCRIBE still works.

It is a small, fast model. This is the compact Whisper model, chosen so it downloads quickly and runs on an ordinary laptop. It is very good, not perfect: expect the odd wrong word, especially with heavy accents, crosstalk or background noise. Always read the transcript before you rely on it.

Longer files take longer. A minute of audio transcribes in seconds on most machines; a long recording can take several minutes, and the tab may pause while it works. Keep the tab open. Shorter clips are snappy.