Local transcription with Whisper, no upload
This tool uses a multilingual Whisper model running entirely in your browser, through WebGPU when available (with an automatic fallback to CPU processing). The audio or video file is never sent to our server: audio extraction, normalization, and inference all happen locally on your device.
Only the AI model files (Whisper weights) are downloaded from Hugging Face the first time you use the tool. After that, the browser caches those files, so future transcriptions do not require a new download.