Skip to main content

@remotion/whisper-webgpuv4.0.518

Transcribe audio locally in the browser or Node.js using timestamped Whisper models and WebGPU through Transformers.js, and convert the result to @remotion/captions.

Installation​

npx remotion add @remotion/whisper-webgpu @huggingface/transformers

Example​

transcribe.ts
import {clearStaleModels, downloadWhisperModel, resampleTo16Khz, toCaptions, transcribe} from '@remotion/whisper-webgpu'; export const transcribeFile = async (file: File) => { await clearStaleModels(); await downloadWhisperModel({model: 'small.en'}); const channelWaveform = await resampleTo16Khz({file}); const transcription = await transcribe({ channelWaveform, model: 'small.en', language: 'en', }); const {captions} = toCaptions({whisperWebGpuOutput: transcription}); return captions; };

Model hosting​

We noticed that hosting the models on R2 leads to faster downloading than through Hugging Face directly.
We mirrored the models to remotion.media, keeping them byte-identical.
We also notified Hugging Face and they are investigating the issue. We may migrate back to the official hosting in the future.

On the server​

See @remotion/whisper-webgpu in Node.js for a complete Node.js example.

APIs​

Requirements​

You obviously need a GPU.
Use canUseWhisperWebGpu() to determine if it is supported.

License​

MIT