Skip to main content

elevenLabsTranscriptToCaptions()v4.0.443

Converts an ElevenLabs Speech to Text API response or a JSON transcript exported from ElevenLabs into an array of Caption objects.

This function can be used in any JavaScript environment, but you should not use the ElevenLabs API in the browser because your API key will be exposed.

Example​

When calling the ElevenLabs Speech to Text API, you must set timestamps_granularity to "word" to include word-level timing in the response.

Example usage
import fs from 'fs'; import {elevenLabsTranscriptToCaptions} from '@remotion/elevenlabs'; const form = new FormData(); form.append('file', new Blob([fs.readFileSync('audio.mp3')])); form.append('model_id', 'scribe_v2'); form.append('timestamps_granularity', 'word'); const response = await fetch('https://api.elevenlabs.io/v1/speech-to-text', { method: 'POST', headers: { 'xi-api-key': process.env.ELEVENLABS_API_KEY!, }, body: form, }); const transcript = await response.json(); const {captions} = elevenLabsTranscriptToCaptions({transcript});

API​

Arguments​

An object with the following property:

transcript​

Pass the parsed JSON object in either of the following formats. The function converts it locally without calling the ElevenLabs API.

Speech-to-Text API response​

The response must include a words array. Set timestamps_granularity to "word" when calling the API, as shown in the example above.

Speech-to-Text response (required fields)
{ "language_code": "en", "words": [ {"text": "Hello", "type": "word", "start": 0, "end": 0.5} ] }

Each "word" entry becomes one caption. start and end are measured in seconds from the start of the audio.

Entries with type: "audio_event" are skipped. Entries with type: "spacing" do not produce captions, but their start times are used as the start times of the following words.

JSON exportv4.0.530​

You can also pass a JSON transcript exported from ElevenLabs. It must contain a language_code field (which can be null) and a segments array:

JSON export without word-level timing
{ "language_code": "en", "segments": [ { "text": "Hello world", "start_time": 0, "end_time": 1 } ] }

Each segment becomes one caption unless it contains a non-empty words array. In that case, each word becomes a caption instead:

JSON export with word-level timing
{ "language_code": "en", "segments": [ { "text": "Hello world", "start_time": 0, "end_time": 1, "words": [ {"text": "Hello", "start_time": 0, "end_time": 0.5}, {"text": " world", "start_time": 0.5, "end_time": 1} ] } ] }

All start_time and end_time values are measured in seconds from the start of the audio, not from the start of the segment.

Invalid input​

The function throws an error if required fields are missing or invalid, timestamps are negative, or an end time is earlier than its start time.

Return value​

An object with the following property:

captions​

An array of Caption objects.

Compatibility​

BrowsersServersEnvironments
Chrome
Firefox
Safari
Node.js
Bun
Serverless Functions

See also​