Scribe v2
ElevenLabs
· scribe-v2
Speech-to-Text
Production
Recommended
State-of-the-art speech-to-text with word-level timestamps, diarization and 99-language coverage.
Best use cases: Subtitling, interview transcription, voiceover QA
9
Quality
8.5
Speed
8
Cost Eff.
9
Brand Safety
7
Control
9
Adherence
Input cost
—
Output cost
—
Generation cost
$0.22 / audio hour
Max context / duration
10h per file
Aspect ratios
—
Resolutions
—
Input types
Audio, video
Output types
Text, timestamps
API endpoint
api.elevenlabs.io/v1/speech-to-text
Commercial use
Allowed
Data retention
Per provider policy
Regions available
Global
API connected
Connected
Last updated
Release date
Documentation