โ€” Audio / Speech-to-Text API

Deepgram

Last updated June 21, 2026 ยท Reviewed by ToolForge Editorial

The fastest, most accurate speech-to-text API for developers. Nova-2 sets the bar for real-time transcription.

โ˜… 4.7/5 ยท 200K+ developers ยท Since 2015 ยท $200 free credit to start
$0.0043/min Pay-as-you-go
Try Deepgram โ†’ Read full review

The dev-first speech-to-text API

Deepgram is what Twilio, Spotify, and NASA reach for when they need to transcribe audio at scale. The Nova-2 model delivers 90%+ accuracy on noisy real-world audio where Whisper and Google STT stumble โ€” and it streams results back in under 300ms, fast enough to power live voice agents.

Who it's for: Developers building voice agents, call center analytics, video captioning, podcast search, or any product that needs to ingest hours of audio.

Key features

Nova-2 Industry-leading accuracy

90.4% WER on the Switchboard benchmark โ€” beats Whisper Large-v3 by 18 points on noisy audio. Trained on 1M+ hours of labeled speech.

300ms Real-time streaming

WebSocket-based streaming API. Start receiving words as the speaker says them. Powers voice agents like Vapi, Retell, and Synthflow.

36 langs Multilingual

36 languages out of the box, plus code-switching detection for mixed-language audio. Custom models for jargon-heavy domains (medical, legal, finance).

$0.0043/min Predictable pricing

Pay-as-you-go at $0.0043/min for Nova-2. Volume discounts to $0.0025/min at 1M+ minutes. $200 free credit on signup โ€” that's 46,500 minutes of transcription.

The honest take

โœ“ What works

  • Fastest streaming STT on the market โ€” sub-300ms latency
  • Beats Whisper on noisy audio (call centers, accents, crosstalk)
  • 36 languages + custom model training
  • Generous free credit ($200) โ€” enough to ship a prototype
  • Solid SDKs for Python, Node, Go, Rust

โœ— What doesn't

  • Not the cheapest at low volumes (AssemblyAI is $0.0025/min flat)
  • No self-hosted option โ€” pure API only
  • Free tier is credit-based, expires after 12 months
  • Voice Agent API is still in beta as of mid-2026

Verdict

If you're building anything voice-first in 2026 โ€” voice agents, transcription, call analytics โ€” Deepgram is the default. It's the fastest, the most accurate on hard audio, and the developer experience is the best in class. The only reason to skip it: you need on-prem deployment, or your volume is so high that the per-minute math hurts (then negotiate or look at self-hosted Whisper.cpp).

๐Ÿ’ก Transparency: This review contains affiliate links. If you sign up through our link, we may earn a commission at no cost to you. We only recommend tools we use ourselves. Full disclosure.

Related Tools

Try Deepgram today

$0.0043/min pay-as-you-go ยท $200 free credit

Get Deepgram โ†’