Skip to main content
DeepInfra hosts Whisper and other speech recognition models. Given an audio file, they produce transcribed text with per-sentence timestamps. Browse all speech recognition models.

Models

Browse all ASR models.

Example

Supported audio formats

  • mp3
  • wav

Response

Additional parameters

Each model exposes different parameters (language, task, etc.). Check the model’s API documentation page for details.

Tutorial

See the Whisper tutorial for a complete walkthrough.