← Back to all generators
thomasmol/whisper-diarization
OfficialView on Replicate →
⚡️ Blazing fast audio transcription with speaker diarization | Whisper Large V3 Turbo & pyannote 4.0 community-1 | word & sentence level timestamps | prompt
Capabilities
No capability data available
Cost
Community model (estimated from hardware time)
Input Parameters
| Name | Type | Description | Default | Constraints |
|---|---|---|---|---|
file | string(uri) | Or an audio file | — | — |
file_string | string | Either provide: Base64 encoded audio file, | — | — |
file_url | string | Or provide: A direct audio file URL | — | — |
language | string | Language of the spoken words as a language code like 'en'. Leave empty to auto detect language. | — | — |
num_speakers | integer | Number of speakers, leave empty to autodetect. | — | min: 1, max: 50 |
prompt | string | Vocabulary: provide names, acronyms and loanwords in a list. Use punctuation for best accuracy. | — | — |
translate | boolean | Translate the speech into English. | false | — |
filestringOr an audio file
file_stringstringEither provide: Base64 encoded audio file,
file_urlstringOr provide: A direct audio file URL
languagestringLanguage of the spoken words as a language code like 'en'. Leave empty to auto detect language.
num_speakersintegerNumber of speakers, leave empty to autodetect.
min: 1, max: 50
promptstringVocabulary: provide names, acronyms and loanwords in a list. Use punctuation for best accuracy.
translatebooleanTranslate the speech into English.
Default:
falseVersion:
744c4f2bffaeUpdated: 7/25/20268.5M runs
cinemasetfree.com