Speech-to-Text
Convert speech to text using AI
Configuration
OpenAI Whisper
Provider*
OpenAI Whisper
Audio/Video File*
Upload an audio or video file
Audio/Video File Reference*
Reference audio/video from previous blocks
Audio/Video URL
Or enter publicly accessible audio/video URL
Language*
Auto-detect
Timestamps*
None
API Key*
••••••••
Model*
Whisper-1
Translate to English
Disabled
OpenAI Whisper (
stt_whisper)Transcribe audio to text using OpenAI Whisper
Input
| Parameter | Type | Required | Description |
|---|---|---|---|
provider | string | Yes | STT provider (whisper) |
apiKey | string | Yes | OpenAI API key |
model | string | No | Whisper model to use (default: whisper-1) |
audioFile | file | No | Audio or video file to transcribe |
audioFileReference | file | No | Reference to audio/video file from previous blocks |
audioUrl | string | No | URL to audio or video file |
language | string | No | Language code (e.g., "en", "es", "fr") or "auto" for auto-detection |
timestamps | string | No | Timestamp granularity: none, sentence, or word |
translateToEnglish | boolean | No | Translate audio to English |
prompt | string | No | Optional text to guide the model's style or continue a previous audio segment. Helps with proper nouns and context. |
temperature | number | No | Sampling temperature between 0 and 1. Higher values make output more random, lower values more focused and deterministic. |
Output
| Parameter | Type | Description |
|---|---|---|
transcript | string | Full transcribed text |
segments | array | Timestamped segments |
language | string | Detected or specified language |
duration | number | Audio duration in seconds |
Deepgram
Provider*
Deepgram
Audio/Video File*
Upload an audio or video file
Audio/Video File Reference*
Reference audio/video from previous blocks
Audio/Video URL
Or enter publicly accessible audio/video URL
Language*
Auto-detect
Timestamps*
None
API Key*
••••••••
Model*
Nova 3
Speaker Diarization
Disabled
Deepgram (
stt_deepgram)Transcribe audio to text using Deepgram
Input
| Parameter | Type | Required | Description |
|---|---|---|---|
provider | string | Yes | STT provider (deepgram) |
apiKey | string | Yes | Deepgram API key |
model | string | No | Deepgram model to use (nova-3, nova-2, whisper-large, etc.) |
audioFile | file | No | Audio or video file to transcribe |
audioFileReference | file | No | Reference to audio/video file from previous blocks |
audioUrl | string | No | URL to audio or video file |
language | string | No | Language code (e.g., "en", "es", "fr") or "auto" for auto-detection |
timestamps | string | No | Timestamp granularity: none, sentence, or word |
diarization | boolean | No | Enable speaker diarization |
Output
| Parameter | Type | Description |
|---|---|---|
transcript | string | Full transcribed text |
segments | array | Timestamped segments with speaker labels |
language | string | Detected or specified language |
duration | number | Audio duration in seconds |
confidence | number | Overall confidence score |
ElevenLabs
Provider*
ElevenLabs
Audio/Video File*
Upload an audio or video file
Audio/Video File Reference*
Reference audio/video from previous blocks
Audio/Video URL
Or enter publicly accessible audio/video URL
Language*
Auto-detect
Timestamps*
None
API Key*
••••••••
Model*
Scribe v1
ElevenLabs (
stt_elevenlabs)Transcribe audio to text using ElevenLabs
Input
| Parameter | Type | Required | Description |
|---|---|---|---|
provider | string | Yes | STT provider (elevenlabs) |
apiKey | string | Yes | ElevenLabs API key |
model | string | No | ElevenLabs model to use (scribe_v1, scribe_v1_experimental) |
audioFile | file | No | Audio or video file to transcribe |
audioFileReference | file | No | Reference to audio/video file from previous blocks |
audioUrl | string | No | URL to audio or video file |
language | string | No | Language code (e.g., "en", "es", "fr") or "auto" for auto-detection |
timestamps | string | No | Timestamp granularity: none, sentence, or word |
Output
| Parameter | Type | Description |
|---|---|---|
transcript | string | Full transcribed text |
segments | array | Timestamped segments |
language | string | Detected or specified language |
duration | number | Audio duration in seconds |
confidence | number | Overall confidence score |
AssemblyAI
Provider*
AssemblyAI
Audio/Video File*
Upload an audio or video file
Audio/Video File Reference*
Reference audio/video from previous blocks
Audio/Video URL
Or enter publicly accessible audio/video URL
Language*
Auto-detect
Timestamps*
None
API Key*
••••••••
Model*
Best
Speaker Diarization
Disabled
Sentiment Analysis
Disabled
Entity Detection
Disabled
PII Redaction
Disabled
Auto Summarization
Disabled
AssemblyAI (
stt_assemblyai)Transcribe audio to text using AssemblyAI with advanced NLP features
Input
| Parameter | Type | Required | Description |
|---|---|---|---|
provider | string | Yes | STT provider (assemblyai) |
apiKey | string | Yes | AssemblyAI API key |
model | string | No | AssemblyAI model to use (default: best) |
audioFile | file | No | Audio or video file to transcribe |
audioFileReference | file | No | Reference to audio/video file from previous blocks |
audioUrl | string | No | URL to audio or video file |
language | string | No | Language code (e.g., "en", "es", "fr") or "auto" for auto-detection |
timestamps | string | No | Timestamp granularity: none, sentence, or word |
diarization | boolean | No | Enable speaker diarization |
sentiment | boolean | No | Enable sentiment analysis |
entityDetection | boolean | No | Enable entity detection |
piiRedaction | boolean | No | Enable PII redaction |
summarization | boolean | No | Enable automatic summarization |
Output
| Parameter | Type | Description |
|---|---|---|
transcript | string | Full transcribed text |
segments | array | Timestamped segments with speaker labels |
language | string | Detected or specified language |
duration | number | Audio duration in seconds |
confidence | number | Overall confidence score |
sentiment | array | Sentiment analysis results |
entities | array | Detected entities |
summary | string | Auto-generated summary |
Google Gemini
Provider*
Google Gemini
Audio/Video File*
Upload an audio or video file
Audio/Video File Reference*
Reference audio/video from previous blocks
Audio/Video URL
Or enter publicly accessible audio/video URL
Language*
Auto-detect
Timestamps*
None
API Key*
••••••••
Model*
Gemini 2.5 Flash
Google Gemini (
stt_gemini)Transcribe audio to text using Google Gemini with multimodal capabilities
Input
| Parameter | Type | Required | Description |
|---|---|---|---|
provider | string | Yes | STT provider (gemini) |
apiKey | string | Yes | Google API key |
model | string | No | Gemini model to use (default: gemini-2.5-flash) |
audioFile | file | No | Audio or video file to transcribe |
audioFileReference | file | No | Reference to audio/video file from previous blocks |
audioUrl | string | No | URL to audio or video file |
language | string | No | Language code (e.g., "en", "es", "fr") or "auto" for auto-detection |
timestamps | string | No | Timestamp granularity: none, sentence, or word |
Output
| Parameter | Type | Description |
|---|---|---|
transcript | string | Full transcribed text |
segments | array | Timestamped segments |
language | string | Detected or specified language |
duration | number | Audio duration in seconds |
confidence | number | Overall confidence score |
Usage Instructions
Transcribe audio and video files to text using leading AI providers. Supports multiple languages, timestamps, and speaker diarization.
Notes
- Category:
tools - Type:
stt