Feature availability
Availability depends on the combination of interaction pattern, model, and deployment, so check the combination rather than the product name. This page covers SaaS on Cloud. For container and virtual appliance deployments, see On-prem availability.
A check mark marks an available combination. An em dash means the combination is either unavailable or not established, and carries no roadmap position either way.
Readiness by combination
Readiness is a separate question from availability: availability says whether something exists for a combination, readiness says whether that combination is usable and documentable. Readiness is a property of the combination, so it is stated here once and not repeated on individual feature pages.
Preview applies to SaaS on Cloud only, because on-prem ships as versioned containers: the question there is which release you need.
Streaming with Melia 1, and agent STT with Linden 1, are available for evaluation and feedback. They are not production-ready and not ready to scale.
Pre-recorded transcription features
Pre-recorded transcription runs on the Standard, Enhanced, and Melia 1 models, in the EU, US, and AUS regions.
Language coverage and selection
Output and formatting
Transcript content and tagging
Speakers and channels
Audio input
Operational
Add-ons
Add-ons are separate products that produce an output derived from a completed transcript. They are selected in addition to transcription.
Audio alignment is available to Enterprise customers only.
Pre-recorded regions
Streaming transcription features
Streaming transcription runs on the Standard, Enhanced, and Melia 1 models. Streaming is available in the EU and US regions, and is not available in the AUS region. Melia 1 for streaming is in Preview.
Streaming language coverage and selection
Streaming output and formatting
Streaming transcript content and tagging
Streaming speakers and channels
Streaming audio input
Responding
These features control when the Realtime API returns a result, and how much of the transcript each result contains.
Streaming operational and add-ons
Agent STT features
Agent STT runs on the Linden 1 model in the EU and US regions, and is in Preview. Because agent STT offers a single model, the following lists name what Linden 1 supports rather than comparing columns.
Language and output
- 56 languages
- Transcription language packs (including bilingual)
- Output locale
- Smart formatting
- Punctuation and casing
- Segment-level timings
Transcript content and tagging
- Custom dictionary
- Medical domain
- Entity detection (basic, legacy)
- Text replacement (find and replace)
Speakers
- Speaker diarization
- Speaker identification
Responding
- Turn detection
- Voice activity detection (VAD)
- Force end of utterance
- Segment-level partials
Audio input
- Audio filtering (volume filtering)
Agent STT regions
Agent STT processes audio in the EU and US regions.