Voice models enable speech-to-text (transcription) and text-to-speech (speech output) in VARIOS AI. Each voice model is assigned a model type that determines the available configuration options and cost fields.
Model Types
Basic Data (Both Types)
Costs by Model Type
Transcription Model Costs
Transcription models process audio inputs and produce text outputs.
Speech Model Costs
Speech models process text inputs and produce audio outputs.
Token costs for voice models differ depending on the model type.
Transcription models have three cost fields (text input, text output,
audio input), while speech models only require two cost fields
(text input, audio output).
Voice models do not have DLP security settings, as data processing is
secured through the associated chat and embedding models.