Skip to content
VaakyoDocs
Navigation
Open console →

Agents

Voice tab

The Voice tab sets how the agent sounds: the text-to-speech provider and model, the voice, the language it speaks and how fast it talks. Three providers are available: Cartesia (the default, voices in many languages), Sarvam AI (voices built for Hindi, Hinglish and 10 more Indian languages) and ElevenLabs (your ElevenLabs account's voices, cloned ones included, in 32 languages).

Fields

All fields live under voice in the agent object.

FieldTypeDefaultAllowedWhat it does
providerstringcartesiacartesia, sarvam, elevenlabsThe text-to-speech provider. Your workspace must be allowed to use it.
modelstringsonic-3.6Cartesia: sonic-3.6, sonic-3.5, sonic-3. Sarvam: bulbul:v3. ElevenLabs: eleven_flash_v2_5 (its default, the quickest), eleven_turbo_v2_5, eleven_multilingual_v2The voice model. If it doesn’t match the provider, Vaakyo replaces it with the provider’s default.
voice_idstring""Cartesia: a voice id from GET /api/v1/voices. Sarvam: a speaker name from GET /api/v1/voices?provider=sarvam. ElevenLabs: a voice id from GET /api/v1/voices?provider=elevenlabsThe voice. If empty, Cartesia uses the first voice it lists for language, Sarvam its default speaker (shubh), and ElevenLabs the first voice of the account’s list. An unknown Sarvam speaker returns 422.
voice_namestring""anyA label for the console. It does not change the sound.
languagestringhiCartesia: a language id from GET /api/v1/catalog (hi, en by default). Sarvam: one of sarvam_tts_languages (hi-IN, en-IN, bn-IN, gu-IN, kn-IN, ml-IN, mr-IN, od-IN, pa-IN, ta-IN, te-IN). ElevenLabs: one of elevenlabs_tts_languages (hi, en, ta, fr, …)The language the voice speaks. hi (or hi-IN) also tells the agent to speak Hinglish in Roman script. For Sarvam, hi and en are accepted and saved as hi-IN and en-IN; for ElevenLabs, hi-IN is saved as hi.
speednumber1.00.6 to 1.5Speaking rate. 1.0 is normal. On Sarvam it is the pace. ElevenLabs takes 0.7 to 1.2, so faster or slower values are capped there.
fallbacksarray[]up to 3Backup voices, tried in order when the voice fails. See Fallback voices.
{
  "voice": {
    "provider": "cartesia",
    "model": "sonic-3.6",
    "voice_id": "a0e99841-438c-4a64-b679-ae501e7d6091",
    "voice_name": "Riya",
    "language": "hi",
    "speed": 1.05
  }
}

Sarvam example

{
  "voice": {
    "provider": "sarvam",
    "model": "bulbul:v3",
    "voice_id": "priya",
    "voice_name": "Priya",
    "language": "hi-IN",
    "speed": 1.0
  }
}

ElevenLabs example

{
  "voice": {
    "provider": "elevenlabs",
    "model": "eleven_flash_v2_5",
    "voice_id": "JBFqnCBsd6RMkjVDRZzb",
    "voice_name": "George",
    "language": "hi",
    "speed": 1.0
  }
}

On Flash and Turbo v2.5 the voice is told the language; Multilingual v2 detects the language from the text. Multilingual v2 sounds the most lifelike but starts speaking later and costs twice as much per character, so Flash v2.5 suits most calls.

Finding a voice

GET /api/v1/voices lists Cartesia voices. Filter with language, gender (as Cartesia names it) and q (a search term). Results are cached for 10 minutes. The ids and names below are examples.

curl "https://api.vaakyo.com/api/v1/voices?language=hi" -H "X-API-Key: $VAAKYO_API_KEY"
[
  {"id": "a0e99841-438c-4a64-b679-ae501e7d6091", "name": "Riya", "description": "Warm, friendly young woman", "gender": "feminine", "language": "hi"}
]

To hear a voice before you pick it, fetch an MP3 preview:

curl -o preview.mp3 \
  "https://api.vaakyo.com/api/v1/voices/a0e99841-438c-4a64-b679-ae501e7d6091/preview?language=hi&model=sonic-3.6&text=Namaste%2C%20main%20Riya%20hoon" \
  -H "X-API-Key: $VAAKYO_API_KEY"

text is up to 300 characters. In the console, the Voice tab shows the voices for the chosen provider and language with a play button on each.

Sarvam speakers

GET /api/v1/voices?provider=sarvam&model=bulbul:v3 lists Sarvam’s speakers. Every speaker speaks every Sarvam language, so language only fills the language field of each result; filter with gender (female, male) or q. bulbul:v3 has 37 speakers. Sarvam recommends priya and ishita (female) and shubh, ratan and mani (male) for a natural, neutral voice; varun is deep and dramatic.

curl "https://api.vaakyo.com/api/v1/voices?provider=sarvam&model=bulbul:v3&gender=female" \
  -H "X-API-Key: $VAAKYO_API_KEY"
[
  {"id": "ishita", "name": "Ishita", "description": "Recommended", "gender": "female", "language": ""},
  {"id": "priya", "name": "Priya", "description": "Recommended", "gender": "female", "language": ""}
]

Preview a speaker with provider=sarvam on the preview route:

curl -o preview.mp3 \
  "https://api.vaakyo.com/api/v1/voices/priya/preview?provider=sarvam&model=bulbul:v3&language=hi-IN" \
  -H "X-API-Key: $VAAKYO_API_KEY"

ElevenLabs voices

GET /api/v1/voices?provider=elevenlabs lists the voices of the ElevenLabs account Vaakyo uses: premade voices, saved library voices and the account’s own (cloned or designed) ones. Every voice speaks every language of its model, so language only sorts the list: voices labelled with that language come first, none are hidden. Filter with gender or q (searches names, descriptions and labels). Each result also has languages, accent, category and preview_url (a hosted sample from ElevenLabs).

curl "https://api.vaakyo.com/api/v1/voices?provider=elevenlabs&language=hi" \
  -H "X-API-Key: $VAAKYO_API_KEY"
[
  {"id": "JBFqnCBsd6RMkjVDRZzb", "name": "George", "description": "british, warm, middle aged", "gender": "male", "language": "en", "languages": ["en"], "accent": "british", "category": "premade", "preview_url": "https://storage.googleapis.com/.../george.mp3"}
]

Preview a voice with provider=elevenlabs on the preview route (model and language as for the agent):

curl -o preview.mp3 \
  "https://api.vaakyo.com/api/v1/voices/JBFqnCBsd6RMkjVDRZzb/preview?provider=elevenlabs&model=eleven_flash_v2_5&language=hi" \
  -H "X-API-Key: $VAAKYO_API_KEY"

Language and Hinglish

When language is hi (Cartesia, ElevenLabs) or hi-IN (Sarvam):

  • the agent’s prompt gets a rule to speak natural Hinglish written in Roman script, never Devanagari, matching how the caller mixes Hindi and English;
  • on Cartesia, text in Roman script is read with Indian English text normalisation (en-IN), and text in Devanagari with Hindi normalisation (hi-IN). Sarvam reads both scripts natively.

For other Indian languages (Tamil, Bengali, Marathi, …), use a Sarvam voice with that language, such as ta-IN, and write the prompt to answer in that language.

Use en for English-only agents. The language list comes from the platform’s settings (languages in GET /api/v1/catalog).

The voice’s language and the transcriber’s language are separate settings. For a Hindi agent on Deepgram, set both voice.language and transcriber.language. See Transcriber tab.

Fallback voices

voice.fallbacks lists up to three backup voices. When the voice fails to connect at the start of a call, or reports an error or drops during the call, the next one takes over for the rest of the call. A reply that was being spoken when the voice failed is spoken again by the backup, so the caller doesn’t miss it.

FieldTypeDefaultWhat it does
providerstringcartesiacartesia, sarvam or elevenlabs. A backup on another provider keeps the call talking if one provider has an outage.
modelstringsonic-3.6The voice model, replaced to fit the provider as for the primary (bulbul:v3 for Sarvam, eleven_flash_v2_5 for ElevenLabs).
voice_idstringrequiredThe backup voice (a Sarvam speaker name for Sarvam).
voice_namestring""A label for the console.
languagestring""The language it speaks. Empty means the same as the primary voice; hi and hi-IN are mapped between providers.

A backup uses the primary voice’s speed. The same model and voice can’t appear twice (primary included); saving returns 422.

{
  "voice": {
    "provider": "cartesia",
    "model": "sonic-3.6",
    "voice_id": "a0e99841-438c-4a64-b679-ae501e7d6091",
    "voice_name": "Riya",
    "language": "hi",
    "speed": 1.0,
    "fallbacks": [
      {"provider": "cartesia", "model": "sonic-3", "voice_id": "f9836c6e-a0bd-460e-9d3c-f7299fa60f94", "voice_name": "Meera", "language": ""}
    ]
  }
}

Each switch is a fallback.used event on the call’s timeline, with the reason. Without fallbacks, a failing voice behaves as before: the call shows an error event, and a voice that can’t connect at all fails the call. See Webhook events.

Fallback voices also help when a voice provider is busy: providers limit how many replies they speak at once, and when the main voice’s provider is full, a fallback voice from another provider speaks that reply (a capacity.fallback event). See Webhook events.

Usage

Every character sent to the voice (Cartesia, Sarvam or ElevenLabs) is counted in the call’s usage.tts_characters, including the welcome message, tool pre_call_messages, the “are you still there?” message and the hang-up message, and replies a fallback voice speaks again.

Esc