Skip to main content
Languages, voices, models, and transcribers define the conversation stack for one agent version. To a caller, that stack becomes a simple experience: did the agent understand me, reply naturally, use the right facts, and finish the task? This page helps you choose those controls as one working system instead of four disconnected dropdowns. DialNexa agent builder controls for voice, model, transcriber, language, and pricing preview. DialNexa voice selector modal showing language, gender, accent, provider, search, voice preview, Nexa voice ID, and per-row language selection.
Do not choose the stack like a buffet. A great voice, wrong language, and impatient transcriber will still produce a strange call.

Before You Begin

Define the caller language, expected accents, background conditions, response-time goal, and budget. Prepare one representative test script and recording criteria before comparing providers.

DialNexa Language Voice Model And Transcriber Selectors

Selection Order That Avoids Rework

1

Pick the caller language first

Start with English, Hindi, Hindi-English, or another enabled language. Language decides which voice rows are useful and which transcribers are valid.
2

Choose the transcriber

For cascaded agents, use Deepgram Flux when quick turn boundaries matter and the selected language is supported, or Soniox for Hindi-English, multilingual, Indian English, or accent-heavy calls that need code-switching strength. If primary STT latency is a concern, open the transcriber settings and configure fallback STT.
3

Choose the voice

Open the voice modal, filter by language, listen to samples, copy the Nexa voice ID if needed, and choose the row language before clicking Use Voice.
4

Choose the model

Pick the LLM model after the listening and speaking layer is sensible. A strong model cannot reason over words the transcriber never heard.
5

Review pricing preview

Check visible INR rates for LLM, transcriber, and voice model where available. Telephony is still separate.

Compatibility Rules The Builder Applies

Voice Selector Details Users Should Know

The voice modal is more than a list of pretty names.

Speech To Speech Uses A Different Stack

Speech to Speech agents use a realtime speech model for both listening and speaking. Choose this pipeline when low turn latency matters more than separate control over STT and TTS providers.

Speech To Speech Agents

Set up OpenAI realtime and Gemini Speech to Speech model paths.

Fallback STT

Fallback STT runs a backup transcriber alongside the primary transcriber for cascaded agents. Enable it when call quality, accents, or provider latency make recognition reliability more important than the lowest possible transcription cost. DialNexa transcriber settings popover showing fallback STT enabled with a fallback transcriber selected and a fallback wait value. The fallback transcriber must differ from the primary transcriber. The dashboard removes the selected primary option from the fallback list, and API requests that send the same transcriber for both return 400 Bad Request. Fallback STT is version-specific. Published versions lock the setting, so create or edit a draft version before changing fallback STT in production. The selector currently shows fallback choices as ₹0.00/min. Review the Billing transaction breakdown after test calls because the saved call charge can include a separate fallback-transcriber line item resolved from the workspace plan.

Stack Recommendations

Add The Action Layer Only After The Stack Works

Once the caller can be heard and answered correctly, connect the call to the system that should receive the result.

What To Compare Before Publishing

Ask callers to use the language mix you expect in production. Do not approve a Hindi-English agent from a polished English-only demo.

Verify The Conversation Stack

Run the same script against each candidate stack. Compare transcription accuracy, response delay, interruption handling, pronunciation, voice consistency, and cost using Call History evidence.

Recap

Choose the stack from the caller’s language and call conditions, then compare candidates with the same script. Publish only after transcription, latency, voice quality, interruption behavior, and cost meet the use case.

Choose Voice AI Providers

Decide what to use when.

Speech To Text

Compare Deepgram and Soniox.

Text To Speech

Tune ElevenLabs and Cartesia voices.

LLMs And Conversation Behavior

Tune model behavior and fallback.

Speech To Speech Agents

Compare OpenAI realtime and Gemini model paths.

Synthesiser Settings Video

Watch the recommended voice configuration walkthrough.

LLM And Transcriber Video

Watch the recommended conversation and transcription setup.