Skip to main content
Speech to text in DialNexa is the part of a Voice AI call that decides what the system believes the caller said. If this layer is wrong, every downstream feature inherits the mistake: the model answers the wrong question, functions receive wrong arguments, summaries sound confident but false, and post-call fields become review work. DialNexa transcriber selector showing the Soniox option for the selected agent language. DialNexa call detail page showing summary, transcription tabs, live transcript, and accurate transcript controls.
If the transcript says the caller asked about carrots when they clearly asked about careers, do not rewrite the whole prompt yet. Start with listening.

DialNexa Speech To Text And Transcription Options

The selector can show different options depending on the agent language. Pricing per minute for each option is shown next to the selector in the dashboard. Cascaded agents can also use fallback STT from the transcriber settings popover. The current dashboard selector does not expose AssemblyAI for new primary or fallback choices. For help comparing transcription providers, see the voice AI provider selection guide. For how transcription evidence appears after the call, see transcripts, recordings, and summaries.

Deepgram Flux Versus Soniox

Start with Deepgram Flux for English calls where fast turn boundaries matter. For supported non-English or mixed-language tests, DialNexa can use multilingual Flux with language hints derived from the agent language.

What DialNexa Configures Behind The Selector

These details explain the behavior users see in the dashboard and call records.

Boosted Keywords For Speech To Text

Boosted Keywords are recognition hints for supported cascaded transcribers. Add short terms that the caller is likely to say and that the transcriber might otherwise miss. The Speech Settings panel validates Boosted Keywords before saving. Terms can contain letters, numbers, and spaces only. DialNexa preserves multi-word phrases, removes duplicate terms case-insensitively, and sends the resulting list to Deepgram or Soniox when the selected transcriber supports it.

Fallback STT

Fallback STT helps when the primary transcriber is slow or unreliable for a specific caller population. When enabled, DialNexa can use a backup transcriber if it finalizes first and the primary result does not arrive within the configured wait. DialNexa transcriber settings popover showing fallback STT enabled with a fallback transcriber selected and a fallback wait value.
Fallback STT applies to cascaded agents. Speech to Speech agents do not use a separate STT provider.
The fallback transcriber must have a different transcriber ID from the primary transcriber. The dashboard hides the primary option from the fallback list. Create or update API requests that send the same ID for both return 400 Bad Request. The current fallback selector displays ₹0.00/min. After a test call, review Billing transactions for a separate fallback-transcriber component because the final line item is resolved from the workspace billing plan. DialNexa processes a valid fallback final result only once. If the fallback result is queued again during timing coordination, it is not dropped just because the first pass was deferred.

Transcript Types

DialNexa can show more than one transcript view for the same call.
Text tells you what the system believed. Audio tells you what happened. When they disagree, trust the recording first.

How To Test Speech To Text

1

Use the same caller script

Compare transcribers with the same greeting, caller answers, interruption, name, city, number, and final outcome.
2

Test names and places

Use real customer names, locality names, company names, product terms, and common abbreviations.
3

Test interruptions

Ask a caller to speak during the greeting, correct themselves mid-answer, pause, and give one-word replies.
4

Test language switching

For Hindi-English calls, include English numbers, Hindi phrases, and mixed casual replies in the same call.
5

Compare transcript with recording

Mark the exact point where the transcript diverges from audio.
6

Review downstream effects

Check function arguments, workflow branches, summaries, and post-call fields. Bad listening quietly becomes bad automation.

How Transcription Affects Integrations

Integrations usually receive data that started as caller speech. If the transcript is wrong, a CRM update, ticket note, WhatsApp message, or spreadsheet row can be wrong too.

Common Speech To Text Mistakes

Deepgram Flux and Soniox can both be useful, but they behave differently by language, accent, and noise. Compare them on real calls before standardizing a stack.
The runtime filters language hints to valid Soniox values. For Hindi-English agents, DialNexa sends English and Hindi hints rather than an invalid combined language code. For unsupported broad choices, it avoids strict hints instead of sending a value Soniox cannot use.
Boosted Keywords bias recognition toward important terms. They do not force exact output, so recordings and transcript review are still required.
If the LLM received the wrong caller words, the LLM is not the first problem.
Plivo, SIP trunking, and web calls can produce different audio with the same agent. Compare through the same route when testing transcribers.
Real callers use speaker mode, traffic, low signal, and short answers. Test those before a large campaign does it for you.

Recap

Choose a transcriber using real caller language, accents, noise, and timing requirements. Compare transcript accuracy and response behavior with the same test calls before publishing the selected stack.

Supported Transcribers

Read the reference table for selectable transcribers.

Speech Settings

Tune Response Eagerness and Denoising Mode.

Latency And Turn Taking

Understand how turn boundaries affect response timing.

Transcripts And Recordings

Review call evidence after live calls.