AI + Data

Voice AI accuracy in regulated services — why the speech model underneath matters more than you might think

Routine call handling is consuming staff time across regulated service organisations. AI voice is a credible answer — but only if the technology can handle your vocabulary.

Published · Taidotech

Routine call handling is consuming staff time across regulated service organisations. AI voice is a credible and low-cost answer — but only if the technology can handle your vocabulary.

Coordinators and administrators across regulated service organisations spend a significant part of their day on calls that follow a predictable pattern. A patient calling a hospital outpatient clinic to reschedule an appointment. A participant querying their remaining NDIS plan balance with a provider's finance team. A support worker calling to confirm a schedule change before a shift. A family member checking whether a home visit is still going ahead.

None of these calls require professional judgement. All of them consume staff time that could be going to work that does.

AI voice platforms can handle these interactions — routing the right information, confirming appointments, logging amendment requests, escalating anything that genuinely needs a person. The cost per interaction is a fraction of staff time, the platform runs around the clock, and every call is automatically documented. For organisations under financial pressure and workforce constraints, the economics are straightforward.

The part that is less straightforward is making the technology actually work in a regulated service environment.

The vocabulary problem nobody warns you about

Speech-to-text models are the ears of a voice AI platform — they convert what the caller says into text the system can act on. General-purpose models handle everyday English well. They handle the vocabulary that regulated service environments generate every day considerably less reliably.

A patient asking to reschedule through a hospital's outpatient management system may have the system name misheard entirely. An NDIS participant querying their Support Coordination budget may find that term not recognised. A support worker calling to amend a visit in a scheduling platform may have the platform name resolved incorrectly — meaning the system cannot find the record and the call fails.

In most software, misrecognition is a minor nuisance. In a regulated service environment it is a different category of problem. The platform cannot complete the task, the caller is frustrated, and a staff member still has to step in. That is not a theoretical risk — it is the failure mode that turns a promising pilot into a shelved project.

What Microsoft's new models change

At Microsoft Build 2026, Microsoft released MAI-Transcribe-1.5 and MAI-Voice-2 — in-house foundation models for speech-to-text and text-to-speech, available through Azure Speech. Both are the models already running inside Copilot, Teams, and PowerPoint at enterprise scale.

MAI-Transcribe-1.5 currently holds the number one position on the FLEURS multilingual benchmark with a word error rate of 3.7% — roughly half that of the previous industry standard. For a voice AI platform handling clinical and service vocabulary, that accuracy gap is material.

Word error rate comparison: MAI-Transcribe-1.5 vs alternatives (FLEURS benchmark, June 2026)
Figure 1 — Word error rate comparison: MAI-Transcribe-1.5 vs alternatives (FLEURS benchmark, June 2026)

The more significant capability for regulated service environments is entity biasing. Before a call is processed, the platform provides the model with a list of domain-specific terms — booking system names, clinical terminology, service types, staff names, regulatory acronyms. The model uses that context to resolve ambiguous audio toward the right term, without forcing incorrect matches. Microsoft reports a 30% reduction in word error rate for domain-specific vocabulary when entity biasing is applied.

In practice: an outpatient booking system name is heard correctly. An NDIS plan management term resolves as intended. A scheduling amendment gets logged without coordinator intervention.

Entity biasing in practice: service-sector vocabulary recognition with and without keyword priming
Figure 2 — Entity biasing in practice: service-sector vocabulary recognition with and without keyword priming

Where these models sit in a voice AI platform

MAI-Transcribe-1.5 and MAI-Voice-2 are not a voice AI platform on their own. They are the speech input and output layers — the ears and voice — in a broader architecture that still requires orchestration, workflow logic, system integration, and governance. The speech model is infrastructure. The platform is the product.

Where MAI models sit in a governed voice AI platform stack
Figure 3 — Where MAI models sit in a governed voice AI platform stack

Evaluating a voice AI platform is not the same as evaluating a speech model. Transcription accuracy matters — but the governance layer above it, how calls are routed, how decisions are made, how escalations are handled under your regulatory framework, is where the real complexity sits.

The Australian data question — and what it actually covers

When a caller interacts with your voice AI platform, the audio is processed somewhere. For Australian providers with Privacy Act obligations, that somewhere matters — and most procurement conversations collapse two distinct questions into one.

Data at rest — transcripts, audio recordings, and call logs — is stored in the region where your Azure service was provisioned. For AU East provisioning, that stored data remains in Australia. This is a genuine, citable Microsoft commitment, and it means the evidence trail that regulators ask for sits on Australian soil.

Inference during the call — the real-time processing of audio while the caller is speaking — is a separate question with a potentially different answer. It may route outside Australia depending on model and service configuration. Written confirmation from your vendor, scoped to the specific model and service components you are deploying, is the right standard to hold them to.

Conflating the two — or accepting a vendor's assurance on stored data as covering real-time processing — creates a gap that a regulator or privacy commissioner may not accept.

Three questions worth adding to any evaluation

  • Can the platform be configured with our specific vocabulary — system names, service types, clinical or regulatory terms? How is that list maintained?
  • Where is audio processed in real time, and can that be confirmed in writing by service component?
  • Where are transcripts and recordings stored at rest — and in which region is the service provisioned?

The providers that ask these questions before signing get better outcomes than the ones who ask them after go-live.

The bottom line

AI voice is a practical, low-cost way to handle the high volume of routine calls that regulated service organisations field every day — appointment changes, balance queries, visit confirmations, schedule amendments. The technology is ready. The question is whether the speech model underneath it is ready for your environment specifically.

MAI-Transcribe-1.5 and MAI-Voice-2 are the strongest in-house voice models Microsoft has released. Entity biasing in particular closes a gap that has made accurate voice AI genuinely difficult in regulated service settings. Both are currently in public preview — not yet recommended for production workloads — but they represent a clear near-term upgrade for platforms already built on Azure. Worth watching, and worth asking your voice AI partner about now.

Let's talk about your operation.

Tell us what you're working on. We'll scope a practical next step and say plainly where we can help.