Accessible voice UX

These principles when followed together support greater numbers of users with diverse and dysfluent speech patterns.

Principles of accessible voice UX

To make voice UX accessible:

1. Provide adaptive pause tolerance

Voice dysfluency can include hesitations, extended pauses and breaks within words or sentences, and these patterns change depending on the situation.

A person may speak fluently in a comfortable setting but experience greater dysfluency in public, unfamiliar or high-pressure situations. For voice interfaces, a pause does not necessarily mean someone has finished speaking.

Acting too quickly can interrupt the person or cause the system to interpret an incomplete statement incorrectly.

Voice interfaces should allow greater tolerance for pauses and recognise the amount of time someone needs may change with the situation. A single fixed pause setting is unlikely to work in every context.

Systems should adapt their listening time, but adaptation should remain understandable and under the user’s control. Voice interfaces should combine adaptive behaviour with clear settings and user overrides.

2. Allow recovery without repeating

A voice interaction can become frustrating when someone has spent significant effort saying something, only to be asked to start again.

If the system misunderstands part of a request, people should be able to repair just that part of the interaction. For example, correcting a word, choosing between possible interpretations or confirming what the system thinks it heard.

Error recovery should not depend entirely on more voice input, especially when voice recognition caused the problem in the first place. Repeated prompts to "say that again" can trap someone in a cycle of correction and misunderstanding.

Voice interfaces should provide alternative ways to recover from errors while preserving what the person has already communicated. It should not force them to repeat themselves or restart the interaction.

3. Allow easy switching between modes

A person may begin an interaction using speech, switch to typing when they reach a difficult word, select an option on screen, and then continue speaking.

Voice interfaces should support these changes naturally rather than assume an interaction which begins with voice must remain voice only.

Progress must be preserved when the person changes how they interact. Switching from speech to text, touch or another input method should not mean starting again or losing context.

4. Provide customisable wake commands

Voice interfaces often rely on a fixed wake word or phrase, but this can create a barrier when it contains sounds that are difficult for someone to say.

A person may block or repeat part of the phrase, meaning the device never hears the complete activation command and does not respond. The problem is not necessarily speech recognition, but the requirement to produce one specific phrase in one specific way.

Voice interfaces should allow people to choose or customise the words they use to activate. Some platforms already provide greater flexibility, while others limit users to a small set of predefined wake phrases.

Giving people control over activation language can reduce unnecessary effort and make voice interaction more reliable, particularly for people whose speech varies depending on particular sounds or words.

5. Provide adaptive dysfluency removal

Dictation can be difficult when natural speech includes filler words, repetitions or other dysfluencies that help someone construct a sentence. Transcribing every sound exactly as spoken may produce text that does not reflect the person’s intended meaning, particularly in formal writing.

Technology can help by identifying likely filler words or unintended repetitions, but this is not as simple as automatically deleting repeated words because repetition can also be grammatically correct and meaningful.

Voice interfaces should distinguish between recognising what was said and refining what the person intended to communicate. Dysfluency should not automatically be treated as noise that needs to be removed.

People should be able to decide whether their speech is transcribed exactly, lightly refined or edited more substantially, with the system making any changes transparent and reversible.

6. Recognise intent, not exact words

Requiring people to say specific words in a specific sequence can create unnecessary barriers. Someone may substitute a difficult word for an easier one while expressing the same meaning.

A voice interface should focus on the person’s intent rather than expecting a precise command. "The weather is really good today" and "the weather is really nice now" communicate essentially the same thing, even though the words are different.

When specific information cannot be substituted, such as a name, address or age. The system should recognise that the requested information may be surrounded by filler words, repetitions or lead-in phrases such as "my name is…" or "my age is…".

Voice interfaces should be flexible to identify the important information within natural patterns of speech rather than requiring people to speak in a prescribed way.

7. Show what’s coming next

Speaking off-the-cuff can be much harder for some people than reading from prepared words, because the person is having to formulate language while also trying to say it.

For some people, this difficulty increases pausing, hesitation and choppy speech.

If a voice interface asks someone to read or repeat content, show all words and content that come next. This gives the person more time to prepare and can reduce the effort involved in speaking.

An example is a teleprompter. Current words are highlighted while upcoming words remain visible. This kind of "ghosting" can make speech tasks easier by helping the person anticipate the next part of the sentence rather than producing it entirely in the moment.

Details

Licensing

This work is licensed under a Creative Commons Attribution-NonCommerical-ShareAlike 3.0 Unported License.

Details

Licensing

This work is licensed under a Creative Commons Attribution-NonCommerical-ShareAlike 3.0 Unported License.

Let’s start the conversation

Tell us a bit about what you’re working on, and we’ll be in touch to explore how we can help you.