Skip to content

Voice interfaces for websites

Voice interfaces let people speak to websites, portals and apps and hear content read aloud.

Background

Voice interfaces matter in three places: accessibility, forms and search fields, and audio or video that should be searchable. Since June 2025, the European Accessibility Act has applied in Germany to many consumer services. Voice adds to a WCAG 2.2 AA base with keyboard and screen reader support, but cannot replace it.

Typical use cases

  • Transcription and captions: for videos, webinars and podcasts. WCAG requires captions for recorded video in any case.

  • Dictation in forms: speaking damage reports or service requests helps on mobile and for people with motor impairments.

  • Voice search: search understands spoken questions.

  • Speech output: text-to-speech reads content aloud in portals or apps.

  • Call analysis: service or advice calls transcribed, summarised and made searchable, with everyone’s consent.

Technology and providers

Speech recognition runs on Azure AI Speech, Google Cloud Speech-to-Text or open-source models such as Whisper on your own servers. Neural voices read technical terms correctly via a pronunciation lexicon.

The browser’s Web Speech API is quick to add. Some browsers send recordings to the vendor’s servers, which rules it out for sensitive input.

Data protection and operations

Voice recordings are personal data. We clarify in advance where they are processed and how long they are kept. Where necessary, models run in EU data centres or on your own servers. Unless storage has been agreed, recordings are deleted after processing.

Approach

  1. Define the use case: who speaks, where, in which language and with what vocabulary.

  2. Compare models: two or three models tested with real recordings on error rate with your vocabulary, latency and cost per minute.

  3. Integration: into your website, portal or app, with a visible status and text input as an alternative.

  4. Testing and operations: checks with screen readers and real users, then monitoring. New model versions pass the same test before any switch.

Project enquiry

Back to top