Quick answer
Speech synthesis (text → speech, window.speechSynthesis) works in every current major browser: Chrome, Edge, Safari, Firefox, Opera, Samsung Internet, and their mobile versions. Speech recognition (speech → text, SpeechRecognition) is much patchier: Chromium browsers and Safari support it (Safari and older Chromium only with the webkit prefix), Firefox does not ship it. The details below are what actually matters when you build on either half — or when you are just trying to get our converter to talk.
Check your own browser
This live check runs in your browser and reports what it exposes right now. Nothing is sent anywhere.
speechSynthesis | checking… |
|---|---|
| Voices reported | checking… |
SpeechRecognition | checking… |
| User agent | – |
If speechSynthesis is present but zero voices are reported, the problem is the operating system's speech engine, not the browser — see install system voices. If the test button produces an error such as not-allowed or synthesis-failed, read the troubleshooting section below.
speechSynthesis support by browser
| Browser | Supported since | Where the voices come from | Notes |
|---|---|---|---|
| Chrome (desktop) | Chrome 33 (2014) | OS voices plus Google network voices (“Google US English”, “Google Deutsch” …) | Network voices stop after roughly 15 seconds of continuous speech; split long text. Network voices need a connection. |
| Edge (desktop) | Edge 14 (legacy), Chromium Edge 79 | Windows voices plus Microsoft “Online (Natural)” voices | The “Online (Natural)” neural voices are an Edge feature; other Chromium browsers (Chrome, Opera, Brave) do not get them. |
| Safari (macOS) | Safari 7 | macOS system voices, including downloaded Enhanced/Premium voices | Some voices installed for Spoken Content may not be offered to web pages; the list is usually long on recent macOS versions. |
| Firefox (desktop) | Firefox 49 | Windows/macOS system voices; on Linux, whatever speech-dispatcher provides | No network voices. On Linux with no speech-dispatcher module installed, the list is empty. |
| Opera, Brave, Vivaldi | Follows Chromium | OS voices (Google network voices are generally Chrome-only) | Fewer voices than Chrome or Edge on the same machine is normal. |
| Safari and every other browser on iOS / iPadOS | iOS 7 | iOS system voices (Settings → Accessibility → Spoken Content) | All iOS browsers use WebKit, so Chrome, Edge and Firefox on iPhone have the same voices as Safari — Edge's natural voices do not appear in web pages there. |
| Chrome on Android | Yes | The Android TTS engine you selected (Google, Samsung, third-party) | pause() is not reliably implemented; tools fall back to stop-and-restart. |
| Samsung Internet, Firefox for Android | Yes | The Android TTS engine | Voice list is often shorter than in Chrome on the same device. |
Version numbers are when a usable unprefixed speechSynthesis first shipped, per MDN's compatibility data. In practice, any browser updated in the last several years supports it; the differences are about voices and quirks, not presence.
SpeechRecognition support (speech-to-text)
Recognition is the half most people mean when they search for “Web Speech API browser support”. In short:
- Chrome, Edge and other Chromium browsers expose
webkitSpeechRecognition(current Chrome also exposes the unprefixedSpeechRecognition). In Chrome, audio is sent to Google's servers by default; Edge uses Microsoft's service. Some Chromium forks (and Electron apps) expose the object but fail with anetworkornot-allowederror because they lack access to the vendor's recognition service. - Safari 14.1+ on macOS and iOS 14.5+ expose
webkitSpeechRecognition, backed by Apple's dictation. - Firefox does not ship speech recognition to users; there is an internal preference, but it is not a working feature in release builds.
The full breakdown — permissions, continuous mode limits, language support, on-device processing and alternatives — is on the SpeechRecognition browser support page.
Autoplay and the user-gesture rule
Browsers block sound that starts without user interaction, and speech synthesis is no exception. Chrome stopped allowing speechSynthesis.speak() without user activation in Chrome 71; the utterance fails with a not-allowed error. Safari on iOS is stricter still: the first speak() call in a page must happen synchronously inside a tap or click handler. If you call it after an await fetch(), the gesture has “expired” and you get NotAllowedError/not-allowed.
Practical pattern: inside the click handler, speak something immediately (even an empty utterance) to unlock speech for the page, then start the real text when your data arrives. That's also why our converter only speaks when you press Play.
Voices that load asynchronously
In Chromium, speechSynthesis.getVoices() returns an empty array on first call; the list arrives later and the browser fires a voiceschanged event. Safari usually returns voices immediately; Firefox can take a moment. Robust code calls getVoices() once, listens for voiceschanged, and stops waiting after a few seconds. Our speechSynthesis developer guide has a copy-paste version.
New voices installed while the browser is open usually only appear after the browser is fully restarted.
Length limits and the 15-second cut-off
The spec does not set a maximum utterance length, but engines do. The best-known case: in desktop Chrome, Google's network voices go silent after roughly 15 seconds of continuous speech, with no error and sometimes no end event. On-device voices are not affected the same way. The reliable fix is to split text into sentence-sized utterances and speak them one after another — which is exactly what our converter does for texts up to 20,000 characters.
Pause, resume and cancel
- Desktop Chrome, Edge, Safari, Firefox:
pause()andresume()work for local voices. With Chrome's network voices, a long pause can leave the engine stuck; cancelling and restarting from the current sentence is safer. - Android:
pause()is effectively unsupported in Chrome; treat pause as “cancel and remember where you were”. cancel()fireserrorwithinterruptedorcanceledon pending utterances in some browsers rather thanend— don't show those to the user as errors.
Things that are not portable
- SSML. The spec allows an utterance to contain SSML, but mainstream browsers do not reliably interpret it — some read the tags aloud, others strip them. Shape speech with punctuation instead; see writing for TTS.
- Word boundaries.
boundaryevents are fired by most engines for local voices, but often not for network voices, and thecharIndex/charLengthdetails vary. Test before building word highlighting on them. - Capturing the audio. The API has no way to hand you the synthesized audio as a stream or file; it only plays it. Tab-capture tricks don't work reliably either, because the OS speech engine often plays outside the tab's audio. If you need a file, see how to save text-to-speech as audio.
- Background tabs. Behaviour when the tab is hidden or the phone screen locks varies; mobile browsers frequently stop speech.
Troubleshooting checklist
- Did the speech start from a click or tap? If not, expect
not-allowed. - Is the tab or system muted? Check the tab's speaker icon and the OS volume mixer.
- Is the voice list empty? Install a voice in the OS, then fully restart the browser — see install system voices.
- Does it stop after ~15 seconds? You're on a Chrome network voice; split text or choose an on-device voice.
- Only “online” voices fail? A firewall, VPN or offline connection can block the vendor's speech service. On-device voices still work.
- Managed (work/school) computer? Open
chrome://policyoredge://policyto see policies applied by your organisation; audio-capture policies (such asAudioCaptureAllowed) block the microphone for speech recognition, and network filters can block online voices. The live check above shows whether the APIs are exposed at all. - Another app speaking? Screen readers and dictation can hold the speech engine; pause them and try again.
FAQ
Does Firefox support the Web Speech API?
Firefox supports speech synthesis (text-to-speech) on desktop and Android. It does not ship speech recognition, so SpeechRecognition and webkitSpeechRecognition are undefined in release Firefox.
Does the Web Speech API work on mobile?
Speech synthesis works in Safari (and all other browsers) on iOS and in Chrome, Samsung Internet and Firefox on Android. The voices come from the phone's own TTS settings. Recognition works in Chrome on Android and in Safari on iOS 14.5+.
Why does Edge have more voices than Chrome or Opera?
Edge adds Microsoft's cloud-based “Online (Natural)” voices to the list; Chrome adds its own “Google …” network voices. Other Chromium browsers usually see only the voices installed in the operating system. There is no setting to give Opera or Brave Edge's online voices.
Is it a W3C standard?
The Web Speech API is specified in a W3C Community Group report, not a full W3C Recommendation. That's part of why implementations differ so much.
When the browser is not the right tool
If you need consistent voice quality for every visitor, a downloadable audio file, or voices licensed for commercial use, a cloud TTS service is the better fit; the browser TTS vs cloud TTS comparison walks through the decision. For listening, proofreading, accessibility and language practice, browser TTS is more than enough.