The 10-line version
if ('speechSynthesis' in window) {
button.addEventListener('click', () => {
speechSynthesis.cancel(); // clear anything queued
const u = new SpeechSynthesisUtterance('Hello from the Web Speech API.');
u.lang = 'en-US';
u.rate = 1; // 0.1 – 10, 1 is normal
u.pitch = 1; // 0 – 2
speechSynthesis.speak(u);
});
}
That works in every current browser. Everything below is about the parts that break in production: loading voices, autoplay rules, long text, pause/resume and errors. This is the same approach our converter uses.
1. Load voices properly (voiceschanged)
getVoices() often returns an empty array on the first call — always in Chromium — and fills in later, firing voiceschanged. Some engines never fire the event, and on some systems there are simply no voices. Handle all three cases:
function loadVoices(timeoutMs = 3000) {
return new Promise((resolve) => {
let voices = speechSynthesis.getVoices();
if (voices.length) return resolve(voices);
const done = () => {
voices = speechSynthesis.getVoices();
if (voices.length) {
speechSynthesis.removeEventListener('voiceschanged', done);
resolve(voices);
}
};
speechSynthesis.addEventListener('voiceschanged', done);
setTimeout(() => resolve(speechSynthesis.getVoices()), timeoutMs); // may be []
});
}
const voices = await loadVoices();
const voice = voices.find(v => v.lang === 'en-GB' && v.localService)
|| voices.find(v => v.lang.startsWith('en'));
- Match on
voice.lang(BCP 47, e.g.pt-BR), not on names — names differ per OS and browser. Some Android builds reporten_USwith an underscore, so normalise. localService: truemeans on-device;falsemeans the browser sends text to a network service (Chrome's "Google …" voices, Edge's "… Online (Natural)" voices).- Store the user's choice by
voiceURI, and re-resolve it aftervoiceschanged; the array can be rebuilt. - Newly installed OS voices usually appear only after a full browser restart.
2. Respect the user-gesture (autoplay) rule
Chrome rejects speak() without prior user activation (since Chrome 71) and iOS Safari requires the first speak() to run synchronously inside a tap/click handler. Calling it after await fetch(...) on iOS fails with a not-allowed error because the gesture has expired. Unlock speech first, then speak the real text when it arrives:
button.addEventListener('click', async () => {
speechSynthesis.speak(new SpeechSynthesisUtterance('')); // unlock inside the gesture
const text = await fetch('/article.txt').then(r => r.text());
speakLong(text);
});
You can check navigator.userActivation?.hasBeenActive to decide whether to show a "Tap to listen" button instead of auto-speaking.
3. Long text and Chrome's ~15-second cut-off
Desktop Chrome's network voices stop after roughly 15 seconds of continuous speech, often without an end event, leaving your UI stuck in "speaking". Other engines have their own soft limits, and the spec caps an utterance at 32,767 characters. The robust fix is to split text into sentence-sized utterances and chain them:
function chunk(text, max = 180) {
const sentences = text.replace(/\s+/g, ' ')
.match(/[^.!?。!?]+[.!?。!?]*\s*/g) || [text];
const out = []; let buf = '';
for (const s of sentences) {
if ((buf + s).length > max && buf) { out.push(buf.trim()); buf = ''; }
if (s.length > max) { // very long sentence: break at spaces
for (const part of s.match(new RegExp('.{1,' + max + '}(\\s|$)', 'g')) || [s]) out.push(part.trim());
} else buf += s;
}
if (buf.trim()) out.push(buf.trim());
return out;
}
let run = 0;
const keep = []; // hold references: Chrome may garbage-collect utterances and skip 'end'
function speakLong(text, voice) {
const id = ++run;
speechSynthesis.cancel();
const parts = chunk(text);
let i = 0;
const next = () => {
if (id !== run || i >= parts.length) return;
const u = new SpeechSynthesisUtterance(parts[i++]);
if (voice) { u.voice = voice; u.lang = voice.lang; }
u.onend = next;
u.onerror = (e) => { if (e.error !== 'interrupted' && e.error !== 'canceled') console.warn(e.error); };
keep.push(u); if (keep.length > 5) keep.shift();
speechSynthesis.speak(u);
};
setTimeout(next, 50); // small gap after cancel() avoids a dropped first utterance in Chromium
}
Speaking one chunk at a time (instead of queueing them all) also gives you accurate progress and lets the user change voice or speed mid-text. The older workaround — calling pause() and resume() every ~14 seconds — works on desktop Chrome but breaks elsewhere; chunking is simpler.
4. Pause, resume and stop
speechSynthesis.pause()/resume()work on desktop browsers for local voices. On Chrome for Android,pause()isn't reliably implemented — emulate it by cancelling and remembering the current chunk, then restarting that chunk on resume.cancel()clears the whole queue. Pending utterances may fireerrorwithinterruptedorcanceledinstead ofend; ignore those, and use a run counter (as above) so stale callbacks don't advance your queue.- Check
speechSynthesis.speakingand.pausedrather than tracking state only in your own variables; another tab can pause or cancel the shared engine.
5. Handle errors by type
event.error | Typical cause | What to tell the user |
|---|---|---|
not-allowed | No user gesture (autoplay policy) | "Tap Play to start speech." |
interrupted, canceled | You called cancel() or started new speech | Nothing — not a real error |
network | Online voice without connectivity | Suggest an on-device voice |
language-unavailable, voice-unavailable | No voice for that language / voice removed | Link to how to install voices |
audio-busy, audio-hardware | Output device busy or missing | Check speakers; close other apps |
synthesis-unavailable, synthesis-failed | Speech engine missing or crashed (common on Linux without speech-dispatcher) | Try another browser or install a voice |
text-too-long | Utterance over the engine's limit | Chunk the text |
invalid-argument | Out-of-range rate, pitch or volume | Clamp your values |
6. Highlighting words with boundary events
u.onboundary = e => highlight(e.charIndex, e.charLength) works with most on-device voices, but network voices often fire no boundary events, and charLength is missing in some browsers. Treat highlighting as progressive enhancement and fall back to highlighting the current chunk.
7. What you cannot do
- Record or download the audio. The API only plays speech. There's no stream to hand to
MediaRecorderor the Web Audio API, and tab capture usually misses it because the OS plays the sound. For files, see save text-to-speech as audio, or use a cloud TTS API. - Rely on SSML. Browsers don't consistently interpret it; some read the tags aloud. Use punctuation — see writing for TTS.
- Guarantee a voice. Every visitor has a different list. Always fall back to "any voice for this language", then to the default.
8. Testing checklist
- Chrome on Windows with a Google network voice and a 2-minute text (the 15-second bug).
- Safari on iPhone, starting speech after an async fetch (the gesture rule).
- Chrome on Android: pause and resume.
- Firefox on Linux with and without speech-dispatcher voices (empty list).
- Edge with an "Online (Natural)" voice while offline (
networkerror). - Switching tabs and locking the phone mid-speech.