Why a web page can't just give you an MP3
Browser text-to-speech (the Web Speech API, which our converter uses) asks your operating system to play speech. It never hands the audio data back to the page, so there is nothing for JavaScript to save. Some sites advertise "free MP3 download" while really sending your text to a cloud service; that can be fine, but it is a different product with different privacy terms.
The good news: every desktop operating system can write speech to an audio file with tools that are already installed or free. Pick your system below. Each method works offline and keeps your text on your computer.
Before you publish: the voices bundled with Windows, macOS and Linux are licensed for use on your own device. Using their output in a commercial video, podcast or product may need permission from the voice's vendor. For commercial work, check the licence or use a cloud TTS plan that explicitly allows it.
macOS: the built-in say command
macOS ships with a command-line tool that speaks text with any installed voice and can write the result to a file.
- Save your text as a plain-text file, for example
input.txton your Desktop. - Open Terminal (Applications → Utilities) and run:
cd ~/Desktop
say -v '?' # list the installed voices
say -v Samantha -f input.txt -o speech.aiff
afconvert -f m4af -d aac speech.aiff speech.m4a # optional: smaller AAC file
-vpicks the voice by name (use a name from the list; Enhanced/Premium voices you downloaded appear there too).-r 180sets the speaking rate in words per minute.afconvertis also built in; the.m4aplays everywhere. For MP3, convert with ffmpeg (see below).
Windows: PowerShell and the built-in voices
Windows PowerShell (the blue powershell.exe that ships with Windows 10 and 11) can use the system speech engine to write a WAV file. Use Windows PowerShell rather than PowerShell 7, which doesn't include the speech library.
Add-Type -AssemblyName System.Speech
$s = New-Object System.Speech.Synthesis.SpeechSynthesizer
$s.GetInstalledVoices() | ForEach-Object { $_.VoiceInfo.Name } # list voices
$s.SelectVoice("Microsoft Zira Desktop") # pick one from the list
$s.Rate = 0 # -10 (slow) to 10 (fast)
$s.SetOutputToWaveFile("$env:USERPROFILE\Desktop\speech.wav")
$s.Speak((Get-Content -Raw "$env:USERPROFILE\Desktop\input.txt"))
$s.Dispose()
This library sees the classic desktop voices (on English systems typically David and Zira, plus any third-party SAPI voices). The newer voices you add in Settings and Narrator's natural voices usually don't appear here; to capture those, record the playback instead (next section). Save input.txt as UTF-8 if it contains accented characters.
Any computer: record the playback with Audacity
If you want exactly the voice you hear in the browser — for example Edge's "Online (Natural)" voices — record your computer's output while the converter speaks.
Windows
- Install Audacity (free, open source).
- In Audacity's toolbar, set the audio host to Windows WASAPI and the recording device to your speakers or headphones marked (loopback).
- Press Record in Audacity, then press Play in the converter. Stop recording when the speech ends.
- Trim the silence and use File → Export to save as MP3 or WAV.
macOS
macOS does not let apps record system output by default. Install a free virtual audio device such as BlackHole, set it (or a multi-output device that includes it) as the output, and choose it as the recording input in Audacity. For most people, say above is simpler and gives a cleaner file.
Linux: eSpeak NG or Piper
eSpeak NG is in every major distribution's repositories, supports over 100 languages, and writes WAV directly. It sounds robotic but is fast and tiny:
sudo apt install espeak-ng # Debian/Ubuntu; use your distro's package manager
espeak-ng -v en-us -s 160 -f input.txt -w speech.wav
Piper is a free, open-source neural TTS engine that runs offline on ordinary CPUs and sounds far more natural. Download a voice model (an .onnx file plus its .json) from the Piper voices collection, then pipe text into it. With the classic command-line release:
cat input.txt | piper --model en_US-lessac-medium.onnx --output_file speech.wav
Piper's command-line options have changed between releases (newer versions are installed with pip install piper-tts), so check the project's README for the exact syntax of the version you install. Piper also runs on Windows and macOS.
Convert WAV or AIFF to MP3 with ffmpeg
ffmpeg converts between almost any audio formats:
ffmpeg -i speech.wav -codec:a libmp3lame -qscale:a 2 speech.mp3
-qscale:a 2 gives high-quality variable bitrate; use 5 for smaller spoken-word files. Audacity can do the same conversion through File → Export.
Phones and tablets
iOS and Android don't offer a built-in "save speech to file" option. The practical route on a phone is the system screen recorder with audio enabled — test a short clip first, because whether speech output is captured depends on the device and app. For anything longer than a few minutes, generating the file on a computer with one of the methods above is easier.
When a cloud service is the better choice
If you need the same voice on every device, word-level timings for captions, SSML control, or clear commercial-use rights, a paid cloud TTS service returns MP3 files directly. The trade-offs — privacy, cost, licensing — are covered in browser TTS vs cloud TTS.
Quick comparison
| Method | Platform | Voice quality | Output | Effort |
|---|---|---|---|---|
say | macOS | Good; excellent with Premium voices | AIFF → M4A/MP3 | One command |
| PowerShell System.Speech | Windows | Basic (classic voices) | WAV | Short script |
| Audacity loopback | Windows (macOS with BlackHole) | Whatever the browser plays, incl. Edge natural voices | MP3, WAV, etc. | Real-time recording |
| eSpeak NG | Linux, Windows, macOS | Robotic, many languages | WAV | One command |
| Piper | Linux, Windows, macOS | Natural (neural) | WAV | Download a model |