
Local Text to Speech AI That Runs Entirely on Your Computer
Local text-to-speech has crossed the quality line: the small models now sound better than the robotic voices most people remember, they need no subscription, and nothing leaves the machine. Piper, Kokoro-82M and XTTS v2 all run fully offline, and each covers a different job. Here's what local synthesis gets you, which engine fits which task, and the honest edges where the cloud still wins.
What Local Text to Speech Actually Gives You
Running speech synthesis locally means no audio or text leaves the machine, no per-character API billing, and no network round trip before the first word plays. The audio starts the moment you ask for it, the setup keeps working on a plane or when the connection drops, and private drafts stay private because the text never touches a vendor. For anything you narrate regularly, those four properties compound.
One confusion worth clearing up first, because half the people searching for this want the opposite direction: text-to-speech reads words aloud, while transcription turns audio into written notes. If your voice memos need to become text, you want local AI transcription, which is a different toolset with the same privacy logic, and on a Mac that job belongs to Shmeetings. This article is about the reading-aloud direction.
The hardware story has changed too. Since an 82-million-parameter model now ranks above far larger ones in blind comparisons, you don't need a big GPU for good results. Small models run fast on ordinary laptops, long narrations generate without strain, and no monthly fee arrives at the end of it.
Three Engines That Run Offline Today

Piper is the speed-and-simplicity pick, built by the Rhasspy team for exactly the small hardware other engines ignore: it speaks in real time on a Raspberry Pi 5 with no GPU, ships more than 30 languages and over 100 downloadable voices, and installs with a single pip command. It's MIT licensed and it's the default voice in Home Assistant, so you've probably heard it already. The honest cost is voice quality, which sits a step below the newer generation.
Kokoro-82M takes the quality-per-watt crown. Its 82 million parameters produce speech that ranked first on the TTS Spaces arena at launch, it runs roughly five times faster than real time on Apple Silicon CPUs, and the Apache license makes it safe for almost any project. The limits are coverage: 54 voices across 8 languages, so rarer languages aren't represented.
XTTS v2 from Coqui is the voice-cloning option: it clones a speaker from a sample of about 6 seconds of audio and supports 17 languages, which makes it the pick for narrating projects in a specific person's voice. The catch is the license, because the weights are non-commercial under the Coqui Public Model License, so a paid product needs Piper or Kokoro instead. For personal use, nothing else in this list sounds like you.
Getting Your First Engine Running
Piper installs with pip install piper-tts, and one command synthesizes a sentence to a file that appears almost instantly. The model and voice files download on first use and land in a cache directory, so the second run needs no network at all. It runs the same way on Windows, Linux and the Pi, which makes it the easiest first engine to test.
For higher quality, pull Kokoro-82M through its CLI tool or run the browser build, which needs no GPU on modern Apple Silicon. Test one sentence in two or three voices before committing to a long project, because voice fit is a taste decision the benchmarks can't make for you. Once an engine is set up, the workflow is the same for all of them: a text file in, a wave file out, voices swappable in seconds without reinstalling anything.
Voice cloning through XTTS runs via the Python API: you hand it a six-second sample and the text, and it generates speech in that voice. Keep the non-commercial license in mind at the start rather than the end of a project, because the terms travel with the weights, not with your code.
When Local Beats the Cloud, and When It Doesn't
Local wins on privacy, recurring cost and independence: nothing leaves the machine, there's no per-character bill, and the setup works when the network doesn't. For reading aloud, home assistants and accessibility narration, that combination usually ends the debate. The Piper project page shows how little hardware the offline path needs.
The cloud still wins three things: the very top of voice quality, the easiest integration with existing apps, and depth in rare languages and dialects. If your product needs a specific premium voice or a language the open models don't cover, the vendor API is the practical answer, and the privacy trade may be worth making knowingly rather than avoiding by default.
The decision rule: personal use and assistants go local, shipped products check the license before the first commit, and anything needing a rare language checks the cloud's coverage first. Most real projects end up hybrid, which is fine, because the tools don't conflict.
Where This Is Heading

The direction is visible in the benchmarks: an 82-million-parameter model outranking systems hundreds of times its size means quality per parameter is climbing fast, and the hardware needed for good audio keeps shrinking. It's the laptop-replaces-the-rig pattern from every other local AI category, arriving in speech.
Licenses are becoming the real sorting problem, because the weights carry usage terms separate from the code, and three guides this year made the same point about XTTS: excellent model, read the terms before you ship. Expect the Apache and MIT engines to keep gaining users for that reason alone. If you want the transcription side of the same story, free transcription from audio to text covers the other direction, and the offline meeting note taker workflow shows the same local-first pattern applied to meetings.
Frequently Asked Questions
What is the best free local text to speech?
Kokoro-82M is the current quality pick, with 54 voices across 8 languages, an Apache license and real-time synthesis on a plain CPU. Piper is the pick for tiny hardware and the widest language list. Both are free, offline, and installable in minutes.
Does local TTS work without internet?
Yes, once the model and voice files are downloaded. Piper, Kokoro and XTTS all synthesize with no network connection after installation, which is the point: no API round trips, no outages in your pipeline, and no data leaving the machine. Download the voices while you have a connection and you're done.
Can local TTS clone a voice?
XTTS v2 clones a voice from about six seconds of sample audio and speaks 17 languages, all offline. The weights are non-commercial, so the clone is for personal use unless you license it otherwise. For a paid product, Kokoro and Piper are the licensed-for-commerce options, without cloning.
Is local TTS good enough for audiobooks?
For personal listening, yes: Kokoro's arena ranking means long-form narration holds up far better than the robotic voices people remember. The tradeoffs are voice variety and emotional range compared to professionally produced audiobooks, and the generation time, which is minutes of compute for an hour of audio.
What hardware do you need for local TTS?
Less than you'd expect. Piper runs real time on a Raspberry Pi 5, and Kokoro runs about five times faster than real time on Apple Silicon CPUs with no GPU. A plain modern laptop covers every engine in this article, and the heavy-rig era of local audio is over.