Access Ito, a non-commercial voice model
Ito's weights are free for hobby, research, education and personal use. Tell us a little about your project: it helps us decide what to build next, and we answer every request for a commercial or collaboration license.
The Ito voice model is licensed under CC BY-NC-SA 4.0 together with Lokutor's terms of use (TERMS.md in this repository). Without a written license from Lokutor you may not use the weights, weights derived from them, or the audio they produce commercially, and you may not use Ito's output to train or improve a text-to-speech or voice model that is offered or used commercially. Ito's output is synthetic speech in the voice of a real (LibriTTS-R) speaker: say that it is synthetic when you share it, and never use it to impersonate or deceive. Small companies, startups, makers, schools and research groups can ask for a no-cost license at contact@lokutor.com.
Log in or Sign Up to review the conditions and access this model content.
Ito: natural-sounding streaming TTS for the ESP32-S3
Ito is English text-to-speech built to run entirely on an ESP32-S3 microcontroller (240 MHz dual-core, 8 MB PSRAM, 16 MB flash), with no cloud and no neural accelerator. It streams: audio starts after a short first chunk (125 ms), and the work before it does not grow with the sentence length. Its 3.3 M parameters fit in 3.8 MB of int8 weights.
Watch the 50 s video: lokutor-ai.github.io/ito/demo.mp4 (with sound) Β· Listen next to sanoTTS: lokutor-ai.github.io/ito Β· Code, firmware and tools: github.com/lokutor-ai/ito (GPLv3, commercial licenses available). From Lokutor, the makers of OΓdo, speech recognition on the same chip.
Status (5 October 2026). The on-chip engine is verified on a laptop and in Espressif's QEMU emulator: the firmware's audio is bit-identical to the host build of the engine, and every sample below is that engine's exact output. Nothing has run on a physical board yet. Time to first audio and real-time factor are estimated from exact QEMU instruction counts and an assumed PSRAM bandwidth. Real-time playback is not established on silicon: the optimistic and central estimates are faster than real time (RTF 0.43β0.47 and 0.63β0.66 with the main weight set), the pessimistic one is only just below it (0.97β0.99). The firmware now ships three weight sets per voice and chooses between them at boot by measuring its own chunk times (see Weight sets below). On 5 October the vocoder went from 256 to 192 channels (22 % fewer weights); in a blind test (#10) one listener heard no difference. An earlier version of this card (and our first GitHub README) said 130β210 ms and real time at 0.5β1 GOPS; that was too optimistic and is corrected below. Board measurements are coming.
Voices
Two voices, each distilled from StyleTTS 2 conditioned on recordings of one LibriTTS-R reader (LibriTTS-R, CC BY 4.0). Each voice is a separate model with the same architecture, size and speed.
| Voice | Speaker | PyTorch | Chip (flash at 0x200000) |
|---|---|---|---|
| female (default) | LibriTTS-R speaker 4970 | ito_female.pt |
ito_female_esp32s3.bin |
| male | LibriTTS-R speaker 5105 | ito_male.pt |
ito_male_esp32s3.bin |
The male voice is new (2 October 2026). Its chip file passes the same host and QEMU checks as the female voice (quantised vs float PESQ 4.55; firmware PCM bit-identical to the host engine). It has not been in a blind listening test yet; the results below are for the female voice.
Listen
The female voice, the chip engine's exact output (192-wide vocoder), for eight sentences Ito never saw in training. The male voice reads the same sentences on the demo page.
| Sentence | Ito (chip-exact) | |
|---|---|---|
| 01 | Hey, are you still coming over for dinner tonight, or should I save you a plate? | |
| 02 | Your package should arrive on Friday, October 9th, sometime before noon. | |
| 03 | When I finally got to the station, the last train had already left, so I ended up sharing a taxi with two strangers who turned out to be surprisingly good company. | |
| 04 | The pharmacist recommended an anti-inflammatory, but honestly, I'd rather try physiotherapy first. | |
| 05 | Thanks so much for calling. I'll check the schedule and get back to you first thing tomorrow morning. | |
| 06 | Could you grab some quinoa and Worcestershire sauce on your way home? | |
| 07 | It's about 23 degrees outside, so you probably won't need a jacket. | |
| 08 | I know it sounds strange, but I actually enjoy the quiet hours before everyone else wakes up. |
The players stream from the public demo page, so they work before you accept the gate. The bit-exact WAVs are in
samples/ in this repository. The demo page plays the same sentences from
sanoTTS and from Ito's teacher, with a blind mode.
Results
Blind listening test #9. One expert listener rated naturalness from 1 to 5, with system names hidden. Each system read four new conversational sentences.
| System | Parameters | Runs on | Mean (4 clips) |
|---|---|---|---|
| Teacher: StyleTTS 2 (LibriTTS model) | large | GPU / laptop | 4.75 |
| Ito (as rated: 256-wide vocoder, 4.4 M; the shipped 192-wide one is 3.3 M, see #10) | 4.4 M | ESP32-S3 (bit-exact in emulation) | 4.00 |
| sanoTTS amy | 1.46 M | browser / desktop (not run on an MCU) | 2.00 |
| sanoTTS heart-nano | 0.29 M | ESP32-S3 | 1.00 |
Blind listening test #10 (5 October 2026), the lighter vocoder: the same listener, the same four sentences, the female voice, 16 clips with hidden system names. Teacher 4.25; Ito with the 256-wide vocoder (the chip engine) 4.00; Ito with the 192-wide vocoder (the chip engine, shipped now) 4.00; the 192-wide vocoder with int4 weights (emulated) 4.00. He heard no difference between the three Ito systems (one listener, four clips per system).
Automatic metrics on the eight sentences above (measured with the 256-wide vocoder of the first release; on 60 held-out utterances the 192-wide one is within 0.02 UTMOS and +0.014 / +0.002 log-mel distance for female / male):
| Teacher | Ito | sanoTTS amy | sanoTTS heart-nano | |
|---|---|---|---|---|
| UTMOS | 4.49 | 4.46 | 3.98 | 2.07 |
| WER, Whisper medium.en / base | 0 / 0 % | 0 / 0 % | 0 / 1.0 % | 1.0 / 1.0 % |
Please read these with their limits:
- One listener (Lokutor's founder, so not a neutral party) and four sentences per system. This is a strong direction, not a MOS study. With n = 4, differences under about half a point are noise.
- The rated Ito clips are the float model through the same streaming path as the chip. The quantised chip engine scores PESQ 4.5 against it; test #10 below rated the chip engine itself.
- UTMOS cannot hear intonation. Treat it as a check, not a verdict.
What we think this supports, and no more: the highest automatically measured naturalness of any complete text-to-waveform neural TTS built for a microcontroller without an NPU (UTMOS 4.46 on 8 sentences; the best previously published on-chip model scores 2.80), and the first streaming neural TTS designed for the ESP32-S3. Both are verified in emulation, not yet on a board. Ito is not the first or the smallest TTS on a microcontroller.
A broader benchmark against other small and embedded TTS systems is being finalised; it will be in
bench/ on GitHub.
Size and compute
| Parameters | 3.34 M: acoustic front 1.62 M + vocoder 1.72 M (was 4.40 M with a 256-wide vocoder) |
| Chip weights | 3.81 MB main set (ito_female_esp32s3.bin): int8 mel head and vocoder, int16 pitch path; plus the fallback sets _int4 (3.20 MB) and _light (3.05 MB) |
| Memory | peak 5.3 of 8 MB PSRAM, 314 of 384 KB internal SRAM (QEMU; about 327 KB on the chip with the I2S buffers) |
| Work before the first audio (125 ms chunk) | 23.0β24.3 M instructions and 3.9 MB of weights read from PSRAM, for any sentence length (exact counts from QEMU) |
| Time to first audio | estimated, not measured: 124β132 ms optimistic, 171β180 ms central, 251β260 ms pessimistic. The first chunk is 125 ms of audio and the chunks behind it are sized so that playback can start with it without a gap |
| Real-time factor | estimated, not measured: 0.43β0.47 optimistic, 0.63β0.66 central, 0.97β0.99 pessimistic (below 1 is faster than real time). The pessimistic case (40 MB/s PSRAM, nothing overlapped) is only just below real time, which is not a margin (it was 1.18β1.27 with the 256-wide vocoder) |
| Start delay for gapless speech | estimated: 125β130 ms optimistic (the same as the time to first audio), 176β182 ms central, 558β663 ms pessimistic (the firmware plans it from its own measurements) |
| On a laptop | RTF β 0.01 on an Apple M4 Max CPU |
All timings are estimated from exact instruction counts, not measured on silicon. The open question is the effective PSRAM bandwidth and how much of it overlaps compute; the pessimistic column depends on both. A weight pass serves 24-frame (300 ms) chunks, and the 192-wide vocoder reads 11 MB of weights per second of audio (15 MB with the 256-wide one). The firmware benchmarks itself at boot
and prints BOARD_SUMMARY. If you flash a board, please
open an issue with that line.
Weight sets and the boot check
Each voice has three chip files, flashed to three partitions: the main set (ito_<voice>_esp32s3.bin, int8), the same model with int4 weights in the ConvNeXt blocks (_int4, 3.20 MB) and a light set
with one block fewer, also int4 (_light, 3.05 MB). At boot the board measures the real time of every chunk of a representative sentence with each set in turn, keeps the first whose measured real-time factor is at or below 0.85,
plans the playback start delay from the measured chunk times, and, if even the fastest set measures 0.95 or more, prints a warning and sets a degraded flag instead of stuttering silently.
Estimated the same way as above (optimistic / central / pessimistic, whole sentences; not measured on silicon):
| main int8 | main int4 | light | |
|---|---|---|---|
| real-time factor | 0.43β0.47 / 0.63β0.66 / 0.97β0.99 | 0.45β0.48 / 0.62β0.64 / 0.92β0.93 | 0.42β0.45 / 0.59β0.61 / 0.88 |
| time to first audio | 124β132 / 171β180 / 251β260 ms | 127β133 / 168β174 / 236β243 ms | 118β124 / 157β164 / 222β229 ms |
| gapless start delay | 125β130 / 176β182 / 558β663 ms | 127β133 / 173β179 / 481β526 ms | 118β124 / 162β169 / 406β429 ms |
int4 weights save a fifth of the PSRAM weight traffic for about 1 % more instructions; they help only where the PSRAM is the limit. The int4 sets are bit-exact against an int8-equivalent blob, and their audio scores PESQ 4.52 (female) / 4.54 (male) against the int8 main set's; the mel distance to the float model grows by about 75 %. One listener heard no difference between the int8 and the (emulated) int4 main set in blind test #10; the light set and the male voice's int4 sets have not been heard. In the pessimistic column no set reaches 0.85, so the boot would keep the light set as "marginal" (0.88). The check is a guarantee by measurement on the board that runs it; it has been tested with simulated timings in QEMU and on the host, not on silicon.
How it works
Text becomes phonemes on the host (espeak-ng), and token ids go to the chip over serial. On the chip, a streaming acoustic front (1.62 M; a forward-only GRU predicts durations, pitch, energy and a mel spectrogram) feeds a Vocos-style vocoder (1.72 M; ConvNeXt blocks, a harmonic pitch source and an inverse STFT) that outputs 24 kHz audio in 100 ms chunks. Ito learned its voice from StyleTTS 2 speaking as a LibriTTS-R reader. The engine is new portable C99 code with the chip's integer arithmetic; the same source builds on a laptop. The training code and recipe are not public.
Use
git clone https://github.com/lokutor-ai/ito && cd ito && pip install -e .
hf download lokutor-ai/ito --include "ito_*" --local-dir models # both voices (accept the terms first)
ito-tts "Good morning! The coffee is ready." -o hello.wav # PyTorch, CPU is fine; female voice
ito-tts --voice male "Good morning! The coffee is ready." -o hello_male.wav # male voice
cd esp32/host && make && cd ../.. # the chip engine, built for your laptop
python esp32/tools/chip_wav.py "Good morning! The coffee is ready." hello_chip.wav # chip-exact audio
esp32/tools/flash.sh /dev/ttyUSB0 # ESP32-S3-DevKitC-1 N16R8 + I2S DAC/amp (female voice)
VOICE=male esp32/tools/flash.sh /dev/ttyUSB0 # ... or the male voice
python esp32/tools/say.py "Hello from a five dollar chip." --port /dev/ttyUSB0
from ito import Ito
tts = Ito.load("models/ito_female.pt") # female voice; Ito.load(voice="male") or "models/ito_male.pt" for the male voice
wav = tts.synthesize("Could you grab some quinoa on your way home?") # float32 numpy, 24 kHz
for chunk in tts.stream("A sentence of any length."): # 100 ms chunks, as on the chip
...
Files
ito_female_esp32s3.bin: the female voice's chip weights, main set (3.81 MB), for the firmware and the host build of the engine.ito_female_esp32s3_int4.bin(3.20 MB) andito_female_esp32s3_light.bin(3.05 MB): its fallback sets for the boot check (flash offsets 0x680000 and 0xB00000).ito_female.pt: the female voice's PyTorch inference checkpoint (14 MB): front, vocoder, the fixed style the chip uses, and an optional text-to-style predictor.ito_male_esp32s3.bin,ito_male_esp32s3_int4.bin,ito_male_esp32s3_light.bin,ito_male.pt: the same for the male voice. The male checkpoint has no text-to-style predictor; it uses the fixed style, as the chip does.samples/01.wavβ¦samples/08.wav: the chip engine's exact output for the sentences above (24 kHz, 16-bit).LICENSE-WEIGHTS,TERMS.md,NOTICE: the license, the terms of use, and third-party attributions.
Limitations
- English only, two voices (one per weights file), one fixed speaking style on the chip.
- Speed on the chip is estimated until board measurements are published.
- The pitch range is slightly narrower than the teacher's (female 0.94, male 0.97).
- Needs an ESP32-S3 with 8 MB PSRAM (N16R8 recommended). Text-to-phoneme conversion runs on the host.
License and attribution
The weights and the Ito audio are licensed CC BY-NC-SA 4.0 together with Lokutor's terms of use. Commercial use needs a written license from Lokutor. That covers products and devices, paid services and APIs, internal business use and commercial content, and also weights derived from Ito. The terms also exclude training commercial TTS or voice models on Ito's output. The engine, firmware and Python package are GPLv3, with commercial licenses available.
Ito learned its voice from StyleTTS 2 (MIT) conditioned on LibriTTS-R speakers
4970 (female) and 5105 (male) (LibriTTS-R, Koizumi et al. 2023, CC BY 4.0;
derived from LibriTTS and LibriVox public-domain recordings). Ito is not affiliated with or endorsed by those
speakers, LibriVox or Google. No third-party weights or recordings are included. Its
output is synthetic speech: please say so when you share it. Full attributions are in NOTICE.
Attribution: "Ito voice model by Lokutor (lokutor.com), CC BY-NC-SA 4.0".
We'd love to hear what you build. Lokutor offers no-cost licenses to small companies, startups, makers selling small batches, schools and research groups, and has more voices, more languages and speech recognition for the same chip. Write to contact@lokutor.com.
