Q-SPT: Learnable Query-Based Compression for Low-Frame-Rate Speech Tokenization
Paper • 2610.01492 • Published
Q-SPT is a speech tokenizer for speech language models. It converts 16 kHz speech into tokens at 6.25 Hz (1 semantic stream with 32768 codes + 7 acoustic streams with 4096 codes each) and reconstructs speech from them.
Set up the environment by following the GitHub repository, then:
import torchaudio
from qspt import QSPT
codec = QSPT.from_pretrained("KU-AGI/Q-SPT").cuda()
wav, sr = torchaudio.load("speech.wav")
codes = codec.encode(wav, sr) # (1, 8, N) tokens at 6.25 Hz
recon = codec.decode(codes) # (1, 1, ~N * 2560) waveform at 16 kHz
torchaudio.save("recon.wav", recon[0].cpu(), 16000)
The semantic encoder SenseVoiceSmall is downloaded automatically on first use.
model.safetensors: Q-SPT weights (fp32, 104M parameters)config.json: model configuration@article{yun2026qspt,
title={Q-SPT: Learnable Query-Based Compression for Low-Frame-Rate Speech Tokenization},
author={Yun, Jeeyoung and Yun, Seohwan and Kim, Sungwoong},
journal={arXiv preprint arXiv:2610.01492},
year={2026}
}
CC BY-NC 4.0