Q-SPT: Learnable Query-Based Compression for Low-Frame-Rate Speech Tokenization

Q-SPT is a speech tokenizer for speech language models. It converts 16 kHz speech into tokens at 6.25 Hz (1 semantic stream with 32768 codes + 7 acoustic streams with 4096 codes each) and reconstructs speech from them.

Usage

Set up the environment by following the GitHub repository, then:

import torchaudio
from qspt import QSPT

codec = QSPT.from_pretrained("KU-AGI/Q-SPT").cuda()

wav, sr = torchaudio.load("speech.wav")
codes = codec.encode(wav, sr)          # (1, 8, N) tokens at 6.25 Hz
recon = codec.decode(codes)            # (1, 1, ~N * 2560) waveform at 16 kHz
torchaudio.save("recon.wav", recon[0].cpu(), 16000)

The semantic encoder SenseVoiceSmall is downloaded automatically on first use.

Files

  • model.safetensors: Q-SPT weights (fp32, 104M parameters)
  • config.json: model configuration

Citation

@article{yun2026qspt,
  title={Q-SPT: Learnable Query-Based Compression for Low-Frame-Rate Speech Tokenization},
  author={Yun, Jeeyoung and Yun, Seohwan and Kim, Sungwoong},
  journal={arXiv preprint arXiv:2610.01492},
  year={2026}
}

License

CC BY-NC 4.0

Downloads last month
22
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for KU-AGI/Q-SPT