Nemotron 3 Diarization — LiteRT INT8 conversion

Community conversion by spybyscript of NVIDIA Nemotron 3 Diarization. No training or fine-tuning was performed. This is not an official NVIDIA, Hugging Face, or Google release.

This is a stateful 1.04 s low-latency streaming integration, not a stateless drop-in model. The preferred v2 bundle contains a feature-stacking graph and one fused encoder/classification-head graph. Application code owns the Arrival-Order Speaker Cache (AOSC), FIFO state, audio framing, and output stabilization. See INTEGRATION.md, the authoritative int8/nemotron3-diarization-manifest.json, and the runnable reference_runtime.py.

Source and license

  • Source revision: f667ed73aee57d40cc39428eb768b4fd87a0a29e.
  • Source model materials remain governed by the NVIDIA Open Model Development Work License 1.1. See LICENSE.md and NOTICE.md.
  • Exact source hashes are in int8/source-manifest.json. The upstream model card is preserved as MODEL_CARD.md; its accuracy claims are not conversion measurements.
  • The conversion pins Transformers revision f5af3202d63d9bb7578a41f0041c9040071e7345 because Nemotron 3 Diarization support was not yet in a stable Transformers release at conversion time.

INT8 format and size

The two LiteRT flatbuffers total 101,905,216 bytes (97.18 MiB). All learned matrix and convolution kernels are stored as per-output-channel INT8 and use dynamic-range quantized kernels where LiteRT supports them. Graph inputs/outputs, activations, caches, biases, normalization constants, and the learned silence embedding remain floating point.

This is therefore an INT8-weight LiteRT release; it is not an integer-input/integer-output or all-activation-INT8 model. A representative-calibrated full-integer attempt did not meet the numerical gate and is not published.

Measured conversion parity

Closed-loop validation used a 182.053 s multilingual, three-speaker recording with the source Transformers implementation as the numerical reference.

Measurement INT8 LiteRT
Speaker-frame activity agreement 99.9821%
Active speaker-frame IoU 99.8470%
False-negative speaker-frames 13
False-positive speaker-frames 13
100 ms stabilized segments 41 vs 41 upstream
Graph-only real-time factor 0.1354

The graph-only timing was measured on a local Apple CPU with four LiteRT threads and excludes feature extraction, cache policy, audio I/O, and application postprocessing. It is not an Android, GPU, or NPU benchmark.

On a Samsung Galaxy S23 Ultra, the silent streaming harness processed the same 182.053 s recording in 112.252 s (0.617 wall RTF), produced 38 stabilized segments across three speaker channels, emitted its first segment after 1.25 s of audio / 0.814 s wall time, and measured 199,252 KiB loaded and 352,057 KiB peak process PSS. The slowest 250 ms append was 735 ms. These are single-run device measurements, not a device-family guarantee; see ANDROID_VALIDATION.md for the exact protocol.

The frame decisions are promising, but this release has not yet been evaluated against human diarization references for DER/JER or across a representative corpus. evaluation_only is therefore true in the manifest. The reported parity is a conversion check—not an independent quality benchmark and not a claim of equivalence to the FP32 source.

Download and run

from huggingface_hub import snapshot_download

snapshot_download(
    "spybyscript/Nemotron-3-Diarization-litert",
    # Pin revision to an immutable release commit for deployment.
    allow_patterns=[
        "int8/*",
        "README.md",
        "INTEGRATION.md",
        "LICENSE.md",
        "LICENSE.OpenMDW-1.1",
        "NOTICE.md",
        "ANDROID_VALIDATION.md",
        "reference_runtime.py",
        "validate_sideload_zip.py",
        "requirements-reference.txt",
        "publication.json",
        "SHA256SUMS",
    ],
    local_dir="nemotron3-diarization-litert",
)
pip install -r nemotron3-diarization-litert/requirements-reference.txt
python nemotron3-diarization-litert/reference_runtime.py \
  nemotron3-diarization-litert/int8 recording.wav

The reference runtime does not load the original checkpoint or Transformers. It validates all runtime file hashes, resamples to mono 16 kHz, executes the cache-aware graph loop, and prints both raw and recommended 100 ms-stabilized segments.

For the Android debug importer, download nemotron3-diarization-litert-int8.zip. It uses the future production-shaped layout (model-manifest.json, configuration, graphs/, silence state, license, and notice) and is generated from the same checked artifacts as the browsable int8/ folder.

from huggingface_hub import hf_hub_download

hf_hub_download(
    "spybyscript/Nemotron-3-Diarization-litert",
    "nemotron3-diarization-litert-int8.zip",
    local_dir=".",
)

Validate the archive before importing it:

python validate_sideload_zip.py nemotron3-diarization-litert-int8.zip

Android integration status

Use adapter ID nemotron3-diarization-low-latency-v2. Keep both graphs and state buffers behind a single inference owner. The model emits eight arrival-ordered speaker-activity channels at 10 ms resolution; overlap is represented by simultaneous active channels. It does not perform speech recognition or attach words to speakers.

Android applications can run this model as an independent speaker-timeline component alongside any ASR engine that supplies compatible word or segment timestamps. ASR text and timing remain owned by the ASR engine; this model supplies only the diarization timeline used for speaker alignment. Correctness, peak PSS, first-result latency, and sustained RTF have been measured on one S23 Ultra. Longer thermal runs, delegate fallback coverage, a broader Android device matrix, concurrent ASR resource tests, and corpus DER/JER remain open production gates.

Related implementation

The two-stage graph boundary is consistent with the community Core ML conversion by smdesai, while tensor shapes and validation in this repository remain pinned to the NVIDIA source revision listed above. This LiteRT conversion does not reuse Core ML weights or binaries.

Downloads last month
200
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for spybyscript/Nemotron-3-Diarization-litert

Finetuned
(10)
this model