SLI Sentences
Continuous Saudi Sign Language recognition from body, face and hand landmarks. Each model reads a whole signed sentence and outputs its glosses with CTC. They power Sentences mode in SLI, the Arabic sign-language interpreter made by the Digital Fingers Team.
Open source. The code is MIT-licensed on GitHub: Digital-Fingers-Team/sli-ai. It includes the training code, and PyTorch checkpoints are included here for fine-tuning. Improvements are welcome: see Improve it.
| File | Model | Params | Size | Dev WER | One hand hidden | 1.5× faster | No face |
|---|---|---|---|---|---|---|---|
sli-sentences-large-v2.onnx |
SLI Sentences Large v2 | 42M | 169 MB | 17.5% | 18.0% | 16.8% | 19.2% |
sli-sentences-phone-v1.onnx |
SLI Sentences Phone v1 | 4.5M | 17.9 MB | 17.4% | 18.0% | 16.7% | 19.4% |
Word error rate (WER) is measured on the signer-independent (SI) dev split of Isharah-1000. The dev signer is never seen in training. Scores come from the exported ONNX files. For reference, the best SI baseline in the Isharah paper is Swin-MSTP (RGB video), with a dev WER of 17.9% (Alyami et al., arXiv:2506.03615, Table 4). The best pose-based entry in the MSLR 2025 challenge reached 7.3% (arXiv:2508.09372), so there is plenty of room to improve.
Stress tests: one hand hidden hides a hand for 0.5–1.5 s at a time (about 1/3 of frames); 1.5× faster speeds up the signing; no face drops the face landmarks entirely (signer far from the camera).
Models
Both are Conformer encoders (self-attention, convolution kernel 15) with a linear CTC head over 673 classes (blank = 0, plus 672 Isharah glosses).
- Large v2: d=384, 12 layers, 6 heads, FF 1536. It is pre-trained on KArSL (502 isolated signs), trained on the Isharah SI train split (10 signers) with strong augmentation for 100 epochs, and self-distilled from an earlier large model. Sentence accuracy on dev is 67.2%.
- Phone v1: d=192, 6 layers, 4 heads, FF 512. It is distilled from a large teacher with the same augmentation. Small enough to run in a browser (onnxruntime-web). It scores slightly better than its teacher.
In the SLI app the large model runs on a server, and the phone model runs in the browser as a fallback.
Input: SF1 features
feats: float32[1, T, 356]at 15 fps. Each frame holds 176 points × (x, y) plus 4 presence flags: pose 6 | face 128 | left hand 21 | right hand 21 | presence (pose, face, left hand, right hand).- Hands are stored by the signer's own side. Each part is normalised to its own anchor and scale. A missing part is all zeros with presence 0.
feats.py(withface_idx.jsonandconventions.json) builds SF1 from MediaPipe landmarks:from_holistic(xyz, conf, w, h)takes MediaPipe Holistic frames (pose-format layout: 33 + 468 + 21 + 21).from_app_raw(raw)takes MediaPipe Tasks pose / face / hands results.resample(frames, t_ms)puts frames on the 15 fps grid.
- The models expect sentences of 8–450 frames (0.5–30 s).
Output
logprobs: [1, T', 673] CTC log-probabilities. Decode greedily: take the argmax per frame,
collapse repeats and drop blanks (0). Class k is the gloss vocab.json["glosses"][k - 1].
vocab.json["lookup"] maps some gloss sequences to written Arabic sentences.
import json, numpy as np, onnxruntime as ort
sess = ort.InferenceSession("sli-sentences-large-v2.onnx")
vocab = json.load(open("vocab.json", encoding="utf-8"))
x = np.load("sentence_sf1.npy").astype(np.float32)[None] # [1, T, 356], 15 fps
lp = sess.run(None, {"feats": x})[0][0]
ids, prev = [], 0
for k in lp.argmax(-1):
if k != prev and k != 0:
ids.append(int(k))
prev = k
glosses = [vocab["glosses"][i - 1] for i in ids]
key = " ".join(glosses)
print(vocab["lookup"].get(key, key))
PyTorch checkpoints
pytorch/ holds the training checkpoints: a dict with state (the state dict), config, arch,
n_classes, dev_wer and epoch. They load with torch.load(path, weights_only=True), on CPU.
| File | What |
|---|---|
sli-sentences-large-v2.pt |
Large v2 (source of the ONNX above) |
sli-sentences-phone-v1.pt |
Phone v1 (source of the ONNX above) |
karsl-pretrain-large.pt |
Large, pre-trained on KArSL's 502 isolated signs (the start of every large round) |
karsl-pretrain-small.pt |
Small, pre-trained on KArSL |
import torch
from training.sentences import model as M # from the GitHub repo
c = torch.load("pytorch/sli-sentences-large-v2.pt", weights_only=True)
m = M.SignCTC(c["n_classes"], **c["arch"])
m.load_state_dict(c["state"])
Improve it
The full guide is CONTRIBUTING.md. In short:
- Attach the public Kaggle datasets (Isharah-1000 pose, Isharah annotations, KArSL keypoints),
and run
training.sentences.prepand thentraining.sentences.train. A free Kaggle GPU trains a large model in 2–3 hours. To fine-tune a checkpoint here, pass--init pytorch/sli-sentences-large-v2.pt --keep-head. - Export to ONNX and score it with
training.sentences.evaluate --stress hand fast noface. Choose settings on SI dev only, and evaluate on SI test once, at the end. - Open a pull request on GitHub with your change and a row in
training/sentences/rounds.md, including runs that did not help. Share the weights in a community pull request here or in your own repo. A model that beats the current one on dev, without getting worse on the stress tests, becomes the next release, with credit.
Also welcome: bug reports from signers, new sentence data, and support for other Arabic sign languages.
Limitations
- The vocabulary is limited to the 672 Isharah glosses, in Saudi Sign Language. Other Arabic sign languages and signs outside the vocabulary are not recognised.
- The models were trained on 10 signers filmed with room around them. Accuracy drops when a hand leaves the frame, and falls by 1–2 points when no face is detected.
- They output glosses, not grammatical Arabic. Only the sentences in
lookupget a written form. - Research prototype: not for medical, legal or other high-stakes interpretation.
Licence and data
The weights are released under CC BY-NC-SA 4.0, following the Isharah dataset licence. You may use, change and share them for non-commercial purposes, with credit, and derived models must use the same licence. The code on GitHub is MIT. Training data: Isharah-1000 (Alyami et al., arXiv:2506.03615) and KArSL (pre-training).