code-daemon-reranker-v1

A cross-encoder reranker for natural-language β†’ code search. It reads a (query, candidate) pair and returns one relevance logit, and re-orders the top candidates of the Code-Daemon hybrid retriever (embeddings + BM25 + code graph).

What makes it different: it was fine-tuned on exactly what it sees when it serves β€” the candidate pools Code-Daemon's own retriever produces, with the retriever's own near-misses as negatives, and the document text in the form the search actually feeds it. A generic reranker learns from a proxy retriever's mistakes; this one learned from the ones it has to fix.

  • ~117M parameters β€” XLM-RoBERTa, 12 layers Γ— 384 hidden, multilingual 250k vocabulary.
  • 2-input ONNX (input_ids, attention_mask, no token_type_ids) β†’ one logit per pair.
  • 256 tokens per pair: the query is kept whole, the document's tail is cut.

The weights under this id were replaced on 2026-09-13. The previous revision, a fine-tune on public CoIR pairs, scored below its own untuned base inside Code-Daemon's search; figures published for it do not describe these weights.

Numbers

Ten repositories never seen in training, 6 141 natural-language queries, run through Code-Daemon's search with only the reranker changed (pool of 20 candidates, same TensorRT FP16 engine path):

reranker hit@1 hit@5 MRR@10 nDCG@10 query latency, p50
mmarco-mMiniLMv2-L12-H384-v1 (untuned base) 0.165 0.450 0.285 0.331 54 ms
this model 0.167 0.470 0.298 0.344 56 ms
  • Better on 9 of 10 repositories, at the same speed β€” the architecture is the base's. Paired over the same queries, Ξ” nDCG@10 is +0.015 (95 % CI +0.011 … +0.018).
  • The gain sits in the middle of the list β€” hit@3 to hit@5 and MRR. It does not move the first result much.
  • The trade is fit, not cost: the model is tuned to Code-Daemon's document shapes. Plain (query, code snippet) pairs work, but that is not what it was tuned on.
  • Absolute numbers are low because the queries are single-positive at file level: a relevant neighbouring file scores zero. Read the difference between the rows, not the level.

Latency is end to end β€” retrieval plus reranking 20 pairs β€” on a laptop RTX 5060, TensorRT FP16.

How to use it

Re-rank the top ~20 candidates of a first-stage retriever. Feed (query, candidate) pairs, sort by logit. The score is a raw logit β€” compare it within one query, not against a fixed threshold; Code-Daemon takes sigmoid(logit) and blends it with the retriever's score.

import onnxruntime as ort, numpy as np
from transformers import AutoTokenizer

tok  = AutoTokenizer.from_pretrained(".")
sess = ort.InferenceSession("model.onnx", providers=["CPUExecutionProvider"])

def rerank(query, docs, max_len=256):
    enc = tok([query] * len(docs), docs, padding=True, truncation="only_second",
              max_length=max_len, return_tensors="np", return_token_type_ids=False)
    logits = sess.run(None, {"input_ids":      enc["input_ids"].astype(np.int64),
                             "attention_mask": enc["attention_mask"].astype(np.int64)})[0]
    return sorted(zip(logits.reshape(-1).tolist(), docs), reverse=True)   # higher = more relevant

Documents. About 900 characters of code per candidate is what was measured above. Code-Daemon often sends a candidate as a file header plus the matched entity's siblings, with a >>> marker on the candidate itself; the model has seen that form.

Shapes. OpenVINO and Core ML ship batch 16 Γ— seq 256 (Core ML also seq 64 and 128 β€” route short pairs there). Everything is FP16: under OpenVINO INT8 this graph hits a known access violation on the Intel iGPU.

How it was made

Warm-started from cross-encoder/mmarco-mMiniLMv2-L12-H384-v1 and fine-tuned with a listwise objective (one positive against up to 16 negatives) on:

  • behavioural queries written per code entity by an LLM that was not allowed to use the entity's identifiers β€” "compare two hash hierarchies and return list of changed paths", not diffMerkleTrees;
  • the candidate pools Code-Daemon's retriever returned for those queries, with negatives taken from the pool's uncertain band, over open-source repositories disjoint from the evaluation set.

The checkpoint was selected on the evaluation repositories' pools replayed offline.

Files

file what it is
model.onnx + model.onnx.data FP32 ONNX β€” the build source and the standalone path
model.safetensors, config.json the same weights as XLMRobertaForSequenceClassification; AutoModelForSequenceClassification loads them as-is; the Apple (MLX) build is prepared from this pair
tokenizer.json, tokenizer_config.json, sentencepiece.bpe.model the tokenizer β€” stock XLM-R
code-daemon-reranker-v1_{win_x64,linux_x64}_trt11.0_sm_120.engine TensorRT FP16 for RTX 50xx
code-daemon-reranker-v1_ov2026.4_{cpu,igpu}_fp16_b16_s256.{xml,bin} OpenVINO 2026.4 FP16, Intel CPU / iGPU
coreml_ane/embed.mlpackage/ Core ML for the Apple Neural Engine, functions b16 s64 / s128 / s256; load with cpuAndNeuralEngine, fall back to MLX for other shapes

TensorRT engines for other GPU architectures are not published for this revision yet β€” build them from model.onnx.

License

Released under the MIT license.

⚠️ The warm-start base cross-encoder/mmarco-mMiniLMv2-L12-H384-v1 derives from mMARCO ← MS MARCO, whose terms are non-commercial research. Whether a fine-tuned model inherits dataset-use terms is legally unsettled; this is not legal advice. No MS MARCO data was used in this fine-tune; retrain from a permissive base if strict compliance is required.

Attribution

Warm-started from cross-encoder/mmarco-mMiniLMv2-L12-H384-v1. Backbone: XLM-RoBERTa.

Downloads last month
150
Safetensors
Model size
0.1B params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for faxenoff/code-daemon-reranker-v1