Clewen

Clewen combines Qwen3.8-27B text and image generation with Cloudflare Clef structured decisions, sharing one backbone. The decision mode uses recovered switchable adapters and the original Clef joint schema head. No additional training was performed. Recovered adapters approximate the learned changes; they are not the original LoRA factors.

Usage

Both modes support text and images. Decision inputs can also contain JSON. Modes are selected explicitly:

  • Qwen: text generation and image understanding.
  • Clef: structured answers with probabilities for choice, noul (yes/no), and score questions.

Required runtime packages:

torch==2.10.0
torchvision==0.25.0
transformers==5.10.2
huggingface-hub>=1.5,<2
accelerate>=1.12,<2
safetensors>=0.6,<1
pillow>=11,<13
numpy>=2,<3

The clewen[cuda] installation below also installs flash-linear-attention[cuda]==0.5.2 and a prebuilt causal-conv1d==1.7.0 wheel for Python 3.11 or 3.12.

Use Transformers 5.10.2. Older versions such as 4.56.0 fail to import PreTrainedConfig from this release. In Jupyter, use %pip instead of pip in the commands below, then restart the kernel and rerun the imports after installation.

Install on Linux with Python 3.11 or 3.12 and a CUDA GPU:

pip install torch==2.10.0 torchvision==0.25.0 --index-url https://download.pytorch.org/whl/cu128
pip install "clewen[cuda] @ git+https://github.com/salyamq/clewer.git"
from PIL import Image
from transformers import AutoModelForCausalLM

model = AutoModelForCausalLM.from_pretrained(
    "salyamq/clewen",  # or "salyamq/clewen-flash"
    trust_remote_code=True,
    device_map="cuda",
    dtype="bfloat16",
)

# 1. Qwen: text generation
reply = model.text(
    [{
        "role": "user",
        "content": "Explain gradient descent in plain language and give one simple example.",
    }],
    max_new_tokens=256,
)
print(reply["answer"])

# 2. Clef: structured decision from text
result = model.decision({
    "state": "I was charged twice for my monthly subscription. Please refund the duplicate charge.",
    "questions": {
        "department": {
            "type": "choice",
            "instructions": "Which support department should handle this request?",
            "criteria": {
                "billing": "Payments, invoices, subscriptions, and refunds.",
                "technical": "Software errors, connection issues, and service outages.",
                "account": "Login, passwords, and account access.",
            },
        },
    },
})
print(result["answers"]["department"])

# 3. Qwen: text generation from a photo
image = Image.open("example.jpg").convert("RGB")
reply = model.text(
    [{
        "role": "user",
        "content": [
            {"type": "image", "image": image},
            {"type": "text", "text": "Describe this photo in two sentences."},
        ],
    }],
    max_new_tokens=128,
)
print(reply["answer"])

# 4. Clef: structured decision from a photo
result = model.decision({
    "state": "Inspect the attached photo and decide whether a cat is visible.",
    "images": [image],
    "questions": {
        "cat_visible": {
            "type": "noul",
            "instructions": "Is at least one cat visible in the photo?",
        },
    },
})
print(result["answers"]["cat_visible"])

For text or JSON decisions, omit images and put your information in state. Both modes use the same loaded model instance.

Results

Decision-mode scores from complete evaluation runs. Scores are percentages; higher is better. Workflow exact-action scores require the complete action set to match the reference labels.

Benchmark / metric Clewen Clewen-Flash Clef Clef-Flash Jev DiffusionGemma Jev Kev 9B Laya
BFCL — case exact accuracy 98.41 98.88 98.5 98.8 95.8 96.5 94.5 38.1
API-Bank — accuracy 91.73 93.11 91.9 93.1 88.2 83.7 56.3 11.5
BANKING77 — macro-F1 94.08 90.80 94.2 90.9 79.7 74.3 84.8 14.3
RAGTruth — hallucination F1 79.30 35.74 79.4 35.6 76.5 70.4 46.2 48.8
When2Call MCQ — accuracy 72.48 65.36 72.4 65.6 81.0 75.4 49.6 11.9
Invoice processing — exact actions 64.67 57.11 64.7 57.1 61.8 — — —
Invoice processing — primary action 86.44 74.22 86.2 73.3 83.1 — — —
Customer service — exact actions 76.31 76.96 76.3 77.0 76.0 — — —
Security incidents — exact actions 63.33 61.67 62.9 61.7 61.7 — — —
Agent trace observability — primary action 68.47 69.82 68.5 69.8 71.6 — — —
Decision latency — median, ms ↓ 142.30 78.06 209.3 38.8 524.1 84.4 51.4 5.8
Decision latency — p95, ms ↓ 274.15 111.26 238.6 122.4 536.0 211.2 187.9 222.5

Latency was measured on one H200 in BF16 across the five Decision Index benchmarks. It includes input encoding and answer decoding, excluding model loading and warmup. These evaluations measure structured decisions, not text-generation quality.

License

Apache-2.0. Qwen components are credited to the Qwen authors; Clef components are credited to Cloudflare. See LICENSE, LICENSE-QWEN, LICENSE-CLEF, and NOTICE for attribution and license details.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 1 Ask for provider support

Model tree for salyamq/clewen

Base model

Qwen/Qwen3.8-27B
Finetuned
Cloudflare/clef
Finetuned
(7)
this model

Collection including salyamq/clewen