8ball

Published by PythiaFinance.

A System One decision model fine-tuned to be a Magic 8-Ball: it answers one 20-way choice question β€” "the asker posed the question in the state; answer it with the classic Magic 8-Ball response that fits best" β€” with a calibrated probability distribution over the canonical 20 answers (10 affirmative / 5 non-committal / 5 negative), each mapped to a signed alignment value in [-1, 1].

421M parameters (ModernBERT-large encoder + decision head), Apache-2.0, runs on CPU (330–390 ms/question) or a single consumer GPU (5–10 ms). Loads like any Laya checkpoint:

import laya

CRITERIA = {  # the canonical 20, verbatim β€” the model was trained on exactly these
    "It is certain": "Strong affirmative: no doubt whatsoever",
    "It is decidedly so": "Strong affirmative: beyond question",
    "Without a doubt": "Strong affirmative",
    "Yes definitely": "Strong affirmative",
    "You may rely on it": "Strong affirmative: dependable yes",
    "As I see it, yes": "Mild affirmative: yes with minor reservations",
    "likely": "Mild affirmative: probable yes",
    "Outlook good": "Mild affirmative: favorable prospects",
    "Yes": "Plain affirmative",
    "Signs point to yes": "Mild affirmative: indicators favor yes",
    "Reply hazy, try again": "Non-committal: rephrase and ask again",
    "Ask again later": "Non-committal: no answer available right now",
    "Better not tell you now": "Non-committal: answer withheld for now",
    "Cannot predict now": "Non-committal: prediction unavailable at this time",
    "Concentrate and ask again": "Non-committal: asker should focus and retry",
    "Don't count on it": "Mild negative: unlikely",
    "My reply is no": "Strong negative: the ball says no",
    "My sources say no": "Strong negative: authoritative no",
    "Outlook not so good": "Mild negative: unfavorable prospects",
    "doubtful": "Mild negative: doubtful",
}

agent = laya.load("PythiaFinance/8ball")
result = agent.predict("Should I deploy on Friday?", {
    "answer": {
        "type": "choice",
        "instructions": "The asker posed the question found in the state. Answer it with the classic Magic 8-Ball response that fits best.",
        "criteria": CRITERIA,
    }
})
probs = result["answers"]["answer"]["probabilities"]  # {answer: p}, sums to 1

Sampling & deflection

The distribution is the point, so sample from it rather than taking the argmax. When the model is unsure (p_max < 0.09, calibrated to this checkpoint), restrict the draw to the five non-committal answers:

import random

NON_COMMITTAL = ["Reply hazy, try again", "Ask again later", "Better not tell you now",
                 "Cannot predict now", "Concentrate and ask again"]

def shake(probs, temperature=0.9, deflect_below=0.09):
    pool = NON_COMMITTAL if max(probs.values()) < deflect_below else list(probs)
    weights = [probs[a] ** (1 / temperature) for a in pool]
    return random.choices(pool, weights=weights)[0]

The signed alignment score is value(answer) * probs[answer], where value runs +1.0 β†’ +0.1 down the affirmative ladder (in 0.1 steps: It is certain, It is decidedly so, Without a doubt, Yes definitely, You may rely on it, Yes, As I see it yes, likely, Outlook good, Signs point to yes), is 0 for all non-committal answers, and βˆ’1.0 β†’ βˆ’0.2 down the negative ladder (in 0.2 steps: My reply is no, My sources say no, Don't count on it, Outlook not so good, doubtful).

Results

The 8-ball persona (held-out eval, n=620)

metric value
Sampled bucket mix 49.7 / 25.1 / 25.2 (canonical: 50 / 25 / 25)
Brier vs persona gold 0.0031 (teacher: 0.0095, uniform: 0.0175)
Mean p_max 0.106 (p5 0.084 Β· p50 0.100 Β· p95 0.145) β€” deliberately humble
Deflection rate @0.09 18.7% of shakes fall back to the non-committal answers

General decision skill (S1MB english-v1, 137 benchmarks, 26,269 judgments)

Evaluated with the official S1MB evaluator (baseline-adjusted scores 0–100; extended-input condition --max-len 9216).

model noul choice score Task Avg
TypeSafe Jev 1.13 64.77 67.31 46.69 59.59
laya-typed-decisions (base, published) 20.06 18.93 5.99 15.00
8ball v1 (persona-only fine-tune) 18.77 17.14 2.14 12.69
8ball v2 (MTL r1, 47k items) 28.87 21.47 19.18 23.17
8ball v3 (this model, 143k items, 6 epochs) 33.85 24.82 21.43 26.70

The multi-task training traded ~1.7 Task Avg vs the base teacher for the persona (mix, alignment scale, humility) β€” and recovered general skill 12.69 β†’ 26.70 along the way. Methodology is summarized under Training below.

Training

  • Recipe: RLCD (upstream Laya loop) β€” proper-scoring-rule policy gradient (spherical 0.75 + ranked probability 1.0) over noisy logit projections, plus soft cross-entropy against gold distributions. 6 epochs total, fp32, per-type temperature calibration after the final epoch.
  • Corpus (143,514 items, 0 skipped):
    • Open-Jev release-v2-redistributable train (79k, CC0) β€” all 12 sources, uncapped
    • 14 public NLP train splits converted to S1MB decision formats (30.1k): SNLI, HANS, PAWS, CREAK, SciTail, ETHOS, Civil Comments, DBpedia, AG News, poem_sentiment, ARC, OpenBookQA, AQuA-RAT, Banking77
    • 8-ball persona gold (11.6k effective): laya-typed-decisions teacher distributions over 6,200 whimsical questions, bucket-rescaled to the canonical 50/25/25 mix, blended 70/30 with uniform (humility baked into targets)
    • Qwen3.8-27B-generated coding/logic/finance decision cases (4.4k)
  • Never trained on S1MB test rows β€” NLP conversion transfers decision formats (instructions/criteria/state shape read from local test schemas) applied to train-split rows only.

Limitations

  • No world knowledge: it judges the state you hand it. Without evidence in the state it is deliberately near-uniform β€” an honest 8-ball, not an oracle.
  • The persona is trained-in: this checkpoint is near-uniform on almost everything. For general typed decisions use convaiinnovations/laya-typed-decisions.
  • Score-primitive skill is weak (21.4) relative to frontier decision models.
  • Trained for the fixed 20-option 8-ball choice; other option sets work (request-time declaration) but are not calibrated for.

Intended use

A calibrated, whimsical decision layer: "should I…" questions in β†’ an 8-ball answer with a signed alignment score between βˆ’1 and +1 (direction from the sampled answer, magnitude from the probability mass). See Sampling & deflection above (deflection threshold calibrated to this checkpoint: 0.09).

Downloads last month
10
Safetensors
Model size
0.4B params
Tensor type
F16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for PythiaFinance/8ball

Finetuned
(8)
this model