AbstractGym tiny controllers

Twelve independently trained small Transformers: role-mapped step and direct-answer controls (A6), a separate role-mapped full-trace model (X4), and a raw-symbol step ablation (X6). Seeds 0, 1, 2 for each family. These are full model weights, not adapters. Each folder has model.safetensors, config.json, training.json, results.json, and a conversion verification record.

Input representation and split

The step, direct, and trace controllers receive symbols already mapped to roles. The harness supplies this abstraction; perfect step execution does not demonstrate learned raw-symbol binding. Exception: raw-step (X6) receives printable raw-symbol IDs and a supplied symbol map; it was explicitly designed to remove role preprocessing.

A6 development uses 367 trajectories on uvwxy, depths 0–4. Step training deduplicates the prefix states to 36 role observations; direct training uses all 367 membership examples with balanced label sampling. Trace training uses the same development trajectories with balanced membership sampling. Raw-step trains on the 36 raw observations of the development symbol map. The original held-out set has 92 cases across two different alphabets at depths 0/1/2/3/4/8/16. Step models also have 32 extra cases at depths 32/64 and an exhaustive 36-state check. Raw-step evaluates 108 states across the three supplied maps.

Architecture and optimization

All: two Transformer layers, width 64, four heads, feed-forward width 128, GELU, dropout zero, pre-norm, sinusoidal positions (maximum 256), FP32. Step/direct use bidirectional attention and CLS classification; trace/raw-step use causal attention and autoregressive outputs. AdamW LR .001, weight decay .01, clipping 1, cosine decay to 1e-5. Step: 800 full-batch updates on 36 states. Direct/trace: 1,500 updates, balanced batch 64. Raw-step: 800 full-batch updates on 36 states. Exact per-seed metadata and vocabularies are supplied.

Exact held-out results

Family Seed Correct / cases
step 0 92/92
step 1 92/92
step 2 92/92
direct 0 76/92
direct 1 80/92
direct 2 76/92
trace 0 60/92
trace 1 60/92
trace 2 60/92
raw-step 0 0/92
raw-step 1 0/92
raw-step 2 0/92

Step/raw-step scores require complete live execution; direct scores are binary membership; trace scores require the entire generated action sequence. These are different tasks and interfaces. All failures remain in the denominator. Results are copied from docs/results/2026-10-03-a6, X4 metadata, and docs/results/2026-10-04-x6. Detailed auxiliary results remain in each folder.

Conversion and loading

Original local state dictionaries were checked against each run’s frozen.json, the tiny runs’ original checkpoint hash manifest. Tensor values remain bitwise identical after safetensors serialization. Reloading into the original architecture reproduced every saved evaluation prediction, including teacher-forced trace outputs, corrupted-history continuations, exhaustive states, and extra-depth step trajectories. No training or evaluation selection was performed during conversion. See conversion_verification.json for exact counts and hashes.

Clone the public code on main, install PyTorch and safetensors, and set PYTHONPATH=src:scripts:

from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
from abstractgym.a6_model import TinyTransformer
path = hf_hub_download("flavianv/abstractgym-tiny-controllers", "step/seed-0/model.safetensors")
model = TinyTransformer()
model.load_state_dict(load_file(path), strict=True)
model.eval()

For direct use TinyTransformer(outputs=2); for trace import TraceTransformer from run_a6_tiny_trace; for raw-step import RawTransformer from run_a6_x6. Use the matching encoder and vocabulary recorded in config.json and implemented in the public source. These are custom PyTorch models, not root-level AutoModel checkpoints.

Limitations and license

Three seeds per family. The role-mapped models depend on privileged preprocessing, and the raw-symbol ablation does not transfer reliably. Synthetic finite tests, bounded positions, and an external state machine do not prove general computation. MIT license. No pickle files or optimizer states are published. Code: flavianv/abstractgym-public, main.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support