By the time the model writes its first token it has already made up its mind; minojev reads that answer instead. One forward pass returns a typed, calibrated probability distribution — no output tokens.
After reading your question, the model has already formed an opinion. Making it write that opinion out as text is the slowest possible way to hear it. So minojev reads the opinion straight from the model's hidden states, and reports it as probabilities over the candidates you declared.
Three steps, no generation loop, no output parser.
Each option becomes a path: state + question + that candidate. Related questions share one state prefix.
Paths are batched through the backbone. Every decision head reads hidden states; no token is sampled, ever.
Choice, boolean, and score questions return distributions. Dev-fitted temperatures keep confidence honest.
Freeze the backbone, cache every candidate path once, then train only the decision head. No backbone drift, no catastrophic forgetting — just minutes of laptop compute for a calibrated decision layer.
What makes this project more than a demo.
A 547k-parameter transformer and a byte tokenizer train on a laptop CPU in minutes — no downloads, no API keys, no GPU.
Datasets, teacher targets, per-question predictions, metrics, and replay bundles are committed in the repo.
Temperature scaling is fitted on dev and stored in the checkpoint. ECE is measured before and after.
Train decision heads from scratch, or read native logits from a pretrained Hugging Face model with zero training.
Many questions per state in one pass; one KV prefill per distinct state, then cheap branches.
The whole model, trainer, calibration, and demos fit in one afternoon of reading. Bilingual docs included.
All replay committed bundles — no server model required.
Every step shows a move distribution, four parallel safety booleans, and the code-enforced final move. Learned from scratch in minutes.
A learned policy eats food while avoiding its own body: move choice, four safety booleans, and a distance score, all from one forward pass per step.
One state, many runtime questions scored together, with teacher distributions overlaid.
Six operational domains — support, moderation, code review, refunds, urgency, leads — answered together by the post-trained model.
One state, three typed questions, one forward pass — watch distributions become a deterministic action through code thresholds.
Drag a confidence threshold and watch coverage, accuracy, and the reliability diagram respond, live in the browser.
Flagship: Qwen3-1.7B head training. Compact from-scratch models power the maze, snake, and console demos. Reproduce with scripts/build_results.sh.
minojev synth --out data/train.jsonl --count 512 --seed 17 --split train minojev train --train data/train.jsonl --dev data/dev.jsonl \ --output-dir runs/synth --steps 800 --calibrate gold --device cpu minojev score --checkpoint runs/synth/checkpoint \ --input examples/decisions.jsonl --output results.jsonl --mode reuse