SYSTEM ONE DECISION MODEL

Decisions, not tokens.

By the time the model writes its first token it has already made up its mind; minojev reads that answer instead. One forward pass returns a typed, calibrated probability distribution — no output tokens.

Chat model

generate token by token, then parse
output tokens 0 time 0 ms

minojev

one forward pass, read the distribution
0.79
0.14
0.07
✓ typed · calibrated · decode_steps = 0
output tokens 0 time ~12 ms

The whole insight, in one sentence

After reading your question, the model has already formed an opinion. Making it write that opinion out as text is the slowest possible way to hear it. So minojev reads the opinion straight from the model's hidden states, and reports it as probabilities over the candidates you declared.

0output tokens per decision
1forward pass for every question
2–255candidates, defined at request time
0.016calibration error on the maze model

How it works

Three steps, no generation loop, no output parser.

Build candidate paths

Each option becomes a path: state + question + that candidate. Related questions share one state prefix.

One forward pass

Paths are batched through the backbone. Every decision head reads hidden states; no token is sampled, ever.

Read & calibrate

Choice, boolean, and score questions return distributions. Dev-fitted temperatures keep confidence honest.

Head training

Freeze the backbone, cache every candidate path once, then train only the decision head. No backbone drift, no catastrophic forgetting — just minutes of laptop compute for a calibrated decision layer.

37 min1.7B head training on a laptop
4 GBpeak memory, MPS
95.0%accuracy vs 80.0% for generation
0output tokens per decision

Different by design

What makes this project more than a demo.

Runs offline, from scratch

A 547k-parameter transformer and a byte tokenizer train on a laptop CPU in minutes — no downloads, no API keys, no GPU.

Every claim has an artifact

Datasets, teacher targets, per-question predictions, metrics, and replay bundles are committed in the repo.

Calibration is first-class

Temperature scaling is fitted on dev and stored in the checkpoint. ECE is measured before and after.

Two engines

Train decision heads from scratch, or read native logits from a pretrained Hugging Face model with zero training.

Parallel + reusable state

Many questions per state in one pass; one KV prefill per distinct state, then cheap branches.

Small enough to read

The whole model, trainer, calibration, and demos fit in one afternoon of reading. Bilingual docs included.

Live demos

All replay committed bundles — no server model required.

Measured on held-out data

Flagship: Qwen3-1.7B head training. Compact from-scratch models power the maze, snake, and console demos. Reproduce with scripts/build_results.sh.

95.8%decision accuracy · 88.3% zero-shot logits
0output tokens per decision
0.024ECE after calibration · 0.014 on 200 requests
~5×better p95 latency than token generation
5/6mazes solved by the compact demo model
73tests passing offline

Train one in a minute

minojev synth  --out data/train.jsonl --count 512 --seed 17 --split train
minojev train  --train data/train.jsonl --dev data/dev.jsonl \
  --output-dir runs/synth --steps 800 --calibrate gold --device cpu
minojev score  --checkpoint runs/synth/checkpoint \
  --input examples/decisions.jsonl --output results.jsonl --mode reuse