← all demos Decision console Domains GitHub Models 中文
This replay is driven by a compact 547k-parameter model trained from scratch on synthetic maze decisions — the flagship Qwen3-1.7B head-training results live on the benchmark page. At each step the model answers nine questions in one forward pass: one choice over four moves, four parallel safety booleans, and one distance score. Code multiplies the move distribution by safety probabilities, filters out blocked moves, and takes the largest remaining value.