← all demos Maze Domains GitHub Models 中文
A compact model trained from scratch on synthetic snake decisions drives this replay. At each step the model answers six questions in one forward pass: one choice over four moves, four parallel safety booleans, and one distance score. Code multiplies the move distribution by safety probabilities, filters out colliding and reversing moves, and takes the largest remaining value — no output token is ever decoded.