๐Ÿ† Caratan โ€” Evals & Matchups

Head-to-head win-rate โ€” full games, trained model vs untrained base

Each trained checkpoint plays the untrained base (Qwen/Qwen3-8B) over the held-out grader_games boards, mirrored seat-swap so win-rate is seat-fair. This measures what training bought in actual games โ€” distinct from the per-decision eval below.

Game-quality stats โ€” measured from the games ยท untrained vs settlement-only vs fully-trained

Per-decision held-out eval โ€” pick quality on frozen scenarios