Each trained checkpoint plays the untrained base (Qwen/Qwen3-8B) over the
held-out grader_games boards, mirrored seat-swap so win-rate is seat-fair. This measures what
training bought in actual games โ distinct from the per-decision eval below.