Each contest is evaluated from four viewpoints. Win rates and game metrics are computed mechanically from the game logs, while the other two assess the quality of dialogue: human judges, and LLM judges that rank the players of each game (relative evaluation).

  • Win rates: wins divided by games, overall and per role, plus a role-composition-weighted rate. Same definition for every contest.
  • Game metrics: how often a team was executed, divined or guarded relative to chance, and how accurately it found werewolves when divining or voting (1.00 = random play).
  • Human evaluation: judges read selected logs and rate or rank the players on criteria A–E (plus F, team play, in 9- and 13-player villages). Ranking since the 2024 winter contest; 5-point scores before that.
  • LLM-as-a-Judge (relative evaluation): LLMs rank the players of each game on the same criteria; ranks are averaged per team. Part of the official evaluation since the 2025 international contest (for the 2025 spring contest it was run afterwards for reference); the models differ by contest (see below). The judge is published as aiwolf-nlp-llm-judge.

Judges by contest

ContestHumanLLM relative
International Contest 2026 (INLG 2026)Not conductedRound 1: GPT-5.4
Gemini 2.5 Pro
Claude Sonnet 4.5
Round 2: GPT-5.6 Terra
Claude Sonnet 5
Gemini 3.1 Pro
gemma-4-31b
Qwen3.6-27b
Mistral-Small-3.2
EXAONE-4.5-33b
Seed-OSS-36b
International Contest 2025 (INLG 2025)4 student judgesGPT-4o
GPT-5
International Contest 2024 (INLG 2024)4 judges (Japanese) / 3 judges (English)Not conducted