Each contest is evaluated from four viewpoints. Win rates and game metrics are computed mechanically from the game logs, while the other two assess the quality of dialogue: human judges, and LLM judges that rank the players of each game (relative evaluation).
- Win rates: wins divided by games, overall and per role, plus a role-composition-weighted rate. Same definition for every contest.
- Game metrics: how often a team was executed, divined or guarded relative to chance, and how accurately it found werewolves when divining or voting (1.00 = random play).
- Human evaluation: judges read selected logs and rate or rank the players on criteria A–E (plus F, team play, in 9- and 13-player villages). Ranking since the 2024 winter contest; 5-point scores before that.
- LLM-as-a-Judge (relative evaluation): LLMs rank the players of each game on the same criteria; ranks are averaged per team. Part of the official evaluation since the 2025 international contest (for the 2025 spring contest it was run afterwards for reference); the models differ by contest (see below). The judge is published as aiwolf-nlp-llm-judge.
Judges by contest
| Contest | Human | LLM relative |
|---|---|---|
| International Contest 2026 (INLG 2026) | Not conducted | Round 1: GPT-5.4 Gemini 2.5 Pro Claude Sonnet 4.5 Round 2: GPT-5.6 Terra Claude Sonnet 5 Gemini 3.1 Pro gemma-4-31b Qwen3.6-27b Mistral-Small-3.2 EXAONE-4.5-33b Seed-OSS-36b |
| International Contest 2025 (INLG 2025) | 4 student judges | GPT-4o GPT-5 |
| International Contest 2024 (INLG 2024) | 4 judges (Japanese) / 3 judges (English) | Not conducted |