Official Results (Awards)
Win-rate awards are based on the role-composition-weighted win rate ("Adjusted" in the tables below).
| Award | Team |
|---|---|
| 5-player village: Best win rate | CanisLupus |
| 5-player village: Best human evaluation | sunamelli |
| 5-player village: Human evaluation, 2nd | yharada |
| 13-player village: Top win rate | sunamelli-c |
| 13-player village: Top win rate | kanolab-nw-B |
Results and awards were presented at the AIWolfDial 2025 workshop (October 30, 2025, INLG 2025, Hanoi).
Win Rates
Computed from the logs of all games that finished normally in the main competition. The official win rate is "Adjusted": the per-role win rates averaged with the weights of the village's role composition, so that uneven role assignments cancel out; teams are ranked by it. "Overall" is simply wins divided by games, and per-role cells show the win rate in that role with (wins / games).
5-player village (120 games)
| # | Team | Games | Adjusted | Overall | Villager | Seer | Possessed | Werewolf |
|---|---|---|---|---|---|---|---|---|
| 1 | CanisLupus | 73 | 66.5% | 67.1% | 66.7% (18/27) | 81.2% (13/16) | 42.9% (6/14) | 75.0% (12/16) |
| 2 | kanolab-nw | 75 | 62.7% | 62.7% | 73.3% (22/30) | 60.0% (9/15) | 66.7% (10/15) | 40.0% (6/15) |
| 3 | sunamelli | 75 | 59.8% | 60.0% | 67.7% (21/31) | 66.7% (10/15) | 50.0% (7/14) | 46.7% (7/15) |
| 4 | CamelliaDragons | 75 | 53.4% | 53.3% | 61.3% (19/31) | 57.1% (8/14) | 37.5% (6/16) | 50.0% (7/14) |
| 5 | GPTaku | 77 | 46.8% | 46.8% | 53.3% (16/30) | 62.5% (10/16) | 33.3% (5/15) | 31.2% (5/16) |
| 6 | yharada | 74 | 46.1% | 46.0% | 50.0% (15/30) | 64.3% (9/14) | 37.5% (6/16) | 28.6% (4/14) |
| 7 | mille | 77 | 44.4% | 44.2% | 60.0% (18/30) | 33.3% (5/15) | 37.5% (6/16) | 31.2% (5/16) |
| 8 | Character-Lab | 74 | 34.7% | 35.1% | 41.9% (13/31) | 46.7% (7/15) | 21.4% (3/14) | 21.4% (3/14) |
13-player village (13 games)
| # | Team | Games | Adjusted | Overall | Villager | Seer | Bodyguard | Medium | Possessed | Werewolf |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | sunamelli-c | 13 | 63.1% | 61.5% | 42.9% (3/7) | - | 100.0% (1/1) | 0.0% (0/1) | 100.0% (1/1) | 100.0% (3/3) |
| 2 | kanolab-nw-B | 13 | 61.5% | 61.5% | 16.7% (1/6) | 100.0% (1/1) | 100.0% (1/1) | 100.0% (1/1) | 100.0% (1/1) | 100.0% (3/3) |
| 3 | sunamelli-b | 13 | 61.5% | 61.5% | 33.3% (2/6) | 100.0% (1/1) | 100.0% (1/1) | 0.0% (0/1) | 100.0% (1/1) | 100.0% (3/3) |
| 4 | mille-B | 13 | 46.2% | 46.2% | 33.3% (2/6) | 0.0% (0/1) | 100.0% (1/1) | 0.0% (0/1) | 100.0% (1/1) | 66.7% (2/3) |
| 5 | kanolab-nw-C | 13 | 46.2% | 46.2% | 33.3% (2/6) | 0.0% (0/1) | 0.0% (0/1) | 100.0% (1/1) | 0.0% (0/1) | 100.0% (3/3) |
| 6 | kanolab-nw-A | 13 | 46.2% | 46.2% | 33.3% (2/6) | 0.0% (0/1) | 0.0% (0/1) | 100.0% (1/1) | 100.0% (1/1) | 66.7% (2/3) |
| 7 | CamelliaDragons | 13 | 46.2% | 46.2% | 50.0% (3/6) | 0.0% (0/1) | 0.0% (0/1) | 0.0% (0/1) | 100.0% (1/1) | 66.7% (2/3) |
| 8 | CanisLupus-B | 13 | 46.2% | 46.2% | 33.3% (2/6) | 0.0% (0/1) | 0.0% (0/1) | 100.0% (1/1) | 100.0% (1/1) | 66.7% (2/3) |
| 9 | CanisLupus-A | 13 | 45.4% | 46.2% | 40.0% (2/5) | 50.0% (1/2) | 0.0% (0/1) | 0.0% (0/1) | 100.0% (1/1) | 66.7% (2/3) |
| 10 | Character-Lab-B | 13 | 30.8% | 30.8% | 33.3% (2/6) | 0.0% (0/1) | 0.0% (0/1) | 0.0% (0/1) | 0.0% (0/1) | 66.7% (2/3) |
| 11 | sunamelli-a | 13 | 30.8% | 30.8% | 33.3% (2/6) | 0.0% (0/1) | 0.0% (0/1) | 0.0% (0/1) | 0.0% (0/1) | 66.7% (2/3) |
| 12 | mille-A | 13 | 15.4% | 15.4% | 0.0% (0/6) | 100.0% (1/1) | 0.0% (0/1) | 0.0% (0/1) | 0.0% (0/1) | 33.3% (1/3) |
| 13 | Character-Lab-A | 13 | 15.4% | 15.4% | 16.7% (1/6) | 0.0% (0/1) | 0.0% (0/1) | 0.0% (0/1) | 100.0% (1/1) | 0.0% (0/3) |
Game Metrics
Behavioural metrics computed mechanically from the game logs, independent of dialogue quality. Each value is a ratio against what random play would produce (1.00 = random). "Executed", "divined" and "guarded" show how often the team was targeted relative to chance; lower is better for being executed or divined. "Divination accuracy" and "vote accuracy" show how often the team found a werewolf relative to chance; higher is better. "Possessed avoidance" is the rate at which, as the possessed, the team voted for someone other than a werewolf.
5-player village
| Team | Executed | Divined | Divination accuracy | Vote accuracy | Possessed avoidance |
|---|---|---|---|---|---|
| CanisLupus | 0.85 | 0.71 | 1.69 | 1.73 | 0.61 |
| kanolab-nw | 0.80 | 0.92 | 1.40 | 1.71 | 0.74 |
| sunamelli | 0.72 | 0.85 | 0.37 | 1.28 | 0.73 |
| CamelliaDragons | 1.04 | 1.03 | 1.35 | 1.20 | 0.58 |
| GPTaku | 1.39 | 1.09 | 1.20 | 1.13 | 0.60 |
| yharada | 0.92 | 0.85 | 1.88 | 1.20 | 0.71 |
| mille | 1.23 | 1.36 | 1.04 | 1.35 | 0.56 |
| Character-Lab | 1.14 | 1.23 | 0.97 | 0.60 | 0.70 |
13-player village
| Team | Executed | Divined | Guarded | Divination accuracy | Vote accuracy | Possessed avoidance |
|---|---|---|---|---|---|---|
| kanolab-nw-B | 0.19 | 0.64 | 1.39 | 1.72 | 1.37 | - |
| sunamelli-b | 0.39 | 0.63 | 0.72 | 1.25 | 1.47 | 1.00 |
| sunamelli-c | 0.67 | 0.53 | 0.83 | - | 1.64 | 0.75 |
| mille-B | 2.89 | 1.50 | 1.96 | 4.00 | 0.92 | 1.00 |
| kanolab-nw-C | 0.58 | 0.57 | 0.76 | 1.84 | 1.39 | 1.00 |
| kanolab-nw-A | 0.72 | 1.61 | 2.08 | - | 1.38 | 1.00 |
| CanisLupus-A | 0.47 | 0.76 | - | 0.65 | 1.59 | 1.00 |
| CamelliaDragons | 3.45 | 0.93 | 2.45 | - | 1.34 | - |
| CanisLupus-B | 1.17 | 1.00 | 0.84 | - | 1.42 | 1.00 |
| Character-Lab-B | 0.56 | 1.64 | 0.91 | - | 0.74 | - |
| sunamelli-a | 0.67 | 1.08 | 0.34 | - | 1.03 | 0.50 |
| mille-A | 2.69 | 0.88 | 0.44 | 2.32 | 0.96 | - |
| Character-Lab-A | 1.30 | 1.40 | 1.63 | 0.70 | 0.20 | 1.00 |
Human Evaluation (Relative)
Human judges read selected game logs and evaluate the players of each game on each criterion. From 2024 winter onward the judges rank the players (average rank, lower is better); in the 2024 spring and 2024 international contests they gave 5-point scores (higher is better). In 13-player villages each entry (e.g. team-A, team-B) is evaluated separately.
5-player village
10 selected games. Average ranks (lower is better).
| # | Team | Natural expression | Contextual dialogue | Logical consistency | Action consistency | Character consistency | Average |
|---|---|---|---|---|---|---|---|
| 1 | sunamelli | 2.17 | 2.08 | 2.00 | 2.25 | 2.54 | 2.21 |
| 2 | yharada | 2.12 | 2.25 | 2.38 | 3.33 | 2.58 | 2.53 |
| 3 | CanisLupus | 3.00 | 2.88 | 2.96 | 2.42 | 2.38 | 2.73 |
| 4 | kanolab-nw | 3.12 | 2.75 | 2.88 | 2.50 | 2.67 | 2.78 |
| 5 | GPTaku | 2.54 | 2.89 | 3.07 | 2.57 | 3.18 | 2.85 |
| 6 | CamelliaDragons | 3.12 | 3.08 | 2.88 | 3.04 | 3.12 | 3.05 |
| 7 | Character-Lab | 3.67 | 3.62 | 3.38 | 3.25 | 3.08 | 3.40 |
| 8 | mille | 4.14 | 4.25 | 4.25 | 4.36 | 4.32 | 4.26 |
13-player village
10 selected games. Average ranks (lower is better).
| # | Team | Natural expression | Contextual dialogue | Logical consistency | Action consistency | Character consistency | Team play | Average |
|---|---|---|---|---|---|---|---|---|
| 1 | sunamelli-b | 3.95 | 3.90 | 4.17 | 4.00 | 5.50 | 4.80 | 4.39 |
| 2 | sunamelli-c | 3.85 | 3.90 | 4.38 | 4.65 | 5.83 | 4.60 | 4.54 |
| 3 | sunamelli-a | 4.00 | 4.38 | 4.67 | 4.67 | 4.90 | 5.00 | 4.60 |
| 4 | CanisLupus-A | 6.08 | 5.00 | 4.88 | 4.28 | 5.17 | 4.55 | 4.99 |
| 5 | CanisLupus-B | 5.53 | 5.40 | 5.67 | 5.15 | 6.42 | 5.20 | 5.56 |
| 6 | kanolab-nw-A | 6.17 | 5.90 | 6.30 | 6.15 | 4.83 | 5.20 | 5.76 |
| 7 | kanolab-nw-C | 7.28 | 6.42 | 7.00 | 6.33 | 4.67 | 5.42 | 6.19 |
| 8 | kanolab-nw-B | 6.58 | 6.90 | 6.88 | 6.92 | 4.90 | 7.17 | 6.56 |
| 9 | Character-Lab-B | 6.75 | 6.80 | 6.72 | 7.30 | 7.17 | 7.00 | 6.96 |
| 10 | Character-Lab-A | 7.47 | 7.83 | 7.45 | 8.35 | 7.20 | 7.90 | 7.70 |
| 11 | mille-A | 10.45 | 10.72 | 10.35 | 10.05 | 10.82 | 10.40 | 10.46 |
| 12 | mille-B | 10.07 | 10.85 | 10.07 | 10.43 | 10.65 | 10.95 | 10.50 |
| 13 | CamelliaDragons | 10.82 | 12.88 | 11.82 | 12.72 | 12.80 | 12.75 | 12.30 |
LLM-as-a-Judge (Relative Evaluation)
LLM judges rank the players of each game on each evaluation criterion; the ranks are averaged per team (lower is better). The judge models differ by contest and are listed under each track. The judge is published as aiwolf-nlp-llm-judge.
5-player village
Judge models: GPT-4o・GPT-5 の 2 モデル平均(本戦全ゲーム)
| # | Team | Natural expression | Contextual dialogue | Logical consistency | Action consistency | Character consistency | Average |
|---|---|---|---|---|---|---|---|
| 1 | sunamelli | 2.27 | 2.33 | 2.43 | 2.17 | 2.73 | 2.39 |
| 2 | yharada | 1.81 | 1.95 | 2.42 | 3.10 | 2.66 | 2.39 |
| 3 | Character-Lab | 3.16 | 2.29 | 2.71 | 2.99 | 2.24 | 2.68 |
| 4 | kanolab-nw | 3.01 | 2.97 | 3.23 | 2.75 | 2.15 | 2.82 |
| 5 | CamelliaDragons | 3.04 | 3.10 | 2.87 | 3.27 | 3.01 | 3.06 |
| 6 | CanisLupus | 3.54 | 3.53 | 3.05 | 2.75 | 3.10 | 3.19 |
| 7 | GPTaku | 3.06 | 3.45 | 3.55 | 3.01 | 3.44 | 3.30 |
| 8 | mille | 4.08 | 4.32 | 3.69 | 3.93 | 4.60 | 4.12 |
13-player village
Judge models: GPT-4o・GPT-5 の 2 モデル平均(本戦全ゲーム)
| # | Team | Natural expression | Contextual dialogue | Logical consistency | Action consistency | Character consistency | Team play | Average |
|---|---|---|---|---|---|---|---|---|
| 1 | sunamelli-b | 4.08 | 4.38 | 5.46 | 5.42 | 6.23 | 5.77 | 5.22 |
| 2 | sunamelli-c | 4.31 | 4.81 | 5.50 | 5.73 | 6.12 | 5.38 | 5.31 |
| 3 | CanisLupus-A | 6.00 | 5.88 | 5.31 | 4.50 | 5.35 | 4.88 | 5.32 |
| 4 | sunamelli-a | 5.23 | 5.27 | 5.62 | 4.73 | 7.38 | 5.54 | 5.63 |
| 5 | kanolab-nw-B | 5.77 | 5.81 | 7.50 | 7.04 | 4.00 | 6.19 | 6.05 |
| 6 | kanolab-nw-A | 6.35 | 5.88 | 7.15 | 6.85 | 5.15 | 6.08 | 6.24 |
| 7 | CanisLupus-B | 6.50 | 6.08 | 6.50 | 6.12 | 6.46 | 6.19 | 6.31 |
| 8 | kanolab-nw-C | 6.73 | 7.12 | 7.23 | 6.42 | 4.96 | 5.77 | 6.37 |
| 9 | Character-Lab-B | 5.96 | 6.69 | 5.96 | 7.96 | 5.62 | 7.19 | 6.56 |
| 10 | Character-Lab-A | 7.08 | 6.81 | 7.23 | 8.19 | 6.77 | 6.50 | 7.10 |
| 11 | mille-B | 10.08 | 10.31 | 9.62 | 7.96 | 10.69 | 9.77 | 9.74 |
| 12 | mille-A | 10.96 | 10.04 | 9.42 | 9.08 | 10.19 | 10.04 | 9.96 |
| 13 | CamelliaDragons | 11.96 | 11.92 | 8.50 | 11.00 | 12.08 | 11.69 | 11.19 |
Game Logs
Logs of the main competition. Only games that finished normally are listed.
| Track | Games | Logs |
|---|---|---|
| 5-player village | 120 | Log list |
| 13-player village | 13 | Log list |