International Contest 2026 (INLG 2026)

Open the contest page →

Official Results (Awards)

The list of awarded teams is being compiled.

Awards will be announced at the AIWolfDial 2026 workshop (October 17–21, 2026, INLG 2026, Utrecht).

Win Rates

Computed from the logs of all games that finished normally in the main competition. "Overall" is wins divided by games. Per-role cells show the win rate in that role with (wins / games). "Adjusted" re-weights the per-role win rates by the village's role composition, cancelling out uneven role assignments.

Round 1, 5-player village (220 games)

#TeamGamesOverallVillagerSeerPossessedWerewolfAdjusted
1yharada10061.0%70.0% (28/40)75.0% (15/20)25.0% (5/20)65.0% (13/20)61.0%
2Oxymoron10061.0%72.5% (29/40)75.0% (15/20)40.0% (8/20)45.0% (9/20)61.0%
3noisy_watcher10060.0%77.5% (31/40)70.0% (14/20)40.0% (8/20)35.0% (7/20)60.0%
4SunameLLi10057.0%67.5% (27/40)65.0% (13/20)30.0% (6/20)55.0% (11/20)57.0%
5wool_100%10057.0%62.5% (25/40)70.0% (14/20)60.0% (12/20)30.0% (6/20)57.0%
6ryanleeai10056.0%75.0% (30/40)60.0% (12/20)35.0% (7/20)35.0% (7/20)56.0%
7yshimoda10051.0%50.0% (20/40)80.0% (16/20)30.0% (6/20)45.0% (9/20)51.0%
8kataken14f10049.0%62.5% (25/40)60.0% (12/20)40.0% (8/20)20.0% (4/20)49.0%
9CamelliaDragons10045.0%57.5% (23/40)65.0% (13/20)20.0% (4/20)25.0% (5/20)45.0%
10gotsumori-NLP10043.0%62.5% (25/40)15.0% (3/20)45.0% (9/20)30.0% (6/20)43.0%
11CanisLupus10040.0%42.5% (17/40)65.0% (13/20)35.0% (7/20)15.0% (3/20)40.0%

Round 1, 9-player village (18 games)

#TeamGamesOverallVillagerSeerBodyguardMediumPossessedWerewolfAdjusted
1CamelliaDragons1866.7%83.3% (5/6)50.0% (1/2)50.0% (1/2)50.0% (1/2)50.0% (1/2)75.0% (3/4)66.7%
2CanisLupus1866.7%83.3% (5/6)50.0% (1/2)50.0% (1/2)50.0% (1/2)100.0% (2/2)50.0% (2/4)66.7%
3Oxymoron1866.7%50.0% (3/6)50.0% (1/2)100.0% (2/2)100.0% (2/2)50.0% (1/2)75.0% (3/4)66.7%
4wool_100%1855.6%66.7% (4/6)0.0% (0/2)50.0% (1/2)100.0% (2/2)100.0% (2/2)25.0% (1/4)55.6%
5yshimoda1855.6%50.0% (3/6)50.0% (1/2)100.0% (2/2)50.0% (1/2)50.0% (1/2)50.0% (2/4)55.6%
6ryanleeai1844.4%33.3% (2/6)100.0% (2/2)50.0% (1/2)50.0% (1/2)0.0% (0/2)50.0% (2/4)44.4%
7noisy_watcher1844.4%50.0% (3/6)50.0% (1/2)0.0% (0/2)100.0% (2/2)50.0% (1/2)25.0% (1/4)44.4%
8SunameLLi1833.3%33.3% (2/6)50.0% (1/2)100.0% (2/2)0.0% (0/2)0.0% (0/2)25.0% (1/4)33.3%
9yharada1833.3%50.0% (3/6)100.0% (2/2)0.0% (0/2)0.0% (0/2)0.0% (0/2)25.0% (1/4)33.3%

Round 1, Speak-Anytime Track, 5-player village (14 games)

#TeamGamesOverallVillagerSeerPossessedWerewolfAdjusted
1yharada1060.0%25.0% (1/4)100.0% (2/2)50.0% (1/2)100.0% (2/2)60.0%
2CanisLupus1060.0%75.0% (3/4)0.0% (0/2)100.0% (2/2)50.0% (1/2)60.0%
3Oxymoron1050.0%50.0% (2/4)50.0% (1/2)50.0% (1/2)50.0% (1/2)50.0%
4CamelliaDragons1050.0%50.0% (2/4)50.0% (1/2)50.0% (1/2)50.0% (1/2)50.0%
5wool_100%1050.0%25.0% (1/4)100.0% (2/2)100.0% (2/2)0.0% (0/2)50.0%
6noisy_watcher1040.0%50.0% (2/4)50.0% (1/2)0.0% (0/2)50.0% (1/2)40.0%
7SunameLLi1040.0%75.0% (3/4)0.0% (0/2)0.0% (0/2)50.0% (1/2)40.0%

Round 2, 5-player village (270 games)

#TeamGamesOverallVillagerSeerPossessedWerewolfAdjusted
1yatolab7569.3%76.7% (23/30)53.3% (8/15)73.3% (11/15)66.7% (10/15)69.3%
2AGKAI7560.0%66.7% (20/30)46.7% (7/15)66.7% (10/15)53.3% (8/15)60.0%
3wool_100%7558.7%53.3% (16/30)80.0% (12/15)46.7% (7/15)60.0% (9/15)58.7%
4CanisLupus7557.3%66.7% (20/30)66.7% (10/15)60.0% (9/15)26.7% (4/15)57.3%
5yharada7556.0%56.7% (17/30)60.0% (9/15)40.0% (6/15)66.7% (10/15)56.0%
6NTT-HAI7554.7%46.7% (14/30)60.0% (9/15)60.0% (9/15)60.0% (9/15)54.7%
7ryanleeai7554.7%40.0% (12/30)86.7% (13/15)40.0% (6/15)66.7% (10/15)54.7%
8SunameLLi7552.0%66.7% (20/30)53.3% (8/15)40.0% (6/15)33.3% (5/15)52.0%
9CamelliaDragons7550.7%60.0% (18/30)40.0% (6/15)46.7% (7/15)46.7% (7/15)50.7%
10kataken14f7550.7%53.3% (16/30)80.0% (12/15)33.3% (5/15)33.3% (5/15)50.7%
11UEC-IL7550.7%53.3% (16/30)60.0% (9/15)40.0% (6/15)46.7% (7/15)50.7%
12Luna7550.7%53.3% (16/30)53.3% (8/15)40.0% (6/15)53.3% (8/15)50.7%
13carrot7546.7%50.0% (15/30)53.3% (8/15)26.7% (4/15)53.3% (8/15)46.7%
14noisy_watcher7545.3%63.3% (19/30)46.7% (7/15)26.7% (4/15)26.7% (4/15)45.3%
15Mackerel7541.3%46.7% (14/30)53.3% (8/15)40.0% (6/15)20.0% (3/15)41.3%
16yshimoda7541.3%50.0% (15/30)40.0% (6/15)40.0% (6/15)26.7% (4/15)41.3%
17Oxymoron7541.3%50.0% (15/30)40.0% (6/15)26.7% (4/15)40.0% (6/15)41.3%
18gotsumori-NLP7538.7%46.7% (14/30)26.7% (4/15)53.3% (8/15)20.0% (3/15)38.7%

Round 2, 9-player village (28 games)

#TeamGamesOverallVillagerSeerBodyguardMediumPossessedWerewolfAdjusted
1yharada1866.7%83.3% (5/6)50.0% (1/2)100.0% (2/2)50.0% (1/2)50.0% (1/2)50.0% (2/4)66.7%
2ryanleeai1861.1%50.0% (3/6)50.0% (1/2)50.0% (1/2)100.0% (2/2)0.0% (0/2)100.0% (4/4)61.1%
3kataken14f1861.1%50.0% (3/6)100.0% (2/2)50.0% (1/2)100.0% (2/2)50.0% (1/2)50.0% (2/4)61.1%
4SunameLLi1861.1%33.3% (2/6)100.0% (2/2)100.0% (2/2)100.0% (2/2)100.0% (2/2)25.0% (1/4)61.1%
5UEC-IL1861.1%66.7% (4/6)50.0% (1/2)100.0% (2/2)50.0% (1/2)50.0% (1/2)50.0% (2/4)61.1%
6noisy_watcher1855.6%83.3% (5/6)50.0% (1/2)0.0% (0/2)50.0% (1/2)50.0% (1/2)50.0% (2/4)55.6%
7wool_100%1855.6%50.0% (3/6)50.0% (1/2)50.0% (1/2)50.0% (1/2)100.0% (2/2)50.0% (2/4)55.6%
8Oxymoron1850.0%66.7% (4/6)100.0% (2/2)0.0% (0/2)50.0% (1/2)50.0% (1/2)25.0% (1/4)50.0%
9Mackerel1850.0%16.7% (1/6)100.0% (2/2)100.0% (2/2)100.0% (2/2)50.0% (1/2)25.0% (1/4)50.0%
10CanisLupus1844.4%50.0% (3/6)0.0% (0/2)100.0% (2/2)50.0% (1/2)0.0% (0/2)50.0% (2/4)44.4%
11yatolab1844.4%66.7% (4/6)0.0% (0/2)50.0% (1/2)50.0% (1/2)0.0% (0/2)50.0% (2/4)44.4%
12yshimoda1844.4%83.3% (5/6)0.0% (0/2)50.0% (1/2)0.0% (0/2)50.0% (1/2)25.0% (1/4)44.4%
13CamelliaDragons1844.4%66.7% (4/6)100.0% (2/2)0.0% (0/2)0.0% (0/2)0.0% (0/2)50.0% (2/4)44.4%
14carrot1833.3%33.3% (2/6)50.0% (1/2)50.0% (1/2)50.0% (1/2)50.0% (1/2)0.0% (0/4)33.3%

Round 2, Speak-Anytime Track, 5-player village (24 games)

#TeamGamesOverallVillagerSeerPossessedWerewolfAdjusted
1yatolab1070.0%75.0% (3/4)100.0% (2/2)100.0% (2/2)0.0% (0/2)70.0%
2ryanleeai1070.0%100.0% (4/4)0.0% (0/2)100.0% (2/2)50.0% (1/2)70.0%
3wool_100%1060.0%75.0% (3/4)100.0% (2/2)50.0% (1/2)0.0% (0/2)60.0%
4SunameLLi1060.0%75.0% (3/4)50.0% (1/2)50.0% (1/2)50.0% (1/2)60.0%
5noisy_watcher1060.0%75.0% (3/4)50.0% (1/2)50.0% (1/2)50.0% (1/2)60.0%
6Oxymoron1050.0%50.0% (2/4)100.0% (2/2)0.0% (0/2)50.0% (1/2)50.0%
7yharada1050.0%50.0% (2/4)50.0% (1/2)0.0% (0/2)100.0% (2/2)50.0%
8kanolab1050.0%50.0% (2/4)50.0% (1/2)50.0% (1/2)50.0% (1/2)50.0%
9kataken14f1040.0%50.0% (2/4)0.0% (0/2)0.0% (0/2)100.0% (2/2)40.0%
10CanisLupus1040.0%25.0% (1/4)100.0% (2/2)50.0% (1/2)0.0% (0/2)40.0%
11UEC-IL1040.0%50.0% (2/4)50.0% (1/2)50.0% (1/2)0.0% (0/2)40.0%
12CamelliaDragons1030.0%25.0% (1/4)50.0% (1/2)0.0% (0/2)50.0% (1/2)30.0%

Game Metrics

Behavioural metrics computed mechanically from the game logs, independent of dialogue quality. Each value is a ratio against what random play would produce (1.00 = random). "Executed", "divined" and "guarded" show how often the team was targeted relative to chance; lower is better for being executed or divined. "Divination accuracy" and "vote accuracy" show how often the team found a werewolf relative to chance; higher is better. "Possessed avoidance" is the rate at which, as the possessed, the team voted for someone other than a werewolf.

Round 1, 5-player village

TeamExecutedDivinedDivination accuracyVote accuracyPossessed avoidance
yharada0.320.861.431.470.40
Oxymoron0.860.891.231.580.69
noisy_watcher0.791.010.841.580.75
SunameLLi0.930.841.001.390.70
wool_100%1.071.221.221.340.67
ryanleeai0.930.871.041.390.48
yshimoda1.000.881.301.480.67
CamelliaDragons0.941.331.641.640.80
gotsumori-NLP1.511.650.970.990.73
CanisLupus1.460.620.681.090.70
kataken1.260.851.711.640.68

Round 1, 9-player village

TeamExecutedDivinedGuardedDivination accuracyVote accuracyPossessed avoidance
CamelliaDragons0.331.861.341.081.200.67
CanisLupus1.931.132.591.521.840.75
Oxymoron1.040.420.340.681.670.75
wool_100%1.281.031.531.271.771.00
yshimoda1.080.530.57-1.700.25
ryanleeai0.740.921.172.151.150.50
noisy_watcher1.321.591.052.271.620.60
SunameLLi1.010.931.171.511.180.60
yharada0.450.83-2.551.850.67

Round 1, Speak-Anytime Track, 5-player village

TeamExecutedDivinedDivination accuracyVote accuracyPossessed avoidance
yharada0.240.83-1.760.75
CanisLupus2.40--2.001.00
Oxymoron0.921.092.000.861.00
CamelliaDragons0.511.001.710.671.00
wool_100%1.371.821.201.001.00
noisy_watcher0.781.121.711.80-
SunameLLi1.330.96-1.00-

Round 2, 5-player village

TeamExecutedDivinedDivination accuracyVote accuracyPossessed avoidance
yatolab0.851.001.571.690.77
AGKAI0.731.080.941.340.82
wool_100%0.661.211.391.330.80
CanisLupus1.340.981.401.230.89
yharada0.550.701.220.970.59
NTT-HAI0.810.760.561.140.64
ryanleeai0.870.841.910.960.71
SunameLLi0.891.190.781.870.58
CamelliaDragons1.111.071.541.450.88
UEC-IL0.870.690.781.010.95
Luna0.831.000.820.980.72
carrot0.930.581.151.060.47
noisy_watcher1.460.840.401.280.38
Mackerel1.210.770.701.200.80
yshimoda1.261.160.980.770.74
Oxymoron0.911.151.040.900.65
gotsumori-NLP2.012.040.900.971.00
kataken0.971.001.331.010.63

Round 2, 9-player village

TeamExecutedDivinedGuardedDivination accuracyVote accuracyPossessed avoidance
yharada0.320.740.151.451.920.50
ryanleeai0.620.851.772.601.670.50
SunameLLi0.550.690.572.741.830.80
UEC-IL0.700.720.711.092.310.83
noisy_watcher0.940.961.911.051.340.75
wool_100%0.761.031.471.101.461.00
Oxymoron1.721.461.79-1.130.67
Mackerel1.550.900.491.081.861.00
CanisLupus1.290.45--1.44-
yatolab0.610.850.591.051.830.67
yshimoda2.050.621.092.662.010.33
CamelliaDragons2.382.120.61-1.760.60
carrot1.301.340.71-1.150.50
kataken0.711.891.751.871.620.67

Round 2, Speak-Anytime Track, 5-player village

TeamExecutedDivinedDivination accuracyVote accuracyPossessed avoidance
yatolab1.331.331.202.001.00
ryanleeai1.871.29-2.001.00
wool_100%1.200.752.002.500.33
SunameLLi0.840.80-1.201.00
noisy_watcher0.821.290.861.671.00
Oxymoron0.861.124.002.670.50
yharada0.310.922.400.860.50
kanolab1.200.75-4.00-
CanisLupus1.201.452.000.671.00
UEC-IL1.670.502.401.331.00
CamelliaDragons0.690.921.200.80-
kataken0.280.97-0.351.00

Human Evaluation (Relative)

Human judges read selected game logs and evaluate the players of each game on each criterion. From 2024 winter onward the judges rank the players (average rank, lower is better); in the 2024 spring and 2024 international contests they gave 5-point scores (higher is better). In 13-player villages each entry (e.g. team-A, team-B) is evaluated separately.

Human evaluation was not conducted in this contest.

LLM-as-a-Judge (Relative Evaluation)

LLM judges rank the players of each game on each evaluation criterion; the ranks are averaged per team (lower is better). The judge models differ by contest and are listed under each track. The judge is published as aiwolf-nlp-llm-judge.

Round 1, 5-player village

Judge models: GPT-5.4・Gemini 2.5 Pro・Claude Sonnet 4.5 の 3 モデル平均

#TeamNatural expressionContextual dialogueLogical consistencyAction consistencyCharacter consistencyAverage
1ryanleeai1.181.452.042.201.471.67
2Oxymoron1.552.002.182.221.501.89
3yharada2.862.552.231.982.282.38
4SunameLLi2.662.872.572.722.892.74
5noisy_watcher2.912.932.722.632.942.83
6kataken14f2.973.052.632.523.112.86
7wool_100%3.113.362.752.993.183.08
8CamelliaDragons4.052.562.652.813.613.14
9yshimoda3.423.782.973.023.593.35
10CanisLupus3.984.323.603.513.723.83
11gotsumori-NLP3.813.904.243.444.203.92

Round 1, 9-player village

Judge models: GPT-5.4・Gemini 2.5 Pro・Claude Sonnet 4.5 の 3 モデル平均

#TeamNatural expressionContextual dialogueLogical consistencyAction consistencyCharacter consistencyTeam playAverage
1ryanleeai1.332.542.483.482.443.562.64
2Oxymoron2.613.703.304.302.224.613.46
3yharada5.304.282.852.634.463.983.92
4CamelliaDragons6.302.613.313.095.243.073.94
5SunameLLi4.524.393.984.374.524.834.43
6noisy_watcher4.416.066.045.395.725.435.51
7yshimoda5.466.415.065.616.065.325.65
8wool_100%6.936.896.265.077.155.506.30
9CanisLupus7.047.746.615.725.917.206.70

Round 1, Speak-Anytime Track, 5-player village

Judge models: GPT-5.4・Gemini 2.5 Pro・Claude Sonnet 4.5 の 3 モデル平均

#TeamNatural expressionContextual dialogueLogical consistencyAction consistencyCharacter consistencyAverage
1Oxymoron1.402.302.032.231.531.90
2SunameLLi2.102.272.132.272.832.32
3yharada2.973.072.431.873.102.69
4noisy_watcher2.672.533.503.432.672.96
5wool_100%3.303.403.003.403.133.25
6CanisLupus3.674.203.173.373.173.51
7CamelliaDragons4.733.203.873.074.533.88

Round 2, 5-player village

Judge models: GPT-5.6 Terra・Claude Sonnet 5・Gemini 3.1 Pro・gemma-4-31b・Qwen3.6-27b・Mistral-Small-3.2・EXAONE-4.5-33b・Seed-OSS-36b の 8 モデル平均

#TeamNatural expressionContextual dialogueLogical consistencyAction consistencyCharacter consistencyAverage
1yharada2.691.381.501.561.251.40
2UEC-IL1.501.752.694.193.502.20
3ryanleeai2.382.883.446.002.383.00
4wool_100%5.195.383.812.506.624.00
5SunameLLi8.126.006.386.196.695.90
6AGKAI9.128.1910.006.757.758.10
7Oxymoron7.319.129.3112.126.388.20
8yatolab13.758.198.005.4412.949.10
9yshimoda4.629.2510.0615.627.259.70
10NTT-HAI12.1211.258.506.1912.5610.10
11kataken14f13.128.8810.069.2511.5010.60
12CamelliaDragons15.509.0013.199.4410.0012.00
13carrot8.3114.7510.4414.5610.3812.20
14Luna9.9414.0012.0011.7515.7512.80
15Mackerel11.3111.8812.3811.9413.1213.00
16CanisLupus13.1215.1214.5613.887.9413.70
17gotsumori-NLP16.6216.3817.5016.3817.2517.40
18noisy_watcher16.2517.6217.1917.2517.7517.60

Round 2, 9-player village

Judge models: GPT-5.6 Terra・Claude Sonnet 5・Gemini 3.1 Pro・gemma-4-31b・Qwen3.6-27b・Mistral-Small-3.2・EXAONE-4.5-33b・Seed-OSS-36b の 8 モデル平均

#TeamNatural expressionContextual dialogueLogical consistencyAction consistencyCharacter consistencyTeam playAverage
1UEC-IL2.001.502.253.385.122.251.75
2yharada3.622.382.252.501.622.942.08
3ryanleeai2.444.003.813.561.755.813.33
4wool_100%2.562.696.386.625.941.694.58
5SunameLLi8.256.385.446.565.755.195.50
6kataken14f9.255.385.255.628.197.126.67
7noisy_watcher6.126.196.387.128.065.506.75
8yatolab9.507.756.254.5611.817.448.17
9yshimoda7.3110.259.448.698.0010.258.58
10CanisLupus10.5011.7510.5610.255.8812.0010.33
11Oxymoron8.8810.2511.3110.626.8812.2510.42
12Mackerel11.5011.0011.3110.0612.129.4411.42
13carrot9.3812.1210.8812.6911.0010.8811.50
14CamelliaDragons13.6913.3813.5012.7512.8812.2513.92

Round 2, Speak-Anytime Track, 5-player village

Judge models: GPT-5.6 Terra・Claude Sonnet 5・Gemini 3.1 Pro・gemma-4-31b・Qwen3.6-27b・Mistral-Small-3.2・EXAONE-4.5-33b・Seed-OSS-36b の 8 モデル平均

#TeamNatural expressionContextual dialogueLogical consistencyAction consistencyCharacter consistencyAverage
1yharada1.251.561.192.501.001.00
2UEC-IL2.192.123.314.504.942.80
3wool_100%4.312.444.253.196.123.40
4SunameLLi5.005.005.814.563.694.40
5ryanleeai4.386.945.446.818.006.20
6kataken14f8.696.504.565.507.696.40
7Oxymoron7.007.067.948.065.317.00
8yatolab9.318.506.944.129.448.00
9CanisLupus8.7510.6211.5011.004.449.60
10noisy_watcher7.759.258.318.6211.569.60
11kanolab8.3810.129.0011.446.259.60
12CamelliaDragons11.007.889.757.699.5610.00

Game Logs

Logs of the main competition. Only games that finished normally are listed.

TrackGamesLogs
Round 1, 5-player village220Log list
Round 1, 9-player village18Log list
Round 1, Speak-Anytime Track, 5-player village14Log list
Round 2, 5-player village270Log list
Round 2, 9-player village28Log list
Round 2, Speak-Anytime Track, 5-player village24Log list

Official Results (Awards)

The list of awarded teams is being compiled.

Awards will be announced at the AIWolfDial 2026 workshop (October 17–21, 2026, INLG 2026, Utrecht).

Win Rates

Computed from the logs of all games that finished normally in the main competition. "Overall" is wins divided by games. Per-role cells show the win rate in that role with (wins / games). "Adjusted" re-weights the per-role win rates by the village's role composition, cancelling out uneven role assignments.

Round 1, 5-player village (220 games)

#TeamGamesOverallVillagerSeerPossessedWerewolfAdjusted
1yharada10061.0%70.0% (28/40)75.0% (15/20)25.0% (5/20)65.0% (13/20)61.0%
2Oxymoron10061.0%72.5% (29/40)75.0% (15/20)40.0% (8/20)45.0% (9/20)61.0%
3noisy_watcher10060.0%77.5% (31/40)70.0% (14/20)40.0% (8/20)35.0% (7/20)60.0%
4SunameLLi10057.0%67.5% (27/40)65.0% (13/20)30.0% (6/20)55.0% (11/20)57.0%
5wool_100%10057.0%62.5% (25/40)70.0% (14/20)60.0% (12/20)30.0% (6/20)57.0%
6ryanleeai10056.0%75.0% (30/40)60.0% (12/20)35.0% (7/20)35.0% (7/20)56.0%
7yshimoda10051.0%50.0% (20/40)80.0% (16/20)30.0% (6/20)45.0% (9/20)51.0%
8kataken14f10049.0%62.5% (25/40)60.0% (12/20)40.0% (8/20)20.0% (4/20)49.0%
9CamelliaDragons10045.0%57.5% (23/40)65.0% (13/20)20.0% (4/20)25.0% (5/20)45.0%
10gotsumori-NLP10043.0%62.5% (25/40)15.0% (3/20)45.0% (9/20)30.0% (6/20)43.0%
11CanisLupus10040.0%42.5% (17/40)65.0% (13/20)35.0% (7/20)15.0% (3/20)40.0%

Round 1, 9-player village (18 games)

#TeamGamesOverallVillagerSeerBodyguardMediumPossessedWerewolfAdjusted
1CamelliaDragons1866.7%83.3% (5/6)50.0% (1/2)50.0% (1/2)50.0% (1/2)50.0% (1/2)75.0% (3/4)66.7%
2CanisLupus1866.7%83.3% (5/6)50.0% (1/2)50.0% (1/2)50.0% (1/2)100.0% (2/2)50.0% (2/4)66.7%
3Oxymoron1866.7%50.0% (3/6)50.0% (1/2)100.0% (2/2)100.0% (2/2)50.0% (1/2)75.0% (3/4)66.7%
4wool_100%1855.6%66.7% (4/6)0.0% (0/2)50.0% (1/2)100.0% (2/2)100.0% (2/2)25.0% (1/4)55.6%
5yshimoda1855.6%50.0% (3/6)50.0% (1/2)100.0% (2/2)50.0% (1/2)50.0% (1/2)50.0% (2/4)55.6%
6ryanleeai1844.4%33.3% (2/6)100.0% (2/2)50.0% (1/2)50.0% (1/2)0.0% (0/2)50.0% (2/4)44.4%
7noisy_watcher1844.4%50.0% (3/6)50.0% (1/2)0.0% (0/2)100.0% (2/2)50.0% (1/2)25.0% (1/4)44.4%
8SunameLLi1833.3%33.3% (2/6)50.0% (1/2)100.0% (2/2)0.0% (0/2)0.0% (0/2)25.0% (1/4)33.3%
9yharada1833.3%50.0% (3/6)100.0% (2/2)0.0% (0/2)0.0% (0/2)0.0% (0/2)25.0% (1/4)33.3%

Round 1, Speak-Anytime Track, 5-player village (14 games)

#TeamGamesOverallVillagerSeerPossessedWerewolfAdjusted
1yharada1060.0%25.0% (1/4)100.0% (2/2)50.0% (1/2)100.0% (2/2)60.0%
2CanisLupus1060.0%75.0% (3/4)0.0% (0/2)100.0% (2/2)50.0% (1/2)60.0%
3Oxymoron1050.0%50.0% (2/4)50.0% (1/2)50.0% (1/2)50.0% (1/2)50.0%
4CamelliaDragons1050.0%50.0% (2/4)50.0% (1/2)50.0% (1/2)50.0% (1/2)50.0%
5wool_100%1050.0%25.0% (1/4)100.0% (2/2)100.0% (2/2)0.0% (0/2)50.0%
6noisy_watcher1040.0%50.0% (2/4)50.0% (1/2)0.0% (0/2)50.0% (1/2)40.0%
7SunameLLi1040.0%75.0% (3/4)0.0% (0/2)0.0% (0/2)50.0% (1/2)40.0%

Round 2, 5-player village (270 games)

#TeamGamesOverallVillagerSeerPossessedWerewolfAdjusted
1yatolab7569.3%76.7% (23/30)53.3% (8/15)73.3% (11/15)66.7% (10/15)69.3%
2AGKAI7560.0%66.7% (20/30)46.7% (7/15)66.7% (10/15)53.3% (8/15)60.0%
3wool_100%7558.7%53.3% (16/30)80.0% (12/15)46.7% (7/15)60.0% (9/15)58.7%
4CanisLupus7557.3%66.7% (20/30)66.7% (10/15)60.0% (9/15)26.7% (4/15)57.3%
5yharada7556.0%56.7% (17/30)60.0% (9/15)40.0% (6/15)66.7% (10/15)56.0%
6NTT-HAI7554.7%46.7% (14/30)60.0% (9/15)60.0% (9/15)60.0% (9/15)54.7%
7ryanleeai7554.7%40.0% (12/30)86.7% (13/15)40.0% (6/15)66.7% (10/15)54.7%
8SunameLLi7552.0%66.7% (20/30)53.3% (8/15)40.0% (6/15)33.3% (5/15)52.0%
9CamelliaDragons7550.7%60.0% (18/30)40.0% (6/15)46.7% (7/15)46.7% (7/15)50.7%
10kataken14f7550.7%53.3% (16/30)80.0% (12/15)33.3% (5/15)33.3% (5/15)50.7%
11UEC-IL7550.7%53.3% (16/30)60.0% (9/15)40.0% (6/15)46.7% (7/15)50.7%
12Luna7550.7%53.3% (16/30)53.3% (8/15)40.0% (6/15)53.3% (8/15)50.7%
13carrot7546.7%50.0% (15/30)53.3% (8/15)26.7% (4/15)53.3% (8/15)46.7%
14noisy_watcher7545.3%63.3% (19/30)46.7% (7/15)26.7% (4/15)26.7% (4/15)45.3%
15Mackerel7541.3%46.7% (14/30)53.3% (8/15)40.0% (6/15)20.0% (3/15)41.3%
16yshimoda7541.3%50.0% (15/30)40.0% (6/15)40.0% (6/15)26.7% (4/15)41.3%
17Oxymoron7541.3%50.0% (15/30)40.0% (6/15)26.7% (4/15)40.0% (6/15)41.3%
18gotsumori-NLP7538.7%46.7% (14/30)26.7% (4/15)53.3% (8/15)20.0% (3/15)38.7%

Round 2, 9-player village (28 games)

#TeamGamesOverallVillagerSeerBodyguardMediumPossessedWerewolfAdjusted
1yharada1866.7%83.3% (5/6)50.0% (1/2)100.0% (2/2)50.0% (1/2)50.0% (1/2)50.0% (2/4)66.7%
2ryanleeai1861.1%50.0% (3/6)50.0% (1/2)50.0% (1/2)100.0% (2/2)0.0% (0/2)100.0% (4/4)61.1%
3kataken14f1861.1%50.0% (3/6)100.0% (2/2)50.0% (1/2)100.0% (2/2)50.0% (1/2)50.0% (2/4)61.1%
4SunameLLi1861.1%33.3% (2/6)100.0% (2/2)100.0% (2/2)100.0% (2/2)100.0% (2/2)25.0% (1/4)61.1%
5UEC-IL1861.1%66.7% (4/6)50.0% (1/2)100.0% (2/2)50.0% (1/2)50.0% (1/2)50.0% (2/4)61.1%
6noisy_watcher1855.6%83.3% (5/6)50.0% (1/2)0.0% (0/2)50.0% (1/2)50.0% (1/2)50.0% (2/4)55.6%
7wool_100%1855.6%50.0% (3/6)50.0% (1/2)50.0% (1/2)50.0% (1/2)100.0% (2/2)50.0% (2/4)55.6%
8Oxymoron1850.0%66.7% (4/6)100.0% (2/2)0.0% (0/2)50.0% (1/2)50.0% (1/2)25.0% (1/4)50.0%
9Mackerel1850.0%16.7% (1/6)100.0% (2/2)100.0% (2/2)100.0% (2/2)50.0% (1/2)25.0% (1/4)50.0%
10CanisLupus1844.4%50.0% (3/6)0.0% (0/2)100.0% (2/2)50.0% (1/2)0.0% (0/2)50.0% (2/4)44.4%
11yatolab1844.4%66.7% (4/6)0.0% (0/2)50.0% (1/2)50.0% (1/2)0.0% (0/2)50.0% (2/4)44.4%
12yshimoda1844.4%83.3% (5/6)0.0% (0/2)50.0% (1/2)0.0% (0/2)50.0% (1/2)25.0% (1/4)44.4%
13CamelliaDragons1844.4%66.7% (4/6)100.0% (2/2)0.0% (0/2)0.0% (0/2)0.0% (0/2)50.0% (2/4)44.4%
14carrot1833.3%33.3% (2/6)50.0% (1/2)50.0% (1/2)50.0% (1/2)50.0% (1/2)0.0% (0/4)33.3%

Round 2, Speak-Anytime Track, 5-player village (24 games)

#TeamGamesOverallVillagerSeerPossessedWerewolfAdjusted
1yatolab1070.0%75.0% (3/4)100.0% (2/2)100.0% (2/2)0.0% (0/2)70.0%
2ryanleeai1070.0%100.0% (4/4)0.0% (0/2)100.0% (2/2)50.0% (1/2)70.0%
3wool_100%1060.0%75.0% (3/4)100.0% (2/2)50.0% (1/2)0.0% (0/2)60.0%
4SunameLLi1060.0%75.0% (3/4)50.0% (1/2)50.0% (1/2)50.0% (1/2)60.0%
5noisy_watcher1060.0%75.0% (3/4)50.0% (1/2)50.0% (1/2)50.0% (1/2)60.0%
6Oxymoron1050.0%50.0% (2/4)100.0% (2/2)0.0% (0/2)50.0% (1/2)50.0%
7yharada1050.0%50.0% (2/4)50.0% (1/2)0.0% (0/2)100.0% (2/2)50.0%
8kanolab1050.0%50.0% (2/4)50.0% (1/2)50.0% (1/2)50.0% (1/2)50.0%
9kataken14f1040.0%50.0% (2/4)0.0% (0/2)0.0% (0/2)100.0% (2/2)40.0%
10CanisLupus1040.0%25.0% (1/4)100.0% (2/2)50.0% (1/2)0.0% (0/2)40.0%
11UEC-IL1040.0%50.0% (2/4)50.0% (1/2)50.0% (1/2)0.0% (0/2)40.0%
12CamelliaDragons1030.0%25.0% (1/4)50.0% (1/2)0.0% (0/2)50.0% (1/2)30.0%

Game Metrics

Behavioural metrics computed mechanically from the game logs, independent of dialogue quality. Each value is a ratio against what random play would produce (1.00 = random). "Executed", "divined" and "guarded" show how often the team was targeted relative to chance; lower is better for being executed or divined. "Divination accuracy" and "vote accuracy" show how often the team found a werewolf relative to chance; higher is better. "Possessed avoidance" is the rate at which, as the possessed, the team voted for someone other than a werewolf.

Round 1, 5-player village

TeamExecutedDivinedDivination accuracyVote accuracyPossessed avoidance
yharada0.320.861.431.470.40
Oxymoron0.860.891.231.580.69
noisy_watcher0.791.010.841.580.75
SunameLLi0.930.841.001.390.70
wool_100%1.071.221.221.340.67
ryanleeai0.930.871.041.390.48
yshimoda1.000.881.301.480.67
CamelliaDragons0.941.331.641.640.80
gotsumori-NLP1.511.650.970.990.73
CanisLupus1.460.620.681.090.70
kataken1.260.851.711.640.68

Round 1, 9-player village

TeamExecutedDivinedGuardedDivination accuracyVote accuracyPossessed avoidance
CamelliaDragons0.331.861.341.081.200.67
CanisLupus1.931.132.591.521.840.75
Oxymoron1.040.420.340.681.670.75
wool_100%1.281.031.531.271.771.00
yshimoda1.080.530.57-1.700.25
ryanleeai0.740.921.172.151.150.50
noisy_watcher1.321.591.052.271.620.60
SunameLLi1.010.931.171.511.180.60
yharada0.450.83-2.551.850.67

Round 1, Speak-Anytime Track, 5-player village

TeamExecutedDivinedDivination accuracyVote accuracyPossessed avoidance
yharada0.240.83-1.760.75
CanisLupus2.40--2.001.00
Oxymoron0.921.092.000.861.00
CamelliaDragons0.511.001.710.671.00
wool_100%1.371.821.201.001.00
noisy_watcher0.781.121.711.80-
SunameLLi1.330.96-1.00-

Round 2, 5-player village

TeamExecutedDivinedDivination accuracyVote accuracyPossessed avoidance
yatolab0.851.001.571.690.77
AGKAI0.731.080.941.340.82
wool_100%0.661.211.391.330.80
CanisLupus1.340.981.401.230.89
yharada0.550.701.220.970.59
NTT-HAI0.810.760.561.140.64
ryanleeai0.870.841.910.960.71
SunameLLi0.891.190.781.870.58
CamelliaDragons1.111.071.541.450.88
UEC-IL0.870.690.781.010.95
Luna0.831.000.820.980.72
carrot0.930.581.151.060.47
noisy_watcher1.460.840.401.280.38
Mackerel1.210.770.701.200.80
yshimoda1.261.160.980.770.74
Oxymoron0.911.151.040.900.65
gotsumori-NLP2.012.040.900.971.00
kataken0.971.001.331.010.63

Round 2, 9-player village

TeamExecutedDivinedGuardedDivination accuracyVote accuracyPossessed avoidance
yharada0.320.740.151.451.920.50
ryanleeai0.620.851.772.601.670.50
SunameLLi0.550.690.572.741.830.80
UEC-IL0.700.720.711.092.310.83
noisy_watcher0.940.961.911.051.340.75
wool_100%0.761.031.471.101.461.00
Oxymoron1.721.461.79-1.130.67
Mackerel1.550.900.491.081.861.00
CanisLupus1.290.45--1.44-
yatolab0.610.850.591.051.830.67
yshimoda2.050.621.092.662.010.33
CamelliaDragons2.382.120.61-1.760.60
carrot1.301.340.71-1.150.50
kataken0.711.891.751.871.620.67

Round 2, Speak-Anytime Track, 5-player village

TeamExecutedDivinedDivination accuracyVote accuracyPossessed avoidance
yatolab1.331.331.202.001.00
ryanleeai1.871.29-2.001.00
wool_100%1.200.752.002.500.33
SunameLLi0.840.80-1.201.00
noisy_watcher0.821.290.861.671.00
Oxymoron0.861.124.002.670.50
yharada0.310.922.400.860.50
kanolab1.200.75-4.00-
CanisLupus1.201.452.000.671.00
UEC-IL1.670.502.401.331.00
CamelliaDragons0.690.921.200.80-
kataken0.280.97-0.351.00

Human Evaluation (Relative)

Human judges read selected game logs and evaluate the players of each game on each criterion. From 2024 winter onward the judges rank the players (average rank, lower is better); in the 2024 spring and 2024 international contests they gave 5-point scores (higher is better). In 13-player villages each entry (e.g. team-A, team-B) is evaluated separately.

Human evaluation was not conducted in this contest.

LLM-as-a-Judge (Relative Evaluation)

LLM judges rank the players of each game on each evaluation criterion; the ranks are averaged per team (lower is better). The judge models differ by contest and are listed under each track. The judge is published as aiwolf-nlp-llm-judge.

Round 1, 5-player village

Judge models: GPT-5.4・Gemini 2.5 Pro・Claude Sonnet 4.5 の 3 モデル平均

#TeamNatural expressionContextual dialogueLogical consistencyAction consistencyCharacter consistencyAverage
1ryanleeai1.181.452.042.201.471.67
2Oxymoron1.552.002.182.221.501.89
3yharada2.862.552.231.982.282.38
4SunameLLi2.662.872.572.722.892.74
5noisy_watcher2.912.932.722.632.942.83
6kataken14f2.973.052.632.523.112.86
7wool_100%3.113.362.752.993.183.08
8CamelliaDragons4.052.562.652.813.613.14
9yshimoda3.423.782.973.023.593.35
10CanisLupus3.984.323.603.513.723.83
11gotsumori-NLP3.813.904.243.444.203.92

Round 1, 9-player village

Judge models: GPT-5.4・Gemini 2.5 Pro・Claude Sonnet 4.5 の 3 モデル平均

#TeamNatural expressionContextual dialogueLogical consistencyAction consistencyCharacter consistencyTeam playAverage
1ryanleeai1.332.542.483.482.443.562.64
2Oxymoron2.613.703.304.302.224.613.46
3yharada5.304.282.852.634.463.983.92
4CamelliaDragons6.302.613.313.095.243.073.94
5SunameLLi4.524.393.984.374.524.834.43
6noisy_watcher4.416.066.045.395.725.435.51
7yshimoda5.466.415.065.616.065.325.65
8wool_100%6.936.896.265.077.155.506.30
9CanisLupus7.047.746.615.725.917.206.70

Round 1, Speak-Anytime Track, 5-player village

Judge models: GPT-5.4・Gemini 2.5 Pro・Claude Sonnet 4.5 の 3 モデル平均

#TeamNatural expressionContextual dialogueLogical consistencyAction consistencyCharacter consistencyAverage
1Oxymoron1.402.302.032.231.531.90
2SunameLLi2.102.272.132.272.832.32
3yharada2.973.072.431.873.102.69
4noisy_watcher2.672.533.503.432.672.96
5wool_100%3.303.403.003.403.133.25
6CanisLupus3.674.203.173.373.173.51
7CamelliaDragons4.733.203.873.074.533.88

Round 2, 5-player village

Judge models: GPT-5.6 Terra・Claude Sonnet 5・Gemini 3.1 Pro・gemma-4-31b・Qwen3.6-27b・Mistral-Small-3.2・EXAONE-4.5-33b・Seed-OSS-36b の 8 モデル平均

#TeamNatural expressionContextual dialogueLogical consistencyAction consistencyCharacter consistencyAverage
1yharada2.691.381.501.561.251.40
2UEC-IL1.501.752.694.193.502.20
3ryanleeai2.382.883.446.002.383.00
4wool_100%5.195.383.812.506.624.00
5SunameLLi8.126.006.386.196.695.90
6AGKAI9.128.1910.006.757.758.10
7Oxymoron7.319.129.3112.126.388.20
8yatolab13.758.198.005.4412.949.10
9yshimoda4.629.2510.0615.627.259.70
10NTT-HAI12.1211.258.506.1912.5610.10
11kataken14f13.128.8810.069.2511.5010.60
12CamelliaDragons15.509.0013.199.4410.0012.00
13carrot8.3114.7510.4414.5610.3812.20
14Luna9.9414.0012.0011.7515.7512.80
15Mackerel11.3111.8812.3811.9413.1213.00
16CanisLupus13.1215.1214.5613.887.9413.70
17gotsumori-NLP16.6216.3817.5016.3817.2517.40
18noisy_watcher16.2517.6217.1917.2517.7517.60

Round 2, 9-player village

Judge models: GPT-5.6 Terra・Claude Sonnet 5・Gemini 3.1 Pro・gemma-4-31b・Qwen3.6-27b・Mistral-Small-3.2・EXAONE-4.5-33b・Seed-OSS-36b の 8 モデル平均

#TeamNatural expressionContextual dialogueLogical consistencyAction consistencyCharacter consistencyTeam playAverage
1UEC-IL2.001.502.253.385.122.251.75
2yharada3.622.382.252.501.622.942.08
3ryanleeai2.444.003.813.561.755.813.33
4wool_100%2.562.696.386.625.941.694.58
5SunameLLi8.256.385.446.565.755.195.50
6kataken14f9.255.385.255.628.197.126.67
7noisy_watcher6.126.196.387.128.065.506.75
8yatolab9.507.756.254.5611.817.448.17
9yshimoda7.3110.259.448.698.0010.258.58
10CanisLupus10.5011.7510.5610.255.8812.0010.33
11Oxymoron8.8810.2511.3110.626.8812.2510.42
12Mackerel11.5011.0011.3110.0612.129.4411.42
13carrot9.3812.1210.8812.6911.0010.8811.50
14CamelliaDragons13.6913.3813.5012.7512.8812.2513.92

Round 2, Speak-Anytime Track, 5-player village

Judge models: GPT-5.6 Terra・Claude Sonnet 5・Gemini 3.1 Pro・gemma-4-31b・Qwen3.6-27b・Mistral-Small-3.2・EXAONE-4.5-33b・Seed-OSS-36b の 8 モデル平均

#TeamNatural expressionContextual dialogueLogical consistencyAction consistencyCharacter consistencyAverage
1yharada1.251.561.192.501.001.00
2UEC-IL2.192.123.314.504.942.80
3wool_100%4.312.444.253.196.123.40
4SunameLLi5.005.005.814.563.694.40
5ryanleeai4.386.945.446.818.006.20
6kataken14f8.696.504.565.507.696.40
7Oxymoron7.007.067.948.065.317.00
8yatolab9.318.506.944.129.448.00
9CanisLupus8.7510.6211.5011.004.449.60
10noisy_watcher7.759.258.318.6211.569.60
11kanolab8.3810.129.0011.446.259.60
12CamelliaDragons11.007.889.757.699.5610.00

Game Logs

Logs of the main competition. Only games that finished normally are listed.

TrackGamesLogs
Round 1, 5-player village220Log list
Round 1, 9-player village18Log list
Round 1, Speak-Anytime Track, 5-player village14Log list
Round 2, 5-player village270Log list
Round 2, 9-player village28Log list
Round 2, Speak-Anytime Track, 5-player village24Log list

How to view logs

Each “Log list” link opens a listing of the logs (.log files) of games that finished normally. A file name starts with the game start time, followed by the names of the participating teams.

  • View in the browser: press “▶ 開く” (open) in the “viewerで見る” (view in viewer) column of the row you want to read. aiwolf-nlp-viewer opens with that log loaded.
  • Save a file: press “⬇ 保存” (save) in the “ダウンロード” (download) column to download the log file. Clicking the file name shows the raw text of the log instead.
  • Save a whole folder: go one level up via “Parent directory/” and press “⬇ zip” in the row of the folder (e.g. success). All files in the folder are zipped in your browser and downloaded (a dialog confirms the number of files).

How to use the viewer is described in the “Log viewer” section of the Links page.