Spearman rank correlation between the human evaluation (overall) and the LLM-as-a-Judge relative evaluation (overall), computed on the teams of each track from the tables published on this site. Positive values mean the LLM ranks teams in the same order as the human judges; 1.00 is perfect agreement. The number of teams per track is small, so individual values fluctuate; the consistent positive sign across contests is the point.

ContestTrackTeamsSpearman ρ
International Contest 2025 (INLG 2025)5-player village80.63
International Contest 2025 (INLG 2025)13-player village130.95

Contests without an LLM relative evaluation (2024) are not listed. Values are recomputed from the published tables whenever the site is built.