How the model is doing
Across 104 matches, the model's most-likely outcome happened 65% of the time — a Brier score of 0.493 vs 0.642 for the base rates alone.
A prediction tool is only as good as its record. Every figure here scores the model's own pre-match calls against what actually happened — no hindsight.
Calibration
When the model says it's X% sure, does it happen X% of the time? Predicted vs actual, by confidence bucket.
The pre-tournament call
The model had Spain — the eventual champion — 1st in its pre-tournament title odds, at 29%.
25 of the 32 teams that reached the Round of 32 were in the model's pre-tournament top 32.
Brier is the mean squared error of the pre-match win/draw/loss probabilities vs the result (a knockout level after 90 counts as a draw). Calibration buckets the model's confidence in its top pick. The pre-tournament call replays the model with no results.