Frozen call Away win · 1-2 · probabilities 31/30/40% · Brier 0.748 · locked 10h 13m before kickoff
How good is it,
really?
The live record shows forecasts actually published before kickoff. Earlier historical hybrid figures were withdrawn after an input error. Their limitation remains visible while the live record grows.
Live 2026/27 record
Locked before kickoff. Never rewritten.
This is the public record of forecasts the site actually showed before Premier League matches—not a replay produced after the result. Historical backtesting starts below.
Log updated
What the record shows
49 of 50 finished fixtures have a settled, pre-kickoff forecast (98% coverage). 49 of 50 settled forecasts and 5 of 5 matchday blocks are recorded. Aggregate scores remain withheld until both minimums are met.
Forecasts frozen
49
Nothing waiting to settle
Finished coverage
98%
49/50 finished fixtures covered
Live outcome accuracy
Withheld
49/50 settled forecasts required
Live Brier / log loss
Withheld
Average lock 8h 16m before kickoff
- Ipswich vs Liverpool · — no frozen forecast
Settled forecasts
Frozen call Home win · 2-1 · probabilities 39/30/31% · Brier 0.720 · locked 7h 43m before kickoff
Frozen call Home win · 2-1 · probabilities 49/29/21% · Brier 0.787 · locked 7h 43m before kickoff
Frozen call Home win · 1-0 · probabilities 54/26/20% · Brier 0.316 · locked 7h 43m before kickoff
Frozen call Home win · 1-0 · probabilities 47/29/24% · Brier 0.887 · locked 11h 34m before kickoff
Frozen call Home win · 2-1 · probabilities 46/28/26% · Brier 0.443 · locked 9h 4m before kickoff
Frozen call Away win · 1-2 · probabilities 27/27/45% · Brier 0.812 · locked 9h 4m before kickoff
Frozen call Home win · 1-0 · probabilities 42/33/25% · Brier 0.504 · locked 9h 4m before kickoff
Frozen call Home win · 1-0 · probabilities 38/29/33% · Brier 0.687 · locked 6h 34m before kickoff
Frozen call Home win · 2-1 · probabilities 47/28/26% · Brier 0.423 · locked 5h 54m before kickoff
Frozen call Home win · 1-0 · probabilities 45/28/27% · Brier 0.459 · locked 3h 48m before kickoff
Frozen call Home win · 2-1 · probabilities 37/28/35% · Brier 0.630 · locked 10h 13m before kickoff
Frozen call Away win · 0-1 · probabilities 31/29/40% · Brier 0.542 · locked 7h 43m before kickoff
Frozen call Away win · 0-1 · probabilities 24/29/48% · Brier 0.413 · locked 6h 40m before kickoff
Frozen call Away win · 0-1 · probabilities 33/28/39% · Brier 0.784 · locked 11h 35m before kickoff
Frozen call Home win · 2-1 · probabilities 48/28/24% · Brier 0.874 · locked 9h 5m before kickoff
Frozen call Home win · 2-1 · probabilities 53/27/20% · Brier 0.856 · locked 9h 5m before kickoff
Frozen call Away win · 0-1 · probabilities 34/28/38% · Brier 0.582 · locked 9h 5m before kickoff
Frozen call Home win · 2-1 · probabilities 38/28/33% · Brier 0.770 · locked 9h 5m before kickoff
Frozen call Home win · 2-1 · probabilities 50/29/22% · Brier 0.804 · locked 9h 5m before kickoff
Frozen call Home win · 2-1 · probabilities 51/27/22% · Brier 0.363 · locked 10h 30m before kickoff
Frozen call Home win · 2-1 · probabilities 41/30/29% · Brier 0.740 · locked 8h before kickoff
Frozen call Draw · 1-1 · probabilities 33/34/33% · Brier 0.651 · locked 11h 37m before kickoff
Frozen call Home win · 1-0 · probabilities 44/30/26% · Brier 0.749 · locked 9h 7m before kickoff
Frozen call Home win · 2-1 · probabilities 52/27/22% · Brier 0.353 · locked 9h 7m before kickoff
Frozen call Home win · 2-1 · probabilities 47/28/25% · Brier 0.807 · locked 9h 7m before kickoff
Frozen call Home win · 1-0 · probabilities 51/27/22% · Brier 0.842 · locked 9h 7m before kickoff
Frozen call Home win · 1-0 · probabilities 45/30/25% · Brier 0.859 · locked 9h 7m before kickoff
Frozen call Home win · 2-1 · probabilities 47/28/25% · Brier 0.808 · locked 6h 37m before kickoff
Frozen call Away win · 1-2 · probabilities 33/26/41% · Brier 0.524 · locked 2h 41m before kickoff
Frozen call Home win · 1-0 · probabilities 49/27/24% · Brier 0.389 · locked 9h 47m before kickoff
Frozen call Home win · 1-0 · probabilities 44/30/26% · Brier 0.469 · locked 7h 17m before kickoff
Frozen call Home win · 2-1 · probabilities 52/26/23% · Brier 0.348 · locked 7h 17m before kickoff
Frozen call Home win · 2-1 · probabilities 41/29/31% · Brier 0.762 · locked 7h 17m before kickoff
Frozen call Home win · 2-1 · probabilities 42/28/30% · Brier 0.754 · locked 9h 11m before kickoff
Frozen call Home win · 1-0 · probabilities 48/26/26% · Brier 0.839 · locked 6h 41m before kickoff
Frozen call Home win · 1-0 · probabilities 44/28/28% · Brier 0.785 · locked 6h 41m before kickoff
Frozen call Home win · 1-0 · probabilities 53/27/21% · Brier 0.855 · locked 10h 51m before kickoff
Frozen call Away win · 0-1 · probabilities 25/28/47% · Brier 0.416 · locked 6h 56m before kickoff
Frozen call Home win · 1-0 · probabilities 43/29/28% · Brier 0.779 · locked 1h 52m before kickoff
Frozen call Home win · 1-0 · probabilities 41/33/26% · Brier 0.677 · locked 6h 31m before kickoff
Frozen call Home win · 1-0 · probabilities 43/33/23% · Brier 0.486 · locked 10h 49m before kickoff
Frozen call Home win · 1-0 · probabilities 41/33/26% · Brier 0.531 · locked 10h 49m before kickoff
Frozen call Home win · 1-0 · probabilities 41/35/24% · Brier 0.528 · locked 10h 15m before kickoff
Frozen call Home win · 2-1 · probabilities 39/31/31% · Brier 0.564 · locked 7h 45m before kickoff
Frozen call Home win · 1-0 · probabilities 42/34/24% · Brier 0.862 · locked 7h 45m before kickoff
Frozen call Home win · 1-0 · probabilities 41/35/24% · Brier 0.532 · locked 7h 45m before kickoff
Frozen call Away win · 1-2 · probabilities 34/32/35% · Brier 0.663 · locked 5h 15m before kickoff
Frozen call Home win · 2-1 · probabilities 59/26/15% · Brier 0.259 · locked 9h 51m before kickoff
Probabilities are ordered home win / draw / away win. “Called correctly” means the highest probability was the result; the predicted score is checked separately. Coverage is settled forecasts divided by all finished fixtures in the provider feed. A missing freeze lowers coverage rather than disappearing from the denominator. Brier is the three-outcome probability error and log loss scores the probability assigned to the realised outcome; lower is better. Aggregate score intervals resample whole matchday blocks so shared matchday conditions are not treated as independent matches.
Historical hybrid record withdrawn
The prior 1,140-fixture GBT/hybrid replay used season-end xG aggregates in earlier fixture rows. Those figures are no longer presented as evidence.
Where it breaks down — 2025/26
Historical Snapshot 5 (2026-05-27), measured on the pre-GBT hybrid. Retained as a dated failure-pattern diagnostic, not a claim about the current hybrid.
Q1 · MD 1–9
53.3%
Q2 · MD 10–19
55.0%
Q3 · MD 20–28
40.0%
Q4 · MD 29–38
42.0%
This dated pre-GBT replay was weaker in the second half of the season. Later corrected hybrid replays show a smaller period-dependent gap, but they have not established its cause. Changes in form or squads are possible explanations to test, not findings.
A candidate we rejected
The weekly optimizer produced a candidate coefficient set fitted on all 380 fixtures of 2025/26. It was worse than the live model on both primary metrics, so it was not shipped.
| Metric | Then-live model | Candidate | Verdict |
|---|---|---|---|
| Outcome Brier | 0.624765 | 0.626315 | Worse |
| Outcome log loss | 1.038122 | 1.039960 | Worse |
| Draw calibration gap | 0.003867 | 0.016039 | Worse |
Decision on 2026-05-27: not shipped. A newer fit is not automatically a better one, and the gate exists to catch exactly this.
How the model works
1. Ratings, not vibes
Each club carries an Elo rating plus separate attack and defence strengths, rebuilt from every completed fixture and weighted towards recent form. Home advantage is fitted, not assumed.
2. Two models, blended
A Dixon-Coles Poisson model turns those ratings into goal expectations and a full scoreline distribution. A gradient-boosted classifier predicts the outcome directly from the same features. The live forecast is a blend of the two — each covers the other’s weak spots.
3. Context, carefully
Season phase, table pressure, fixture congestion and squad availability adjust the baseline. These features are deliberately weighted low: they help in-season and hurt on seasons they were not calibrated against, so they are never allowed to dominate the ratings.
4. Simulated seasons
Title, top-four and relegation probabilities come from a validated Monte Carlo sweep of every remaining fixture. The sweep runs offline on a schedule; pages only ever read the resulting snapshot and its recorded iteration count.
5. Replayed and gated
Every candidate model is replayed out-of-sample against completed seasons. It ships only if it improves Brier score and log loss without a season quarter regressing. Candidates that fail are documented and discarded.
What we are not claiming
The failure modes we know about, published alongside the numbers rather than after someone finds them.
- The Snapshot 6 GBT/hybrid historical figures are withdrawn because season-end xG leaked into earlier pre-match rows. A corrected chronological calibration audit is committed, but the public hybrid replay will remain unavailable until its live-hybrid and quarter-stability rerun is complete.
- The frozen 2026/27 Premier League log is the only current public result record. It records forecasts before kickoff, never back-fills a missed fixture, and publishes coverage alongside accuracy.
- The corrected GBT calibration candidate improved Brier and log loss on its untouched 2025/26 holdout, but it is not promoted: live-hybrid comparison and quarter-stability gates remain outstanding.
- On historical seasons the plain baseline beats the fully-enriched context model. The context layer (pressure, congestion, manager, form weights) adds net noise against seasons it was not calibrated on; the hybrid blend recovers most of that, but the context layer needs regularisation.
- A run-in fatigue feature was tested and rejected — it made the model worse in both 2023/24 and 2024/25. It is not enabled.
- Squad-sentiment weighting cannot be validated before 2026-04-28, when the sentiment archive starts. It runs at deliberately conservative weights pending a full-season ablation.
- Exact scorelines are much harder than outcomes and are evaluated separately — treat any single projected score as the middle of a wide distribution, not a call.
This is a forecasting tool, not a tipping service. Every number here is a model output — a statement about how often something happens in simulation, not advice about what to do with it.
Want the current inputs and diagnostics? Explore the model page.