Skip to main content
Sensible Stats / Recent research

First establish that the tackle actually happened

AI can help turn football video into event data. A 2026 SoccerNet report shows why the gap between counting actions and understanding a player still matters.

By Sensible Stats · Published

Revised

Imagine watching a midfielder stretch for a loose ball. The camera cuts, two shirts overlap and possession changes. Was that a tackle, an interception, a block or a pass that went astray? Now imagine building a season-long account of the player’s defending from thousands of such moments. A small labelling mistake has acquired a very large spreadsheet.

Automated football analysis depends on getting those moments right, including who performed them. A 2026 SoccerNet challenge report by Parthsarthi Rawat offers an instructive case: its event-detection system improves on its test benchmark, then loses ground on a separate challenge set. The interesting story is what happens between those two results, and what it means for anyone tempted to turn a tidy event count into a tactical verdict.

References:[1] arXiv · SoccerNet challenge report

What is player-centric football action spotting?

The FOOTPASS project supplies a benchmark of 54 complete match broadcasts from European competitions in 2023–24, with 102,992 annotated events across eight action categories. Its records connect an action to a frame, team and shirt number; supporting information includes tracking and player identity. The project documentation describes the dataset and its baselines. It is historical research material, not evidence of current Premier League coverage.

Why should a football reader care? Consider an illustrative scouting question: does a midfielder win the ball by stepping out to challenge, or mainly collect it after a teammate’s pressure? A count of possession changes cannot answer that alone. You need the right action attributed to the right player, then enough context to interpret it. Otherwise, a defensive partnership can become an individual hero story through an administrative misunderstanding.

Even perfectly labelled ball actions would leave some work undone. In that example, the teammate’s off-ball pressure might be the decisive contribution. A reliable event log is a foundation for analysis, not a complete description of the match. The temptation is to jump straight from a machine-readable row to a judgement about a player. The missing middle is the football.

References:[2] Jérémie Ochin and collaborators

How reliable are the SoccerNet 2026 results?

Rawat’s June 2026 preprint builds on FOOTPASS baselines, combining visual evidence, player relationships, class weighting and event clean-up. Its reported macro F1 improves from a 0.493 test baseline to 0.548, but the final system reaches 0.446 on the separate challenge set. Macro F1 averages the per-class scores equally, so frequent passes cannot simply drown out rare tackles.

F1 balances precision and recall. Precision asks how many predicted events were correct; recall asks how many real events were found. As an illustrative calculation, a detector that finds eight genuine events, invents two and misses two has 80% precision and 80% recall, giving F1 of 0.8. That score describes the specified detection task. It is not a probability of predicting a match correctly.

Equal weighting makes the choice of metric consequential. If the intended use is a defender report, rare defensive actions matter even when a system handles routine passes well. Before trusting a headline score, we would ask which classes contribute to it and whether the difficult classes are precisely the ones needed for the football question.

References:[1] arXiv · SoccerNet challenge report

The tackle that appeared rather too often

The report’s challenge-set tackle results are unusually revealing: two true positives, 46 false positives and 24 ground-truth tackles. From those published counts, precision is 2/48, about 4.2%, and recall is 2/24, about 8.3%. Their harmonic mean is approximately 0.056, matching the reported tackle F1. This calculation concerns that class and split, not the system’s overall accuracy.

The author identifies poorly transferring, test-tuned thresholds as a major failure mode. That should matter to a reader because the tuning evidence and the final examination are doing different jobs. A threshold can look sensible where it was chosen and behave badly elsewhere. An improved test result is therefore a reason to investigate transfer, not permission to assume it.

In our opening hypothetical, such an error could produce an entirely plausible but false account of industrious defending. The prose would be confident, the chart would be immaculate, and the tackle might never have happened. When a metric feeds a player judgement, checking a sample of the underlying clips is part of the analysis, not an optional courtesy to the video department.

References:[1] arXiv · SoccerNet challenge report

New research, different scoreboards

There is further work to watch. Ruifeng Wang, Di Yang and Jiangtao Wang’s August 2026 preprint preserves player identity through sequence modelling and reports micro F1 of 0.778 on FOOTPASS validation data. It is a newer research direction, but its micro average and validation setting do not provide a like-for-like comparison with Rawat’s macro F1 on the challenge set. Both remain preprints.

Micro averaging pools event counts; macro averaging gives each class equal weight. That distinction can favour different strengths. We would need the same data split, matching rules, class breakdown and evaluation procedure before deciding which system better serves a specific scouting task. A larger number printed beside a newer date is not a controlled comparison.

The more useful question is whether preserving player identity helps on unseen footage, especially when players overlap or move out of view. That is a proposed practical test, rather than an outcome established by the comparison above. Progress can be real while the route from benchmark to reliable match analysis remains unfinished.

References:[3] Ruifeng Wang[1] arXiv · SoccerNet challenge report

How we would use it without fooling ourselves

Start with the intended decision. Searching video for possible tackles can tolerate a different error profile from publishing a definitive player ranking. For clip retrieval, an analyst may check the candidates. For a numerical claim, missed events and mistaken identities can change the conclusion. The same detector can be useful for one job and insufficient for another.

Our proposed audit would reserve unseen matches, check action classes separately and review clips where the system is uncertain. It would also track the consequences of mistakes. In a labelled hypothetical, two players could each make 15 genuine tackles, while a detector finds 12 for one and six for the other because their footage is harder to interpret. The resulting ranking measures detection conditions as well as defending.

For Sensible Stats, automated events would need a legitimate current source, stable definitions, visible coverage gaps and independent evidence that any derived input improves forecasts. Neither reviewed preprint supplies that live feed. We can study the methods without quietly treating their benchmarks as operational data.

The next time a tactical chart seems to explain everything, ask to see the events behind it. A convincing account of pressing needs more than columns that add up. First establish who touched the ball. Then we can have a properly informed argument about whether they ought to have done so.

References:[2] Jérémie Ochin and collaborators[1] arXiv · SoccerNet challenge report[3] Ruifeng Wang

Sources and research status

Tags:SoccerNetFootball event dataComputer visionPlayer analysis

From research to forecasts

These findings are not automatic inputs to our model. Explore the accuracy record and methodology or the Premier League forecasts.

Found a mistake? Send a correction.

Established research ·

TacticAI: Can AI improve football corner kicks?

The corner committee has acquired a computer

Liverpool’s TacticAI collaboration offers coaches new ways to arrange a corner. What would turn an appealing suggestion into evidence of better football?

Source: Nature Communications · Peer reviewed

Read the article →