Skip to content
Zackcast

Zackcast Insights

What advanced stats actually measure

EPA, xG, WAR and the rest are not magic and they are not nonsense. Each one answers a narrow question — and most arguments about them are arguments about the question.

Zackcast Editorial Desk 3 min read

They all do the same thing

Advanced statistics look unrelated until you notice they share a structure. Each one takes an event, asks what usually happens in that situation, and records the difference between what usually happens and what happened this time.

That is the whole idea. A shot from eight yards out normally becomes a goal a quarter of the time; scoring it is credit for roughly three quarters of a goal. Second-and-7 at midfield is normally worth a certain number of future points; the play that follows either adds to that or subtracts from it. Everything else is bookkeeping over the same move.

Once you see the pattern, the acronym stops mattering and only two questions remain: what does this model think is normal, and where did it learn that?

Expected goals (xG), and its cousins

Expected goals scores each shot by the chance an average player would convert it, using distance, angle, body part, the kind of pass that created it, and the pressure on the shooter. A match that finishes 1-0 but 2.6 to 0.4 on xG tells you the loser was badly outplayed and got away with it for eighty minutes.

Used over a season, xG is one of the better predictors in sport, because scoring is rare and final scores are noisy. Used on a single match, it is a description of chance quality, not a verdict — and it is blind to the thing you watched, which is whether the shooter was a striker who genuinely finishes better than average.

Expected points added (EPA), and football's version of the same trick

Every down-distance-field-position combination has an average number of points the offense eventually scores from there. EPA credits each play with the change in that number. A four-yard gain on third-and-3 is a good play; the same four yards on third-and-8 is a bad one. Yardage cannot tell those apart and EPA can, which is why it has taken over football analysis.

Its weakness is that it hands the whole outcome to whoever touched the ball. Every EPA number for a quarterback contains his line, his receivers' hands, the coordinator's call and the defense's coverage. Treat it as a measure of what the offense produced, attributed to a player for convenience.

WAR, PER and the all-in-one numbers

Wins Above Replacement in baseball, and its equivalents in basketball and hockey, go one step further: they convert everything a player did into a single currency, wins, against a defined baseline — the freely available player a team could sign tomorrow.

The appeal is obvious and so is the risk. A one-number summary hides every judgement that went into it, and different publishers' versions of WAR disagree, sometimes by a win or two, because they value defense differently. Use them to sort players into broad tiers. Do not use them to argue about a gap of half a win, which is inside the error bars of the method itself.

Where these models break

Every one of them learns “normal” from a pile of past games, and inherits whatever was true of that pile:

  • Small samples. Almost all of these numbers need a season or more before they mean much. A month of xG overperformance is not finishing skill.
  • Unusual players. A model built on the league is worst at describing the players who are least like the league — which is often exactly who you are arguing about.
  • Missing inputs. Most public models cannot see the defense's alignment, the goalkeeper's position, or whether the shooter was falling over.
  • Rule and style changes. A model trained on how a sport was played five years ago quietly mis-scores how it is played now.
  • Attribution. The model measures what happened on the play; assigning that to one of the twenty-two people on the field is a separate assumption, usually an arguable one.

How to use them without being used by them

The productive stance is neither reverence nor dismissal. These numbers are compressed descriptions of thousands of events you did not watch, and that is worth a great deal — as long as you keep asking what was compressed away.

In practice: prefer rate stats to totals, prefer a season to a night, look up whose version of the metric you are reading, and treat any single number that claims to rank human beings as a starting point for a conversation rather than the end of one.

The best analysts use these to find the questions worth watching for, then go back to the video to answer them. That order — number first, eyes second, conclusion last — is most of the skill.