edge lab · factor writeup · 7 min read

Confidence Scoring: Two Attempts, Two Failures

We tried twice to build a score telling you which favourites are safest. Version one came back inverted, version two came back non-monotonic. Neither shipped.

This is the feature people ask for most: a single number saying how safe a favourite really is, beyond the odds. We built it twice. It failed twice, and the second failure was more interesting than the first.

What we tested

Version one: when our own ELO power ratings agreed with the betting market, the pick should be safer than when they disagreed. Version two: when bookmakers agree tightly with each other — low dispersion across many books — the pick should be safer than when they are scattered.

How we tested it

Both versions were scored the same way. Games were bucketed by confidence, and each bucket's real favourite win rate was compared with its de-vigged market probability. For a confidence score to be worth shipping, the buckets have to line up in order: higher confidence, higher outperformance. Every game was priced from real bookmaker moneylines pulled from the historical odds archive, then de-vigged so the two sides add to 100%. That de-vigged number is the market's honest win probability. A factor only counts as an edge if games matching it beat that probability by more than chance would explain.

  • Version one used ELO ratings built game-by-game from 1999 onward, with home-field advantage and inter-season regression, checked as zero-sum across every game.
  • Version two used real per-book moneylines, measuring the standard deviation of implied probability across books and the number of books quoting each game.
  • Both were run across the full 2021–2025 odds backfill.

What the numbers said

Window
2021–2025
v1 result
inverted
v2 result
non-monotonic

Version one did not merely fail — it ran backwards. Games our model flagged as low confidence, meaning ELO disagreed with the market, saw the favourite win 76.9% of the time. That is the market being right and our rating being blind: ELO knows results, it does not know that a starting quarterback was ruled out on Friday. Disagreement was not a warning sign, it was ELO missing the news.

Version two failed more quietly. Bookmaker dispersion did not order itself: mid-confidence buckets outperformed high-confidence ones, then flipped again. A score that does not increase monotonically is not a score, it is noise with a label.

Why we are telling you this

A shipped confidence number would be the most clickable thing on the site. It is not on the site because it did not work, and shipping it anyway would have been selling a decoration as an instrument.

What it means for your picks

The de-vigged market probability is the confidence score. When several books disagree about a game, that is worth knowing as context — you can see it yourself in the book-depth panel on the odds board — but it does not reliably tell you the favourite is in trouble.

How our ratings are actually built, and where they are and are not used, is documented in the methodology. The register of every test is in the Edge Lab.

Track your real pool free — 10-day trial, no card →