Home Technological Edge Benchmarking Prediction Systems: Reading Accuracy Claims

Benchmarking Prediction Systems: Reading Accuracy Claims

Every forecasting system claims an accuracy figure. Almost none of those figures are comparable with each other, and the reasons why are worth understanding before you take any of them seriously.

Why “75% accurate” means almost nothing

The number depends entirely on choices that usually go unstated:

  • Which market. A two-way tennis market and a three-way football market are not the same problem. Football with a draw caps sensible accuracy far lower.
  • Which events. Restricting to heavy favourites inflates accuracy without adding any skill.
  • Which period. One good season is a sample, not a result.
  • Whose baseline. Beating a coin flip is trivial; beating the closing market price is the real bar.

An accuracy figure without a stated market, sample and baseline is a marketing claim, not a measurement.

The baseline problem

Scored

Scored

First Deposit Sports Bonus: 100% up to €250

Ivibet

Ivibet

First Deposit Bonus for Sports Betting: 100% up to €150

BetRepublic

BetRepublic

First Deposit Bonus: 100% up to €250

Slotsgem

Slotsgem

Welcome package up to €1,450 + instant bonus round + 225 free spins

22bet

22bet

Welcome bonus up to €122 for sports betting

Slotrave

Slotrave

Welcome package: 450% up to €2,500 + 150 free spins + 1 bonus game

Winhero

Winhero

150% welcome freebets up to €300 for the first three deposits

Playbaze

Playbaze

50% welcome risk-free bet

Spinbetter

Spinbetter

Up to €500 for the first five deposits

GG.BET

GG.BET

Up to 200% bonus + €100 freebet

CrazyTower

CrazyTower

First deposit bonus: 100% up to €200

18+. Advertising. Betting carries risk — gamble responsibly and never stake more than you can afford to lose.

Comparing a model to random guessing is the most common way to make it look good. Three more demanding baselines:

Always back the favourite

A rule requiring no model at all, and one that already reaches respectable accuracy in most leagues. Any system that cannot beat it has demonstrated nothing.

The closing price

The market at the moment it closes aggregates everything known, including the information that arrived last. It is a genuinely hard benchmark, and beating it consistently is rare.

A simple statistical model

A Poisson model on goals, or an Elo rating, takes an afternoon to build. If an elaborate neural system cannot outperform one, the elaboration is decorative.

What a fair comparison requires

  • The same events. Both systems predict every fixture in a defined set, with no discretion to skip.
  • The same timing. Predictions locked at the same point before kick-off, since later is easier.
  • The same metric. Ideally a probabilistic one — log loss or Brier — rather than a hit count.
  • A stated sample size. With confidence intervals, because a few hundred matches cannot resolve small differences.

How benchmarks get gamed

Selective reporting

Publishing the competitions where the model did well. The corrective is simple and rarely applied: define the coverage first, then report all of it.

Retrospective prediction

Any statistic recorded after kick-off will improve a model dramatically and be unavailable when it matters. Timestamping predictions before the event is the only real defence.

Tuning against the test set

Adjusting a model repeatedly while checking the same held-out data leaks that data into the model, one decision at a time. After enough iterations the test set has become a training set.

Silent revision

Editing a published prediction after the result is known. Locking the record is what makes it evidence rather than narrative.

Reproducibility

A result nobody can reproduce is not a result. In practice that requires four things recorded alongside the prediction:

  • Which model version produced it
  • What data it had access to at that moment
  • When it was generated, to the minute
  • What it actually said, before anyone knew the answer

This is unglamorous engineering rather than modelling work, and it is the difference between a system that can be audited and one that can only be advertised.

How to read someone else’s numbers

Four questions cover most of it:

  • Over how many events, and which ones?
  • Compared with what baseline?
  • Were the predictions recorded before the events?
  • Are the losing calls still published?

A system that answers all four is making a claim you can check. A system that answers none is asking to be trusted, which is a different request entirely.

Conclusion

Scored

Scored

First Deposit Sports Bonus: 100% up to €250

Ivibet

Ivibet

First Deposit Bonus for Sports Betting: 100% up to €150

BetRepublic

BetRepublic

First Deposit Bonus: 100% up to €250

Slotsgem

Slotsgem

Welcome package up to €1,450 + instant bonus round + 225 free spins

22bet

22bet

Welcome bonus up to €122 for sports betting

Slotrave

Slotrave

Welcome package: 450% up to €2,500 + 150 free spins + 1 bonus game

Winhero

Winhero

150% welcome freebets up to €300 for the first three deposits

Playbaze

Playbaze

50% welcome risk-free bet

Spinbetter

Spinbetter

Up to €500 for the first five deposits

GG.BET

GG.BET

Up to 200% bonus + €100 freebet

CrazyTower

CrazyTower

First deposit bonus: 100% up to €200

18+. Advertising. Betting carries risk — gamble responsibly and never stake more than you can afford to lose.

Benchmarking is where technological claims either survive or collapse. Fixed coverage, honest baselines, probabilistic metrics and locked timestamps are what make one system comparable with another.

The strongest signal of a serious forecasting operation is not a high accuracy figure. It is a complete record, including the parts that did not work.