Methodology
How the tennis prediction model works
This page describes how the published probabilities are produced, how they are validated, and where the model's limits lie. It is written to be checkable rather than reassuring.
Last updated: 14 September 2026
What is predicted
One thing only: the probability that each player wins the match. A likely set score is published too, but it is a secondary output and less reliable than the win probability.
The model covers the ATP, WTA, Challenger men's and Challenger women's tours, in singles and doubles. ITF tournaments are not predicted, but their results feed the strength ratings of the players who compete in them.
The data behind it
The model is a gradient boosting model trained on professional match history since 2020, in four steps:
History since 2020
Every archived professional match
Match variables
Elo, form, fatigue, head-to-head, market odds
Model
Gradient boosting trained on the history
Published output
Win probability + likely set score
- an overall Elo and a surface Elo, the latter seeded from the overall rating rather than from a neutral value, without which it converges poorly and ends up less predictive than the overall rating alone;
- recent form and a fatigue measure drawn from matches played in the preceding days;
- the head-to-head record, weighted by recency: a meeting five years ago does not count as much as one last month;
- the market's implied probability where odds are available, with the bookmaker's margin removed.
The rule that governs everything else
No information from after the match can enter the calculation. It is the one rule that, when broken, produces a model that looks excellent in testing and is useless in production, without ever raising an error.
Before the match — each player's state is read
The match is played
After the match — the state is updated
In practice, each player's state is read and updated in that order, by a single shared implementation used both for training and for prediction. A match's detailed statistics, which are only known once it has finished, are never used as they are: they enter only as rolling averages computed over the player's earlier matches.
How the model is validated
Chronological split
Training, then validation, then test, always in that time order — never a random shuffle, which would amount to training on matches played after the ones being scored.
Walk-forward validation
The model is retrained at each window on all the history available so far, then scored only on the next window, never seen before.
The accuracy shown in the application is computed differently, and it is the figure that really counts: it compares the predictions actually published before the matches with the results that followed. It is not reconstructed after the fact and cannot be filtered retroactively.
The metrics used
Accuracy alone is not enough to judge a model that publishes probabilities: it ignores confidence. Saying 51% or 95% for the same correct call gives the same accuracy, though the two are not worth the same.
Log loss
Penalises misplaced confidence far more heavily than a plain accuracy miss.
Brier score
Measures the gap between the announced probability and the real outcome, across every match.
Significance test
A gap between two versions only counts once it is shown not to be plain statistical noise.
Finally, the probabilities produced are capped: no prediction shows absolute certainty. A calibration left unbounded produces values of exactly 0 or 1 in thinly covered regions, and a single one of those proved wrong costs as much as hundreds of moderate errors.
Where the data comes from
Schedules, results, live scores and odds come from the provider api-tennis.com. An interruption or an error on its side can leave some information temporarily inaccurate. See the legal notice and the terms of use.
See the model at work
Today's predictions, the track record and the real accuracy are visible right away.