Tennis prediction algorithm: how an AI model actually works
· Pronocast

“AI-powered prediction” has become a phrase you read everywhere, and it says nothing about what actually happens inside the machine. There is rarely any magic behind it: a statistical model receives a few dozen numbers describing two players and returns a single one, the probability that the first wins. All the work lies in choosing those numbers, making sure they were known before the match, and honestly checking what the model is worth once it faces matches it has never seen.
This article walks through that mechanism end to end, using the model behind Pronocast as the example. It explains what data goes in, how the model is prevented from “seeing the future” while it learns, what a well-calibrated probability means, and why the accuracy of such an algorithm plateaus, whatever you do, around that of bookmakers’ odds. The goal is not to sell a success rate, but to give you the keys to judge any automated prediction, this one included.
What a tennis prediction model really predicts
One thing only: the probability that each player wins the match. Not the score, not the number of games, not the duration. A likely score in sets can be derived afterwards, but it is a secondary output, noticeably less reliable than the win probability.
A probability, not a verdict
Saying “Alcaraz wins at 71%” is not saying “Alcaraz wins”. It means that over a hundred matches with exactly this profile, he is expected to win seventy-one. The other twenty-nine are not model errors: they are the share of uncertainty in tennis, which a good probability should quantify rather than hide. That is why the success rate alone is a poor measure: a model that says 51% and a model that says 95% for the same correct prediction have the same success rate, but not at all the same value.
The scope covered
The Pronocast model covers the ATP, WTA, men’s Challenger and women’s Challenger tours, in singles and doubles. ITF tournaments are not predicted, but their results feed the strength calculation of the players who compete there: a player climbing up from the ITF circuit arrives in Challenger with a history, not with a blank rating.
The data that goes into the calculation
The model itself is a gradient boosting, a family of algorithms that builds hundreds of small successive decision trees, each correcting the errors of the previous ones. It is trained on the history of professional matches since 2020. What matters is not so much the algorithm as what it is given to read for each match.
Elo, global and by surface
The first signal is an Elo rating, a measure of a player’s strength derived from their results and the strength of their opponents. A global Elo is complemented by a surface Elo (hard, clay, grass), which captures the fact that a player is not the same on clay and on grass. One detail changes everything: the surface Elo is initialised on the player’s global rating, not on a neutral value. With three to four times fewer matches per surface than overall, a rating that starts from scratch converges poorly and ends up predicting worse than the global rating alone. We measured this effect directly before choosing the initialisation.
Form and fatigue
Recent form is the share of wins over the last ten matches. The fatigue indicator counts matches played in the preceding days: a player coming off a tournament won in six matches arrives with a load that Elo does not see.
Head-to-head
The head-to-head history is weighted by recency: a win from five years ago does not weigh as much as one from last month. The weighting follows an exponential decay with a half-life in the range of one to two years. A plain count of wins, with equal weights, would give too much importance to matches played by players who no longer really exist in that form.
The market, when it exists
Finally, the implied probability of the odds is fed to the model when a price is available, after removing the bookmaker’s margin. This choice has a consequence that must be owned: a model that reads the market can no longer be used to measure whether it beats it. We come back to this below.
The rule that governs everything else: never see the future
A sports prediction model trains on past matches, whose result is known, to predict future matches, whose result is not. The central trap of the exercise fits in one sentence: if a feature captures, even indirectly, information that did not yet exist at the time of the match, the model looks excellent in testing and collapses in production, without ever raising an error.
Read before, write after
The safeguard is an implementation discipline more than a mathematical trick. Each player’s state (their Elo, form, head-to-heads) is read before the match, then updated after, in that order, match after match, walking through the history chronologically. And it is one and the same implementation that builds the training set and predicts today’s matches. Two “equivalent” implementations diverge sooner or later, and the divergence is invisible: the model keeps running, just worse.
The case of match statistics
Detailed match statistics (aces, first-serve percentage, points won on serve) are only known once the match is over. They are therefore never used as they are. They only enter as rolling averages computed on the player’s previous matches. And, a counter-intuitive result measured on this project, those averages brought no gain at all: a player who serves well wins more matches, so already has a better Elo. The information was already there.
How you check that a model is worth anything
This is where the difference between a serious algorithm and a shop window is decided. Three principles.
A strictly chronological split
Training matches precede validation matches, which precede test matches. Never a random draw, which would amount to training the model on matches later than those it is evaluated on. The validation set is genuinely used to choose the model and its settings; the test set is looked at only at the end, once.
Walk-forward validation
A single fixed split can give a good result by chance. So the model is retrained at each window, for instance every quarter, on all the history available up to then, and evaluated only on the next, never-seen window. On this project, the model beat its baseline on nine windows out of ten; the only exception was the very first one, when there was not yet enough data. A fixed split would never have shown that nuance.
The right metrics
The success rate is readable, but it ignores confidence. The metrics that matter for a model that displays a probability are log-loss and the Brier score, which penalise misplaced confidence. A model that announces 95% and is wrong pays dearly; that is exactly what we want. Add a common-sense rule: always establish a simple baseline first (“the higher-rated player wins”) and check that the model really beats it, tour by tour.
Calibration, or why no probability should ever show 100%
A well-calibrated model is one whose 70% predictions actually win seven times out of ten. Raw model outputs are often adjusted with a post-hoc calibration. The trap, met on this project: fitted on a modest validation set, such a calibration can produce probabilities of exactly 0 or 1 in sparsely populated regions. A handful of those “certain” predictions being wrong is then enough to blow up the log-loss: a single error at absolute confidence costs as much as hundreds of moderate ones.
Measured directly: a calibration that improved the log-loss on validation degraded it sharply on test, through that mechanism alone. The conclusion is simple and applies to any automated prediction: the output is capped, for instance between 3% and 97%. No tennis match is decided in advance, and a model that shows 100% is lying, at best clumsily.
The odds market as the final judge
Since tennis has an active betting market, consensus odds, averaged across several operators and converted into probabilities after removing the margin, are the best public predictor available. On this project, a model based only on Elo, form and head-to-heads plateaued around 65 to 67% accuracy; the market reached 68.5%. Adding the market probability as an input made it possible to reach that level, not to durably exceed it.
Compare on the same scope
A classic trap is to compare a model on 100% of matches with a market that covers only 90%. A comparison made that way suggested a gain of 2.3 points that did not exist; redone on the strict intersection, the gap fell to 0.06 points, that is, nothing. Any claim of the kind “our AI beats the bookmakers” should state the scope and the significance test. Otherwise it is marketing.
What the model brings, then
Full coverage, including the hundreds of Challenger matches without a price, a calibrated and verifiable probability for each match, and a transparency the market does not offer: every prediction is computed before the match, frozen, and published with the result. The accuracy is recomputed from those pages, not from a declared figure.
The limits you should know
- Injuries, last-minute withdrawals and the day’s conditions are not in the data. A diminished player keeps their Elo until their results bring it down.
- New players and long-absence comebacks have an unreliable rating, for lack of recent history. The K-factor, higher for them, speeds up the correction without making it instant.
- Doubles is less predictable than singles: pairs change, team history is thin and the market is often absent.
- A model retrained every month slightly changes its behaviour. The model version is shown on each prediction for that reason.
Today’s Pronocast predictions
The model described here computes a prediction every morning for every match on the tours covered. Upcoming matches are listed on the predictions of the day page, with the ATP, WTA and Challenger hubs. Once a match is over, the prediction and the result are published on the match page, and the real accuracy is recomputed.
Frequently asked questions
- Can a tennis prediction algorithm predict every match?
- No. A good model is right on roughly two matches out of three across all tours, and more often on lopsided matches. The remaining third is not a calculation error: it is the real uncertainty of tennis, which the displayed probability quantifies.
- What data does the Pronocast model use?
- A global Elo and a surface Elo, recent form, a fatigue indicator, the head-to-head history weighted by recency, and the implied probability of the odds when available. No data from after the match enters the calculation.
- Why is the real accuracy published match by match?
- Because a success rate announced without the matches behind it cannot be checked. Each prediction is computed before the match, frozen, then published with the result on its own page: the accuracy is recomputed from those pages, not from a declared figure.
- Is the model better than the bookmakers?
- On matches where odds exist, the model matches the market’s accuracy without durably beating it: consensus odds are the best known public predictor. What the model adds is full coverage, Challenger included, and a calibrated probability for every match.
Other articles
- Elo in tennis: understanding the rating that predicts matchesWhy a rating designed for chess predicts a tennis match better than the ATP ranking, and which adjustments it needs to really work.
- Tennis odds and implied probability: what they really sayHow to read bookmaker odds, strip out the margin, and why the market stays the benchmark most models can’t durably beat.
- How to bet on tennis: bankroll, value betting and riskHow much to stake, what value really means, and why our own model lost 9% betting on its disagreements with the market.
- Clay, hard or grass: what the surface really changes in tennisAcross 36,210 matches, a player’s overall level predicts better than their record on the day’s surface, and grass is where favourites fall most often.
- Head-to-head in tennis: does the H2H really predict the winner?Across 38,123 matches, the head-to-head leader wins less often than the Elo favourite. But a record against the favourite really does count.
- Tennis fatigue and rest: do tired players really lose more?Stringing matches together is barely measurable. Returning after a month without playing is: nearly ten points of win probability less than Elo predicts.
- Indoor tennis: are favourites more reliable under a roof?Indoors, the favourite wins 64.9% of matches, against 65.0% on outdoor hard courts. A good indoor record mostly reflects a good player.