BayerRank is a statistical model built from 65,000+ professional pickleball matches. Every rating, odds board, and title probability on this site comes out of one pipeline: there is no hand-tuning between the model and the page.
Under the hood: a statistical engine that reads each player's full match record, plus a second engine that reads their recent results in order, so a player surging and a player fading look different even when their averages match. Each season projection is the remaining schedule simulated forward 100,000 times. Forecasts are evaluated against a locked set of held-out matches played after everything the model learned from, and calibration is checked the boring way: when it says 70%, that should happen about 70% of the time.
It does not know about injuries, lineups before they are announced, or what anyone said on a podcast. Every forecast is frozen before first serve and graded after, misses included, and limitations get documented here as they are found.
How accurate is it?
The Events tab is the answer: accuracy, log loss, and calibration published for every graded match, misses included. A single flattering accuracy number quoted here, with the conditions chosen after the fact, is exactly what this site exists to avoid.
The track record just started. Why should I trust it today?
You are not being asked to yet. The public ledger opened in July 2026, and a calibration table built on one event is an anecdote, not evidence: judging whether the 70% predictions really land 70% of the time takes hundreds of graded matches, which takes a season rather than a weekend. The model was developed against thousands of past matches it had never trained on, which is why these forecasts are worth publishing at all, but private testing should not buy public trust. What is on offer is a commitment: every forecast frozen before the result, every miss kept, and a sample that only grows. Skepticism now is the correct read; the ledger is how it gets answered.
Why publish probabilities instead of just picking winners?
Because the probability is the claim. Naming a winner only says which side of 50% a match sits on; the probability says exactly where, and that number is what gets graded. This is a stricter standard than picking winners, not a softer one. A winner prediction can only be right or wrong, once per match; a probability is scored on every match by how far it sat from the result, so overconfidence and hedging both cost. More ambitious, and much easier to catch being wrong.
The model said 70% and that side lost. Was it wrong?
Not by itself. A 70% forecast never said the favored side would win; it said sides like this one win about seven times in ten. So the model working correctly does not look like every 70% prediction landing; it looks like about 70% of them landing and the other 30% missing. A spotless record would itself be a red flag: predictions that never miss deserved a number far above 70, so the 70 was the error. One match settles nothing about a probability; the calibration table on the Events tab is where hundreds of them settle it, band by band.
What do accuracy, log loss, Brier score, and calibration actually measure?
Four grades for four different questions. Accuracy is the simple one, how often the side the model favored won, and it is the easiest to fool: lopsided matchups flatter it on their own. Log loss grades the probability itself: a confident miss costs far more than a cautious one, so bluster and hedging both lose points, and lower is better. The anchor is a coin: guessing 50% on every match scores about 0.69, and the question is how far below that a model can stay. The Brier score grades the same probabilities with squared error instead of logarithms: lower is still better, the same coin scores 0.25, and because each match's penalty is capped, one terrible miss cannot dominate it the way it can dominate log loss. Calibration checks honesty in bulk: across all the matches where the model said 70%, that side should have won about 70% of the time. All four are published because each catches a failure the others miss.
What does the model actually look at?
More than the ratings. The foundation is the per-player rating built from every recorded result, but a rating alone is not a forecast: a real share of the model's accuracy comes from turning two ratings into a well-calibrated probability for one specific match. That layer reads context the ladder deliberately ignores, including each player's recent form and trajectory, margins of victory and not just wins, the format and length of the match, and the strength of the field it is played in. Every signal earned its place by improving predictions on held-out matches, and most candidates tested did not (see What is overfitting? at the end of this section). The exact signals, their weights, and how they combine stay in the lab. That is the one part of this site you cannot check, so it is never the part you are asked to believe: the claim on offer is that the forecasts hold up, and that claim is graded in public with or without the recipe.
Is this just DUPR with extra steps?
No. Every number on this site is computed from match results alone: opponents, dates, events, and scores. No DUPR ratings, no seedings, no external rankings of any kind go in. DUPR answers a different question: one rating for every player, from rec courts to the tour, that has to hold up across all of them. This model has exactly one job, predicting pro matches, and it gets graded on that job on the Events tab. So nothing here asks you to adopt a new scale: DUPR rates your own play; the Bayer Rating exists so the forecasts on this site can be checked.
Where does the match data come from?
The public record of professional play, collected and cross-checked into one corpus going back to January 2022. The major tours are all in, MLP, PPA, and APP, along with the wider pro record: national championships, international events, and other professional draws. That is where a count like 65,000+ comes from: every division, more than four seasons, one continuous record. Nothing private goes in and nothing external rides along: no ratings, no seedings, no rankings, just who played whom, when, and what the score was. The unglamorous half of the work is keeping each player's history stitched together across sources that disagree on how to spell a name.
What are the model's biggest limitations?
It reads results, so anything that has not shown up in results yet is invisible to it. It does not predict trades or roster moves: season projections assume current rosters and get re-run when a move is announced. It does not know who is carrying an injury, who switched paddles, or who spent the off-season rebuilding a serve until those things start showing up as wins, losses, and scorelines. Players with thin records carry more uncertainty than any single displayed number can show, which is why young ratings are held off the boards (see The ratings, below). And pro pickleball's public record is young and scattered, so the corpus behind the ratings is smaller than what models in older sports get to learn from. When any of this costs accuracy, the Events tab shows it.
What is overfitting?
Overfitting is when a model memorizes the quirks of the data it was trained on instead of learning what generalizes, like a student who memorizes last year's answer key rather than the subject. Adding parameters just hands that student a longer answer key: the practice score goes up, and this year's exam does not care. That is why a bigger model is often a worse one. The defense here is strict: every candidate signal must improve predictions on held-out matches the model never trained on, played after everything it learned from. Two more guards run against luck: the held-out window rotates on a schedule, so there is never a stable answer key to memorize, and admission takes a clear, statistically significant improvement, not a lucky decimal. Most candidates fail, which is the system working. Wikipedia has a fuller explanation.
When are forecasts locked?
At first serve: the placement board and every match probability freeze as one snapshot, and nothing is edited after that. An event board is published early and can move while roster and lineup news is still landing, but pre-freeze movement comes only from new inputs entering the same pipeline (results, rosters, lineups), never from a hand on the output. The last edits land the night before the event, once day-one lineups are known; from there the board rides untouched to first serve. Mid-event injuries, surprise lineups, and plain bad predictions are graded as misses, never patched, and matches on later days grade against that same snapshot. If holding a Thursday freeze through Sunday costs accuracy, it shows up in the grades, where it belongs.
How often does everything update?
Ratings, rankings, and the season board refresh once the weekend's results are recorded, usually by Monday morning. Season projections also re-run when a roster move is announced, never speculatively. Between events nothing drifts: a rating never decays on its own, though a long-absent player eventually ages off the boards until they compete again. Event forecast cards run on their own clock: published early, final the night before play, frozen at first serve (see When are forecasts locked?). Nothing on this site updates live during a match.
A team played a different lineup than the forecast assumed. Does it still count?
Yes, it counts. Every published probability is graded as published, no exceptions. Roster and lineup surprises are part of what a real forecast has to survive, and a track record with hand-picked exclusions is not a track record. The only matches waiting on a grade are the ones without an official final score yet.
Why can a team with lower title odds have a higher predicted finish?
Because they answer different questions. The Pred column ranks teams by expected finish: the average of where a team lands across every simulated run of the event. Title odds care about one outcome only, first place. With unbalanced pools the two can split. A clear second-best team in the favorite's pool reaches the semifinals almost every time and usually loses there, a very reliable fourth place. A stronger team in a loaded pool might miss the semifinals half the time, yet win the title more often when it gets through. The first team has the better predicted finish; the second has the better title odds. The card shows both because both claims get graded.
Which events does BayerRank cover?
Coverage is a declared rule, never a weekend-by-weekend choice: every MLP event, no exceptions, and starting with the 2027 season every PPA event awarding at least 1,000 ranking points. From then on everything above that line is covered, never selectively and never retroactively, so choosing easy events can never quietly pad the track record. Coverage governs which events get frozen, graded forecasts; results from every recorded pro event feed the ratings either way. The ledger opened with the first frozen forecast, the Mid-Season Tournament in July 2026. Earlier 2026 boards were never frozen before first serve and are not reconstructed after the fact: a forecast produced after the results are known is not a forecast. If coverage ever slips, the gap gets recorded on the Events tab like any other miss, never papered over.
What does the model version mean?
The number left of the dot changes when an update can move ratings and rankings for reasons other than play, such as a model improvement or expanded match coverage. The number right of the dot changes when predictions sharpen without touching the ratings. Either way, forecasts are never re-graded: every event card keeps the version that made it and is graded on what that version actually said.
What is the Bayer Rating?
The model's estimate of a player's current strength in a division. A new player starts at 50, and the rating moves with results, weighted by the quality of the opposition and the scoreline, so beating a top-five player moves it far more than edging out a qualifier. The scale is fixed: a ten-point gap converts to roughly a 64% chance of winning a single head-to-head game, and larger gaps compound quickly. Each division is ranked on its own ladder, so a rating in men's doubles and one in women's singles are separate claims and not comparable. The ratings come out of the same model that makes every forecast on this site, with no separate ranking to tune, and a match forecast also reads current form and matchup context, so it can see more than the gap between two ladder numbers.
Why isn't a player I follow on the board?
Three rules keep players off a board. Ratings built on fewer than 20 recorded games in a format are provisional and not shown (games, not matches: a best-of-three can supply up to three): a young rating mostly reflects uncertainty, and it would sit next to numbers built on hundreds of games while looking just as authoritative. Staying listed also takes recent play, ten matches in that division within the past twelve months, so long-absent players age off until they compete again. And each board shows the top 30 of its division, so a rated player can simply be below the cut. Provisional players appear automatically once they reach 20 games. The search box below the boards covers every rated player, not just the top 30, and shows each one's exact rank; a name it cannot find is off the ladder for one of these three reasons.
The rating disagrees with my eye test on a player. Who is right?
Neither automatically. The model does not watch matches; it reads results, weighing each one by opposition quality and scoreline, and nothing else. Sometimes the eyes are ahead: a player leveling up mid-season is visibly better than their last fifty games, and the number takes a few events to catch up (the flame and snowflake badges on the Rankings tab exist for exactly this). Sometimes the eyes are a highlight reel: a player who looks dominant while dropping close games to mid-ladder opponents will rate like it. The useful part is that the disagreement gets settled in public: ratings feed forecasts, forecasts get graded on the Events tab, and if the eye test keeps winning there, that is a model problem, and it gets documented.
Why do ratings move so much after one weekend?
Mostly it is the ranks moving, not the ratings. The middle of a ladder is packed: ranks 20 through 50 can sit within a dozen rating points of each other, so a strong weekend worth eight points can move a player twenty places, while the same eight points near the top, where the gaps are wide, moves one or two. The rating change is modest either way; the rank change amplifies it wherever the field is dense. As for how quickly the ratings themselves respond, that is not a dial turned by taste. Recent form genuinely predicts upcoming results in this sport, and every mechanism that makes the model responsive had to prove it on matches the model had never seen, against a less responsive version of itself, before it shipped. Damping the movement to look more stable would make the forecasts measurably worse. The check that responsiveness has not tipped into overreaction is public: forecasts are graded on the Events tab, and a model that overreacted to single weekends would show up there as miscalibration, with favorites winning less often than their stated odds.
Why aren't there separate gendered and mixed doubles ratings?
Half of that split already exists: men's and women's boards are separate ladders in both singles and doubles. The half that does not is mixed. A mixed result counts on each player's own board, his on the men's ladder, hers on the women's, because a player's doubles rating is built from their same-gender and mixed play together. A mixed-only board would cut every player's history into thinner slices, and thin slices are where a model starts fitting noise instead of skill (see the overfitting question under The model). Every mixed split tested has predicted held-out matches no better, usually worse, so it stays out until one earns its way in. The same bar keeps out other appealing ideas, like a partnership bonus for pairs who play together often: nothing ships unless it improves predictions on matches the model has never seen.
Can I get a Bayer Rating?
Only by playing the pros. The model is built entirely from professional match records, so it has nothing to say about rec or amateur play; that is DUPR's job, and DUPR is built for it. Make a pro draw, and once your results enter the public record a rating starts accruing automatically, provisional until 20 recorded games in a format.
Two players show the same rating. Who ranks higher?
Whoever is ahead on the full-precision rating underneath. Displayed ratings are rounded to whole numbers, so two players sharing a number is common, but genuine ties are vanishingly rare: the tie lives in the rounding, not in the model.
Who's behind this? I'm Ross Bayer. Behind the model: a Master's from Stanford specializing in Artificial Intelligence, built on a strong math and statistics background going back to representing New Zealand at the International Mathematical Olympiad, and a career spent making large-scale data systems at Facebook, Microsoft, and Airtable. Full disclosure: I play pickleball badly but enthusiastically. BayerRank began as a personal question, whether a model could predict pro pickleball better than the takes, and grew into 65,000+ matches, a few hundred controlled experiments, and this site. It is a personal project, built on my own time, with no employer involved. Treat the credentials as a reason to look; the Events tab is the reason to believe.
Why free, why no ads? Because the point is the scoreboard, not a business. The model gets graded in public, including the misses. That only means something if nobody is paying for the answer.
Seen these numbers before? The model's season projections posted on X under the ErneCast name before this site had a home. Same model, new address. The ledger counts nothing from before its own rules existed, in either direction: no credit claimed for the hits, no grades assigned to the misses.