R at Bet Better
Draw a calibration curve for betting odds in R
28 September 2026. R, dplyr, ggplot2.
A calibration curve asks one question of a set of probabilities: when the market said 60 per cent, did the outcome happen about 60 per cent of the time? Here is the whole thing in R, using the open market calibration dataset: 36,476 outcome observations, both sides of 18,238 settled two-way markets, vig removed, no bookmaker names and no bettor selections.
1. Load
Download market_calibration.csv from the Kaggle mirror (CC BY 4.0), then:
library(dplyr); library(ggplot2)
cal <- read.csv("market_calibration.csv")
# columns: event_month, sport, market_family, market_implied_probability, outcome
2. Bin and compare
curve <- cal %>%
mutate(bin = cut(market_implied_probability, seq(0, 1, 0.1), include.lowest = TRUE)) %>%
group_by(bin) %>%
summarise(predicted = mean(market_implied_probability),
observed = mean(outcome),
n = n())
curve
3. Plot against the diagonal
ggplot(curve, aes(predicted, observed)) +
geom_abline(slope = 1, intercept = 0, linetype = 2) +
geom_line() + geom_point(aes(size = n)) +
coord_equal(xlim = c(0, 1), ylim = c(0, 1)) +
labs(x = "Market probability (vig removed)", y = "Observed frequency",
title = "Betting markets are well calibrated",
caption = "Bet Better market calibration dataset, CC BY 4.0")
Every decile lands within about two points of the diagonal. Look closely and the low bins sit slightly above the line: outcomes priced under 50 per cent happen a touch more often than their vig-free price says, which is the favourite-longshot tilt showing through after the margin is stripped out.
4. Split by market family
cal %>% count(market_family)
# repeat the binning with group_by(market_family, bin):
# player props tilt about twice as much as game lines
Data licensed CC BY 4.0, free to reuse with a credit. Model estimates, not advice. 18+. If gambling is causing you harm: gamblinghelponline.org.au.