<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/">
<channel>
  <title>R at Bet Better</title>
  <link>https://betbetter.world/rstats/</link>
  <description>Short, runnable R posts about open sports betting data: bookmaker margins, odds to probabilities, calibration. Data CC BY 4.0.</description>
  <language>en</language>
  <atom:link href="https://betbetter.world/rstats/feed.xml" rel="self" type="application/rss+xml" />
  <lastBuildDate>Tue, 29 Sep 2026 09:00:00 GMT</lastBuildDate>
  <item>
    <title>Pull the daily bookmaker margin index into R</title>
    <link>https://betbetter.world/rstats/pull-the-bookmaker-margin-index-in-r</link>
    <guid isPermaLink="true">https://betbetter.world/rstats/pull-the-bookmaker-margin-index-in-r</guid>
    <pubDate>Mon, 28 Sep 2026 05:00:00 GMT</pubDate>
    <dc:creator>Edward Glush</dc:creator>
    <category>R</category>
    <description>Load the Bet Better Margin Index CSV in R with readr, filter the latest day by sport, and plot the season trend with ggplot2.</description>
    <content:encoded><![CDATA[
<p class="meta">28 September 2026. R, readr, dplyr, ggplot2.</p>
<p class="lede">The <a href="https://betbetter.world/studies/margin-index">Bet Better Margin Index</a> is a daily CSV of the margin bookmakers build into their odds, by sport, market class and market shape, measured across tens of millions of pre-match markets. Here is how to load it, tidy it and plot it in R.</p>
<h2>1. Read the file straight from the web</h2>
<pre><code>library(readr)
library(dplyr)
mi &lt;- read_csv("https://betbetter.world/studies/margin-index?format=csv",
show_col_types = FALSE)
glimpse(mi)
# as_of_date, sport, market_class, shape, markets_n, books_n,
# mean_pct, median_pct, min_pct, max_pct</code></pre>
<p>Each row is one day, one sport, one market class (head to head, spread, total, player prop) and one shape (two-way or three-way). <code>mean_pct</code> is the average margin in per cent of stake and <code>markets_n</code> is how many markets were measured that day.</p>
<h2>2. The latest day, by sport</h2>
<pre><code>latest &lt;- mi %&gt;%
filter(as_of_date == max(as_of_date),
market_class == "h2h", shape == "2way",
sport != "ALL", markets_n &gt;= 50) %&gt;%
arrange(mean_pct) %&gt;%
select(sport, mean_pct, markets_n, books_n)
latest</code></pre>
<p>On most days the NFL and college football sit at the cheap end, about 4.4 cents in the dollar, and T20 cricket at the dear end, near 7.</p>
<h2>3. The trend across the season</h2>
<pre><code>library(ggplot2)
mi %&gt;%
filter(sport == "ALL", market_class == "h2h", shape == "2way") %&gt;%
ggplot(aes(as_of_date, mean_pct)) +
geom_line() +
labs(x = NULL, y = "Average margin (% of stake)",
title = "What bookmakers keep on a two-way bet, every day",
caption = "Bet Better Margin Index, CC BY 4.0")</code></pre>
<h2>Reuse</h2>
<p>The file is CC BY 4.0. Cite it as: Glush, E. (2026). The Bet Better Margin Index: daily measured bookmaker margins across market types. https://betbetter.world/studies/margin-index</p>
<p class="rg">Data licensed CC BY 4.0, free to reuse with a credit. Model estimates, not advice. 18+. If gambling is causing you harm: gamblinghelponline.org.au.</p>
<p>Originally published at <a href="https://betbetter.world/rstats/pull-the-bookmaker-margin-index-in-r">https://betbetter.world/rstats/pull-the-bookmaker-margin-index-in-r</a>. Data is CC BY 4.0, free to reuse with a credit to Bet Better.</p>
]]></content:encoded>
  </item>
  <item>
    <title>Turn betting odds into probabilities in R, and see what the bookmaker keeps</title>
    <link>https://betbetter.world/rstats/odds-to-probabilities-in-r</link>
    <guid isPermaLink="true">https://betbetter.world/rstats/odds-to-probabilities-in-r</guid>
    <pubDate>Mon, 28 Sep 2026 05:00:00 GMT</pubDate>
    <dc:creator>Edward Glush</dc:creator>
    <category>R</category>
    <description>Implied probability, overround and fair odds from decimal prices in base R, then the same for a whole CSV of bookmaker prices.</description>
    <content:encoded><![CDATA[
<p class="meta">28 September 2026. Base R only.</p>
<p class="lede">Decimal odds hide two numbers: the bookmaker's probability for each outcome, and the cut they take. Ten lines of R separate them.</p>
<h2>1. Implied probability</h2>
<pre><code>odds &lt;- c(home = 1.85, away = 2.05)
implied &lt;- 1 / odds
implied
#  home  away
# 0.541 0.488</code></pre>
<p>The two implied probabilities add to 1.029, not 1. That extra 2.9 per cent is the overround, the margin built into the prices.</p>
<h2>2. The margin in cents per dollar</h2>
<pre><code>overround &lt;- sum(implied) - 1
margin_pct &lt;- 100 * overround / sum(implied)
round(margin_pct, 2)
# 2.8   (per cent of every dollar staked, across the market)</code></pre>
<h2>3. Remove it to get fair probabilities</h2>
<pre><code>fair &lt;- implied / sum(implied)
fair
#  home  away
# 0.526 0.474
fair_odds &lt;- 1 / fair
round(fair_odds, 2)
# 1.90  2.11</code></pre>
<p>This is the simple proportional method. It assumes the margin is spread evenly across outcomes, which it usually is not: bookmakers load more of it onto the longer price. The <a href="https://betbetter.world/studies/market-calibration">calibration study</a> measures that tilt across 18,238 settled markets.</p>
<h2>4. Do it for a whole file</h2>
<p>A clean file of bookmaker prices, one row per priced outcome, is on the Internet Archive under CC BY 4.0:</p>
<pre><code>prices &lt;- read.csv("https://archive.org/download/betbetter-market-prices-2026/betbetter-market-prices-2026.csv")
str(prices)
library(dplyr)
margins &lt;- prices %&gt;%
group_by(sport, game_date, home_team, away_team, market) %&gt;%
summarise(overround = sum(1 / bookmaker_price) - 1,
margin_pct = 100 * overround / sum(1 / bookmaker_price),
sides = n(), .groups = "drop") %&gt;%
filter(sides == 2)
summary(margins$margin_pct)</code></pre>
<p>Group by market, apply the three steps, and you have a margin per market that you can compare by sport, bookmaker or date. That is exactly how the <a href="https://betbetter.world/studies/margin-index">Margin Index</a> is built each day.</p>
<p class="rg">Data licensed CC BY 4.0, free to reuse with a credit. Model estimates, not advice. 18+. If gambling is causing you harm: gamblinghelponline.org.au.</p>
<p>Originally published at <a href="https://betbetter.world/rstats/odds-to-probabilities-in-r">https://betbetter.world/rstats/odds-to-probabilities-in-r</a>. Data is CC BY 4.0, free to reuse with a credit to Bet Better.</p>
]]></content:encoded>
  </item>
  <item>
    <title>Draw a calibration curve for betting odds in R</title>
    <link>https://betbetter.world/rstats/calibration-curve-in-r</link>
    <guid isPermaLink="true">https://betbetter.world/rstats/calibration-curve-in-r</guid>
    <pubDate>Mon, 28 Sep 2026 05:00:00 GMT</pubDate>
    <dc:creator>Edward Glush</dc:creator>
    <category>R</category>
    <description>Bin 36,476 vig-free market probabilities against outcomes with dplyr and plot the calibration curve with ggplot2.</description>
    <content:encoded><![CDATA[
<p class="meta">28 September 2026. R, dplyr, ggplot2.</p>
<p class="lede">A calibration curve asks one question of a set of probabilities: when the market said 60 per cent, did the outcome happen about 60 per cent of the time? Here is the whole thing in R, using the open <a href="https://betbetter.world/studies/market-calibration">market calibration dataset</a>: 36,476 outcome observations, both sides of 18,238 settled two-way markets, vig removed, no bookmaker names and no bettor selections.</p>
<h2>1. Load</h2>
<p>Download <code>market_calibration.csv</code> from the <a href="https://www.kaggle.com/datasets/eddieglush/market-calibration-dataset">Kaggle mirror</a> (CC BY 4.0), then:</p>
<pre><code>library(dplyr); library(ggplot2)
cal &lt;- read.csv("market_calibration.csv")
# columns: event_month, sport, market_family, market_implied_probability, outcome</code></pre>
<h2>2. Bin and compare</h2>
<pre><code>curve &lt;- cal %&gt;%
mutate(bin = cut(market_implied_probability, seq(0, 1, 0.1), include.lowest = TRUE)) %&gt;%
group_by(bin) %&gt;%
summarise(predicted = mean(market_implied_probability),
observed  = mean(outcome),
n = n())
curve</code></pre>
<h2>3. Plot against the diagonal</h2>
<pre><code>ggplot(curve, aes(predicted, observed)) +
geom_abline(slope = 1, intercept = 0, linetype = 2) +
geom_line() + geom_point(aes(size = n)) +
coord_equal(xlim = c(0, 1), ylim = c(0, 1)) +
labs(x = "Market probability (vig removed)", y = "Observed frequency",
title = "Betting markets are well calibrated",
caption = "Bet Better market calibration dataset, CC BY 4.0")</code></pre>
<p>Every decile lands within about two points of the diagonal. Look closely and the low bins sit slightly above the line: outcomes priced under 50 per cent happen a touch more often than their vig-free price says, which is the favourite-longshot tilt showing through after the margin is stripped out.</p>
<h2>4. Split by market family</h2>
<pre><code>cal %&gt;% count(market_family)
# repeat the binning with group_by(market_family, bin):
# player props tilt about twice as much as game lines</code></pre>
<p class="rg">Data licensed CC BY 4.0, free to reuse with a credit. Model estimates, not advice. 18+. If gambling is causing you harm: gamblinghelponline.org.au.</p>
<p>Originally published at <a href="https://betbetter.world/rstats/calibration-curve-in-r">https://betbetter.world/rstats/calibration-curve-in-r</a>. Data is CC BY 4.0, free to reuse with a credit to Bet Better.</p>
]]></content:encoded>
  </item>
</channel>
</rss>
