A free teaching dataset: 6,131 finished betting markets
One row per settled market. Each row has the price that was on offer, the
chance that price implies, the chance the market implies once the bookmaker's cut is
taken out, what our model thought, and whether it actually won. Free to use with a link
back to us.
Cut 18 September 2026 · 12 sports · 134 market types ·
games from 18 May 2025 to 18 September 2026
6,131
settled markets, all graded
12
sports and competitions
15
columns, all plain English
2.9%
average bookmaker's cut
Who it is for
It was built for a teacher who needs a real, messy, honest dataset for a first or second
course in statistics or data science — one where the questions are genuinely open
and the answers are checkable against finished games.
- Probability and statistics. Prices convert straight to
probabilities, so students can test whether stated chances match what happened.
- Data science and machine learning. A clean binary outcome
(
WIN or LOSS), a strong baseline to beat, and an obvious
train-and-test split by date.
- Economics and behavioural science. The price of a bet is a visible,
measurable fee, and it differs by sport and by market.
No sign-up, no key, no account. Download it and hand it out.
What is in a row
| Column | Meaning |
sport, game_date, home_team, away_team | Which game, and when it was played. |
market, selection, line | What was bet on, which side, and the number it had to beat. |
result | WIN or LOSS. Nothing ungraded is in the file. |
bookmaker_price | The decimal price, averaged across the books quoting that market. |
our_probability | The chance our model gave that selection before the game. |
price_implied_probability | The chance the price implies, cut included. |
fair_probability_no_cut, fair_price_from_market | The same thing with the bookmaker's cut taken back out. |
bookmaker_cut_pct, cut_source | The size of the cut on that selection, and whether it was measured or estimated. |
Three exercises, ready to set
A. Do prices tell the truth?
Bucket the rows by the market's fair chance, then compare each bucket's average chance
with the share that actually won, and plot it against a 45-degree line. Expect the big,
heavily traded markets to sit close to the line and the long-odds buckets to be noisy.
The follow-up question is the real lesson: how wide is the uncertainty band, and does any
gap survive it?
B. How much does the bookmaker keep, and where?
Average the cut by sport, then by market type, then repeat using only the rows where the
cut was measured rather than estimated. Expect around 2.9 percentage points overall,
modest differences between sports and much bigger ones between market types —
simple win or lose markets are the cheapest, player markets the dearest.
C. Beat the market, if you can
Fit a simple model on the games before a cut-off date, test on the games after it, and
compare three sets of probabilities on the test set: the student's model, the market's,
and ours. Most first models will not beat the market. The point is the size of the gap,
and the harder question underneath it — whether anything beats the market by more
than the bookmaker's cut, because anything less than that still loses money.
Free to use with attribution to Bet Better
(betbetter.world). Cite it as: Bet Better (2026), Teaching dataset:
betting markets, prices and results, betbetter.world/studies/teaching-dataset.
If you use it in a course, we would like to hear about it.
How it was collected, and what it is not
We record prices for the games our models price, we store the model's own probability at
the same moment, and after the game finishes a settling job marks each selection won or
lost from the official result. This file is a balanced sample of that settled record:
markets are collapsed to one row per game, market and selection, prices are averaged
across the books that quoted them, and each sport is capped so no single sport takes over
the file. Anything ungraded, unpriced or missing a model probability was dropped.
Please pass these limits on to students.
- No bookmaker is named and no live price is served. Every row is a market that has
already finished.
- The rows per sport are a sampling choice, not a measure of how much each sport is
played or bet on.
- NHL (8 rows) and Cricket (43 rows) are too small to draw a per-sport conclusion.
- Our model's probability is a live product output, not a teaching benchmark.
Finding it poorly calibrated is a legitimate result, and we would like to see it.
A companion game-level file — final scores, our projected margin and home win chance
— is published at
betbetter-teaching-dataset-2026.csv
with its own codebook.
Companion research: Do betting odds tell the truth? ·
What does a bet actually cost? ·
Free charts you can embed
18+. Gambling involves risk and most people lose money.
This file is published for teaching and research. Nothing here is betting advice, and
nothing in it makes betting profitable. If gambling is causing you harm, support is
available at
Gambling Help Online (AU)
or 1-800-GAMBLER (US).