# BetBetter teaching dataset 2026 — codebook **File:** `betbetter-teaching-dataset-2026.csv` **Rows:** 4,740 finished games (one row per game) **Leagues:** NFL, MLB, NBA, NHL, AFL, NRL — current seasons **Built:** 18 September 2026 **Licence:** CC BY 4.0 — free to use, teach with, and republish. Please credit "BetBetter" with a link to https://www.betbetter.world/studies/market-calibration The first six lines of the CSV start with `#`. Skip them when you load the file (`pandas.read_csv(path, comment='#')` or `readr::read_csv(path, comment = "#")`). --- ## Columns | Column | Type | What it is | | --- | --- | --- | | `league` | text | `nfl`, `mlb`, `nba`, `nhl`, `afl`, `nrl` | | `season` | number | Season year as the league labels it | | `game_date` | date `YYYY-MM-DD` | Date the game was played (UTC) | | `home_team` | text | Home team name | | `away_team` | text | Away team name | | `home_score` | number | Final score, home team | | `away_score` | number | Final score, away team | | `home_margin` | number | `home_score − away_score`. Positive means the home team won | | `total_points` | number | `home_score + away_score` (runs in MLB, goals in NHL) | | `home_win` | 0 or 1 | 1 if the home team won. Blank on a draw | | `model_home_margin` | number | What our model expected `home_margin` to be, before the game | | `model_total` | number | What our model expected `total_points` to be, before the game | | `model_home_win_prob` | 0–1 | Our model's pre-game chance that the home team wins | | `closing_home_handicap` | number | The handicap on the home team in the last price we saw before kick-off. Negative means the market made the home team the favourite | | `closing_total_line` | number | The market's last total line before kick-off | **Blank means missing, never zero.** Blank cells are games where we had no model run or no stored closing price for that game. --- ## How full each league is | League | Games | `model_home_margin` | `model_total` | `model_home_win_prob` | `closing_home_handicap` | `closing_total_line` | | --- | ---: | ---: | ---: | ---: | ---: | ---: | | MLB | 2,703 | 1,275 | 1,251 | 640 | 1,006 | 1,006 | | NBA | 838 | 75 | 28 | 58 | 0 | 0 | | NHL | 773 | 355 | 0 | 9 | 0 | 0 | | NRL | 166 | 46 | 45 | 28 | 47 | 47 | | AFL | 165 | 48 | 44 | 28 | 61 | 55 | | NFL | 95 | 0 | 0 | 0 | 16 | 16 | | **All** | **4,740** | 1,799 | 1,368 | 763 | 1,130 | 1,124 | MLB is the league with enough of everything to do real work. NFL rows are scores and closing lines only — the current NFL season had just started when this file was built, and our stored NFL model runs are all for games that had not been played yet. NBA and NHL have no stored closing lines in this file. --- ## Where the numbers come from - **Scores** — our own score store, built from public scoreboard feeds, one row per team per game, finished games only. We take the final score, not period scores. - **Model numbers** — the pre-game output of BetBetter's per-sport models, saved at the time the tip was published, so nothing here was written after the result was known. - **Closing lines** — the last price we recorded from licensed bookmakers before the game started, averaged across the books we had. - **Joining** — scores and model rows are matched on league, home team and game date (allowing one day either side, because our score dates are UTC). Where a team is written differently in the two sources (NRL and AFL nicknames versus full club names) we match on the shorter name inside the longer one. A handful of games will therefore be unmatched rather than wrongly matched. ### Honest limits - Team names are not standardised across leagues. NRL rows use club nicknames ("Panthers"); AFL and the US leagues use full names. Normalise before you join anything of your own. - This is one season, not a long history. It is built for teaching and for hackathons, not for settling arguments about long-run performance. - Nothing here is betting advice. Gamble responsibly; in Australia call 1800 858 858. --- ## Three exercises ### 1. Calibration check — when the model says 65%, does it happen 65% of the time? Keep rows where `model_home_win_prob` is filled. Sort them into buckets (50–55%, 55–60%, and so on). In each bucket, compare the average `model_home_win_prob` with the share of rows where `home_win` is 1. Plot one against the other with a 45-degree line. A model that tells the truth sits on the line. Start with MLB — it has the most rows. ### 2. Home advantage — how big is it, and where? For each league, take the average of `home_margin` and the share of games where `home_win` is 1. Then split by season month, or by whether the home team was the market favourite (`closing_home_handicap` below zero). How much of home advantage is just better teams playing at home? ### 3. Does the market beat the model? Keep rows where both `model_home_margin` and `closing_home_handicap` are filled. The market's expected home margin is `−closing_home_handicap`. For each, compute the average absolute error against the real `home_margin`. Whichever number is smaller was closer to the truth. Do the same for `model_total` versus `closing_total_line` against `total_points`. Report the counts alongside the errors — with a few hundred games, small differences mean very little.