Season 1 · 2025 · v1.0 Historical baseline. Weeks 9 and 11 only, 15 games, 86 bets. Models: 2025 launch-era ChatGPT, Claude, and Gemini. Kept public because the corrections history is the point.
Season 1 · 2025 Archive

Where the study started.

The 2025 sample was a senior-year experiment, run manually and after the fact for 15 games across two weeks of the season. It caught model-assignment errors, then a second correction pass caught bet-grading and metadata errors against the actual ESPN box scores. Both correction rounds are preserved here because the fixes are what turned the project into a real study.

Games Graded
Weeks 9 & 11
Bets Recorded
All models combined
Corrections
11
Preserved with original CSV
Models
3
2025 launch-era
Season 1 games
The 15 archived matchups
Full bet database →
Season 1 model P/L
How the 2025 launch-era models finished
Prop win rate by category
Corrections
The 11 grading fixes

Every correction is a bet whose outcome was wrong on first pass, later fixed after checking the actual ESPN stat line. The original CSV value is preserved next to the fix so nothing is silently overwritten. This log is the reason the site can be trusted with its own numbers.