Runs via Codex CLI, which gives ChatGPT local repo access. Capital preservation, break-even math, sizes to edge over conviction.
The 2026 season is live.
Same rules, blind test, real box scores. ChatGPT, Claude, and Gemini each get a $20 hypothetical bankroll per game, decide their own allocation, and get graded against ESPN every week. This site is the public record of what they picked, why, and whether they were right.
Every model reads the repo
Local filesystem where the model supports it, raw GitHub URL fetch where it does not, plus open-web research on every matchup. Every model is expected to reflect on its own past picks and adjust.
$20 per game
Each model decides its own allocation across straight bets, parlays, SGPs, or reserve. Reserve is a valid answer. Sizing itself becomes a graded signal.
Locked before kickoff
Every pick is timestamped and pushed to this repo pre-game. No post-hoc editing. Raw model responses are saved verbatim in `Docs/Responses/`.
Graded against ESPN
Final grading uses the ESPN box score, not memory or headlines. Every prop bet is audited against a real stat line. Reasoning is graded separately from outcome.
Reads the repo directly, reviews its own past picks in NFL_BETS, cites factors from what it saw.
Same repo, same reflection on past picks, fetched through raw GitHub URLs.
Yesterday's picks retired before kickoff
The 2026 Week 1 lockup on Sep 8 did not honor the blind-test intent. New prompt system published today, three lanes, strength-honoring, fair-comparison baseline preserved.
GradingReasoning graded separate from outcome
A model can pick wrong and still grade well if it named a real factor. A model can pick right on a vibe and grade low.
IterationEvery prompt change logged
Templates carry a version. Migrations preserve old records. Nothing gets silently overwritten.
Season 1The 2025 baseline, warts and all
15 games, 86 bets, 11 corrected props. Kept public because the corrections story is the point.
AboutIndependent project, no affiliation
Not affiliated with the NFL or any model provider. Not financial advice. Study and portfolio purposes.
Season 2 begins the encyclopedia. Each future year captures a fresh generation of models against the same rules, so this site grows into a record of how AI research assistants evolve.