SPOTTER.
How the published grades have done · 2026-07-29 – 2026-09-17
Methodology. One row per game and market: the final pre-pitch grade if one was written, otherwise the first. Every grade written enters this record and the record is never pruned — nothing is dropped for having been wrong. Grades are frozen before first pitch and settled by machine against the box score. Win % counts decided cards only; pushes are shown and sit out of the percentage. Samples under n=60 are labeled insufficient.
OVERALL
SEGMENTNWLPUSHWIN %NOTES
Every settled grade13666666247651.6%
of which graded — A, B or C6093012773152.1%
of which no lean — no band7573653474551.3%
The two indented rows sum to the first. Graded is the pick record — every card where the model claimed an edge, and the only population the band table below describes. No lean is every card where it agreed with the market and made no call.
BY MARKET
SEGMENTNWLPUSHWIN %NOTES
Full-game total6833283203550.6%
First-5-innings total6833383044152.6%
BY BAND
SEGMENTNWLPUSHWIN %NOTES
A grades813541546.1%
B grades21611394954.6%
C grades3121531421751.9%
AB combined — C set aside2971481351452.3%
The last row is a subtotal, not a fourth band. It is the A and B rows above added together — the population that would remain if C stopped being published. Read it beside those rows, never instead of them: a combined row cannot show whether the two bands rank in the order the grades claim, and that ordering is the question a ruling on the set-aside band turns on. Nothing has been dropped from this page — every band still has its own row above, and every band is still published.
BY MARKET AND BAND· the same grades, crossed — does the order hold?
This is the only table here that can show whether the bands rank. A band is a claim about confidence, so A should beat B and B should beat C — and that claim is about a market, not about the two markets averaged. The two tables above are the margins of this one: by market pools the bands together, by band pools the markets together, and either can hide an ordering that fails inside a book. The ruled row under each market is what setting a band aside would actually do in that market, which is a different question from what it does to the pooled record — and the one that a decision to stop publishing a band turns on. Read the sample sizes: these cells are the smallest on the page, and most of them are labeled insufficient for that reason.
SEGMENTNWLPUSHWIN %NOTES
Full-game total · A431922246.3%insufficient sample
Full-game total · B874341351.2%
Full-game total · C1437560855.6%
Full-game total · AB combined — C set aside1306263549.6%
First-5-innings total · A381619345.7%insufficient sample
First-5-innings total · B1297053656.9%
First-5-innings total · C1697882948.8%
First-5-innings total · AB combined — C set aside1678672954.4%
AGAINST THE CLOSE· the morning grade's line vs the same book's close
Against the close. The same games and markets as the pick record, scored a second way: did the line move toward the pick between the grade and the close? Measured from the morning grade of each game — its side, its band, its line — against the same book's close; a close from another book is not a move and is left unpriced. N is picks with a usable close; beat % is of the picks whose line moved at all — flat picks are shown and sit out of it, the way pushes sit out of win %. Avg move is runs toward the pick, and the two markets are never pooled: a run in five innings is not a run over nine. The morning band can differ from the final band above, so these cells do not reconcile to the pick record and are not meant to. Measured from the final grade instead — written about an hour before first pitch, with the close captured 35 minutes before — 606 of 643 full-game total picks and 641 of 678 first-5-innings total picks closed at the very same line, so that anchor can only say the line stops moving once lineups are in. The morning grade is the one with hours to move.
SEGMENTNBEATFLATWORSEBEAT %AVG MOVE
Full-game total · A35723558.3%+0.03 runsinsufficient sample5 unpriced
Full-game total · B705471821.7%-0.14 runs13 unpriced
Full-game total · C14522972645.8%-0.01 runs3 unpriced
Full-game total · all graded picks250341674941.0%-0.04 runs21 unpriced
First-5-innings total · A424281028.6%-0.07 runsinsufficient sample6 unpriced
First-5-innings total · B8611651052.4%+0.01 runs10 unpriced
First-5-innings total · C176291222553.7%+0.01 runs15 unpriced
First-5-innings total · all graded picks304442154549.4%-0.00 runs31 unpriced
SEGMENTNBEATFLATWORSEBEAT %AVG MOVE
Full-game total · no lean355622237047.0%-0.01 runs57 unpriced
First-5-innings total · no lean316332443945.8%-0.00 runs32 unpriced
BY MODEL VERSION· the same picks, split by the model that graded them
The same picks, split by the model that graded them. Every table above pools every model that has published — that is the honest record of what went out, and it is never pruned. But a pooled cell cannot say how the current model is doing on its own, and after a model change that is the question. Each block here is one model's public grades over the slates it graded; the newest block starts at zero on its first slate and fills nightly. Read the sample sizes: a fresh block is labeled insufficient for weeks, and that label is the point — nothing here is a verdict until it clears the floor. Shadow grades never appear on this page.
Current model (since Sep 4 — First-5 umpire fix) · slates 2026-09-04 – 2026-09-17
SEGMENTNWLPUSHWIN %NOTES
Full-game total · A1339125.0%insufficient sample
Full-game total · B271215044.4%insufficient sample
Full-game total · C512125545.7%insufficient sample
First-5-innings total · A724133.3%insufficient sample
First-5-innings total · B351617248.5%insufficient sample
First-5-innings total · C392613066.7%insufficient sample
All graded picks1728083949.1%
Aug 8 – Sep 3 model (v2) · slates 2026-08-08 – 2026-09-03
SEGMENTNWLPUSHWIN %NOTES
Full-game total · A22138161.9%insufficient sample
Full-game total · B462618259.1%insufficient sample
Full-game total · C925435360.7%
First-5-innings total · A251013243.5%insufficient sample
First-5-innings total · B673826359.4%
First-5-innings total · C994152644.1%
All graded picks3511821521754.5%
Jul 29 – Aug 7 model (v1) · slates 2026-07-29 – 2026-08-07
SEGMENTNWLPUSHWIN %NOTES
Full-game total · A835037.5%insufficient sample
Full-game total · B1458138.5%insufficient sample
First-5-innings total · A642066.7%insufficient sample
First-5-innings total · B271610161.5%insufficient sample
First-5-innings total · C311117339.3%insufficient sample
All graded picks863942548.1%
UNBANDED
SEGMENTNWLPUSHWIN %NOTES
No lean — the market's right7573653474551.3%
This is not a track record. These are the no-lean cards: the model looked and agreed with the market, so nothing here was ever a recommendation. Settlement still writes a W or an L against the model's nominal side, which is bookkeeping, not a call. It is published because leaving it out would let the graded tables stand for the whole population — but do not read it as a record of anything.
BY SIDE — GRADED PICKS· over and under, split by market
SEGMENTNWLPUSHWIN %NOTES
Full-game total · Over1377358655.7%
Full-game total · Under1366465749.6%
First-5-innings total · Over2761311301550.2%
First-5-innings total · Under603324357.9%
BY SIDE — A AND B ONLY· the bands the model rates highest · C set aside
The grades the model rates highest, over and under. A and B are the bands it claims the most for, so this is the block to read for how the picks worth acting on have gone, market by market and side by side. It is also the answer to a second question — what publishing would look like with C dropped — because it is exactly the block above minus each row's C grades. A subset of the pick record, not a second record: none of it was ever bet as its own book, since every card in it went out on a slate that published C alongside it. Two things it cannot tell you: whether A and B rank against each other, or whether the cut behaves the same way in both markets — the market and band table above is where both of those live. And the cut is thin: read the sample column before you read a rate off a row.
SEGMENTNWLPUSHWIN %NOTES
Full-game total · Over612634143.3%
Full-game total · Under693629455.4%
First-5-innings total · Over1437263853.3%
First-5-innings total · Under24149160.9%insufficient sample
BY SIDE — NO LEAN· over and under, split by market
Not picks. On every card below the model looked and agreed with the market, so it made no call. The side shown is the model's nominal lean — which way it leaned while claiming no edge — and it was never a recommendation. This block is split by side because the nominal lean has so far been predictive, which is an open question under test, not a finding. Read the rows as evidence about that question. Do not read them as a record.
SEGMENTNWLPUSHWIN %NOTES
Full-game total · Over2181001071148.3%
Full-game total · Under19291901150.3%
First-5-innings total · Over218109971252.9%
First-5-innings total · Under12965531155.1%
BY STARTING-PITCHING AGREEMENT· did the starters argue for the pick? · registered watch
A registered watch, not a rule. On 2021-2024 out-of-sample grades, a lean the starting-pitching factors argued FOR won about 54% in both markets, and a lean they argued AGAINST won 47-49% — a gap well beyond chance in each market (reports/CONVICTION_READ_2026-09-03.md). The first five live weeks did not show it; First-5 ran the other way. Nothing is filtered on this. The rows are here so the October read has a live record to hold against the bars declared in advance: for-minus-against of at least 3.0 points in both markets, 300 grades per arm, on the 2025 look. A pick whose five drivers do not include the starters is counted separately, not as either side.
SEGMENTNWLPUSHWIN %NOTES
Full-game total · starters argued for the pick1598072752.6%
Full-game total · starters argued against it281314148.1%insufficient sample
Full-game total · starters not among the five drivers864437554.3%
First-5-innings total · starters argued for the pick2581211241349.4%
First-5-innings total · starters argued against it492721156.2%insufficient sample
First-5-innings total · starters not among the five drivers29169464.0%insufficient sample
SERIES_UNDER — THE PERSONAL RULE· not a Spotter grade, and not a bet log
A different thing from everything above. SERIES_UNDER is the owner's own rule, registered and screened before any 2026 data was looked at: when a game finishes at least 4 runs under its closing total, back the under in the same series' next game. The model does not produce it, no flag carries it, no grade is written for it, and nothing on the slate board acts on it — the pipeline sends the owner a nightly note and stops. It is here because it fires on Spotter's own games and the triggers are reconstructible, so its record can be shown honestly instead of remembered.
POPULATIONNUNDEROVERPUSHUNDER %ROI
Every trigger552924254.7%+5.9%
of which severe — 5+ runs under301612257.1%+10.0%
What the rows are. 57 triggers since 2026-08-05 · 55 graded · 2 still awaiting a closing total. Every graded row is the under at that game's own closing total, one flat unit at the actual closing price, settled off the box score — the same arithmetic the screen ran. They are not bets. Nothing in the bet log is a SERIES_UNDER position; no money moved on any row in this table, and the ROI column is what a flat unit on every trigger would have returned, not what was won.
And the sample is far too small to mean anything yet. On 53 decided triggers the 95% interval around 54.7% runs 41–67% — it contains a coin flip, it contains the screened edge, and it cannot tell them apart. The registration's own minimum was 400 triggers, and the screen needed 1,261 to clear its bars: 54.61% under and +5.63% ROI (the severe subgroup 56.06% / +8.32% on 737), against a 52.38% break-even. The only claim this table supports is the weak one: so far it is tracking the registration rather than diverging from it. That is not evidence the rule works — it is the absence of early evidence that it does not.
Degraded window · 2026-07-29, 2026-07-30, 2026-07-31. Every grade on those dates was produced while the boxscore-context collector was dark — no fielding row had landed since 2026-07-12, so the 30-game fielding fold stayed frozen there. The 2026-08-01 re-parse refilled the fold; 2026-08-01 is the first slate graded against it. model_grades is append-only and nothing re-grades, so these rows keep the numbers they were written with — the repair fixes the slates after it, not these. They are counted above, labeled here.
Degraded window · 2026-07-29, 2026-07-30, 2026-07-31, 2026-08-01, 2026-08-02. Every grade on those dates was produced while the IL collector was dark — no stint had been fetched since 2026-07-07, and a stint that was never fetched reads as no stint, so a starter just back from the injured list scored as healthy. It misread 11 of the 71 games on these dates — 12 starter-sides that should have carried an IL-return flag and did not. The 2026-08-03 repair refilled the stints; 2026-08-03 is the first slate graded against them. model_grades is append-only and nothing re-grades, so these rows keep the numbers they were written with — the repair fixes the slates after it, not these. They are counted above, labeled here.
What this counts: every grade Spotter wrote, settled against the box score — winners and losers in the same table, and no row ever removed. Model fit 2021–2025, leave-one-season-out; F5 grades ride a shorter history (2023–2025). No grade is ever re-graded; only its result is filled in after the game. · Research, not betting advice. 21+. Problem? 1-800-GAMBLER.