Enter the Lab →
Methodology

How ET Parlays works, and where it is wrong

By ET · Updated 2026-09-09
The voice behind ET Parlays. Grades the props and teaches the math the books don’t itemize.

ET Parlays rates player props against its own model, grades every rating against the official box score, and publishes the result whether it flattered us or not. This page is the receipt. It covers where the data comes from, how a read is produced, how it is graded, how much of the board never gets graded, and every place our numbers are currently wrong.

We are not a sportsbook. You cannot place a bet here, we hold no money, and no sportsbook pays us. We publish research, and research is worth exactly what its track record says it is worth.

What ET is, and what it is not

ET is a research tool. It reads a slate, rates props, and explains the reasoning in plain English. It does not tell you what to bet.

We call our outputs reads, not picks, and the distinction is the point. A pick is somebody’s bet. A read is a piece of research with a number attached and a record behind it. You are meant to disagree with it sometimes.

Three things we will not do, so you can hold us to them:

Where the numbers come from

Two kinds of data go in, and they are not equally interesting.

For grading, we use the official league feed for that sport, and nothing else. For MLB that is the same box score the league publishes. We never grade a read against our own copy of a result, our own scraper, or a third party’s summary. If our record says a player got two hits, the league says he got two hits, and you can check it in ten seconds.

That is the part that decides whether the tables on this page are worth anything, so that is the part we pin down.

For producing a read, we use a mix of public league data and a commercial odds feed. We do not itemise those sources, for the same reason we do not publish the model: naming the ingredients is most of the recipe. Treat that as a real limitation rather than a footnote. You can audit our outputs, not our method, and this page is built so the outputs are enough.

How a read is produced

Every read is a probability first and a letter second.

  1. A model estimates the chance that one player clears one line in one specific game.
  2. That probability is compared against the market price for the same prop, where we captured one.
  3. The probability is banded into a letter, S through C, so you can scan a board without reading twenty numbers.
  4. The read is written to the database with its probability, its line, its side and the game date, before the game starts.

That last step is what makes the record checkable, and it is the only step here we would defend at length. A read cannot be edited after the fact, because grading looks up the row that was written before first pitch. Nothing in the tables below is a bet we decided we had made after we saw how it went.

How a read is graded

For MLB props, a nightly job pulls every box score for the slate and compares the actual stat to the line on the row.

That last rule matters more than it sounds. A did-not-play is not a wrong read. Scoring it as a loss would make us look worse than we are, and scoring it as a win would make us look better. It is removed.

What counts as graded, and what does not

The whole board, read from the database when this page last rebuilt.

Reads written since 2026-04-14111,114
Reads on games already played110,790
Graded79,259
Voided, player did not play3,957
Played, ungraded, no explanation on the row27,600

24.9% of the board we have already played has never been graded. That is the least flattering number on this page, so it sits above the calibration tables rather than below them.

Why those rows are missing

One-off audit · September 9, 2026 · not live

Most of them are players who never played. We rate a slate before lineups are posted. When a player we rated does not end up in the lineup, and the automatic void rule does not catch it, the row sits there unresolved.

On the September 9, 2026 audit, of 24,684 ungraded player prop rows, 19,031 belonged to a player with no graded row at all that day. For 87.9% of those, no lineup spot was ever recorded, against 52.3% for rows that did grade.

The remaining 5,653 are different, and worse. Those players did play. They graded in another market the same day. Their rows are gradeable and missing.

Two families are graded, but not in the record this page uses

One-off audit · September 9, 2026 · not live

This is a bookkeeping defect, and it is the least defensible thing the audit found. No runs first inning had 1,124 graded outcomes in its own table and zero in the unified record. Game winner had 896 of 896 graded in its own table and 804 of 2,221 in the unified record.

The outcomes exist. We know what happened in those games. When the grader finishes a market it copies the outcome into the unified record as a second, best-effort write, and it throws away the result of that write without checking it. For no runs first inning, that copy has never landed once in any month since April, and nothing in the system could have told us, because the code discards the answer.

We found it by auditing this page rather than by an alert firing. That is the honest version of how it surfaced.

The part that should worry you most

One-off audit · September 9, 2026 · not live

The ungraded rows are not a random sample of the board. Grading is more complete on the reads we rated highest. On the audit, S reads were 8.8% of the graded set against 3.1% of the ungraded set, and C reads were 54.4% against 70.9%. The missing rows are more C-heavy and slightly lower-probability than the graded record.

What that means for the tables below: they are computed on a subset that over-represents our confident reads. The missing rows have no outcome, so we cannot know directly whether they would have landed at the same rate.

What we can say is more specific, and it comes from asking why a row goes ungraded rather than what the missing rows look like. The tables are banded by the probability we stated, not by the letter, so a shift in the mix of letters cannot move a band. For our record to be flattered, the missing rows would have to be losers within their own probability band. The mechanism that makes a row go missing is that the player was not in the box score, and whether a man was in the lineup is a fact about the lineup, not about whether the read would have won.

That is a mechanism rather than a resemblance, and it is the reason we are willing to publish the tables at all.

Where that argument does not hold: the 5,653 rows belonging to players who did play. We expect that group to skew toward losses, because the usual reason a hitter finishes with one at bat is that he came in late. If that is right, grading them will make the numbers on this page slightly worse, not better. We would rather tell you that before we do it than explain it afterwards.

What the record actually says

Every table below is read from the graded record when this page rebuilds. Each one bands the reads by the probability we stated, then shows what happened in that band. A band appears only once it has at least 30 graded outcomes. Season window, calibration version 2026-09-09.

Home runs, 23,943 graded

We saidLandedGapReads
3.5%4.5%+1.0599
7.2%7.6%+0.43,343
12.0%11.3%−0.73,861
17.1%14.6%−2.53,999
22.0%15.9%−6.02,735
27.2%17.0%−10.23,205
31.9%16.8%−15.12,719
36.7%18.4%−18.31,441
42.0%21.0%−21.1930
46.8%20.7%−26.1484
50.0%23.1%−26.9627

One or more hits, 16,634 graded

We saidLandedGapReads
42.6%45.3%+2.7106
47.5%54.3%+6.8521
52.4%53.6%+1.21,524
57.2%55.1%−2.13,076
62.0%59.1%−2.94,141
66.8%61.2%−5.63,473
71.9%63.1%−8.81,998
76.6%63.3%−13.31,280
81.3%61.8%−19.5429
85.7%81.4%−4.386

1 further band is withheld for having fewer than 30 graded outcomes.

Total bases, 16,232 graded

We saidLandedGapReads
17.7%29.3%+11.682
22.6%28.1%+5.5342
27.4%28.6%+1.31,048
32.3%32.6%+0.43,519
36.9%35.5%−1.33,888
41.9%38.2%−3.73,016
47.1%39.9%−7.23,938
51.6%32.9%−18.8280
55.9%43.8%−12.164
67.0%29.1%−37.955

7 further bands are withheld for having fewer than 30 graded outcomes.

Hits plus runs plus RBIs, 13,916 graded

We saidLandedGapReads
27.5%32.4%+4.868
32.6%37.5%+4.9283
37.4%39.4%+2.11,009
42.3%41.6%−0.72,316
47.0%45.3%−1.63,328
52.0%49.0%−3.03,182
56.9%51.0%−5.92,154
61.7%47.0%−14.71,154
66.6%48.5%−18.1355
71.2%56.7%−14.467

2 further bands are withheld for having fewer than 30 graded outcomes.

Hits plus runs plus RBIs at the higher rung, 1,520 graded

We saidLandedGapReads
72.3%63.9%−8.4133
77.2%72.4%−4.8344
82.1%69.8%−12.3616
86.5%71.3%−15.2362
90.8%73.8%−17.065

2 further bands are withheld for having fewer than 30 graded outcomes.

Moneylines, 770 graded

We saidLandedGapReads
52.1%49.9%−2.2527
57.2%59.0%+1.8195
62.0%60.4%−1.648

3 further bands are withheld for having fewer than 30 graded outcomes.

Pitcher outs, 615 graded

We saidLandedGapReads
62.5%38.2%−24.368
67.1%47.0%−20.1100
72.4%59.6%−12.8114
77.0%33.3%−43.696
82.2%63.9%−18.372
87.4%50.0%−37.436
92.9%58.1%−34.931
99.1%53.1%−46.098

1 further band is withheld for having fewer than 30 graded outcomes.

First five innings, 301 graded

We saidLandedGapReads
52.3%56.4%+4.1149
56.9%38.0%−18.9108
61.9%47.7%−14.244

2 further bands are withheld for having fewer than 30 graded outcomes.

4 further markets have a record too thin to publish, so they show nothing here at all: k · NPB · Tennis (ATP) · Tennis (WTA). The floor is the same 30 outcomes, and it applies to us when it is inconvenient.

Read the pattern, not the rows

Where we say something is unlikely, we are close to right. Home run reads in the 5% to 15% range land within about a point of the stated number, across thousands of reads. That is the largest and most reliable part of our board.

Where we say something is likely, we are too confident, and the gap grows with the confidence. A home run read we called a coin flip lands about a quarter of the time.

The shape is the same in every family with enough data to check. This is a known statistical failure mode. When you rank thousands of candidates and keep the top ones, the top of the list is partly genuine quality and partly luck that has not reverted yet. The luck reverts. We have now measured the same effect in three separate places in our own product.

Anyone telling you their high-confidence selections outperform their stated probability, without showing you a table like the ones above, has not checked.

What we know is wrong right now

  1. Confidence runs hot in every family, and worse the higher the number goes. Shown above.
  2. The pitcher outs model is the worst offender. In its top band, 98 reads stated about 99% and landed 53%. Treat outs reads as directional only. We are rebuilding it.
  3. 24.9% of the played board is ungraded, and the missing part is not random. Shown above.
  4. Two markets are graded but never reach this record. A silent write that has never been checked. Being fixed.
  5. Strikeout reads have no published record right now. We found the model running about 9% hot, rebuilt it, and it went live on 2026-09-08. A recalibrated model’s record starts the day it first scored a board. Rows from the old model describe a model that no longer exists, so we deleted its record rather than let it flatter or damn the new one.
  6. Our one apparent edge over the sportsbook was selection. Measured 2026-09-09 and written up above. Price capture covers 22.3% of graded reads and is heavily skewed toward reads we rated likely.
  7. We do not publish our method. Outputs are auditable, inputs are not.

Why there is no single accuracy number on this page

We could compute one. It would be a good number. It would also be wrong.

Pool every family together and you get a flattering score that mostly measures something trivial: that home run reads are rarer than hit reads. Split the same data by family and the apparent skill largely disappears. The pooled figure describes our board’s composition, not our judgment.

This is a well-known trap, and it is exactly the number a founder reaches for. So we wrote a rule for ourselves: no pooled performance metric, ever, on any surface. If you see one on a competitor’s site with no per-market breakdown underneath it, that is what it is.

What happened when we checked whether we beat the book

One-off audit · September 9, 2026 · not live

We had one number that looked like an edge. On hits, our Brier score was 0.2416 against the book’s 0.2490 across 10,096 graded reads that carried a captured price. Lower is better. That is the number a founder puts on the homepage.

We ran the check anyway. It does not survive.

Prices are captured on 22.3% of our graded reads, and that fraction is not a random sample. Coverage runs from 0.1% of the reads we rated under 10% to 66% of the reads we rated above 90%. The priced set has a mean stated probability of 57.8%. The unpriced set, 38.3%. We were not comparing ourselves to the book across our board. We were comparing ourselves on the part of the board where we had already decided the answer was likely.

Hold the stated probability fixed and the edge goes away. Inside matched bands we are better in four and the book is better in four, but two of our four are differences of two ten-thousandths. Our average winning margin is 0.0011. The book’s is 0.0171, fifteen times larger. Across the nine markets with enough priced reads to check, the book’s Brier is better in seven.

The comparison is also rigged in our favour, and it still loses. The book’s implied probability includes its own margin, which makes the book’s forecast look worse than the book’s actual belief. We could not remove it, because that needs both sides of the same prop captured together and ours are not paired. The real gap is wider than the numbers above.

There is a third problem and it is the one that would have been hardest to spot. Whether a read got priced is correlated with whether it won, even inside a single probability band. In the band where we stated above 90%, priced reads landed 52.3% and unpriced reads landed 76.5%. We do not yet know why, and until we do, no comparison drawn from the priced subset means very much.

That is why “we beat the sportsbook” appears nowhere on this site. The check is rerunnable and it lives in our repository rather than in someone’s notebook, so the day it passes will be a fact rather than a feeling.

What we do about it

The measurement above is not filed away. It runs the product.

Every night we rebuild a calibration table from the graded record, banding reads by the probability we stated and recording what happened in each band. When Slip Check or the public checker at /check shows you a probability, it has been mapped through that table first. The number you see is the number our own record supports, not the number the model produced.

Two rules constrain that layer, and both exist to stop it being used as a marketing device. They are promises you can hold us to:

Ten questions to ask any betting site, including this one

Borrowed from David Spiegelhalter’s The Art of Statistics and answered here about ourselves. Take them to any site making claims about accuracy.

  1. How rigorously was this done? Complete count of the graded record, no sampling. But a quarter is ungraded.
  2. What is the uncertainty? Every band above carries its sample size. Bands under 30 are not shown.
  3. Is the summary appropriate? Banded by market, never pooled. See the section above.
  4. Is the source reliable? We grade ourselves, using official league box scores. You should weigh that.
  5. Is it being spun? The most flattering framing available to us is that our low-probability reads are accurate within a point. It is true. It is also the least useful part of our board for most bettors, and we said so.
  6. What are you not being told? That the ungraded part is not random, and that we do not publish our method.
  7. Does it fit what else is known? Yes. Overconfidence at the top of a ranked list is expected and well documented.
  8. Is an explanation being claimed? No. The tables describe what happened, not why any single read hit.
  9. Is it relevant to me? Only if you were going to trust a confidence label. That is what it measures.
  10. Is the effect important? The ordering is real and useful. The stated probabilities at the top are not, yet.

How to check us

What changes this page

The tables are read live. The defect list changes immediately in both directions: when one is fixed, it moves to a dated entry rather than disappearing.

We publish the failure before we publish the fix. It is the only order that lets you check whether the fix worked.

Nothing here is betting advice, and no research changes the fact that these are negative-expectation wagers. Bet only what you can afford to lose. If it stops being fun, take a break.

Questions

Is ET Parlays a sportsbook?

No. You cannot place a bet on ET Parlays. We do not hold funds and we are not paid by sportsbooks. We publish research about props you bet somewhere else.

How accurate is ET Parlays?

It depends on the market and the confidence level, which is why this page carries a table rather than a number. Our low-probability reads land within about a point of what we state, across thousands of graded outcomes. Our high-confidence reads land well below what we state. Both are on this page, computed from the live record.

Do you delete losing reads?

No. Every graded read stays in the record. Reads where the player did not play are voided rather than scored, and that rule is stated on this page so you can judge it.

Is this a tout service?

No. We do not sell a slate of picks or claim a winning record. We rate props, publish how those ratings have performed, and let you decide.

How often is this page updated?

The tables are read from the live record every time the page rebuilds, at most an hour old. The defect list is updated whenever one is found or fixed.

Why will you not publish your model?

Competitive reasons, and we would rather say that than dress it up. It is a real limitation on how far you can audit us, which is why every output on this page is checkable even though the method is not.