How props are scored

This page is loading — it needs JavaScript. The written page below is served without it.

Every number on this site is built the same way: a player’s official game log is counted against a specific line, shrunk toward the market in proportion to how little history stands behind it, and compared to the price a sportsbook is posting. This page walks that pipeline end to end, and states plainly what it is measured to be unable to do.

What the measurement showed

Across 96,155 graded props the model tracked the de-vigged consensus of five books to within the margin those books charge. Historical testing showed that calibration alone is not enough to overcome sportsbook pricing, and no configuration tested cleared that margin — which is why this product is built on price, context, sample quality and transparency rather than a claimed edge over the closing line.

The board hits about 50.9% where 52.4% is break-even at −110, against a measured margin of 3.49 points per side. It is tracking the market to within the cut the market charges. Nothing here is a lock or a guaranteed winner.

The pipeline, end to end

  1. Ingestion

    Scheduled jobs pull prices from seven sportsbooks and final box scores from each league’s own feed, then stage both into one database.

    Everything downstream is a comparison between a price and a game log, so both halves have to arrive on a schedule and land in the same place. Prices for near-term games refresh every two minutes; the day’s events are rebuilt hourly; box scores load four times a day once games are final.

  2. Normalization

    Every book’s own spelling of a market, a player and a side is mapped onto one internal vocabulary, so one prop is one row rather than five.

    One book calls it "Player Points", another "points"; the same player appears with and without a suffix. Without a single vocabulary the same prop would be scored five times as five different props, and a game log would never join to the line it belongs to.

  3. Historical game logs

    Each player’s per-game statistics are stored exactly as the official box score recorded them.

    A hit rate is only as trustworthy as what it counted. These are settled, official numbers — not projections, not live in-game feeds, and not anybody’s estimate of what a player was on pace for.

  4. Hit-rate calculation

    For a given line, we count the games in which the player actually cleared it, and publish that count over the number of games.

    A bet does not settle at the line, it settles one whole unit past it: Over 0.5 needs at least 1, not at least 0. Joining a game log at the line rather than at the settling threshold is how a routine prop comes out at a 100% hit rate and an enormous fake edge. The count is always shown with the denominator behind it, because a percentage without a sample is the same defect in a different font.

  5. Sample-size adjustment

    Two floors decide whether a number may be computed at all and whether it may be shown as a suggestion, and a third reading reports how much history is behind it.

    A rate over three games can only take four values, so a player who went 3-for-3 in a fortnight reads as a certainty. Rather than one blunt cutoff, the floors rise with the strength of the claim being made: a probability may not be computed below 5 games, and may not be recommended below 10. Above that the sample is reported rather than enforced — it sets a confidence reading shown beside the letter grade, because how much evidence exists and whether the price is good are two separate questions and answering both with one letter charges a prop twice for one fact.

  6. Bayesian shrinkage

    The raw hit rate is pulled toward the market’s own probability, by an amount set by how little evidence sits behind it.

    This is the single most important correction in the pipeline. The estimate is a blend of what the player has done and what the book thinks, weighted by sample: at k = 20, a player with 20 games of history is weighted half history and half market. Regressing toward the market rather than toward a coin flip is deliberate — the book has priced the starting pitcher, the injury report and the rest day, none of which a game log can see. Unshrunk, the same board averaged a +26% edge and topped out near +95% off two or three games.

  7. Sportsbook implied probability

    The posted American odds are converted into the probability that price implies.

    A price and a probability are the same statement in different units, and you cannot compare an estimate to a price until both are in the same units. An unusable or missing price returns nothing at all rather than a neutral 50% — a fabricated coin flip does not read downstream as missing data, it reads as analysis.

  8. Vig normalization

    The two sides of a market imply more than 100% between them; the excess is the book’s margin, and removing it is what turns a price into a fair probability.

    At −110 each way both sides imply 52.38%, summing to 104.76%. The 4.76 points of surplus are the book’s margin; removing them leaves a fair probability, and the book’s hold on that market is 4.76/104.76, or about 4.5%. This correction is now computed and published — every book quoting BOTH sides of a line is de-vigged and the median across them is shown on the card as the market’s fair price, and it is what the board is ordered on. It needs at least two books on both sides, so it is present on a majority of the board and simply absent on the rest rather than filled in with a single book’s opinion. What it deliberately does NOT do is move the letter grade: the edge that sets the grade is still measured against the raw, margin-inclusive price, which charges the model the book’s full cut on the side being valued. Two numbers with two authors, kept apart on purpose — see the limitation below.

  9. Price edge

    The edge is the shrunk estimate minus the probability the price implies, in probability points.

    This is the number the letter grade reads, and it is a description, not a verdict: it says how measured history compares to what is being charged for it. Positive does not mean the book is wrong. It is published to four decimal places under a stated rounding rule, because two implementations that round differently disagree about props sitting exactly on a boundary.

  10. Model ranking

    Rows that cannot support a recommendation are removed, and what survives is ordered by the market’s own de-vigged fair price against the price on the card — not by a Parlay Builder opinion, and not by the grade.

    Eligibility comes first: a prop has to clear a minimum edge, a minimum sample, a maximum price age and a plausibility cap before it is ranked at all. The cap matters most — past about 20 points, a stale price or a bad line join is a likelier explanation than an opportunity, so those rows are treated as data artifacts instead of being ranked at the top. What orders the survivors changed on 5 September 2026. It used to be the price edge above, which put a Parlay Builder reading at the top of the page; it is now the field’s own, arrived at by de-vigging the books that quote both sides and taking the median. Rows with no fair price sink to the bottom of that ordering rather than being scored zero, because "no two-sided consensus exists" and "the consensus says nothing" are different statements. The grade has never been the sort key and still is not.

  11. Correlation handling

    Legs from the same game are limited, and any combined probability quoted from multiplying legs is labelled as assuming independence.

    Multiplying leg probabilities assumes they are independent, and legs from one game rarely are. A three-leg slip built from a single game is one game’s worth of risk sold as three, and its advertised hit probability is simply wrong. So the builder caps legs per game and per player, and where a correlation-adjusted figure exists it is shown beside the independent one rather than replacing it silently.

  12. Final scoring

    The edge alone sets a letter from A to F. The sample behind the estimate and the count of agreeing form signals are reported beside that letter, on their own, and neither one moves it.

    Three readings, measured separately and never allowed to move one another, because a prop can be well priced on thin evidence with cold form and a reader is better served by three legible numbers than by one letter that has quietly netted them off. The letter thresholds are absolute rather than a curve, so a weak slate is supposed to grade badly — a percentile scale would mint the same number of A’s every night regardless of what was on offer. Form is counted rather than applied because grading on it would invert the scale: measured across the board, the form composite anti-correlates with the price edge. Until 26 August 2026 this description was not what the code did — a strong edge was demoted a letter when fewer than two form signals agreed or the sample was short, which counted the same evidence twice and contradicted this page. The code was changed to match the description, not the other way round.

  13. Historical grading

    Every prop the app surfaces is logged with its line, price and model numbers frozen at the moment it was shown, then settled later against the final box score.

    This is what makes the track record a record rather than a replay. Grading against the numbers as they were served is the only version that cannot flatter itself; a backtest run against today’s lines is a projection, and is labelled as one wherever it appears. A line that lands exactly on the number is a push, and a player who did not play is void — neither is counted as a loss.

Where this falls short

Calibration is not an edge over the closing line, and we measured that rather than assumed it
The replay covered 96,155 graded props, and every configuration tested landed inside the de-vigged five-book consensus rather than ahead of it: the board hits about 50.9% where 52.4% is break-even at −110, against a measured margin of 3.49 points per side. Historical testing showed that calibration alone is not enough to overcome sportsbook pricing. That is the ceiling this product is built to, and it is why what we sell is calibration, coverage, speed and the workings behind every number — not an edge over the market.
A grade is a statement about price, not a prediction about tonight
A high grade doesn’t guarantee a win; it means the model’s historical estimate sits above the probability the posted price implies. It is a reading of the price, and the letter is a step function of that one gap — the sample and the form agreement are printed beside it and never move it. It is also not a ranking of bet quality: when the ordering was tested against outcomes, no band of it returned a profit, and the largest claimed gaps were the ones that fell furthest short of their own estimate — the same result the replay above reports from the other direction. An A can lose and an F can win, and neither outcome says the grade was wrong. Nothing on the board is a lock, a guaranteed winner, or a claim that the book has made a mistake.
The edge is measured against the margin-inclusive price
Two different numbers on a prop card are computed against two different prices, and the difference is deliberate. The market’s fair price IS de-vigged — that is the whole of what it is, and it is what the board is ordered on. The EDGE, and therefore the letter grade, is not: it compares our shrunk estimate to the raw posted price, which charges the model the book’s full cut on the side being valued. That makes every edge on the board read worse than a fair-probability comparison would, and it is the main reason most of the board is negative. It is not that the correction is unavailable — the feed now carries both sides on a majority of served rows — it is that moving the grade onto a fair price would restate every published threshold and every measured result on this page at once, so it is a versioned change to the scoring contract rather than an edit.
The price a prop is SCORED against comes from one book
Seven books are collected and line-shopped, but the number the edge and the grade are computed against is one book’s price rather than a median across all seven. A single book’s price is that book’s opinion, so calling it consensus would overstate what is behind the number. The market fair price shown beside it IS a consensus and is the one figure on the card that is — a median across the books quoting both sides, which is why it is missing wherever fewer than two of them do.
A pulled market can sit on the board for up to six hours
Nothing in the feed marks a market as suspended. A market nobody is pricing any more simply stops refreshing and ages out on a six-hour staleness rule. That window matters because a market is most likely to be pulled exactly when it has become mispriced — an injury scratch, a late lineup change — so check the news before acting on anything here.
Whole-number lines are valued without a push term
Settlement handles pushes correctly, but the valuation maths treats every market as two-outcome. On player props this currently costs nothing, because the lines all end in .5 and a push is impossible. On whole-number spreads and totals it is a real distortion, and those figures should be read as approximate.
The two feeds do not always spell a player the same way
Prices and box scores arrive from different upstreams with different name conventions. Where the two spellings fail to join, a player can carry a hit rate assembled from the wrong games or no history at all. Accented names are the known weak point. This is a data-quality defect with a fix in progress, not an intended behaviour.
This is research tooling, and it is not advice
Parlay Builder does not take bets, hold money or settle anything. Nothing on this site is a recommendation to place a wager, and no output should be treated as financial advice. If betting has stopped being entertainment, help is available at 1-800-GAMBLER.

How the chat answers, and where a model is involved

The chat is not a language model with a database bolted to it. At most two calls to a model are made in a turn, both named below, and between them sits ordinary code: choosing what to look up, running the query, ranking the rows and writing the sentence are all functions, reading the same materialised views the rest of this site reads. That is why a number in an answer can be traced to the row it came from — there is no generation step between the table and the text.

The two model calls, in full

Working out what you asked for — claude-haiku-4-5-20251001 · AGENT_USE_LLM_ROUTER · 3s

One call reads your message and returns a label from a fixed list of intents, plus any names it picked out — a player, a market, a league, a number of legs.

What comes back is checked against that list and discarded if it is not on it, and the names are sanitised before they are allowed near a query. So this call chooses which piece of code runs next and what that code searches for; it does not produce a hit rate, a price or a sentence. If it is switched off, answers slower than three seconds, or comes back below the confidence floor, a keyword classifier takes over and the turn is answered the same way — the chat works end to end with no model call at all.

Reading a bet slip you upload — claude-opus-5 · AGENT_VISION_MODEL · 45s

If you attach a screenshot of a slip, one call transcribes it — the players, markets, lines and prices visible in the image — into structured fields.

This is the one place a figure in the conversation is produced by a model rather than read from a row, and it is worth being plain about: it is a transcription of your own screenshot, not a number we publish. Where the transcription is unsure of a load-bearing field it stops and asks rather than guessing, and what it returns is then treated as data — looked up against the same tables a typed question would hit, with anything instruction-shaped inside it classified rather than followed.

Every figure in an answer was read out of a row
A hit rate the chat quotes was counted from final official box scores in the same materialised views the board reads. A price came from the same odds feed the prop cards use. The chat runs the same queries the pages run — it is a different way in, not a second data set — so a figure it prints and the same figure on a prop card cannot disagree.
The sentences are written by code, not generated
Each kind of answer is a function that formats its own result, so the same question against the same board comes back in the same words every time. That is what makes the chat testable at all: an answer can be asserted character for character, which is not something a generated reply supports.
It is tested against the real function, not a description of it
A suite loads the actual edge function, points it at fixed board fixtures with the router switched off, and asserts what comes back. It has caught bugs users found first — a ranking that answered "nobody" in every league, a query against a table that never existed in production, and prop cards whose add-to-slip button was disabled by an eligibility check reading a field the feed does not send.
The claims an answer may not make are checked before release
An answer may never call a prop guaranteed, never call it a lock, never say it cannot lose, and never say it beats the books. Those patterns are run over every answer the suite produces and a match fails the build. The rule is scoped to assertions rather than to words, so copy that names one of those claims in order to refuse it still passes. To be exact about what this is: a gate on the code that writes the answers, run before release — not a filter inspecting live traffic.
An empty board is said out loud, not filled in
When there is nothing priced — an off-season league, a slate not yet posted — the answer says so and offers somewhere to go. It does not reach for an older number or describe a game nobody is pricing, and every case in the suite is checked for exactly that: a reply, never a dead end, and never a prop that is not on the board.

Traceability is the claim this design supports, and it is a precise one rather than a grand one: an answer can be no better than the row behind it, and every row is one you can go and check. Everything listed above under what falls short reaches the chat exactly as it reaches the board — the edge is still measured against a margin-inclusive price, a prop is still scored off one book, a pulled market can still sit there for hours, and a player whose two feeds spell him differently still carries the wrong history.

Questions about the grades

What does an A grade on a prop mean?

That the shrunk estimate of how often the player clears this line sits +3 points or more above the probability the current price implies. That is the whole rule — the sample behind the estimate and how much of the recent form agrees are printed beside the letter and do not move it. Read an A as a statement about the price, not about the bet: the grade describes where measured history and the market disagree, and when that ranking was tested against results it did not beat the market.

Why do most props grade C or below?

Because a two-way prop priced at −110 on both sides carries roughly a 4.5% hold for the book — that is the book’s cut on the market, not a per-side share of it — so a prop priced about where the market prices everything is genuinely a C. The thresholds are absolute rather than a curve: a weak slate is supposed to grade badly, and a percentile scale would mint the same number of A’s every night no matter what was on offer.

Is the grade the same as the signal score?

No, and they can disagree. The signal score measures how strongly a prop’s supporting signals agree — consistency, trend, streak, rest, matchup. The grade is set by the price. Measured across the board those two anti-correlate: props with the highest signal scores had, on average, the worst prices. So a green A above an amber signal score is not a contradiction; it says the price is good and the supporting form is thin.

Can a grade tell me a bet will win?

No. A grade describes how a price compares to measured history, and history is not a forecast. An A can lose and an F can win. The grade is there to tell you what you are paying for the chance, not what the chance will do.

What is Bayesian shrinkage and why is it applied to a hit rate?

It pulls a raw hit rate toward the probability the market is already implying, by an amount set by how few games stand behind it. At k=20 a player with 20 games of history is weighted half history and half market price. Without it a player who cleared a line in three of three games reads as a certainty and produces an enormous false edge — unshrunk, the same board averaged a +26% edge and peaked near +95% off two or three games.

Where do the hit rates and the odds actually come from?

Hit rates are counted from final official box scores — the same settled numbers a league publishes, never projections or in-game estimates. Odds are the prices actually posted at DraftKings, FanDuel, BetMGM, Caesars and BetRivers, refreshed every two minutes for games starting soon. A prop is joined to its game log at the settling threshold, so an Over 0.5 line is counted against games with at least one, not at least zero.

Can the AI chat make up a statistic?

No figure it prints about a player, a prop or a price is generated by a language model — each one is read out of a row in the same materialised views the board reads. Retrieval, scoring and phrasing are all deterministic code, so there is no generation step between the table and the sentence, and the same question against the same board comes back in the same words. The one carve-out is stated plainly on this page: if you upload a screenshot of a bet slip, a model transcribes it, and those numbers are a reading of your own image rather than anything published here.

How much of the AI chat is actually a language model?

Two calls in a turn at most. One classifies your message into a fixed list of intents and picks out the names in it, which decides what gets looked up; it is capped at three seconds and can be switched off entirely, and a keyword classifier answers the turn the same way when it is. The other runs only when you upload a screenshot, and transcribes it. Everything else — the queries, the ranking, the wording — is ordinary code, which is what lets the chat be tested against fixed board fixtures before release rather than sampled after it.

Does the model remove the sportsbook’s margin before calculating edge?

Not for the edge, and that is now a deliberate split rather than a missing capability. The de-vig itself is live: wherever at least two sportsbooks quote both sides of a line, each book’s pair has its margin stripped out and the median across those books is published on the card as the market’s fair price — and that figure, not the edge, is what the ranked board is ordered on. It is present on a majority of the board and simply absent where fewer than two books quote both sides, rather than being filled in from one book’s opinion. The edge, and therefore the letter grade, is still measured against the raw posted price with the margin left in, which charges the model the book’s full cut on the side being valued; that is why most of the board still reads negative. Moving the grade onto the fair price would restate every published threshold and every measured result on this page at once, so it is a versioned change to the scoring contract rather than a copy edit.

Odds and lines come from licensed data feeds and refresh on a schedule; hit rates, edges and grades are Parlay Builder’s own calculations.