Reading a crowd out of a prize table
Nobody publishes which numbers people picked. A prize table publishes something close: for each outcome, how many tickets reached it.
Write T for the number of lines in play and πm for the chance one line matches m of the drawn numbers. If lines were chosen uniformly, the winners at tier m would be Poisson with mean Tπm, and every tier would give the same estimate of T. They do not. The mid tiers run high on some draws and low on others, and the pattern follows which numbers came out.
Recovering the ticket count. One tier is immune to the crowd: on a game with a separate bonus pool, the tier that pays for the bonus ball and no main number is reached by a line's bonus choice alone, which carries no preference over the main pool. Its expected winners are Tπ0 regardless of what people like, so
T^=πbonus onlyWbonus only
is a ticket count uncontaminated by number preference. Games without such a tier fall back to the median across tiers, and everything downstream of that fallback is marked as carrying the crowd's own bias.
Weighting the balls. Give ball n a weight θn, mean one over the pool, and let a line's popularity be the product of its balls' weights. The expected share of tickets holding exactly the drawn set at tier m is then the elementary symmetric polynomial em of the drawn weights, normalised by over the pool. Fitting by ball is hopeless — on Powerball the first seven years' per-ball weights correlate with the last four, which is noise — so the weights are built instead from six named properties a number either has or does not:
logθn=∑jβjx
with xjn the indicator that ball n is a day of the month, a month, seven, one of 3/8/9, thirteen, or a multiple of ten. Six coefficients against thousands of tier counts, rather than sixty-nine.
Fitting. The likelihood is quasi-Poisson in the winner counts. T profiles out in closed form for each draw, so only the β remain, and they are found by coordinate ascent with a scan before each refinement rather than a Newton step: the curvature of a shared coefficient is not the sum of its per-ball curvatures, and a step built that way overshoots by up to a factor of thirty.
The overall level of θ is not absorbed by the ticket count. Scaling every weight by c multiplies tier m's expectation by cm, so the level is identified by the spread of tiers, and pinning it with a geometric rather than an arithmetic mean reported a true 1.30 as 1.62.
Honesty checks. Every standard error is widened by the square root of the dispersion, which runs near 250 rather than 1. A fit is published only if it beats a flat crowd on later draws it never saw, and only if at most one coefficient stopped at the edge of its range. Where it fails either test the page shows the published literature figures and says so.
Sales against the jackpot. Implied lines and advertised jackpot are regressed in logs, so a power law is a straight line and the exponent is its slope. Powerball measures 0.456 where the site had assumed 0.7.
What cannot be measured. A prize table counts tickets by how many of the drawn numbers they held. A property of the whole line — six in a row, an even step, every number a date — only distinguishes tickets at the jackpot tier, which is empty on almost every draw. Those three multipliers stay as published, and no quantity of extra draws will change that.
The only edge there is
There is exactly one thing about a lottery a player can change, and it is not the chance of winning.
Six numbers from forty-nine come up with the same probability whichever six you hold. What differs is who else is holding them. A prize is divided among the tickets that reach it, so a line nobody else played is worth more when it wins than a line half the country played — the same odds, a different prize.
Two draws show the size of it.
In March 2008 the Canadian Lotto 6/49 drew 40, 41, 42, 43, 44, 45. The obvious guess is that a run like that is played by nobody, being too tidy to feel random. Two hundred and thirty-nine tickets matched five of them plus the bonus, against the handful the ticket count predicted. People do not avoid a designed-looking line. They queue up for it.
In December 2020 South Africa's Powerball drew 5, 6, 7, 8, 9 with a powerball of 10. Twenty tickets held all six. A single draw with twenty jackpot winners is not a measurement of anything, and it is the reason this page will not put a number on what a whole line's shape is worth: that information lives only in the jackpot row, and the jackpot row is empty almost every week.
What the prize tables can measure is the crowd's taste in individual numbers, and here they are unambiguous. Group every draw by how many of its numbers a calendar date could have produced — nothing fitted, nothing modelled, just counting — and the winners climb with the count, rung after rung. Dates are what people reach for, and the top of the pool is where the crowd is thin.
None of which is a system. Playing above 31 does not make a win likelier by one part in a million; it makes a win, in the rare event of one, less likely to be split. That is a smaller claim than the industry usually makes, and it has the advantage of being true.