Loading…
Reading the draws and working out what chance predicts.
Loading…
Reading the draws and working out what chance predicts.
The audit, live
These are the tests statisticians use to audit national lotteries and certify random number generators, running on every draw Cash Pop Late Night has made in this era. Seven instruments, and the reading you want from all of them is nothing at all.
Reading the morning draw only — 1,212 draws. This game also runs 4 separate series (late_night, matinee, afternoon, evening), left out on purpose: runs and autocorrelation only mean something inside one draw series.
All 8 needles at rest. Boring is the correct reading.
And here is what that rules out. Over 1,212 draws a ball favoured enough to appear 41% more often than its share — 114 times instead of 81 — would have shown by now, four times in five. Anything subtler is still invisible, and “no bias found” means exactly that and nothing more.
Consistent with a fair machine. The variation between numbers is exactly the size randomness produces.
Draw totals alternate around their midpoint as often as chance predicts — no streaks, no drift.
The most recent window spreads across the pool exactly as widely as a fair machine does.
No draw carries information about the next. The strongest correlation across every lag tested is what noise looks like.
No weekly rhythm. No lunar cycle. No message. The strongest cycle present is the size noise always supplies.
Repeats from one draw to the next arrive at exactly the rate the hypergeometric predicts — the machine is not avoiding its last result, nor favouring it.
Final digits appear as often as the pool allows — including the zeros, which are genuinely rarer.
No single number keeps a rhythm of its own. The strongest cycle in any one ball’s history is the size the search itself produces.
Five of the eight instruments are judged against 200 simulated fair histories of this game’s own length and format, rather than against a textbook tail. The memory, final-digit and repeat instruments each had an approximation standing in for a null — a look-elsewhere correction that assumed independent lags, a dependence factor borrowed from a different test, and an asymptotic tail on dependent cells — and the simulation retires all three.
Cash Pop Late Night1/15 era · since Apr 20221,212 draws
43 p-values — every instrument, every lag, every number. Under an honest machine they are spread evenly across the whole range, and this histogram is flat to within p = 0.824 Which is the reading a working lottery and a working analysis both produce.
Cash Pop Late Night1/15 era · since Apr 20221,212 draws
Cash Pop Late Night1/15 era · since Apr 20221,212 draws
Cash Pop Late Night1/15 era · since Apr 20221,212 draws
The report
Randomness audit
1/15 era · since Apr 2022 · report as of 2026-09-22
| test | statistic | p | p adjusted | reading |
|---|---|---|---|---|
| Number frequencies | 13.94 χ² | 0.908 | 0.908 | consistent |
| Runs above and below the median | -1.25 z | 0.210 | 0.657 | consistent |
| Spread of the numbers (entropy) | 2.72 nats | 0.488 | 0.780 | consistent |
| Memory between draws | -2.28 z | 0.687 | 0.864 | consistent |
| Rhythm in the calendar | 0.01 g | 0.281 | 0.657 | consistent |
| Numbers repeating from the last draw | 0.18 χ² | 0.756 | 0.864 | consistent |
| Final digits | 4.53 χ² | 0.299 | 0.657 | consistent |
| Rhythm in any one number | 0.02 g | 0.328 | 0.657 | consistent |
Across 8 tests on 1,212 draws, no statistically significant departure from fair, independent drawing was found after controlling for the number of tests performed. The results are consistent with fair draws.
This analysis cannot prove fairness, only fail to find deviation. Absence of a detected deviation is not proof that none exists — a small bias needs a long record to surface, and detection time grows as the square of how small it is.
Provenance: 1,212 draws come from a single feed and have not been independently cross-checked. A test battery can only ever be as sound as the record it reads.
Method: Pearson χ² on number frequencies with the Joe (1993) within-draw dependence correction; Wald–Wolfowitz runs test on draw totals; rolling Shannon entropy with the Miller–Madow bias correction against a seeded simulated band; autocorrelation to lag 20 against ±1.96/√D; Fisher's test on the periodogram's largest peak; χ² on the hypergeometric repeat distribution; χ² on final digits against the pool's exact digit baseline. Battery-level false discovery rate controlled by Benjamini–Hochberg.
Now rig it yourself
The instruments above found nothing, which is only worth something if they would have found something. Here they are watching a machine you control — including two rigs they cannot catch.
The 1980 Pittsburgh fix: all but a few balls weighted so they cannot rise. A gross rig, and the instruments find it almost immediately.
4 of 8 instruments sit outside the calm zone after adjusting for how many tests are on this panel. That is a statistic, not a finding — read the caveats below before it becomes one.
The numbers are not coming up equally often — this is the pattern a biased machine would leave, and worth investigating.
Draw totals alternate around their midpoint as often as chance predicts — no streaks, no drift.
The latest window is more concentrated or more even than a fair machine usually manages.
No draw carries information about the next. The strongest correlation across every lag tested is what noise looks like.
No weekly rhythm. No lunar cycle. No message. The strongest cycle present is the size noise always supplies.
Repeats arrive at a rate further from expectation than usual.
Final digits fall further from their pool baseline than usual.
No single number keeps a rhythm of its own. The strongest cycle in any one ball’s history is the size the search itself produces.
The two layers underneath
Every needle on this panel answers one question, and it is not "is this lottery fair?". It is: if the machine were honest, how often would a reading this extreme turn up? That number is the p-value on the bezel.
This matters more than it sounds. A test can find a deviation, or it can fail to find one. It can never confirm there is nothing to find, because "nothing to find" and "something too small to see with 1,212 draws" produce the same calm needle. The Integrity Report is worded to say exactly that and nothing warmer.
The frequency test compares how often each number appeared against . The obvious statistic is Pearson's:
read against on degrees of freedom. That is wrong here, and wrong in the direction that flatters the lottery.
The numbers within one draw are not independent trials — they are drawn without replacement, so a number appearing makes the others slightly less likely. Each count is Binomial rather than multinomial, giving
against the a expects. So the statistic runs systematically low, and rescaling by puts it back.
This is the practical form of the modified statistic Joe (1993) derives for lotto draws, and skipping it is not a rounding error. Measured over 800 simulated honest eras of 15 numbers: the uncorrected test rejects at 1.4% where it should reject at 5% — under-powered by more than threefold — and declares honest lotteries "too good to be true" 10.9% of the time instead of 5%. Both mistakes make a lottery look better than it is.
The sandbox's 5% preset is there to be disappointing. Detection time scales like : halving a bias quadruples the record you need to catch it. A gross rig like a painted ball set is obvious within dozens of draws; a five percent lean survives hundreds and needs tens of thousands.
The consequence is a sentence worth carrying away: "nothing was found in fifty draws" is very weak evidence of anything. Regulators keep long records because short ones cannot answer the question.
Under an honest machine, a p-value is uniform on — not clustered near 1, which is the common intuition, but genuinely flat. Half of all honest readings fall below 0.5. One in twenty falls below 0.05.
That makes the histogram of every p-value the battery produces a test of the battery itself. A miscalibrated null shows up as a lean; a real deviation shows up as a spike against the left wall. It is the one chart here that would betray this page's own analysis if the analysis were wrong.
Because the panel produces so many p-values, they are also reported with a Benjamini–Hochberg adjustment, which controls the share of flagged results that are false alarms rather than demanding no false alarms at all — the right trade for a monitor that must stay sensitive enough to be worth reading.
The last-digit test is where a naive implementation gives itself away. Final digits are not uniform for a bounded pool: a 1–49 pool holds five numbers ending in each of 1 through 9, but only four ending in 0. Testing against "10% each" reports a shortage of zeros on every honest lottery on earth.
The baseline used here is counted from the pool itself, so it is exact for any , and uniform only in the case where ten divides it.
Before anyone could trust a table of random numbers, someone had to work out how to check one. Maurice Kendall and Bernard Babington Smith built a machine to generate random digits and then, facing the obvious question, invented the tools to interrogate it: a frequency test, a serial test, a gap test, a poker test.
Their 1938 paper is the ancestor of every randomness certificate since. Knuth carried the battery into computing in The Art of Computer Programming — the gap test that appears elsewhere in this playground is theirs — and the modern suites that certify casino generators are the same idea with more arithmetic.
The panel above is a direct descendant. There is nothing novel about it, and that is its recommendation.
In 2006, Jeffrey Rosenthal was asked whether Ontario lottery retailers were winning too often. He did the arithmetic that the retailers' own regulator had not: given how many tickets clerks handled, roughly how many wins should insiders have had?
Fewer than 60. The records showed 200 or more.
He presented it on CBC television, and the arithmetic was simple enough that nobody could wave it away. Investigations followed, then refunds, then a restructured regulator. No new mathematics was required — only somebody willing to compute an expectation and compare it with a count.
That is what this page automates. The full case belongs to another feature; the instruments are his toolkit, running continuously instead of after a scandal.
It is worth saying plainly that audits of national lotteries overwhelmingly find nothing. The UK National Lottery has been examined repeatedly by academic statisticians — the same chi-square that sits first on this panel — and it has passed each time.
That is not a boring result. Nothing in the physics of a ball machine guarantees fairness; it has to be engineered, maintained, and checked. A needle at rest is a machine keeping a promise that could have been broken.
Then there is Eddie Tipton, who wrote the code that drew Hot Lotto and inserted a routine that, on a few days each year, narrowed the draw to a small set of combinations he could predict.
He was not caught by statistics. He was caught because a $16.5 million ticket went unclaimed for nearly a year and then was claimed through a lawyer on behalf of an anonymous trust, and because the surveillance video of the purchase was eventually recognised. The forensic recovery of his code came afterwards.
The sandbox's Tipton preset shows why. Tampering with two percent of draws, and only ever producing legitimate-looking combinations, leaves the instruments undisturbed. Run it for a realistic era and every needle rests, because there genuinely is not enough signal to see.
Statistics did do real work in that case — once suspicion existed, it quantified how implausible the associated win patterns were. But the honest lesson is that a panel like this has limits, and a page that only ever demonstrated its own successes would be selling something.
Every needle at rest is a machine keeping its promise. The Observatory exists for the day one isn't — and for the quiet reassurance of every day one is.