Showing posts with label predictions. Show all posts
Showing posts with label predictions. Show all posts

Friday, May 15, 2015

Consequence of Morey's Law: Lucky vs Unlucky teams

In a previous post, I looked at a 1994 paper by Daryl Morey (current Houston Rockets GM) who investigated how a team's winning percentage was related to the number of points they scored and allowed, deriving the "modified Pythagorean theorem":

expected win percentage =
  pts_scored ^ 13.91 / (pts_scored ^ 13.91 + pts_allowed ^ 13.91)

At the end of his paper, Daryl explores teams who had the biggest delta between their actual and predicted wins. In 1993-1994, the Chicago Bulls and Houston Rockets top the list and Daryl refers to them as lucky teams. But why is lucked involved?


The rationale is that if you have two teams A and B with almost identical points scored and points allowed, we would expect them to have very similar win percentages. The only way to create a discrepancy (without changing points scored and points allowed... too much), is by changing the outcome of the very close games. So for all the games team A won by a point, flip the scores so that they lose by 1, and reversely for team B who now wins all the games they previously lost by 1. With this hypothetical construction, we will have two teams still with very similar points scored and allowed but potentially different records. It would make common sense that for very close games the probability of each team winning is around 50%, so winning or losing amounts to "luck", whether a desperation buzzer-beater is made or bounces off the back of the rim. And so it would make sense that teams with high discrepancies between actual and predicted wins were either much better or much worse than 50% in close games. Let's confirm.

Here's the table of teams with discrepancies greater or equal to 6 between their actual and projected records, ranked by year:

Team Year Scored Allowed Wins (proj) Wins (actual) Win %
NJN 2000 98.0 99.0 38 31 37.8
DEN 2001 96.6 99.0 34 40 48.8
NJN 2003 95.4 90.1 56 49 59.8
CHA 2005 94.3 100.2 24 18 22.0
NJN 2005 91.4 92.9 36 42 51.2
IND 2006 93.9 92.0 47 41 50.0
TOR 2006 101.1 104.0 33 27 32.9
UTA 2006 92.4 95.0 33 41 50.0
BOS 2007 95.8 99.2 31 24 29.3
CHI 2007 98.8 93.8 55 49 59.8
DAL 2007 100.0 92.8 61 67 81.7
MIA 2007 94.6 95.5 38 44 53.7
SAS 2007 98.5 90.1 64 58 70.7
NJN 2008 95.8 100.9 27 34 41.5
TOR 2008 100.2 97.3 49 41 50.0
DAL 2010 102.0 99.3 49 55 67.1
GSW 2010 108.8 112.4 32 26 31.7
MIN 2011 101.1 107.7 24 17 20.7
PHI 2012 93.6 89.4 43 35 53.0
BRK 2014 98.5 99.5 38 44 53.7
MIN 2014 106.9 104.3 48 40 48.8


So how did these teams fare in close games? I've labelled a team/year as High if they won 6 or more games than expected (8 teams from the previous list), Low if they lost 6 or more  games than expected (13 teams from the previous list), and Normal otherwise. I then look for each group their win percentage in closely contested games (final scores within 1, 2 and 3 points).

Final scores within 1 point:

Type # Wins # Games Win %
Normal 721 1439 50.1
Low 8 21 38.1
High 2 2 100.0

Final scores within 2 points:

Type # Wins # Games Win %
Normal 1760 3515 50.1
Low 20 52 38.5
High 10 13 76.9

Final scores within 3 points:

Type # Wins # Games Win %
Normal 2820 5617 50.2
Low 24 80 30.0
High 14 19 73.7

Our intuition was correct and so were Daryl's closing comments: teams can indeed be qualified as lucky and unlucky, some winning almost 3 out of 4 close match-ups, others losing 2 out of 3 tight games. This intangible "luck" factor is sufficient to explain why certain teams have much better or worse records than their offense/defense would typically lead to. It doesn't take much for to flip the outcome of an entire game.


As a quick aside, much has been said about the San Antonio Spurs this year and their drop from a potential 2nd seed to 6th seed entering the Playoffs. Most articles focused on their loss on the final day of the regular season which led to that seeding free-fall, but was excessive focus placed on that last game? Had they been particularly lucky/unlucky during the season? It turns out their record is a couple games lower than what the modified Pythagorean theorem would have predicted, and that they weren't particularly lucky or unlucky in their close games, winning 2 of 5 games decided by 1 point, and 6 of 13 decided by 3 points or less.

Tuesday, July 8, 2014

World Cup: Don't get all defensive!



Saying that the 2014 World Cup is currently underway in Brazil is as obvious as statements come. Even if you're not a huge soccer fan (heck, even if you hate soccer!) it's everywhere, ads, fan tweets, TV results, schedules, preferred links, friends' Facebooks updates with posts as insightful as "Go [insert country name] !![adapt number of exclamation points based on time until next game]".

Given that half of my posts on this blog so far have dealt with NBA basketball you've probably guessed what sport I typically tune in to. But the NBA Finals are over ("Go Spurs!!!!"), and I have to admit I got somewhat caught up in the World Cup excitement.


I have only watched one game from start to finish, but have kept track of results and scores (like everyone on Earth there is a world cup pool going on at work), and something stood out: the first few games were rather exciting (aka high-scoring) (I would sit for short periods of time through games and almost always caught a goal). But the past few games have actually been more on the boring (aka low-scoring) side. Again, associating the number of goals to excitement is quite a shortcut, but I don't know enough of soccer to appreciate the strategic subtleties in a game ending at 0-0 (remember I come from the basketball world were scores average 100). But forget the excitement, is there real a drop in goals scores or is this just a fluke?

Quick soccer recap for the non-experts like me (from Wikipedia):
The current final tournament features 32 national teams competing over a month in the host nation(s). There are two stages: the group stage followed by the knockout stage.
In the group stage, teams compete within eight groups of four teams each. Each group plays a round-robin tournament, in which each team is scheduled for three matches against other teams in the same group. This means that a total of six matches are played within a group.
The top two teams from each group advance to the knockout stage. Points are used to rank the teams within a group. Since 1994, three points have been awarded for a win, one for a draw and none for a loss (before, winners received two points).
The knockout stage is a single-elimination tournament in which teams play each other in one-off matches, with extra time and penalty shootouts used to decide the winner if necessary. It begins with the round of 16 (or the second round) in which the winner of each group plays against the runner-up of another group. This is followed by the quarter-finals, the semi-finals, the third-place match (contested by the losing semi-finalists), and the final.


My naive observation was that scores in the first stage (group stage) were typically higher than in the second stage (knockout stage).
Let's look at the numbers for the 2014 World Cup:

First stage:

  • Average Goals: 2.833
  • First quantile: 1
  • Median: 3
  • 3rd quantile: 4
  • Max: 7

Second stage:

  • Average Goals: 1.917
  • First quantile: 1
  • Median: 2
  • 3rd quantile: 3
  • Max: 3

I won't run any statistical test to look into the difference, but even the worst (best?) devil's advocate will agree that goal production has gone down. The most reasonable explanation is that group stage games don't automatically eliminate you from the competition, everyone is guaranteed three games, so there is a little less pressure to win as opposed to the knockout stage where it is win-or-go-home to use the famous NBA wording. Also, as the competition progresses, the teams get better, it's harder to score 8 goals on your opponent.

But one could have made the case that scores in the knockout stage should have been higher on average: recall that draw (any other sport calls this a tie) are acceptable in the group stage, but a winner is required in the knockout stage and so extra time is added in case of a draw after regular time. The need to break the draw and provide extra time should theoretically generate higher scores.

But is this drop specific to this world cup or has it always been the case? If we break out the scores by stage, does the trend continue with scores getting lower and lower throughout the tournament? Let's look at the evolution of scores for the World Cups from 1990 onwards, broken out by stage:

No strong trends surface from the graph, other World Cups share the same strong decline through the first stages (2002, 2006), but others had increases throughout the same phases of the tournament, as recently as the 2010 World Cup.

Life is full of ironies: as I am writing this post, Germany is leading Brazil 5-0 at the half of the first semi-finals! So much for teams focusing more and more on defense....

So, enough pessimism, the tournament doesn't necessarily generate less goals throughout the tournament, crazy things can happen anytime. Some will say it's the beauty of the sport!

Oh, and based on the graphs, I predict two goals in the Finals....

Friday, September 13, 2013

Tuesday Oct 29 2013: Bulls at Miami, who will win?


The 2013-2014 NBA season kicks-off Tuesday October 29th and the spotlight will be the return of two of the injured megastars: the Bulls' Derrick Rose in Miami, and the Kobe Bryant's Lakers host the Clippers.

On this blog we love drama, we love high intensity games, but we also like tough questions. Such as: who will win those two games? And while we're at it, who will come out victorious of the other 1228 games of the season?

In this post I will present my efforts to predict the outcome of a game based on metrics related to the two opposing teams.




The data

The raw data consisted of all regular-season NBA games (no play-offs, no pre-season) since the 1999-2000 season. That’s right, we’re talking about 16774 games here. For each game I pulled information about the home team, the road team and who won the game.


The metrics

After pulling the raw data, the next step was to create all the metrics relate to the home and road team’s performance up until the game I want to predict the outcome of. Due to the important restructuring that can occur in a team over the off-season, each new season starts from scratch and no results carried over from one season to the next.

Simple Metrics
The simple metrics I pulled were essentially the home, road and total victory percentages for both the home team and the road team. Say the Dallas Mavericks, who have won 17 of their 25 home games and 12 of their 24 road games visit the Phoenix Suns who have won 12 of their 24 home games and 13 of their 24 road games, I would compute the following metrics:

  • Dallas home win probability: 17 / 25 ~ 68.0%
  • Dallas road win probability: 12 / 24 ~ 50.0%
  • Dallas total win probability: 29 / 49 ~ 59.2%
  • Phoenix home win probability: 12 / 24 ~ 50.0%
  • Phoenix road win probability: 13 / 24 ~ 54.2%
  • Phoenix total win probability: 25 / 48 ~ 52.1%


Discounted simple metrics
However, these statistics seemed a little too simplistic. A lot can happen in the course of a season. Stars can get injured or return from a long injury. A new team might struggle at first to play with each other before really hitting their stride. So I included some new metrics which have some time discounting. A win early in the season shouldn’t weigh as heavily as one in the previous game. In the non-discounted world we kept track, for home games and road games separately, of the number of wins and number of losses, incrementing one or the other by 1 depending on whether the team won or lost. We do exactly the same here with a discount factor:

new_winning_performance = discount_factor * old_winning_performance + new_game_result

new_game_result is 1 if they won the game, 0 if they lost.

When setting the discount factor to 1 (no discounting), we are actually counting the number of wins and are back in the simple metrics framework.

To view the impact of discounting, let us walk through an example:
Let’s assume a team won it’s first 3 games then lost the following 3, and let us apply a discount factor of 0.9.

  • After winning the first game, the team’s performance is 1 (0 * 0.9 + 1)
  • After winning the second game, the team’s performance is 1.9 (1 * 0.9 + 1)
  • After winning the third game, the team’s performance is 2.71 (1.9 * 0.9 + 1)
  • After losing the fourth game, the team’s performance is 2.44 (2.71 * 0.9 + 0)
  • After losing the fifth game, the team’s performance is 2.20 (2.44 * 0.9 + 0)
  • After losing the sixth game, the team’s performance is 1.98 (2.20 * 0.9 + 0)

But now consider a team who lost their first three games before winning the next three:

  • After losing the first game, the team’s performance is 0 (0 * 0.9 + 0)
  • After losing the second game, the team’s performance is 0 (0 * 0.9 + 0)
  • After losing the third game, the team’s performance is 0 (0 * 0.9 + 0)
  • After winning the fourth game, the team’s performance is 1 (0 * 0.9 + 1)
  • After winning the fifth game, the team’s performance is 1.9 (1 * 0.9 + 1)
  • After winning the sixth game, the team’s performance is 2.71 (1.9 * 0.9 + 1)

Although both teams are 3-3, the sequence of wins/losses now matters. A team might start winning more if a big star is returning after an injury of due to a coach change or some other reason so our metrics should reflect important trend changes such as those. Unsure of what the discounting factor should be, I computed the metrics for various values.

Discounted opponent-adjusted metrics
A third type of metric I explored was one where the strength of the opponent is incorporated to compute a team’s current performance. In the above calculations, a win was counted as 1 and a loss as 0, no matter the opponent. But why should that be? Just like in chess with the ELO algorithm, couldn’t we give more credit to a team beating a really tough opponent (like the Bulls snapping the Heat’s 27 consecutive wins), and be harsher when losing to a really weak team?
These new metrics were computed the same way as previously (with a discounting factor) but using the opponent’s performance instead of 0/1.

Let’s look at an example. The Thunder are playing at home against the Kings. The Thunder (pretty good team) have a current home win performance of 5.7 and a current home loss performance of 2.3. This leads to a “home win percentage” of 5.7 / (5.7 + 2.3) = 71%. The Kings (pretty bad team) have a current road win performance of 1.9 and a current road loss performance of 6.1. This leads to a “road win percentage” of 1.9 / (1.9 + 6.1) = 24%.

If the Thunder win:

  • The Thunder’s home win performance is now: 0.9 * 5.7 + 0.24 = 5.37
  • The King’s road loss performance is now: 0.9 * 6.1 + (1 - 0.71) = 5.78

If the Thunder lose:

  • The Thunder’s home loss performance is now: 0.9 * 2.3 + (1 - 0.24) = 2.83
  • The King’s road loss performance is now: 0.9 * 1.9 + 0.71 = 2.42
If you win, you get credit based off of your opponent’s win percentage. If you lose, you get penalized according to your opponent’s losing percentage (hence the 1- in the above formulas). The worst teams hurt you most if you lose to them.
As seen from the example between a very good team and a very bad one, winning does not guarantee that your win performance will increase and losing does not guarantee your losing performance will increase. The purpose is not to have a strictly increasing function if you win, it is to get, at a given point in time, an up-to-date indicator of a team’s home and road performance.


The models

For all the models, the training data was all games except those of the most recent 2012-2103 season which we will use to benchmark our models.


Very simple models were first used: how about always picking the home team to win? Or the team with the best record? Or compare the home team’s home percentage to the road team’s road percentage?

I then looked into logistic models in an attempt to link all the above-mentioned metrics to our outcome of interest: “did the home team win the game?”. Logistic models are commonly used to look at binomial 0/1 outcomes.

I then looked into machine learning methods, starting with Classification and regression trees (CART). Without going into the details, a decision tree will try to link the regressor variables to the outcome variable by a succession of if/else statements. For instance, I might have tracked over a two week vacation period whether my children decided to play outside or not in the afternoon, and also kept note of the weather conditions (sky, temperature, humidity,...). The resulting tree might look something like:



If it rains without wind tomorrow, I would therefore expect them to be outside again!

Finally, I also used random forest models. For those not familiar with random forests, the statistical joke behind it is that it is simply composed of a whole bunch of decision trees. Multiple trees like the one above are “grown”. To make a prediction for a set of values (rain, no wind), I would look at the outcome predicted by each tree (play, play, no play, play….) and pick the most frequent prediction.





The results

As mentioned previously, the models established were then put to the test on the 2012-2013 NBA season.

As most of the models outlines above use metrics that require some historical data, I can’t predict the first game of the season not having any observations for past games for the two teams (yes, I do realize that the title of the post was a little misleading :-) ). I only included games for which I had at least 10 observations of home games for the home team and 10 road games for the road team.

Home team
Let’s start with the very naive approach of always going for the home team. If you had used this approach in 2013, you would have correctly guessed 60.9% of all games.

Best overall percentage
How about going with the team with the best absolute record? We get a bump to 66.3% of all games correctly predicted in 2013.

Home percentage VS Road percentage
How about comparing the home team’s home percentage with the road team’s road percentage? 65.3% of games are correctly guessed this way. Quite surprisingly, this method provides a worse result than simply comparing overall records.

Logistic regressions
Many different models were tested here but I won’t detail each. Correct guesses range from 61.0% to 66.6%. The best performing model was the simplest one which only included the intercept and the overall winning percentage for the home team and the road team. The inclusion of the intercept explains the minor improvement to the 66.3% observed in the “Best overall percentage” section.
Very surprisingly, deriving sophisticated metrics discounting old performances in order to get more accurate readings on a team’s performance did not prove to be predictive.

Decision tree
Results with a decision tree was in the neighborhood of the logistic regression models, with a value of 65.8%.


Interestingly, the final model only looks at two variables and splits: whether the home team's home performance percentage (adjusted with a discounting factor of 0.5) is greater than 0.5 or not, and whether the road team's performance incorporating opponent strength (discounting factor of 1, so no discounting actually) is greater than 0.25 or not.

Random Forests
All our hopes reside in the Random Forests to obtain a significant improvement in predictive power. Unfortunately we obtain a 66.2% value, right where we were with the decision tree, the regression model and more shamefully the model comparing the two teams overall Win-Loss record!
When looking at which variables were most important, home team and road team overall percentages came up in the first two positions.


Conclusions

The results are slightly disappointing from a statistical point of view. From the simplest to the most advanced techniques, none are able to break the 70% threshold of correct predictions. I will want to revisit the analyses and try to break the bar. It is rather surprising to see how overall standings matter compared to metrics that are more time sensitive. The idea that the first few games of the season matter to predict the last few games of the season is a great insight. This can be interpreted by the fact that good teams will eventually lose a few consecutive games in a row, even against bad teams, but that should not be taken too seriously. Same with bad teams winning a few games against good teams, they remain bad teams intrinsically.

From an NBA fan perspective, the results are beautiful. It shows you why this game is so addictive and generates so much emotion and tension. Even when great teams face terrible opponents, no win is guaranteed. Upsets are extremely common and every game can potentially become one for the ages!

Friday, April 19, 2013

2013 NBA Playoffs Predictions



It's that time of year again! It is what an entire season has lead to, the famous "win or go home" period.
We all just love the Playoffs, when superstars are made and forgotten, when every single match-up has its own little history behind it, when teams without any hard feelings against each other a Game 1 tip-off end up hating each other by Game 3. Unparalleled intensity, game after game!

This year's Playoffs trigger mixed feelings, as it seems the Champions have already been elected. Can anybody beat Miami? Can anybody make it at least somewhat of a challenge for them?

But even so, there are bound to be plenty of upsets we thrive on,  and of course we will all have eyes on the Howard-led Lakers. They can't afford to go through the Playoffs the same way they roller-coasted through the season, but if they continue on their current trend they can be a potential threat.

So what are my predictions? Well, based on season performance and home/away results, as well as taking each team's specific homecourt advantage into account, here's what I obtained for the most likely outcome at each phase with associated probability:




















So it does seem that Miami has the clearest path to the trophy, but there no guarantees. And if we look at each team probability of wining the championship, we obtain the following table:

NBA team Champion Probability
Miami Heat 22.8%
Oklahoma City Thunder 13.1%
San Antonio Spurs 11.3%
New York Knicks 9.5%
Denver Nuggets 7.9%
Los Angeles Clippers 7.2%
Memphis Grizzlies 6.5%
Brooklyn Nets 4.8%
Indiana Pacers 4.8%
Chicago Bulls 2.8%
Golden State Warriors 2.5%
Atlanta Hawks 2.4%
Los Angeles Lakers 1.5%
Houston Rockets 1.3%
Boston Celtics 1.0%
Milwaukee Bucks 0.8%


While the table confirms that Miami has the greatest probability of winning the trophy, the probability is still less than 25%, offering challengers a great opportunity. The probability is still incredibly high considering it is almost as big as the sum of probabilities for the two next contenders (Oklahoma and San Antonio). This is in part due to Miami definitely enjoying a huge Conference advantage battling weaker foes, and if it falls it will most likely be in the Finals. An extra dimension favoring the Heat to consider and not (currently) taken into account in my model is wear-and-tear: the Heat will be more rested and less likely to be injured when they meet the Western Conference finalist...

I will regularly post updates as more games get played.




Monday, February 18, 2013

Who will make the Playoffs?


The All-Star weekend is a nice break in the middle of the regular season. Dunk contest, three-point contest, rising stars game and naturally the All Star game itself. A good break indeed from the intense regular season schedule.

But for me it offered a unique opportunity: no games for four consecutive nights! The perfect opportunity to run simulations. Many simulations.

A. Downpour. Of. Simulations.

You be the judge: the outcome of over 2 BILLION games were simulated.

I previously worked on some code to forecast the outcome of the rest of the season based on each team's latest performance, code which I have already used to look at the Lakers probability of making the Playoffs, and whether Dallas or LA (Lakers) had a better chance of making the Playoffs. But I was eagerly waiting for the All Star weekend to run the code for all days starting on December 1st 2012, in order to look at the trends and shifts in each team's probability of making the Playoffs, and each team's expected final standing.

Here are the results by conference.

Eastern Conference

Let's start by looking at the evolution of team standings:

And now for the evolution of the probability of making it to the Playoffs:

The last values as of February 15th 2013 give the following values:

Team Position Playoff Probability
MIA 1.196 100%
NYK 3.157 100%
BRK 4.455 99.8%
IND 4.076 99.8%
CHI 4.445 99.7%
ATL 4.942 99.6%
BOS 6.706 95.4%
MIL 7.230 91.4%
PHI 9.666 8.2%
TOR 10.195 4.2%
DET 10.588 1.8%
WAS 13.069 0.1%
CHA 14.470 0%
CLE 12.564 0%
ORL 13.241 0%

Although Philadelphia still has a glimmer of hope of making the Playoffs provided Milwaukee runs into a very bad stretch, the eastern teams making the Playoffs seem decided. The huge uncertainty however lies around positions 3-6 where all teams are extremely close. With the Brooklyn Nets facing Chicago and Indiana in the last two weeks of the regular season, the eastern final brackets will not be decided until the very end.


Western Conference

Similarly to what we did for the Eastern Conference, let's start by looking at the standings evolution:

And now for the "Playoff Probability" evolution:

The last values as of February 15th 2013 give the following values:

Team Position Playoff Probability
LAC 2.749 100%
OKC 2.371 100%
SAS 1.329 100%
DEN 4.899 99.9%
MEM 4.682 99.5%
GSW 5.824 97.5%
UTA 7.190 86.9%
HOU 7.550 80.7%
POR 9.752 14.8%
LAL 9.877 12.1%
DAL 10.376 8.1%
MIN 12.519 0.4%
NOH 12.870 0.1%
PHO 14.224 0%
SAC 13.788 0%

The story is a little different on the western front, where suspense is not around the center 3-6 positions and the final Playoff bracket but on the final eight spot. Houston has a good chance of keeping that spot which it currently holds, but Portland and Los Angeles (Lakers of course!) are on their heels from a safe-but-not-THAT-safe distance. Any slip and the two will pounce on the position!



I will provide updated probabilities as the season progresses!



Friday, July 20, 2012

Can the lottery code be cracked?



I just came across this article about a Canadian who broke the Scratch Lottery Code.

The article s about Mohan Srivastava who came across two unscratched Tic-Tac-Toe tickets near his desk. Not a big lottery fan but with nothing else to do, he scratched them. Lost on the first, won $3 on the second. He went to cash in his winning at the nearby gas station, but all the while he started thinking about how these tickets are created. Because the lottery corporation needs to keep careful track of how many winning tickets get printed, the computers can't just generate random numbers onto each card, winning and losing tickets need to be carefully created while giving the illusion of being completely random.

The more he thought about it, the more he became convinced there could be a way to determine whether a ticket was a winning ticket or not without having to scratch it. Here's what a Tic-Tac-Toe ticket looks like:



You have 8 3-by-3 grids with visible numbers. You then scratch the 24 numbers on the left. If any of t3 of the 24 numbers are lined up in any direction in any of the 8 grids, you've won!

Well Mohan bought many tickets, scratched them all, and ultimately found a relatively simple rule to determine ahead of time whether a lottery ticket is a winning ticket or not.

The trick? In your lottery ticket you have 72 numbers visible in the grids (8 * 3 * 3). Because the numbers go only up to 39, so will have to occur multiple times. Mohan kept track for each number how many times it occurred on the ticket:



He was especially interested in singletons that appeared only once. His finding was that if three singletons lined up, then you had a winning ticket!



I am not a 100% about why that is and how Mohan came to this conclusion, but intuitively I think the algorithm works something like this (at a very high level)

  • randomly separate numbers from 1 to 39 into two sets, 24 singletons and the 15 duplicates
  • fill in the grids with duplicate numbers, and with at most 2 singleton numbers
  • list the 24 singleton numbers on the left hand side

This will guarantee that the ticket is a losing one. You can make it a winning one by adding three singletons in a row, column or diagonal in any grid. This will guarantee that the ticket is a winning one with only one gain.

This is overly simplistic but it does allow a simple way to generate the tickets. Creating random grids and then testing whether they are winners or losers before printing them would take to much time, the above process creates them in one shot without testing.

On the example given in the article we can see that the 24 numbers given on left come out a total of 35 times in the grid (average 1.45), whereas the other 15 numbers not listed come out 37 times (average 2.47). This hints that the duplicates are indeed something like background noise, and the 24 numbers selected on the left are actually much less likely to occur in the grid. And this is where we are fooled: If we reveal 24 numbers, we should find a lot more of them in the grid made out of 39 different numbers shouldn't we?

The article also talks about the lottery industry in general and how the mob uses the lottery for money laundering.

And as to why Mohan revealed the trick instead of making a fortune? He simply calculated that if he were to spending all his time identifying winning tickets in each store across town, scratching them, and redeeming them, he would make less than his current full time job!

All in all, a rather interesting read, although I already gave out the end...