Showing posts with label home court advantage. Show all posts
Showing posts with label home court advantage. Show all posts

Friday, October 16, 2015

A brief history of NBA runs: Do teams really 'get hot'?

"Cleveland with a 10-0 in the last 2:05...."
"Warriors answering with an 8-0 run over the end of the first and beginning of the second quarter..."

I've watched a number of games during this 2014-2015 season, Playoffs included, and couldn't help but notice the amount of runs being announced on screen. This was particularly true when Chicago went on long scoring droughts against Cleveland in the Eastern semifinals.


But a question that nagged me all this time and which I wanted to investigate a little further is whether these droughts - or runs depending on whose side you're on - are natural and expected, or on the contrary are influenced by external factors.

In a previous post I took a closer look at overtimes in the NBA and showed that because they are an equilibrium caught between two highly unstable states (team A losing by a handful of points on one side, team B losing by a handful of points on the other), overtimes are about three times more likely to occur than one would naively expect. Could the same be said for runs? Once a team has gone on an 8-0 run, is it more likely to push it to 10-0? The 8-0 run could be the result of one team having a much better lineup on the floor, or a player with a particularly hot hand (although the notion of hot hand is debatable, one I will probably look into in a future post). Or is the team on the bad end of the run more likely to score, perhaps by calling a timeout to stop the first team's 'mojo' or to set up a specific play with higher scoring probability?

Data
The first step was to collect as much data as possible. I pulled from nba.com all available games (regular season and playoffs but excluding pre-season), all the way back to the 2009-2010 season.
For each game, I split it into a succession of 'runs' and for each I computed the number of possessions and points.
Consider for instance the first few minutes of Game 1 of this year's Western Conference Finals between the Houston Rockets and Golden State Warriors:


After parse the information, we would get something like:
  • Rockets had a 2-0 run over 45s
  • Warriors then had a 2-0 run over 15s
  • Rockets then had a 7-0 run over 1m51s
  • ...

Although points is what is always being reported and what everyone ultimately cares about, I decided to focus on number of scoring possessions instead. A 7-0 which is the results of seven consecutive trips to the freethrow line, with 1 out of 2 freethrow being made each time is very different from a three-pointer followed by another three-pointer on which the shooter is fouled and completes the four-point play. In the former case the scoring team needs to get (at least) 6 defensive stops, in the latter they need only one.

So the data would actually look like this:
  • Rockets score 2 points on 1 scoring possession
  • Warriors score 2 points on 1 scoring possession
  • Rockets score 7 points on 4 scoring possession
  • ...

So the question we are interested in is how many consecutive times can a team score uninterrupted?


Preliminary Graphs
Before we jump into any modeling, let us first look at the frequency of uninterrupted scoring possessions:


That's quite a nice shape! It seems the occurrence of every run is a little under half of the previous number. We can verify this ratio visually:



Indeed, the frequency of each run is a remarkably stable ~45% of the previous run frequency.

A natural question is whether there is a difference between regular season games and Playoff games? Defense is supposed to be cranked up on notch so are longer runs less frequent? In the following graph, the proportion of runs in Playoff games is represented via the red histogram, that of regular season games via the blue histogram, and the overlap is purple.


The little blue tip would suggest a little more runs of 1 scoring possession in the regular season (and hence a little more 2+ scoring possession runs in the Playoffs), but a Chi-Square test reveals no statistical significance in the difference between the two histograms.

We can also split runs by home and road team. Can homecourt provide an additional boost and extend runs?


It appears as if the home team is slightly more likely than the road team to have longer runs, and this time the difference (as small as it appears visually) is significant. Once the home team gets it going and gets the crowd involved, good things happen!

Yet another splitting option is by quarter. Perhaps a team hasn't got its rhythm in the first quarter and is more likely to suffer a run, whereas the defense is a little tougher in the fourth quarter thus limiting scoring opportunities. To avoid overlaying 5 histograms over each other (the fifth being for overtimes), I used lines instead:


Although the lines appear nearly identical on the graph, the Chi-Square did pick up a significant difference across the associated table, even when dropping runs in overtime (harder to get long runs in a 5 minute overtime than a 12 minute quarter).
But it turns out that longer runs are more likely in the fourth quarter than the first. Either the defense gets a little tired, or the losing team realizes that they need to step things up quickly to avoid picking up the L.


Model
Having a better sense of how the runs behave, we can apply a little bit of modeling.
Let us assume that the two teams are of similar strength, and that when they have possession of the ball they both have the same probability p of scoring.
Team A just scored, team B now has possession of the ball. What is team B's probability of interrupting A's run? There are theoretically an infinite number of ways for that to happen, the pattern being quite obvious:
  • B scores (probability = p)
  • B misses, A misses, B scores (probability = (1-p)(1-p)p)
  • B missesA missesB missesA misses, B Scores (probability = (1-p)(1-p)(1-p)(1-p)p)
  • ...
  • (B misses, A misses) n times, B Scores (probability = (1-p)^(2n) * p)

Reminder: p is the probability of scoring on a team's possession, so it incorporates missing a shot but getting the offensive rebound and shooting again for instance.

Adding all the pieces yields the probability of team A's run to be interrupted:

Let's now look at the probability of extending the run by exactly n more possessions, which we will denote P(n). We will break up this probability as the probability of scoring one more time and then exactly (n-1) times to get the recurrence with P(n-1):
  • B missesA scoresA scores exactly (n-1) more times (probability = (1-p)p * P(n-1))
  • B missesA missesB missesA scores, A scores exactly (n-1) more times (probability = (1-p)(1-p)(1-p)p * P(n-1))
  • ...
  • (B misses, A misses) n times, B missesA scores, A scores exactly (n-1) more times (probability = (1-p)^(2n) * p(1-p) * P(n-1))
Adding everything up yields:

And so, realizing that P(interrupted) is actually P(0), we get the general formula:

Let's get a few curves for various values of p:



I've overlaid the empirical curve (blue curve in bold). It's a little difficult to spot as it is extremely close to the curve with p = 25%.

Here's the plot with just those two curves:


The similarity in the two curves is really impressive!
We can even refine the true value by fitting a model to identify the value of p which best fits our empirical data. The result is 24.2%.

But back to our initial problem? Recall that we were trying to determine whether runs occur as frequently as one would expect, or if there are external factors that make them more/less likely? In our model, we assume no such external effects, the probability of any team to score when it has possession of the ball is a constant p, and does not depend on the past (whether team A has scored 0, 1 or 10 consecutive times already, the same way heads will come up 50% of the time with a fair coin even if we've just had a run of 10 heads or ten tails right before). And the fact that under this assumption theoretical and empirical values match so well would suggest that there are no external effects (or perfectly compensate each other!), and that when we observe 8-0, or 11-0 runs we were simply bound to see them occur.


Let's end all the modeling with a fun fact: any idea what the greatest run from these past years has been (with one team remaining scoreless)? 15-0? 19-0? Turns out it was 29-0, by the Cleveland Cavaliers led by LeBron James.... before his Miami days. Over the first two quarters of the game and almost 9 minutes, the Cavs scored 29 consecutive points on 19 possessions over the Milwaukee Bucks on Dec 6th 2009 (a day short of the Pearl Harbor Anniversary!).






Thursday, June 21, 2012

Homecourt and rest time advantage

In a previous post, I looked at the true impact of homecourt advantage in the NBA, for the league in general and for each individual team. The model was simple, only considering whether the game was at home or away.

The main take-away was that playing at home bumped your probability of winning by almost 20 percentage points, from 40% to 60%. Quite a significant jump, although not every team observed the same jump.

I did however feel that the model was a little over-simplistic in ignoring another phenomenon which could impact a team's performance: rest time between games. Especially over the 2011-2012 condensed season with certain teams playing back-to-back-to-back games, one can definitely wonder how rest days come into play. If a team is playing on the road, can the fact that they have had three days of rest as opposed to their opponents back-to-back games mitigate the opponent's home court advantage?

The data and methodology are almost identical to the post I mentioned earlier: I looked at all 2009-201 and 2010-2011 games, and for each match-up looked at which team played at home and how many days of rest each team had.

Since we are now looking at multiple variables instead of just the homecourt impact, I will only provide the breakdown of results for the league in general, providing them for each team would just take up too much space.


Impact of rest days

The following table provides the victory probability based on where the game is played and the number of rest days for both teams.

Team A at home Rest days (Team A) Rest days (Team B) Win probability
Yes 1 1 59.1%
No 1 1 40.9%
Yes 2 1 65.2%
No 2 1 47.4%
Yes 3+ 1 62.6%
No 3+ 1 44.6%
Yes 1 2 52.6%
No 1 2 34.8%
Yes 2 2 59.1%
No 2 2 40.9%
Yes 3+ 2 56.4%
No 3+ 2 38.3%
Yes 1 3+ 55.4%
No 1 3+ 37.4%
Yes 2 3+ 61.7%
No 2 3+ 43.6%
Yes 3+ 3+ 59.1%
No 3+ 3+ 40.9%


Some interesting highlights are that:
  • independently of the number of rest days each team has had the difference homecourt advantage is always around 17-18%
  • the homecourt effect is much more predominant than the number of rest days: even in the best case scenario, the win probability on the road is 47.4%, so essentially a +7% percentage uplift due to rest days, as opposed to the +20% we saw in the previous post for the homecourt advantage impact.
  • it turns out that resting 2 days improves probability of victory compared to one day only, and three or more days is also more beneficial than one day only, two days is actually preferable to 3 or more days. This is also a debate that comes around often especially during playoff time, where one team comes out of a game 7 to meet a team that finished a sweep over a week before. Is too much rest a bad thing? From this data it does appear that 2 days provides the optimal balance between hitting your stride while you're hot and resting your sore legs.

Team's optimal rest days

What is true for the league isn't necessarily true for individual teams. I wanted to check if all teams preferred to rest 2 days instead of 1 or 3+ days. Were younger teams eager to have back-to-back games? Were older teams dreadful of tight schedules?

Team Significant Optimal rest days
NBA Yes 2
ATL Yes 2
BOS No 3
CHA Yes 2
CHI Yes 2
CLE No 2
DAL No 3
DEN Yes 3
DET Yes 3
GSW Yes 1
HOU No 2
IND Yes 2
LAC Yes 2
LAL Yes 1
MEM Yes 2
MIA No 2
MIL Yes 1
MIN Yes 1
NJN Yes 2
NOH Yes 3
NYK No 3
OKC No 3
ORL No 2
PHI No 2
PHO Yes 3
POR Yes 1
SAC No 1
SAS No 2
TOR Yes 3
UTA No 3
WAS Yes 3


Upon close inspection there does not seem to be any strong correlation between the team's age and the preferred number of rest days. Sure Boston is an old team preferring over three days and Golden State is one of the youngest team performing best on back-to-back games, but the Lakers are an old team also preferring back-to-back teams and the Wizards are a young team with best odds after 3+ days of rest.

To conclude, while rest days do influence performance in different ways for different teams, homecourt advantage remains the most impactful variable for outcome prediction of a game.



Wednesday, May 30, 2012

Home-court advantage in the NBA

We have heard the concept countless times and almost take it for granted.

Home-court advantage. "They should win the next since they're playing at home." "They managed to steal one on the road."

But doesn't it all come down to Team A versus Team B? If A is a better team it should win no matter the location. The court has the same dimensions, the baskets are identical, the shots have the same likelihood of falling in. It's not like being server or receiver in a tennis game.

Or is it? There has actually been quite some studies around home-court advantage in an attempt to tease out the external factors that could cause it. Many potential causes have been brought forward: the home crowd of course, cheering when the hometeam gains momentum. The fact that the players in the home team can sleep at home instead of being in a hotel. Familiarity with the locker rooms, the facilities in general. Mostly psychological explanations difficult to accurately measure.

Or just consider the distractions when shooting free-throws:


I don't have a degree in psychology so I will tackle from the data point of view, and try to see how playing at home can impact the game's outcome.


Data

I looked at all NBA games (regular season only) for the past two years. I didn't go further back as other factors could have come into play such as the differences in team composition.


Methodology

For each team, and for the league in general, I computed the empirical probabilities of winning at home and on the road. The values can be directly computed form the creation of a two-by-two table win/loss VS home/away. I choose to approach the problem via a logistic model which provides exactly the same point estimates but in addition provides confidence intervals which in turn allow me to determine whether the homecourt advantage is significant or not.


Results

If an NBA team plays against another NBA team, how does the probability of victory change depending on where the game is played?

With a simple model where the only variable is where the geme is played I obtained that for the NBA in general, playing at home offers a +19.8% winning probability (59.9% VS 40.1%). This was a very significant uplift. There we have it, homecourt advantage is very present in the NBA.

Does this +20% hold for all NBA teams?

Here the variability is much greater, ranging from +37.8% for the Denver Nuggets (81.7% at home vs 43.9% away) to only 2.4% for the Dallas Mavericks (69.5% at home VS 67.1% away).
Full team-by-team table is at the end of the post.

Does any team player better away than at home?

No, all teams play better at home, even if by only a little such as the Dallas Mavericks where the uplift is only 2.4%.

Is homecourt advantage statistically significant for all teams?

Homecourt advantage was significant for all teams except 8: Dallas Mavericks (+2.4%), Miami Heat (+3.7%), Boston Celtics (+9.8%), Philadelphia 76ers (+9.8%), Oklahoma City Thunder (+11.0%), Sacramento Kings (+11.0%), Houston Rockets (+13.4%), New York Knicks (+13.4%). These are not necessarily all good or bad teams, just teams that are just as good (or bad) away as at home.


Closing thoughts

Going back to Denver and its overwhelming homecourt advantage (+37.8%) compared to the second-best advantage +32.9% for the Los Angeles Clippers, what might wonder if the altitude isn't the Nuggets biggest fan...


In a later post I will explore how the model can be tweaked to take rest time between games into account. Does more rest improve your winning probability?


Appendix: Full table

Team Home win % Away win % Delta % Significance
DAL 69.5% 67.1% 2.4% No
MIA 65.9% 62.2% 3.7% No
BOS 69.5% 59.8% 9.8% No
PHI 46.3% 36.6% 9.8% No
OKC 69.5% 58.5% 11.0% No
SAC 35.4% 24.4% 11.0% No
HOU 58.5% 45.1% 13.4% No
NYK 50.0% 36.6% 13.4% No
MIN 26.8% 12.2% 14.6% Yes
CLE 57.3% 40.2% 17.1% Yes
LAL 78.0% 61.0% 17.1% Yes
POR 68.3% 51.2% 17.1% Yes
UTA 64.6% 47.6% 17.1% Yes
ORL 76.8% 58.5% 18.3% Yes
PHO 67.1% 47.6% 19.5% Yes
NBA 59.9% 40.1% 19.8% Yes
CHI 73.2% 52.4% 20.7% Yes
NJN 32.9% 11.0% 22.0% Yes
ATL 70.7% 47.6% 23.2% Yes
DET 46.3% 23.2% 23.2% Yes
MIL 61.0% 37.8% 23.2% Yes
SAS 79.3% 56.1% 23.2% Yes
MEM 64.6% 40.2% 24.4% Yes
TOR 50.0% 25.6% 24.4% Yes
NOH 63.4% 37.8% 25.6% Yes
WAS 42.7% 17.1% 25.6% Yes
IND 57.3% 26.8% 30.5% Yes
CHA 63.4% 31.7% 31.7% Yes
GSW 53.7% 22.0% 31.7% Yes
LAC 53.7% 20.7% 32.9% Yes
DEN 81.7% 43.9% 37.8% Yes