Showing posts with label kobe bryant. Show all posts
Showing posts with label kobe bryant. Show all posts

Thursday, February 5, 2015

Are we seeing All-Stars at the All-Star?


The starters for the Western and Eastern teams of the upcoming NBA All-Star game were just announced Jan 22nd. The selection was uniquely based on fan votes.

In the West we have the vote-leading player Steph Curry, along with Marc Gasol, Blake Griffin, Kobe Bryant and Anthony Davis. Their Eastern counterparts will be Pau Gasol (not sure how often two brothers have faced each other in an All Star Game...), LeBron James, Kyle Lowry, John Wall and Carmelo Anthony.

The selection did raise quite a few eyebrows to say the least. Kobe? Sure he's an NBA legend, future hall-of-famer and all, but look at the Lakers record this season, look at his abysmal shooting percentage of 37.3%. Carmelo is also somewhat of a surprise given how the Knicks are performing this year. Sure the All Star is not about the team but the player, but his stats aren't eye-popping either. And then consider all the ones who didn't get in, James Harden, Klay Thompson, the entire Atlanta Hawk roster... Even if not for those reasons but purely on the voting volume, Mark Cuban declared the voting system broken.


fivethirtyeight.com had a very interesting post on the topic, attempting to correlate players' performance with the number of votes received. Performance was measured in terms of Win Above Replacement (WAR), the number of team wins attributable to that player (computed as the difference between the number of wins the team got with that player in the game, versus a hypothetical world where the player is replaced by an average player). It does seem that the above a certain threshold, high-impact players get the votes they deserve, but under that threshold it's all more or less random.

Now I think the real question is: what do we want in an All Star game? Players naturally view it as an honor, a testimony of a great year they're having. But are fans voting for players deserving recognition? Or do they want pure 100% showtime? Imagine a natural born dunker, explosive, athletic and artistic at the rim. Even if that player had below average EFG%, below average WAR, RPM, RAPM or any of the other advanced metrics to measure player performance, wouldn't fans still want to see him in the All Star game?


So while I'm not saying it's fair to the players, I can understand why a Kobe would get voted in, and why a Paul Millsap or Kyle Korver wouldn't. If we really want to understand how fans vote, it would be interesting to see if we could find a metric that better correlates with player votes than WAR. Or perhaps first start including WAR for past seasons as well? I'm sure that if we did that we would have a better understanding as to why Kobe got voted. But how about a combination of team wins + number of dunks in the season? Or number of fast break points?

The debate does seem old and familiar, perhaps because it's so closely related to the one we have every single year about who should be MVP and how MVP is defined? The player with the stellar stats? The player who was most impactful on his team's success?



Monday, December 15, 2014

Clutch or not clutch, that is the stat question


There has always been debate on what clutch is, how it is measured, who performs best in these moments, the list goes on...
I'm definitely not going to settle the debate once and for all in one blog post, but wanted to share a few thoughts and ideas instead.

When we talk about clutch in the NBA, a few names immediately come to mind, Jordan, Bryant, Bird, Miller and Horry. Countless lists can be found with the simplest search, all as subjective as the next ("that was an incredible play in that game!").

But can we actually measure it, and rank players by it?

nba.com has a whole section dedicated to clutch stats on its website. A great first step but amidst all the numbers it's hard to compare players.

SBNation also tackled the issue, clustering players into recipients, creators and scorers (not mutually exclusive) during clutch time (Is Kobe Bryant Actually Clutch? Looking At The NBA's Best Performers In Crunch Time). The article stresses the importance of efficiency by placing all performances in perspective using possessions per 48 minutes on the x-axis. Here are the results for the 2010-2011 season:


Efficiency is a trademark at SB Nation, and Kobe is their primary scapegoat given this viewpoint:


SBNation's perspective is interesting and allows swift comparisons across players, but I feel that it lacks some rigor and robustness around these numbers. How large are the sample sizes? Are the effects significant? Which players are shouldering the most pressure and confronting it head-on? The author underlines these issues himself:

But all of that said ... how reliable are these numbers? There's a school of thought that firmly believes that "clutch" is in the eye of the beholder. They contend that as fans, we see things that may not actually be there. We see Kobe hit a step-back 20-footer and credit his clutch ability, when perhaps we simply should have attributed it to the fact that he's amazing at basketball (in the 1st or 4th quarter).
There are rigorous methods of testing for statistical significance. Rather than dive into those, however, a glance at some yearly efficiency trends can be just as telling.


I also came across this very nice post, Measuring Clutch Play in the NBA, on the Inpredictable blog which offers an interesting and elegant alternative. In a nutshell, the idea is to look at how each player's actions impacted his team's probability of winning the game, referred to as Win Probability Added WPA. Made shots, rebounds, steals increase your team's probability, while missed shots and turnovers hurt it. Some adjustments are required to clean up the cumulative WPA for each player (essentially comparing the impact of the same play under normal circumstances), but it does at the end provide an intuitive metric that makes sense and allows quick comparisons.

I do however have some slight concerns with this metric. The first is that, unless I misread, the metric is cumulative, so that players with more minutes in the clutch have more opportunities to modify their team's WPA. The second is best illustrated with a small example: with a few seconds remaining, if a player makes a two-point shot with his team down by 2 or down by 1, it will make a huge impact on the WPA: in the first case they're tied, likely to go to overtime with 50/50% for each team to win the game, in the other case his team leads by 1 and have a good chance of winning the game. But is it fair to credit the player with very different WPA in both cases? What really matters is that, under tremendous pressure, the player made the shot.


This in turn leads to another question: what was the likelihood of that shot going in in the first place? How frequently does that player make that shot under normal circumstances without the game on the line? How frequently do other players make the shot? How much does clutch pressure reduce the average player's chance of making the shot, and was the player able to rise to the occasion and overcome the pressure?

According to Stephen Shea in his book Basketball Analytics, "90% of teams performed worse (in terms of shooting percentages) in the clutch than in non-clutch situations." Can this me modeled? How significant is the effect?

I will try to explore this path further, looking into statistical models that would offer some elements of response to these questions.

But looking at all hat has been said it seems the debate originates from the fact that "being clutch" is never well-defined. Suppose we could at any point during a game give a score from 0 to 100 as to how good a player is. Suppose player A is at 90 throughout non-clutch times, but drops to 80 in clutch situations. Whereas player B is at 60 in non-cutch situations, but steps his game up to 70 when the game is on the line. Which is clutchier? The one with highest absolute value, or the one stepping up his game and taking the pressure head-on. Answering this would already be a giant step in the right direction.

In the meantime, please enjoy this youtube compilation of clutch shots:

Thursday, April 18, 2013

Who's searching for Steve Nash?



In a previous post we analyzed the trend and seasonality in Google searches for Kobe Bryant. We looked at another famously injured NBA star, Derrick Rose, as well as Lebron James. Let's turn our attention to another star who made lots of headlines for switching teams: Steve Nash.

Here's the raw data from Google Trends:



Using the same decomposition analysis detailed in the Kobe Bryant post, we get the following results:








































We're starting to become experts analyzing these decompositions now.
One interesting insight is that, similarly to Derrick Rose and unlike Kobe and Lebron, Steve's peaks are indeed during the Playoffs, but in May as opposed to June, indicating that while he often makes the Playoffs he's never played a game in the Finals...

Of course we see a giant spike in the 2012 offseason, right after the 2012 Playoffs. Well Steve decided to take his talent from Phoenix to LA. We all know since then that the Lakers have not really impressed us this last season, but at the time the Steve-Kobe sent shockwaves worldwide!


Monday, April 15, 2013

Who's searching for Kobe Bryant?




Unless you live on another planet without any WiFi connection to Facebook or Twitter, it was pretty hard to miss out on Kobe Bryant's injury in a crucial game against the Golden State Warriors last week. A torn Achilles has just ended his season (although to be fair I am not sure he will be missing that many games, even if Lakers do make the Playoffs...).

But this got me thinking: how many people searched for "Kobe Bryant" on the day and next of his injury? Could a huge spike be seen in Google Trends?




The above chart is at the monthly level so we only have partial April data. Nonetheless, Google Trends predicts that April 2013 will be Kobe's all star spike in searches, his MVP (most valuable peak) if you will.

But we also seem to notice some older spikes in the data, and a closer look reveals that these tend to regularly occur in June. What type of seasonal pattern lurks in the data?

In order to investigate seasonality, and also to tease out the underlying trend in searches for Kobe, I processed the data using R's stl() function (seasonal decomposition of a time series with loess) to split a time series into its trend, seasonal component and residuals.

Here's what I obtained:



We can't see the April spike yet because the downloaded data doesn't have Google Trends' estimates for the month, but we can predict a big spike given that halfway through April we already have almost as many queries as in a full month.

Going back to the figure, the top panel has the raw monthly data since 2004.

The next panel displays the smoothed loess trend. After a dip throughout 2004, Kobe has slowly but surely gained momentum since.

The next panel shows the seasonal component. We easily identify a strong pattern with a low volume in the off-season, an uplift when the season starts which remains quite constant during the regular season, another uplift during the playoffs with the maximum attained in June as we had observed previously.
It's also interesting to see how the seasonality's amplitude has steadily increased over the years.

Finally, the last panel shows the residuals, what cannot be explained by the trend and the seasonality. The cluster of yellow bars in the summer of 2004 could be due to the fact that Shaquille O'Neal was traded away from Los Angeles and there were many rumors around Kobe's future with the Lakers.

But how big the spike will be in April 2013 is anybody's guess!




Tuesday, January 22, 2013

From "Will they be champs?" to "Will they make the playoffs?"


I have not posted any basketball-related articles since sharing my data-driven predictions for the 2012 NBA Finals (unfortunately for me, the data suggested the Oklahoma City Thunder had a slight advantage against the Miami Heat).

But there has been SO much talk recently on the Los Angeles Lakers' performance that I just had to take a stab at predicting their season outcome.


Background

For those of you not too familiar with the situation, it has all the drama of a Hollywood script, and I will briefly summarize the situation.

The Lakers have had great results in the past years, winning the championship in 2009 and 2010, and being tough opponents in the other years. Over the 2012 summer and a disappointing 2012 year, they signed a couple of superstars: Steve Nash from Phoenix, desperate to win a championship and one of the league's top point guards, as well as Dwight Howard from Orlando, who can be quite a beast under the basket. On paper they were an All-Star team, nobody contested that. Almost everyone had them reaching the Finals, the only question was whether they would beat Miami or not.

But their performance so far surprised even the most pessimistic: they started the season by accumulating losses, and have lost more games than they have won since the start of the season. Things get worse as they are in the Western Conference which is much more competitive than the Eastern conference. Add in a fired coach, a surprising replacement coach, and you have all the drama necessary for a lot of ink to be poured!

Total confusion on the court has been a common sight in the Staples Center:

And so the questio has become: will the Lakers even make it to the NBA playoffs? Only the best eight teams of the Western conference make the playoffs. Many analysts have looked at historical data, noticing that the eight team in the West has an average of 48 wins, and if the Lakers want to reach 48 wins this season they need to start winning fast.

I've decided to take a slightly different approach, by tring to forecast what the rest of the season would look like for all teams based on the first half of the season. Also, the previous methodology ignores what the rest of the teams are doing in your conference, as well as the impact of playing another team also fighting for a playoff spot. When you win against one of those teams, it basically counts double!


Methodology

Here's the methodology I've taken:
  • extract all of the 2012-2013 season scores and schedule data for all teams, namely home and away winning percentages
  • for all upcoming games I compute the home team's probability of victory as follows:
    home team winning percentage / (home team winning percentage + away team winning percentage).
    For example: Let us say that Chicago with a 12-5 road record (70.8% road winning percentage) is visiting Miami with a 16-3 home record (84.2% home winning percentage), I would estimate Miami's probability of winning the game as:
    84.2% / (84.2% + 70.8%) = 54.4%.
  • Based on this probability, I simulate the game's outcome and then update Chicago's and Miami's records (if Chicago wins, their road record is now 13-5 and Miami's home record becomes 16-4).
  • I continue proceeding that way until all the season's games have been simulated. This then becomes one possible outcome for the season.
  • I then repeat the entire process outlined above thousands of time to get an idea of all the scenarios that can play out.
Based on the simulations, I can then determine which teams are most likely to have the best record in the NBA, and, going back to our original question, how likely the Lakers are to make the playoffs.


Results

Here are the results (based on scores up to 21/01/2013):

Top three NBA records:
  • First place:
    Oklahoma City Thunder (32.9%), San Antonio Spurs (29.2%), Los Angeles Clippers (28%)
  • Second place:
    San Antonio Spurs (26.6%), Oklahoma City Thunder (24%), Los Angeles Clippers (23.4%)
  • Third place:
    San Antonio Spurs (18.7%), Oklahoma City Thunder (17.6%), Los Angeles Clippers (17.1%)
Clearly, the odds are that these three teams will enter the playoffs with the best records and so most likely a Western Conference team will have homecourt advantage in the Finals!

As for the Lakers, here are the probabilities of their final standing in the Western conference:

Position Probability
6 0.4
7 1.4
8 4.1
9 8.9
10 12.6
11 18.1
12 24.2
13 16.9
14 10
15 3.4

Wow! Only a 5.9% probability of the Lakers making the Playoffs!

But let's still monitor them closely, and I'll update these results in the weeks to come.

Please leave a comment if you have any questions/suggestions on the methodology and/or results!


  

Thursday, May 17, 2012

Lakers - Thunder Series

This post is actually an expanded comment to Sekou Smith's Hang Time Blog on nba.com concerning the Lakers - Thunder series.

This series is one everybody has been waiting for since the start of the season.

Experience VS youth.
Kobe VS Kevin.


VS



A blowout in game 1.
An incredible comeback in game 2.

What's in store for the next 2 + X games?

5 nba.com's experts on Sekou's blog give their predictions after 2 games: one says 4, two say 5, and 2 say 6.

But what do the stats say?

I recently updated my model from the last two posts (here and here) in two ways: homecourt advantage is now incorporated (another post soon on this topic, namely how we can quantify it, whether all teams have a significantly higher probability of winning at home than on the road, and which teams have the greatest delta in home ganes vs away games), and by providing more details on each series with not only the probability of one team winning it but also the breakout in how many games the series will play out.

Which is exactly what I did here for the Lakers - Thunder series.

And now for the results:

Winner Number of games Probability
Thunder4 22.2%
Thunder 5 30.4%
Lakers 6 5.8%
Thunder 6 17.2%
Lakers 7 9.5%
Thunder 7 14.9%

So Thunder in 5 is actually the most likely scenario, followed by Thunder in 4 and in 6. Overall, if you're a Laker fan you should feel depressed with Lakers having only a 15.3% probability of facing the Spurs. But I have to admit that I haven't factored Kobe-back-against-the-wall variable in my models :-)

Let me know your thoughts!

Monday, April 23, 2012

NBA player rankings


So in addition to boardgames, I am also a big NBA fan, and luckily NBA and stats mix really well.
A topic that has often been covered is how to rank players? Which player has the greatest impact? Which is the greatest player of all time (well, that one's easy ;-) )?


The debates rage because the questions are so vague and open to interpretation. What does it mean for a player to be a better basketball player than another? Does it mean having better stats? If I score more, rebound more, assist more, steal more, turnover less, clearly I am better than you? Another approach is to compare win/loss records when the player plays or sits out, although for most players it will be difficult to have a good sample size for the "sitting out" observations. A new metric that has emerged and solves this "sitting out" problem is the plus/minus statistic, which keeps track of the score before and after a player enters the game. So say a player enters the game with game tied at 10, and leaves it with his team ahead by 5. That's +5 for him. He re-enters the game with his team ahead by 10, and leaves it (without coming back) with his team only ahead by 2. That's -8. With the earlier +5 that's an overall -3 for that player in that game.

Today I wanted to look at a different approach, a more statistical approach. I have no clue where it will lead me, but after different trials and errors and tweaking here and there I hope to come up with a new interesting way to rank players.

Ultimately, what we care most about is wins. Sure it's great to score 100 in a game, but if you lose that game that's just wasted effort. So the idea is to find a relationship between a player's efforts and the impact it has on the game. In other words, how does a player's stats in a game change the probability of winning the game?


In terms of the data, I looked at the past six seasons (not including 2011-2012), and for each player looked at his stats with the game outcome for all games played. I only considered players with at least 50 wins and 50 losses, playoffs not included.

As our variable of interest is a probability (of winning the game), we naturally turn towards a logistic regression. We are not directly modeling the probability as a linear combination of the covariates but rather the log odds: log(P(win) / (1 - P(win))). The interpretation of the coefficients will not be entirely straightforward but will still allow us to rank players. Which player has the greatest coefficient, and has the greatest impact on the log odds and thus the probability of winning the game?

Well it depends on our covariates. Since we do want to find an easy way to rank, it's best to only consider one covariate.

Points

Let's naively only consider points scored. How does scoring an extra point improve the log odds?
The top 5 impactful players are (in order): Calvin Booth, Greg Ostertag, Antonio Davis, Anderson Varejao and Bruce Bowen.

All metrics

Points might be too restrictive, since a player can have an impact without scoring. So let's consider (points + rebounds + steals + assists + blocks - turnovers), referred from here onwards as "all metrics" as the covariate.
The top 5 impactful players are (in order): Bruce Bowen, Calvin Booth, Eddie Griffin, Kevin Durant, Antonio Davis.

Minutes played

If we were to suspect that a player's impact is difficult to track with simple metrics only, let us take minutes played as a proxy for everything observed and not observed (good defense, good picks...)
In this last case, the top 5 impactful players are (in order): Kevin Durant, Othella Harrington, Gerald Wallace, Eddie Griffin, Zach Randolph.

Where are the superstars?

It's interesting to see the same names come up, and to notice that aside from Kevin Durant, none of the players have superstar status.

Talking of superstars, where are they in the rankings?

Out of the 479 players considered, here is how some superstars ranked respectively for points, all metrics, and minutes played:
Kobe Bryant: 363, 333, 479
LeBron James: 247, 101, 475
Kevin Garnett: 396, 416, 478

Wow.

Caveats

As I was mentioning, I am discovering these results in almost real time with you, and still a little unclear how to interpret them myself. There are a lot of things that could hurt the analysis, namely the fact that the coefficients are hear interpreted as "change in log odds for an additional unit increase in the covariate". But an additional point for Kobe isn't exactly the same thing as an additional point for Eddie Griffin.

We also have a case pointed out in "Superfreakonomics" about ranking good surgeons. Looking at patient death rate for instance can be misleading because of selection bais. People with more critical conditions will go see the better surgeon but because of there condition increase the risk of increasing the surgeons death rate because of the very critical condition. Bad doctors only seeing healthy patients will have impeccable track records. Similarly in basketball, it could be argued that when the game is on the line you will go to your superstars that will have to play exceptionally well to win the game, whereas you might put all your bench in the game when the game has already been won for a while.

There is definitely room for improvement, but I will continue to explore this approach to try to identify lesser known players that have strong yet unnoticeable impacts on the game.
Stay tuned!