Wednesday, April 26, 2017

Guy Molyneux and Joshua Miller debate the hot hand

Here's a good "hot hand" debate between Guy Molyneux and Joshua Miller, over at Andrew Gelman's blog.

A bit of background, if you like, before you go there.

-----

In 1985, Thomas Gilovich, Robert Vallone, and Amos Tversky published a study refuting the "hot hand" hypothesis, which is the assumption that after a player has recently performed exceptionally well, he is likely to be "hot" and continue to perform exceptionally well.

The Gilovich [et al] study showed three results:

1. NBA players were actually *worse* after recent field goal successes than after recent failures;

2. NBA players showed no significant correlation between their first free throw and second free throw; and

3. In an experiment set up by Gilovich, which involved long series of repeated shots by college basketball players, there was no significant improvement after a series of hits.

-----

Then, in 2015-2016, Joshua Miller and Adam Sanjurjo found a flaw in Gilovich's reasoning. 

The most intuitive way to describe the flaw is this:

Gilovich assumed that if a player shot (say) 50 percent over the full sequence of 100 shots, you'd expect him to shoot 50 percent after a hit, and 50 percent after a miss.

But this is clearly incorrect. If a player hit 50 out of 100, then, if he made his (or her) first shot, what's left is 49 out of 99. You wouldn't expect 50%, then, but only about 49.5%. And, similarly, you'd expect 50.5% after a miss.

By assuming 50%, the Gilovich study set the benchmark too high, and would call a player cold or neutral when he was actually neutral or hot.

(That's a special case of the flaw Miller and Sanjurjo found, which applies only to the "after one hit" case. For what happens after a streak of two or more consecutive hits, it's more complicated. Coincidentally, the flaw is actually identical to one that Steven Landsburg posted for a similar problem, which I wrote about back in 2010. See my post here, or check out the Miller paper linked to above.)

------

The Miller [and Sanjurjo] paper corrected the flaw, and found that in Gilovich's experiment, there was indeed a hot hand, and a large one. In the Gilovich paper, shooters and observers were allowed to bet on whether the next shot would be made. The hit rate was actually seven percentage points higher when they decided to bet high, compared to when they decided to bet low (for example, 60 percent compared to 53 percent).

That suggests that the true hot hand effect must be higher than that -- because, if seven percentage points was what the participants observed in advance, who knows what they didn't observe? Maybe they only started betting when a streak got long, so they missed out on the part of the "hot hand" effect at the beginning of the streak.

However, there was no evidence of a hot hand in the other two parts of the Gilovich paper. In one part, players seem to hit field goals *worse* after a hit than after a miss -- but, corrected for the flaw, it seems (to my eye) that the effect is around zero. And, the "second free throw after the first" doesn't feature the flaw, so the results stand.

------

In addition, in a separate paper, Miller and Sanjurjo analyzed the results of the NBA's three-point contest, and found a hot hand there, too. I wrote about that in two posts in 2015. 

-------

From that, Miller argues that the hot hand *does* exist, and we now have evidence for it, and we need to take it seriously, and it's not a cognitive error to believe the hot hand represents something real, rather than just random occurrences in random sequences. 

Moreover, he argues that teams and players might actually benefit from taking a "hot hand" into account when formulating strategy -- not in any specific way, but, rather, that, in theory, there could be a benefit to be found somewhere.

He also uses an "absence of evidence is not evidence of absence"-type argument, pointing out that if all you have is binary data, of hits and misses, there could be a substantial hot hand effect in real life, but one that you'd be unable to find in the data unless you had a very large sample. I consider that argument a parallel to Bill James' "Underestimating the Fog" argument for clutch hitting -- that the methods we're using are too weak to find it even if it were there.

------

And that's where Guy comes in. 

Here's that link again. Be sure to check the comments ... most of the real debate resides there, where Miller and Guy engage each other's arguments directly.






Labels: , , ,

Tuesday, August 06, 2013

Basketball robot shoots worse than some humans

Today, at the Carnegie Science Museum in Pittsburgh, I watched their resident basketball robot shoot some free throws.  Here's a video of what it looks like.

How accurate do you think the robot is?  Take a guess before you read on.  (Or, just read on -- who am I to give you orders?)

The answer is ... not that accurate.  Well, at least, a lot less accurate than I thought.  When I was there, the robot's FT% was only 83 percent (405 for 488).

I was a bit shocked.  I expected close to perfect.  After all, it's the same throw under exactly the same circumstances, every time.  (Actually, it's two different shots: sometimes the robot throws underhand, and sometimes overhand from behind his back.  But, I witnessed the robot missing shots from both positions.)

It seems wrong, doesn't it, that a human can outperform an expensive robot at a repetitive physical task?  In his career, Rick Barry routinely shot over 90 percent (albeit underhanded).   So, the machine misses almost twice as many shots as Barry. 

What's going on?  I don't know.  

For what it's worth, here's my theory:

The robot lets the ball roll down the ramp that's his "hand" before actually doing the throw with his "arm".  Maybe the position it reaches varies randomly, based on random differences in friction.  Maybe if a dirty part of the ball contacts a dirty part of the arm, the ball doesn't quite reach the expected point, and the throw misses.

There could be other friction-related issues that cause variation, like, perhaps, the axis of rotation of the ball when it's released.  (The robot hits the backboard every time.)

Any physicists reading who can deliver a more informed hypothesis?

------

If a human can outperform a robot, then it must be that he does *something* better than the machine does.  What? 

My guess is: when the human shoots, he can notice if something's a little off, like the ball slips a bit.  In that case, he can adjust his motion on the spot to try to counter that.  The robot, of course, doesn't do that.

That's the only thing I can think of.  

-------

I'd love to hear other opinions, because I'm very, very surprised.  I would have bet good money that you could easily make a robot that shoots, say, 98 percent.  Could it really be that tiny differences caused by friction could make such a big difference in outcomes?



Labels: , ,

Tuesday, February 12, 2013

Gut reactions about foul shooting and luck

This is just something I've been thinking about lately.  I'm not sure what my point is, or even if I have a point, but I'm finding it interesting.

For the past few seasons, Dirk Nowitzki has been one of the best free throw shooters in the NBA, with success rates ranging around 90 percent.  Now, let me give you six scenarios.

1.  One game, Nowitzki has a bad night and goes 3 for 6.  And, the Mavericks play roughly as expected but lose by a point.  Would you say Dirk was unlucky?  Would you say the team was unlucky?  

2.  In an alternate universe, the NBA decides that free throws are boring for fans to watch.  So, before the season starts, every player shoots a few hundred free throws.  The scorekeepers have the results.  Any time there's a foul, they go to the sheets, take the player's next shots off the list, and adjust the score accordingly.

Nowitzki still shot 90 percent overall.  But this game, they were at the part of his sheet where had three misses in six shots.  The Mavs lose.

Again, how unlucky was Nowitzki?  How unlucky was the team?  More, or less than the previous scenario?

3.  Same situation, but, this time, instead of progressing on the sheet in the order in which Nowitzki took his 800 preseason shots, the scorekeepers choose randomly.  There's a big urn labled "Nowitzki 90%," with 720 white balls and 80 black balls.  Whenever Nowitzki gets fouled, they reach into the urn.  

This game, they happen to pull three white balls and three black, and the Mavericks lose by a point.  Again: how much luck would you say determined the outcome?

4.  In another alternate universe, all free throws are saved up and taken at the *end* of the season, before the playoffs.  Nowitzki winds up shooting 90 percent, as usual.  But, when it comes time for the shots that are going to apply to *this* game, they tell Dirk, "you need to sink four to tie, five to win."  To everyone's surprise, Nowitzki sinks only three, and the Mavericks lose.  

Does that make a difference?

5.  This time, Nowitzki is not told the score, or what game his end-of-season free throws are going to apply to.  But the session is televised, and the announcers know.  They tell the viewers.  There is tension in the booth.  Nowitzki steps up and hits 3 of 6 again.  

If Nowitzki didn't know the situation, does that make it feel more unlucky?  

6.  Finally ... almost the same thing, but, this time, Nowitzki starts by just taking the shots, without anyone knowing which game they apply to.  It's only afterwards that they determine the game, by randomly spinning a big wheel.  They get to the part where Nowitzki's next six shots happen to be 3 makes and 3 misses.  The NBA spins the wheel, and it comes up that particular game, and they add Nowitzki's 3-for-6, and the Mavericks lose by 1.

This one is the most unlucky so far, right?  

And what if Nowitzki had taken the 800 foul shots *before* the season, but the wheel was still spun after?  To me, that scenario seems the most like bad luck.

------

As I said, I don't have a specific point here.  I just find it interesting how my gut reaches different conclusions about each situation, even though my brain thinks they're really almost the same.  







Labels: , ,

Tuesday, January 01, 2013

Foul shooting and luck

I always think that performance in sports is a combination of luck and skill.  When an 80 percent free-throw shooter steps to the line, I think what happens is almost all luck. To me, the shot is like the player flipping a coin that lands heads 80 percent of the time. 

Where's the skill?  The skill is in the player having practiced enough to bring his coin up to 80 percent heads.  Personally, I'm probably a 25 percent shooter, at best.  But if I practice and practice, I'll get better; maybe eventually I'll hit 50 percent.  Once I do, my shots are like flips of a fair coin.

Some people disagree.  They reject the idea that there can be any luck at all.  It's just the player, and the ball, and the basket.  It's all under the player's control.  Where's the luck?

I think our disagreement is that they're thinking of luck only as "external" luck, something outside the player -- like a bad bounce in the infield, or getting dealt good cards in poker.  I think there can also be "internal" luck, even though if everything is, in theory, within the player's control.

-------

We can probably all agree that coin flipping is a good example of luck.  But in a way, coin flipping and foul shooting are almost the same.  If I'm doing the flipping, and I'm trying to land a head ... it's all within my control, the same way as shooting a basket.

The difference, you might argue, is that it's impossible to influence the outcome of a coin flip.  I may *want* it to land heads, but I just can't do it.

Except that ... I think I *can* do it.  Not all the time, but better than 50 percent.  I bet you if I practiced, I could make it land heads more than tails.  Suppose that, flipping a coin that starts head up, I try to make the coin rotate exactly 10 times, and land flat in my hand.  That's hard, but ... who's to say that I can't do it often enough to get 50.0001 percent heads?  That would mean out of one million flips, I manage to convert one tail into a head because of the skill I developed.  Doesn't seem unreasonable. 

So after I've achieved that amount of skill, you hand me a coin, and I flip it, and it lands heads.  Was it luck, or skill? 

Everyone should agree that it's mostly luck.  My skill is being able to increase my chances from 500,000 out of a million, to 500,001 out of a million.  After that, the individual flip is just luck. 

It has to be, doesn't it? 

-------

In the comments to the previous post, someone wrote that foul shooting *seems* like it's all skill because the success rate is so high.  That seems right to me; we have a default assumption in our brains that either we can do something, or we can't.  If we can, it seems like we should be able to do it all the time, and when we don't, it's our fault, we made a mistake.  At least, to me, that's how it *feels*.  When I screw something up, I feel like, geez, I did that before so many times, why couldn't I do it now?  I feels like I choked, or did something wrong.

And it might have been a choke, actually.  Even though we shoot at 80 percent normally, maybe this time we chose not to concentrate enough.  Maybe we should have bent our knees more.  Maybe we were too impatient, or too nervous.

The thing is, it *could* have been all those.  But it probably wasn't.  It's just our nature to think, we could have done it differently!  And, yes, we could have, just like we could have flipped the coin harder, or softer.  But, so what?  There was no way of knowing in advance that *this time* was one of the 20 percent what we did wasn't going to work. 

And, in a way, it's a good thing we think it was our fault.  Because it's still *possible* that we did something wrong, and it's *possible* that if we figure out what it is, mental or physical, we could get our success rate up from 80 percent to 81 percent.  Or, it could be that we've slipped up and forgotten something we knew, so that if we don't fix it, we'll fall to 78 or 79 percent. You don't want to just think, "Hey, I won't worry about it, it was just luck." 

But it probably was. 

-------

None of this is to say that preparation and concentration and patience aren't important.  But, you'd expect, professionals would have learned to do that all the time.  That, I think, is one of the underrated skills that professional athletes have: the ability to focus and concentrate much better than we do. 

I can probably hit my trash can with a wad of paper most of the time.  It's probably 60 percent when I don't care, and 80 percent when I aim. 

Now, if I miss, it might certainly be that I didn't care, and you might be reasonable in assuming that it wasn't just luck, that it was mostly within my control -- that if I cared more, I'd have made it.  In fact, if half the time I don't care, then, I'm missing 30 percent of the time, rather than just 20 percent if I always tried my hardest. 

In a sense, then, you might have the right to say that there's more than luck involved: that maybe my misses are 2/3 luck, and 1/3 bad decisions.

The thing about the pros, though, is: I think they have the ability concentrate and prepare *almost all the time*.  I have no evidence for this, but I bet an 80 percent NBA player is seldom, if ever, even a 75 percent player, for reasons under his conscious control. 

In any case, there still has to be a substantial amount of randomness.  Because, there's no way an 80 percent shooter is able to be a 100 percent shooter, if he just takes the task more seriously.  That's just not plausible.

-------

If you agree with me that my coin flipping is luck (despite my 50.0001 success rate), but disagree that foul shooting is luck, I have one more argument for you.

Suppose in my coin flipping, I'm trying to get the coin to turn over exactly 20 times: that is, do ten 360s, for a total rotation of exactly 36,000 degrees.  And suppose you measure a large number of my tosses, and you find that my average is indeed 36,000, and my distribution is a bell curve with a large standard deviation.  You analyze the distribution and figure out that based on my accuracy, I should indeed wind up with heads 50.0001 percent of the time. 

Now, if you agree that my coin flipping is largely luck, you have to agree that my deviations from 36,000 are largely luck. 

Now, what if you did the same thing for free throws?  You measure all my free throws.  You find that my angle of release is normally distributed with a mean of 52 degrees, and a certain standard deviation.  And that the ball goes in when my angle is within 1.3 SDs either way, which works out to 80 percent success.

Isn't that the same thing as the coin?  If my deviations of coin flipping rotation are luck, then why aren't the deviations of release angle luck? 

--------

A summary of the overall argument is: there are natural limitations to how much humans can control their muscles.  With practice, you can develop "muscle memory," which increases the amount of control you have.  That is, you get the standard deviation to shrink.  But, humans just aren't capable of getting so precise that we can be 100% shooters of basketball free throws. 

We're all going to have our individual accuracy curves.  Good shooters will have a narrower and better-placed curve than poor shooters.  But, once you have the curve for your own personal muscle control, where you land on that curve, for any given shot, is luck. 

If you don't like using the word "luck," because it's internal to the brain, and it feels like it's under your control ... well, fine.  But then you can't call coin flipping luck, either.  Because, aside from the probability of success, it's almost exactly the same phenomenon.




Labels: , ,

Monday, June 06, 2011

Interpreting regression interaction terms

Last post, after talking about the results from the "choking foul shooter" study, I mentioned that there was one additional assumption I had to make. That assumption was that, in the regression, the coefficients for "last 15 seconds" and "down 1-4 points" were close to zero.

The easiest way to explain that is to go through what an interaction term means in a regression. (Warning: This is boring statistics stuff, no sports content until the end.)

------

Suppose I want to figure out if stimulants help a student do better on an exam. So I run a regression to predict the exam score. I use a bunch of variables, like age, time studying, performance on other exams, grades on assignments, number of classes missed, and so on, but I also include a dummy variable for whether the student had (both) coffee and Red Bull before the exam.

After the exam, I run the regression, and I find the coefficient for "both coffee and Red Bull" is -3, and statistically significant. I conclude that if I were a student, I might consider not taking both coffee and Red Bull.

Fair enough, so far.

But, now, suppose I do the same experiment again, but, this time, I add a couple of new dummy variables -- whether or not the student had coffee (with or without Red Bull), and whether or not the student had Red Bull (with or without coffee). I don't remove the original "had both" variable -- that stays in.

I run the regression again, and, again, the coefficient for "both coffee and Red Bull" comes out to -3 -- exactly the same as last time. What am I able to conclude this time about the desirability of drinking both coffee and Red Bull?

The answer: almost nothing. That coefficient, *on its own*, does not give much useful information at all about how performance is affected by the coffee/Red Bull combination.

Let me explain why.

-------

(First, a quick not on terminology. In a regression, the "both coffee and Red Bull" variable would be referred to as "the interaction of coffee and Red Bull". That would be written as "coffee x Red Bull," or a suitable abbreviation (In fact, I'm going to start referring to "coffee" as C, "Red Bull" as R, and "Coffee x Red Bull" as CxR). The "x" is a multiplication sign -- it's there because you can get the coefficient by multiplying together the dummy values for C and R. That is, if either C or R is zero, CxR equals zero; if both coffee and Red Bull are 1, then CxR equals one. That's exactly what we want.)

-------

In a regression result, the simplest way to interpret the coefficient of a dummy variable is, "what happens when you change the value from 0 to 1 and leave all the other variables the same." In the first regression, that works fine. But in the second regression, it can't work. Because if you change CxR and leave everything else constant, your data and regression become inconsistent. You wind up with CxR being 1 (meaning both coffee and Red Bull), but you'll have either C=0 (no coffee) or R=0 (no Red Bull). Those three variables are tied together, so you can't just change CxR and leave the other two constant.

Put another way, there are four possible combinations for C, R, and CxR:

C = 0, R = 0, CxR = 0
C = 1, R = 0, CxR = 0
C = 0, R = 1, CxR = 0
C = 1, R = 1, CxR = 1

You can't change CxR from 0 to 1, and still have a combination that's on the list. So the "change CxR but leave all other variables the same" strategy no longer works. If you change CxR from 0 to 1, you'll have to change one of the other variables, too.

-------

Which ones should you change? It depends what question you're trying to answer. For example, suppose you do the regression and you get these coefficients:

C = -5
R = -10
CxR = -3

If you're trying to ask, "what's the effect of taking coffee alone versus nothing at all," it's like asking, "what is the effect of changing (C=0, R=0, CxR=0) to (C=1, R=0, CxR = 0)?" The answer is -5.

If you're trying to ask, "what's the effect of taking both coffee and Red Bull versus nothing at all?", it's like asking, what's the effect of changing (C=0, R=0, CxR=0) to (C=1, R=1, CxR =1)?" The answer is -18.

And so on. But none of those kinds of questions lead to the answer of -3 points, because none of these questions can be answered by changing CxR alone.

So what does the -3 represent? The non-linearity of the coffee and Red Bull variables. Or, put another way, the "increasing or diminishing returns" to combining coffee and Red Bull. Or, put a third way, the effects of the *interaction* of coffee and Red Bull, independent of their individual effects. Or, put a fourth way, the amount of effects *duplicated* from both coffee and Red Bull, that you can't count twice even if you take both drinks.

The -3 is NOT any indication of whether it's a good thing to take coffee and Red Bull together. Even though the coefficient of the interaction is negative, coffee and Red Bull together might be a positive thing. Suppose the regression coefficients had looked like this:

C = +10
R = +20
CxR = -3

The CxR coefficient is still -3, but now look what happens:

Take coffee, score 10 points higher
Take Red Bull, score 20 points higher
Take both coffee and Red Bull, score 27 points higher!

In this case, you're still going to want to take both coffee and Red Bull. What the -3 is telling you is, there are diminishing returns to taking both. You might think that, since coffee improves you by 10, and Red Bull improves you by 20, that, if you take both, you'll improve by 30. That's not right. There are diminishing returns of -3, so, if you take both, you'll only improve by 27.

-------

Of course, if the coefficients of C and R are both zero, then the CxR variable is indeed the entire effect. So if coffee does nothing, and Red Bull does nothing, but, when you take them together, you lose 3 points ... in that case, the CxR variable actually IS the effect of taking both C and R.

-------

This is fairly standard stuff, I would think ... I looked for an explanation on the web, so I wouldn't have to type all this, but I couldn't find one.

Anyway, going back to the choking study ... there, I looked at a variable called "Last15 x Down1_4", which was the interaction of shots that happen in the last 15 seconds (a dummy variable called "Last15"), and with the shooting team up by 1 to 4 points (dummy variable "Down1_4").

It turned out that the coefficient for that was -0.058 (as compared to 11-point+ blowouts). I implied that meant that shooters were 5.8 percentage points worse in those clutch situations than in blowouts.

But that wasn't right, because "Down1_4" and "Last15" were also in the regression. It's like the "Coffee / Red Bull / Both" case. If I want to compare the effects of shooting in the last 15 seconds down by 1-4 points, against shooting where *neither* of those is true, I have to add up all three coefficients:

Down1_4 = A
Last15 = B
Down1_4 x Last 15 = -0.058

To get the true clutch effect, I have to compute A + B - 0.058. It could turn out that A and B are huge: maybe they're +7 points each! In that case, the effect would be 0.07 + 0.07 - .058, which would be +0.082 -- which would mean shooters were GREAT in the clutch.

The study doesn't give us A and B. However, the authors do tell us (and author Dan Stone reiterated in the comments to the previous post) that almost all the omitted coefficients are less than 0.01.

Still, suppose they are as high as exactly +0.01. That means that A + B - 0.058 would be -0.038, which would less significant a choke effect than I thought. Or, suppose they were as low as negative 0.01. In that case, players would be even chokier -- at -0.078.

That's why I added a note to the end of my post, saying I had to make one additional assumption. That assumption is that A and B were both close to zero. If they're exactly zero, the -0.058 stands.

Labels: , , , ,

Monday, May 30, 2011

A new basketball free throw choking study

"Performance Under Pressure in the NBA," by Zheng Cao, Joseph Price, and Daniel F. Stone; Journal of Sports Economics, 12(3). Ungated copy here (pdf).

------

There's a new paper demonstrating evidence that NBA players choke when shooting free throws under pressure. The link to the paper is above; here's a newspaper article discussing some of the claims.

Using play-by-play data for eight seasons -- 2002-03 to 2009-10 -- the authors find that players' percentage on foul shots goes down significantly late in the game when they're behind. They present their results through a regression, but it's more obvious just by using their summary statistics. Let me show you the trend, which I got by doing some arithmetic on the cells in the paper's Table 1:

In the last 15 seconds of a game, foul shooters hit

.726 when down 1-4 points (922 attempts)
.784 when tied or up 0-4 points (5505)
.776 when up or down 5+ points (4510)

With 16-30 seconds left, foul shooters hit

.748 when down 1-4 points (727 attempts)
.775 when tied or up 1-4 points (2652)
.779 when up or down 5+ points (6174)

With 31-60 seconds left, foul shooters hit

.752 when down 1-4 points (922 attempts)
.742 when tied or up 0-4 points (1634)
.767 when up or down 5+ points (8969)

In all other situations, foul shooters hit about .751 regardless of score differential (400,000+ attempts).

------

Take a look at the first set of numbers, the "last 15 seconds" group. When down 1 to 4 points, it appears that shooters do indeed "choke," shooting almost 2.5 percentage points (.025) worse than normal. In 5+ point "blowout" situations late in the game, they shoot more than 2.5 percentage points *better* than normal.

But neither of these numbers is statistically significantly different from the overall average (which I'm guessing is about .751). The difference of .025 is about 1.7 SDs.

The real statistical significance comes when you compare the "down by 1-4" group, not to the average, but to the "5+ points" group. In that case, the difference is double: the "down 1-4" is .025 below average, and the "5+" group is .025 *above* average. The difference of .050 is now significant at about 3 SDs.


UPDATE: The above paragraphs are incorrect in one aspect. Dan Stone, one of the paper's authors, corrected me in the comments. What I didn't notice was that in Table 1, the overall free throw percentage of each group was provided. Those percentages are .779 (down 1-4 group), .795 (up 0-4 group), and .782 (5+ group). So the average for those particular players is move like .787 than .751. All three groups shot below expected, but the "down 1-4" group shot WAY below expected.

So the "down 1-4" group is, on its own, statistically significant from expected, without regard to the other two groups. My apologies for not noticing that earlier.

So, if you look only at the last 15 seconds of games, it looks like players down by 1-4 points choke significantly compared to players who are up or down by at least five points.

There are similar (but lesser) differences in the 16-30 seconds group, and the 31-60 seconds group. I haven't done the calculation, but I'm pretty sure you also get statistical significance if you combine the three groups, and compare the "down 1-4" to the "5+" group.

-------

So that's what we're dealing with: when you compare "down 1-4 late in the game" to "up or down 5+ late in the game", the difference is big enough to constitute evidence of choking. The most obvious explanation is that the foul shooters in the two groups might be different. However, that can't be the case, because the authors controlled for who the shooter was, and the results were roughly the same. Indeed, they controlled for a lot of other stuff, too: whether the previous shot was made or not, which quarter it is, whether it's a single foul shot or multiple, and so on. But even after all those controls, the results are pretty much the way I described them above.

Again, I repeat: the authors (and the data) do NOT say that the "choke" group shoots significantly worse than average. They can only say that the "choke" group is significantly worse than one specific group of players: the "don't care" group, shooting late in the game when the result is pretty much already decided, with a gap of 5+ points.

But this fits in with the authors' thesis: that the higher the leverage of the situation, the more choking you see. They later break down the "5+" group into "5-10" and "11+", and they find that even that breakdown is consistent -- the 11+ group shoots better than the (slightly) higher leverage 5-10 group. Indeed, for most of the study, they compare to "11+" instead of "5+". For some of the regressions, they post two sets of results, one relative to the "5-10 points" group, and one relative to the "11+" group. The "11+" results are more extreme, of course.

-------

As I said, the authors don't present the results the way I did above ... they have a big regression with lots of tables and results and such. The result that comes closest to what I did is the first column of their Table 5. There, they say something like,

"In the last 15 seconds of a game, a player down 1-4 points will shoot 5.8 percentage points (.058) worse than if the game were an 11+ point blowout. The SD of that is 2.1 points, so that's statistically significant at p=.01."

-------

Oh, and I should mention that the authors did try to eliminate deliberate misses, by omitting the last foul shot of a set with 5 seconds or less to go. Also, they omitted all foul shots with less than 5 minutes to go in the game (except those in the last 15/30/60 seconds that they're dealing with). I have absolutely no idea why they did that.

-------

Although the authors do mention the "down 1-4" effect above, it's almost in passing -- they spend most of their time breaking the effect down in a bunch of different ways.

The biggest effect they find is for this situation:

-- shooting a foul shot that's not the last of the set (that is, the first of two, or the first or second of three);
-- in the last 15 seconds of the game;
-- team down exactly one point.

compared to

-- shooting a foul shot that's not the last of the set (that is, the first of two, or the first or second of three);
-- in the last 15 seconds of the game;
-- score difference of 11+ points in either direction.

For that particular situation, the difference is a huge 10.8 percentage points (.108), significant at 2.5 SDs.

Also: change "down by one point" to "down by two points", and it's a 6.0 percentage point choke. Change "not the last of the set" to "IS the last of the set," and the choke effect is 6.6 points when down by 1, and 6.0 points when down by 2.

This highly specific stuff doesn't impress me that much ... if you look at enough individual cases, you're bound to find some effects that are bigger and some that are smaller. My guess is that the differences between the individual cases and the overall "down 1-4" case are probably random. However, the authors could counter with the argument that the biggest sub-effects were the ones they predicted -- the "down by 1" and "down by 2" case. On the other hand, late performance is actually *better* than blowouts when the score is tied (by around 0.2 points), a finding the authors say they didn't expect.

So my view is that maybe the "1-4 points" result is real, but I'm wary of the individual breakdowns. Especially this one: in this situation:

-- last 15 seconds of the game
-- for a visiting team
-- where the most recent foul shot was missed
-- down by 1-4 points

the player is 9.6 percentage points (.096) less likely to make the shot than

-- last 15 seconds of the game
-- for a visting team
-- where the most recent foul shot was missed
-- score 11+ points in favor of either team.

Despite the large difference in basketball terms, this one's only significant at .05.

------

Another thing about the main finding is ... we actually already knew it. Last year, I wrote about a similar study (which the authors reference) that found roughly the same thing. Here, copied from that other post, are the numbers those researchers found, for all foul shots in the last minute of games, broken down by score differential:

-5 points: -3% [percentage points]
-4 points: -1%
-3 points: -1%
-2 points: -5% (significant at .05)
-1 points: -7% (significant at .01)
+0 points: +2%
+1 points: -5% (significant at .05)
+2 points: +0%
+3 points: -1% ("also significant")
+4 points: +1%
+5 points: -1%

There are some differences in the two studies. The older study controlled for player career percentages, instead of player season percentages. It didn't control for quarter (which is why commenters suspected it might just be late-game fatigue). It didn't control for a bunch of other stuff that this newer study does. And it used only three seasons of data instead of eight.

But the important thing is: the newer study's eight seasons *include* the older study's three seasons. And so, you'd expect the results to be somewhat similar. It's possible that the three significant seasons are enough to make all eight seasons look significant, even if the other five seasons were just average.

Let's try, in a very rough way, to see if we can tease out the new study's result for those other five seasons.

In the older study, if we average the -1, -2, -3, and -4 effects, we get -3.5. So, in the last minute, down by 1-4 points, shooters choked by 3.5 percentage points.

How do we get the same number for the newer study? Well, in the top-left cell of Table 5, we get that, in the last 30 seconds and down by 1-4 points, shooters choked by 3.8 percentage points.

That's our starting point. But the new study's selection criteria are a little different from the old study's, so we need to adjust.

First, the "-3.8" in the new study is from comparing to games in which the point differential is 11 or more. The "5-10" is probably a more reasonable comparison to the previous study. The difference between "11+" and "5-10", at 30 seconds, appears to be about one percentage point (compare the second columns of Tables 3 and 4). So we'll adjust that 3.8 down to 2.8.

Second, the new study is for the last 30 seconds, while the old study is for the last minute. From earlier in this post, we see that the 31-60 difference between the "down 1-4 group" and the "5+" group is only about -1.5 percentage points. Averaging that with the -2.8 from the above step (but giving a bit more weight to the -2.8 because there were more shots there), we get to about -2.4.

So we can estimate, very roughly, that for the same calculation,

Old study (three seasons): -3.5
New study (eight seasons): -2.4

Let's assume that if the new study had confined itself to only the same three seasons as the older study, it would have come up with the same result (-3.5). In that case, to get an overall average of -2.4, the other five seasons must have averaged -1.74. That's because, if you take five seasons of -1.74, and three seasons of -3.5, you get -2.4.

So, as a rough guess, the new study found:

-3.5 -- same three seasons as the old study;
-1.7 -- five seasons the old study didn't cover;
-------------------------------------------------
-2.4 -- all eight seasons combined.

So, in the new data, this study finds only half the choke effect that the other study did. Moreover, I estimate it's only 1 SD from zero.

-------

That's for "down 1-4 points." Here's the same calculation, broken down by individual score. Here "%" means percentage point difference:

-2 points: First three: -5%. Next five: -2.3%. All eight: -3.3%.
-1 points: First three: -7%. Next five: -1.7%. All eight: -3.7%.
+0 points: First three: +2%. Next five: -1.8%. All eight: -1.0%.
+1 points: First three: -5%. Next five: -0.5%. All eight: -2.2%.
+2 points: First three: +0%. Next five: -0.3%. All eight: -1.3%.


Generally, five new seasons are closer to zero than the three original seasons. That's what you would expect if the original numbers were mostly luck.

-------

So, in summary:

-- The study finds that in the last seconds of games, players behind in close games shoot significantly worse than in blowouts.

-- In the last 30 seconds, they're maybe about 2.8 percentage points worse. In the last 15 seconds, they're maybe about 4.8 percentage points worse.
-- The effect is biggest when down by 1 in the last fifteen seconds.

-- However, they are not statistically significantly better or worse than *average,* just statistically significantly worse than blowouts (although they certainly are "basketball significantly" worse than average).

-- The effect is mostly driven by the three seasons covered in the earlier study. If you look at the other five seasons, the effect is not statistically significant (but still has the same sign).

What do you think? I'm not absolutely convinced there's a real effect overall, but yeah, it seems like it's at least possible.

However, I do think the most extreme individual breakdowns are overstated. For instance, the newspaper article says that in the last 15 seconds, down by 1, players will shoot "5 to 10 percentage points worse than normal." (They really mean "worse than 11+ blowouts," but never mind.) Given that that's the most extreme result the study found, I think it's probably a significant overestimate. I'd absolutely be willing to bet that, over the next five seasons, that the observed effect will be less than five percentage points.

--------

P.S. One last side point: the newspaper article says,

"Shooters who average 90 percent from the line performed slightly better than that under pressure, while 60 percent shooters had a choking effect twice as great as 75 percent shooters. That suggests that a lack of confidence begets less confidence, and vice versa."

This is a correct summary of what the authors say in their discussion, but I think it's wrong. The regressions that this comes from (Table 5, columns 2 and 6) don't include an adjustment for the player. So what it really means is that the 60 percent shooter will be *twice as far below the average player* as the 75 percent shooter. That makes sense -- because he's a worse shooter to begin with, even before any choke effect.

------

UPDATE: After posting this, I realized that I may have missed one aspect of the regression ... but I think my analysis here is correct if I make one additional assumption that's probably true (or close to true). I'll clarify in a future post.


Labels: , , ,

Saturday, February 26, 2011

Why is there no home-court advantage in foul shooting?

There's a home-site advantage in every sport.

Why is that? Nobody knows. One hypothesis is that it's officials favoring the home team. One piece of data that appears to support that hypothesis is that when you look at situations that don't involve referee decisions, the home field advantage (HFA) tends to disappear. In "Scorecasting," for instance, the authors report that, in the NBA, the overall home and road free-throw percentages are an identical .759. Also, in the NHL, shootout results seem to be the same for home and road teams, and likewise for penalty kick results in soccer.

However, there's a good reason for the results to look close to identical even if HFA is caused by something completely unrelated to refereeing.

The reason is that free-throw shooting involves only one player. At the simplest level, you could argue that foul shooting is offense. All other points scored in basketball are a combination of offense and defense. Not only is the offense playing at home, but the defense is playing on the road, which, in a sense, "doubles" the advantage. Therefore, if the home free-throw shooting advantage is X, the home field-goal shooting advantage should be at least 2X.

That's an oversimplification. A better way to think about it is that a foul shot attempt is the work of one player. A field goal attempt, on the other hand, is the end result of the efforts of *ten* players. Not every player is directly involved in the eventual shot attempt, but every player has the potential to be. A missed defensive assignment could lead to an easy two points, and the offense will take advantage regardless of which of the five defensive players is at fault. The same for offense: if a player beats his man and gets open, he's much more likely to be the one who gets the shot away. The weakest or strongest link could be any one of the ten players on the court.

So it might be better to guess that the HFA for a possession is 10X, rather than just X. We can't say that for sure -- it could be that the things a player has to do on a normal possession are so much more complex than a free throw, that the correct number is 20X. Or it could be that a normal possession is less complex than a free throw, so perhaps 5X is better. I don't know the answer to this, but 10X seems like a reasonable first approximation.

------

What would the actual numbers look like?

The home court advantage in basketball is about three points. That means that instead of (say) 100-100, the average game winds up 101.5 to 98.5.

Three points per game, divided by 10 players, is 0.3 points per game per player. Over (say) 200 possessions, that's 0.0015 points per possession per player.

If home-court advantage were made up only of serious mistakes, mistakes that turn a normal 50 percent possession into a 100 percent or zero percent possession, then that works out to exactly one point per mistake. In that case, the average player would make one such extra mistake every 667 possessions. That's a little less than one every three games. If you assume that a mistake is worth only half a point, then it's one mistake per player for every 333 team possessions.

In reality, of course, it's probably not nearly as granular as "mistakes" or "good plays". It's probably something like this: a player plays his role with an overall average of 50 effectiveness units, random between possessions, plus or minus some variation. But that's an average of home, where he plays with an average of 51 effectiveness units, and road, with an average 49 effectiveness units.

Still, that doesn't matter to the argument: the important thing is HFA is one point per player for every 667 total team possessions, regardless of how it manifests itself.

------

Now, let's go back to free throws. I'm going to assume that a player's HFA on a single possession should be about the same as a player's HFA on a single free throw. Is that OK? It's a big assumption. I don't have any formal justification for it, but it doesn't seem unreasonable. I'd have to admit, though, that there are a lot of alternative assumptions that also wouldn't seem unreasonable.

But the point of this post is that it is NOT reasonable to assume that a player's HFA on a free throw should be the same as the overall HFA for an entire game. That wouldn't make any sense at all. That would be like seeing that the average team wins 50 percent of games, and therefore expecting that the average team should win 50 percent of championships. It would be like seeing that the Cavaliers are winning 17 percent of their games, and expecting that they score 17 percent of the total points.

In any case, the overall argument stays the same even if you argue that the HFA on a single possession should be twice that of a single free throw, or half, or three times. But I'll proceed anyway with the assumption that it's one time.

If the HFA on a free-throw is the same 0.0015 points per player as on a possession, then you'd expect the difference between home and road free throw percentages to be 0.15%. Instead of the observed .759 home and road, it should be something like .75975 home, and .75825 road.

Why don't we see this? Well, here's one possible explanation. Visiting teams are behind more often, so will commit more deliberate fouls late in the game. They will try to foul home players who are worse foul shooters. Therefore, the pool of home foul shooters is worse than the pool of road foul shooters, which is why it looks like there's no home field advantage in foul shooting.

Since we're talking about such a very small HFA in the first place, this doesn't seem like an unreasonable explanation. It would be interesting to run the numbers, but controlling for who the shooter is. I suspect if you have enough data, you'd spot a very small home-court advantage in foul shooting.



Labels: , , ,

Friday, July 16, 2010

Do foul shooters choke in the last minute of close games?

Searching Google Scholar for studies about "choking," I came across an interesting one, a short, simple analysis of free-throw shooting in NBA games.

It's called "Choking and Excelling at the Free Throw Line," by Darrell A. Worthy, Arthur B. Markman, and W. Todd Maddox. (.pdf)

The authors looked at all free throws in the last minute of games in the three seasons from 2002-03 to 2004-05. They broke their sample down by score differential, and compared the success percentage to the players' career percentages.

They found that for most of the scores, the shooters converted fewer than expected. Here's the data as I read it off the graph (but see the PDF for yourself). The score differential is from the perspective of the shooting team, and the "%" column is actually percentage points.

-5 points: -3%
-4 points: -1%
-3 points: -1%
-2 points: -5% (significant at 5%)
-1 points: -7% (significant at 1%)
+0 points: +2%
+1 points: -5% (significant at 5%)
+2 points: +0%
+3 points: -1% ("also signficant")
+4 points: +1%
+5 points: -1%

The authors conclude that choking occurs, especially when down by 1 point.

It may not be obvious at first glance from the chart, but there's a tendency to "choke" all the way down: there are 8 negatives and only 3 positives (and the negatives are generally more extreme than the positives). Do players actually shoot worse in the last minute of close games?

----

I couldn't think of any statistical reason the results might be misleading ... but in an e-mail to me, Guy came up with a good one.

Suppose that career shooting percentage is not always a good indicator of a player's percentage that game. Maybe it varies throughout a career, somehow -- higher at peak age and lower elsewhere, or, even, increasing throughout a career. (It doesn't matter to the argument *how* it varies, just that it does.)

You might expect that the differences would just cancel out. However, the overestimated shooters would be appear in each category more than the underestimated shooters. Why? Because they would miss the first shot more often, and take a second shot *within the same score category*.

As an example, suppose two players have 75% career percentages, but, on this day, A is a 100% shooter and B is a 50% shooter. Suppose they each go to the line twice with the game tied. On their first shot, A makes two and B makes one. So far, their percentage is 75%, as expected. Perfect.

But, only B gets to take a second shot with the game still tied. He does that once, the one time in two he missed the first shot. And he makes it half the time.

So, on average, you have these guys taking five shots, and making 3.5 of them. That's 70 percent -- 5 percentage points less than the career average would suggest.

Now, the numbers I used here are not very realistic -- nobody's a 100% shooter, and hardly anyone is a 50% shooter. What if I change it to 80% and 70%?

Then, following the same logic, and if my arithmetic is correct, those two players combined would make 74.8% of their shots instead of 75%. It's still something, but not nearly enough to explain the results. Still, I really like Guy's explanation.

----

So, there you go: it does look like, for those three seasons, players shot worse in the last minute than expected. Can anyone think of an explanation, other than "choking" and luck, for why that might be the case? Has anyone done this kind of analysis for other seasons to confirm these results?


UPDATE: Maybe it's just fatigue! See comments.



Labels: , , ,

Thursday, March 05, 2009

Why hasn't foul shooting improved?

Free-throw shooting percentages haven't changed much over the past 50 years, according to this New York Times article. Between 1950 and 1970, the conversion rate was around 72 percent. Since 1970, it's fluctuated between 72 and 77%.

Here's the NYT graph:



So it looks like free throwing hasn't really improved over the decades. That makes foul shooting an anomaly, because most other skills have improved: marathon times are better, football kicking is better, and "swimming records seemingly fall at each international event."

Why hasn't foul shooting improved? According to the article:

Ray Stefani, a professor emeritus at California State University, Long Beach, is an expert in the statistical analysis of sports. Widespread improvement over time in any sport, he said, depends on a combination of four factors: physiology (the size and fitness of athletes, perhaps aided by performance-enhancing drugs), technology or innovation (things like the advent of rowing machines to train rowers, and the Fosbury Flop in high jumping), coaching (changes in strategy) and equipment (like the clap skate in speedskating or fiberglass poles in pole vaulting). ...

“There are not a lot of those four things that would help in free-throw shooting,” Stefani said.


And that's fair enough. But what about, say, bowling? The article says explicitly that "bowling a 300 game is not as unlikely as it once was," and there are strong similarities between bowling and foul shooting. Physiology doesn't seem like it would help either way; technology and innovation don't seem like issues; and it's hard to see how coaching would be of more help in bowling than in foul shooting.

I'd propose another explanation: foul shooting is an ancillary skill in basketball – players are chosen for their overall ability, not just their free-throw potential. And so "natural selection" won't weed out mediocre shooters or reward the best shooters, at least not very much compared to other skills.

Compare this to other sports: bowling strikes is the primary goal of the game, the most important skill of all. And, in football, field-goal kickers are chosen for one thing: their ability to kick field goals. Any kicker below average in accuracy is out of the league instantly. But any NBA player who can't hit free throws can make it up in other aspects of the game (like Shaq). (A version of this argument was also made in the first comment of a discussion on Tango's blog, here). And coaches don't force their players to shoot underhand, which would make many players more accurate; that provides support for the idea that the NBA thinks free throw percentage doesn't matter that much.

If you want to *really* see if the skill is improving, don't look to NBA players, who may not be the best in the world at the skill. You'd have to look at free-throw specialists. I Googled "free throw shooting contest results," and got a link to an Iowa State contest where the winner made 49 out of 50 throws. That's 98%, and about 4 standard deviations away from the NBA average of 75%. Even considering that the contest had 72 entries, that's pretty significant.

And here's another argument: if foul shooting isn't considered a major skill, young players won't practice it as much, and it stands to reason that you won't get as much improvement over time if there's not as much energy expended to get better at it.

One last point: if you consider the graph's increase from 71 to 77 percent to be real, then that's actually pretty good evidence of an increase in skill. When you're already at a 71% level, it's harder to improve than if you start from, say, a 34% level (as field-goal percentage did). In 1950, players were missing 29% of their foul shots. In 2008, they were missing only 23%. That means that over the past 58 years, players learned to convert 20% of their misses into hits. That's pretty good. The field goal percentage improvement, from 34% to 46%, looks more impressive, but results from converting 18% of misses into hits – almost an identical improvement (although they probably shouldn't be compared directly, because field goals are influenced by where they're taken from, and the quality of the defense).

In summary:

-- there are good reasons you wouldn’t expect foul-shooting to improve as much as other skills over time;
-- if you look at the numbers more closely, there actually *is* a significant amount of improvement.

So I don't think there's as huge a mystery there like the Times does.


Labels: , ,

Monday, January 15, 2007

An NBA free throw coach

A post from the Freakonomics blog points to a New York Times article about the free throw coach hired by the Dallas Mavericks.

Apparently, coaching works well, so well that Dubner calls it an "arbitrage opportunity" and wonders why other teams haven't done likewise.

If you want to read the full NYT article, go now -- it doesn't take long for the Times to make their archives subscription-only.

My previous post on underhanded foul shooting is here.

Labels: ,