Wednesday, May 17, 2017

The hot hand debate vs. the clutch hitting debate

In the "hot hand" debate between Guy Molyneux and Joshua Miller I posted about last time, I continue to accept Guy's position, that "the hot hand has a negligible impact on competitive sports outcomes."

Josh's counterargument is that some evidence for a hot hand has emerged, and it's big. That's true: after correcting for the error in the Gilovich paper, Miller and co-author Adam Sanjurjo did find evidence for a hot hand in the shooting data of Gilovich's experiment. They also found a significant hot hand in the NBA's three-point shooting contest

I still don't believe that those necessarily suggest a similar hot hand "in the wild" (as Guy puts it), especially considering that to my knowledge, none has been found in actual games. 

As Guy says,


"Personally, I find it easy to believe that humans may get into (and out of) a rhythm for some extremely repetitive tasks – like shooting a large number of 3-point baskets. Perhaps this kind of “muscle memory” momentum exists, and is revealed in controlled experiments."

-------

Of course, I keep an open mind: maybe players *do* get "hot" in real game situations, and maybe we'll eventually see evidence for it. 

But ... that evidence will be hard to find. As I have written before, and as Josh acknowledges himself, it's hard to pinpoint when a "hot hand" actually occurs, because streaks happen randomly without the player actually being "hot."

I think I've used this example in the past: suppose you have a 50 percent shooter when he's normal, but he turns in to a 60 percent shooter when he's "hot," which is one-tenth of the time. His overall rate is 51 percent.

Suppose that player makes three consecutive shots. Does that mean he's in his "hot" state? Not necessarily. Even when he's "normal," he's going to have times where he makes three consecutive shots just by random luck. And since he's "normal" nine times as often as he's "hot," the normal streaks will outweigh the hot streaks.

Specifically, only 19 percent of three-hit streaks will come when the player is hot. In other words, four out of five streaks are false positives.

(Normally, he makes three consecutive shots one time in 8. Hot, he makes three consecutive shots one time in 4.63. In 100 sequences, he'll be "normal" 90 times, for an average 11.25 streaks. In his 10 "hot" times, he'll make 2.16 streaks. That's about a 4:1 ratio.)

Averaging the real hotness with the fake hotness, the player will shoot 51.9 percent after a streak. But his overall rate is 51.0 percent. It takes a huge sample size to notice the difference between 51 percent and 51.9 percent.

Even if you do notice a difference, does it really make an impact on game decisions? Are you really going to give the player the ball more because his expectation is 0.9 percent higher, for an indeterminate amout of time?

-------

And that's my main disagreement with Josh's argument. I do acknowledge his finding that there's evidence of a "muscle memory" hot hand, and it does seem reasonable to think that if there's a hot hand in one circumstance, there's probably one in real games. After all, *part* of basketball is muscle memory ... maybe it fades when you don't take shots in quick succession, but it still seems plausible that maybe, some days you're more calibrated than others. If your muscles and brain are slightly different each day, or even each quarter, it's easy to imagine that some days, the mean of your instinctive shooting motion is right on the money, but, other days, it's a bit short.

But the argument isn't really about the *existence* of a hot hand -- it's about the *size* of the hot hand, whether it makes a real difference in games. And I think Guy is right that the effect has to be negligible. Because, even if you have a very large change in talent,  from 50 percent to 60 percent -- and a significant frequency of "hotness", 10 percent of the time -- you still only wind up with a 0.9 percent increased expectation after a streak of three hits. 

You could argue that, well, maybe 50 to 60 percent understates the true effect ... and you could get a stronger signal by looking at longer streaks.

That's true. But, to me, that argument actually *hurts* the case for the hot hand. Because, with so much data available, and so many examples of long streaks, a signal of high-enough strength should have been found by now, no? 


-------

This debate, it seems to me, echoes the clutch hitting debate almost perfectly.

For years, we framed the state of the evidence as "clutch hitting doesn't exist," because we couldn't find any evidence of signal in the noise. Then, a decade ago, Bill James published his famous "Underestimating the Fog" essay, in which he argued (and I agree) that you can't prove a negative, and the "fog" is so thick that there could, in fact, be a true clutch hitting talent, that we have been unable to notice.

That's true -- clutch hitting talent may, in fact, exist. But ... while we can't prove it doesn't exist, we CAN prove that if it does exist, it's very small. My study (.pdf) showed the spread (SD) among hitters would have to be less than 10 points of batting average (.010). "The Book" found it to be even smaller, .008 of wOBA (a metric that includes all offensive components, but is scaled to look like on-base percentage). 

To my experience, a sizable part of the fan community seizes on the "clutch hitting could be real" finding, but ignores the "clutch hitting can't be any more than tiny" finding. 

The implicit logic goes something like, 

1. Bill James thinks clutch hitting exists!
2. My favorite player came through in the clutch a lot more than normal!
3. Therefore, my favorite player is a clutch hitter who's much better than normal when it counts!

But that doesn't follow. Most strong clutch hitting performances will happen because of luck. Your great clutch hitting performance is probably a false positive. Sure, a strong clutch performance is more likely to happen given that a player is truly clutch, but, even then, with an SD of 10 points, there's no way your .250 hitter who hit .320 in the clutch is anything near a .320 clutch hitter. If you did the math, maybe you'd find that you should expect him to be .253, or something. 

Well, it's the same here, with the hot hand:

1. Miller and Sanjurjo found a real hot hand!
2. Therefore, hot hand is not a myth!
3. My favorite player just hit his last five three-point attempts!
4. Therefore, my player is hot and they should give him the ball more!

Same bad logic. Most streaks happen because of luck. The streak you just saw is probably a false positive. Sure, streaks will happen given that a player truly has a hot hand, but, even then, given how small the effect must be, there's no way your usual 50-percent-guy is anything near superstar level when hot. If you had the evidence and did the math, maybe you'd find that you should expect him to be 52 percent, or something.

-------

For some reason, fans do care about whether clutch hitting and the hot hand actually happen, but *don't* care how big the effect is. I bet psychologists have a cognitive fallacy for this, the "Zero Shades of Grey" fallacy or the "Give Them an Inch" fallacy or the "God Exists Therefore My Religion is Correct" fallacy or something, where people are unwilling to believe something into existence -- but, once given license to believe, are willing to assign it whatever properties their intuition comes up with.

So until someone shows us evidence of an observable, strong hot hand in real games, I would have to agree with Guy:


"... fans’ belief in the hot hand (in real games) is a cognitive error."  

The error is not in believing the hot hand exists, but in believing the hot hand is big enough to matter. 

Science may say there's a strong likelihood that intelligent life exists on other planets -- but it's still a cognitive error to believe every unexplained light in the sky is an alien flying saucer.



Labels: , ,

Tuesday, December 02, 2014

Players being "clutch" when targeting 20 wins -- a follow-up

In his 2007 essay, "The Targeting Phenomenon," (subscription required), Bill James discussed how there are more single-season 20-game winners than 19-game winners. That's the only time that happens, that the higher number happens more frequently than the lower number. 

This is obviously a case of pitchers targeting the 20-win milestone, but Bill didn't speculate on the actual mechanisms for how the target gets hit. In 2008, I tried to figure it out. But, this past June, Bill pointed out that my conclusion didn't fit with the evidence:

"... the Birnbaum thesis is that the effect was caused one-half by pitchers with 19 wins getting extra starts, and one-half by poor offensive support by pitchers going for their 21st win, thus leaving them stuck at 20. But that argument doesn't explain the real life data. 

"[If you look closely at the pattern in the numbers,] the bulge in the data is exactly what it should be if 20 is borrowing from 19 -- and is NOT what it should be if 20 is borrowing both from 19 and 21."

(Here's the link.  Scroll down to OldBackstop's comment on 6/6/2014.)

So, I rechecked the data, and rethought the analysis, and ... Bill is right, as usual. The basic data was correct, but I didn't do the adjustments properly.

-----

My original study covered 1940 to 2007. This study, though, will cover only 1956 to 2000. That's because I couldn't find my original code and data. The "1956" is what I happened to have handy, and I decided to stop at 2000 because Bill did. 

First, here are the raw numbers of seasons with X wins:

17 wins: 159
18 wins: 132
19 wins:  92
20 wins: 113
21 wins:  56
22 wins:  35
23 wins:  20
24 wins:  20

You can see the bulge we're dealing with: there are way too many 20-win pitchers. And it can't be that the excess comes from the 21-win bucket, because, then, the average of 20 and 21 would stay the same, and wouldn't be much lower than 19. That can't be right. And, as Bill pointed out, even if only *half* the excess came from the 21 bucket, 20 would still be too big relative to 19.

So, let me try fixing the problem.

In the other study, I checked four ways in which 20 wins could get targeted:

1. Extra starts for pitchers getting close
2. Starters left in the game longer when getting close
3. Extra relief apparances for pitchers getting close
4. Better performance or luck when shooting for 20 than when shooting for 21.

I'll take those one at a time.

-------

1. Extra starts

The old study found that pitchers who eventually wound up at 19 or 20 wins did, in fact, get more late-season starts than others -- about 23 more overall. In this smaller study (1956-2000 instead of 1940-2007), that translates down to maybe 18 extra starts. 

That's about 9 extra wins. Let's allocate four of them to pitchers who wound up at 19 instead of 18, and the other five to pitchers who wound up at 20 instead of 19. If we back that out of the actual data, we get:

18 wins: 132 136
19 wins:  92  93
20 wins: 113 108
21 wins:  56  56

(If you're reading this on a newsfeed that doesn't support font variations: the first column is the old values, which should be struck out.)

What happens is: the 18 bucket gets back the four pitchers who won 19 instead. The 19 bucket loses those four pitchers, but gains back the five pitchers who won 20 instead of 19. The 20 bucket loses those five pitchers.

(In the other study, I didn't bother doing this, backing out the effects when I found them, so I wound up taking some of them from the wrong place, which caused the problem Bill found.)

So, we've closed the gap from 21 down to 15.

--------

2. Starters left in the game longer

After I had posted the original study, Dan Rosenheck commented,
"You didn't look at innings per start. I bet managers leave guys with 19 W's in longer if they are tied or trailing in the hope that the lineup will get them a lead before they depart."

I checked, and Dan was right. In a subsequent comment, I figured Dan's explanation accounted for about 10 extra twenty-game winners. Those are all taken from the 19-game bucket, because the effect occurred only for starters currently pitching with 19 wins.

For this smaller dataset, I'll reduce the effect from 10 seasons to 7. 

So:

18 wins: 136 136
19 wins:  93 100
20 wins: 108 101
21 wins:  56  56

Now, the bulge is down to 1.  We still have a ways to go, if the 19 is to be significantly higher than the 20, but we're getting there.

---------

3. Extra Relief Apparances

The other study listed every pitcher who got a win in relief while nearing 20 wins. Counting only the ones from 1956 to 2000, we get:

3 pitchers winding up at 19
5 pitchers winding up at 20
2 pitchers winding up at 21

Backing those out:

18 wins: 136 139
19 wins: 100 102
20 wins: 101  98
21 wins:  56  54

The gap now goes the proper direction, but only slightly.

------

4. Luck

This was the most surprising finding, and the one responsible for the "getting stuck at 20" phenomenon. Pitchers who already had 20 wins were unusually unlikely to get to 21 in a subsequent start. Not because they pitched any worse, but because they got poor run support from their offense.

When Bill pointed out the problem, I wondered if the run-support finding was just a programming mistake. It wasn't -- or, at least, when I rewrote the program, from scratch, I got the same result.

For every current starter win level, here are the pitchers' W-L records in those starts, along with the team's average runs scored and allowed:

17 wins:   483-311 .557   4.30-3.61
18 wins:   350-250 .608   4.30-3.61
19 wins:   260-182 .588   4.24-3.56
20 wins:   150-136 .524   3.81-3.54
21 wins:    94- 61 .606   4.49-3.44
22 wins:    59- 23 .720   4.26-2.80

The run support numbers are remarkably consistent -- except at 20 wins. Absent any other explanation, I assume that's just a random fluke.

If we assume that the 20-win starters "should have" gone 171-115 (.598) instead of 150-136 (.524), that makes a difference of 21 wins.

The mistake I made in the previous study was to assume that those wins were all stolen from the "21-win" bucket. Some were, but not all. Some of the unlucky pitchers eventually got past the 20-win mark; a few, for instance, went on to post 23 wins. In their case, it becomes the 23-win bucket stealing a player from the 24-win bucket.

I checked the breakdown. For every starter who tried for his 21st win but didn't achieve it that game, I calculated where he eventually finished the season. From there, I scaled the totals down to 21, the number of wins lost to bad luck. The result:

20  wins:  9 pitchers
21  wins:  5 pitchers
22  wins:  3 pitchers
23  wins:  1 pitcher
24  wins:  2 pitchers
25+ wins:  less than 1 pitcher

So: the 20-win bucket stole 9 pitchers from the 21-win bucket. The 21-win bucket stole 5 pitchers from the 22-win bucket. And so on. 

Adjusting the overall numbers gives this:

18 wins: 139 139
19 wins: 102 102
20 wins:  98  89
21 wins:  54  50
22 wins:  35  33

-------

And that's where we wind up. It's still not quite enough, to judge by Bill's formula and even just the eyeball test. It still looks like there's a little bulge at 20, by maybe five pitchers. If 20 could steal five more pitchers from 19, we'd be at 107/84, which would look about right.

But, we've done OK. We started with a difference of +21 -- that is, 21 more twenty-game winners than nineteen-game winners -- and finished with a difference of -13. That means we found an explanation for 34 games, out of what looks like a 39-game discrepancy.

Where would the other five come from? I don't know. It could be luck and rounding errors. It could also be that the years 1956-2000 aren't a representative sample of the original study, so we lost a bit of accuracy when I scaled down.  Or, it could be some fifth real factor I haven't thought of.

In any case, here's the final breakdown of the number of "excess" 20-game winners:

-- 5 from getting extra starts;
-- 7 from being left in games longer than usual;
-- 3 from getting extra relief appearances;
-- 9 from bad run support getting them stuck at 20;
-- 5 from luck/rounding/sources unknown.

By the way, one important finding still stands through both studies. Starters didn't seem to pitch any better than normal with their 20th win on the line, so you can't accuse them of trying harder in the service of a selfish personal goal.




Labels: , , ,

Wednesday, February 15, 2012

Absence of evidence vs. evidence of absence

People tell me that Albert Pujols is a better hitter than John Buck. So I did a study. I watched all their at-bats in August, 2011. I observed that Pujols hit .298, and Buck hit .254.

Yes, Pujols' batting average was better than Buck's, but the difference wasn't statistically significant. In fact, it wasn't even close: it was less than 1 standard deviation!

So, clearly, August's performance shows no evidence that Pujols and Buck are different in ability.

Does that sound wrong? It's right, I think, at least as I understand how things work in the usual statistical studies. If you fail to reject the null hypothesis, you are entitled to use the words "no evidence."

Which is a little weird, because, of course, it *is* evidence, although perhaps *weak* evidence. I suppose they could have chosen to say "not enough" evidence, or "insufficient" evidence, but that carries with it an implication that the null hypothesis is correct. If I say, "the study found no evidence that whites are smarter than blacks," that sounds fine. But if I say, "the study found insufficient evidence that whites are smarter than blacks," that sounds racist.

The problem is, if you don't really know what "no evidence" really means, you might get the wrong impression. You might have 25 different studies testing whether Pujols is better than Buck, each of them using a different month. They all fail to reject the hypothesis that they're equal, and they all say they found "no evidence". (That's not unlikely: to be significant at .01 for a single month, you'd have to find Pujols outhitting Buck by about 200 points.)

And you think, hey, "25 studies all failed to find any evidence. That, in itself, is pretty good evidence that there's nothing there."

But, the truth is, they all found a little bit of evidence, not *no* evidence. If you multiply *no* evidence by 25, you still have *no* evidence. But if you multiply a little bit of evidence by 25, now you have *enough* evidence.

------

There's an old saying, "absence of evidence is not evidence of absence." The idea is, just because I look around my office and don't see any proctologists or asteroids, it doesn't mean proctologists or asteroids don't exist. I may just not be looking in the right place, or looking hard enough. Similarly, if I look at only one month of Pujols/Buck, and I don't see a difference, it doesn't mean the difference isn't there. It might just mean that I'm not looking hard enough.

This is the point Bill James was making in his "Underestimating the Fog." We looked for clutch hitting, and we didn't find it. And so we concluded that it didn't exist. But ... maybe we we just need to look harder, or in different places.

What Bill was asking is: we have the absence of evidence, but do we have the evidence of absence?

------

Specifically, what *would* constitute evidence of absence? The technically-correct answer: nothing. In normal statistical inference, there's actually no evidence that can support absence.

Suppose I do a study of clutch hitting, and I find it's not significantly different from zero. But ... my parameter estimate is NOT zero. It's something else, maybe (and I'm making this up), .003. And maybe the SD is .004.

If I think clutch hitting is zero, and you think it's .003, we can both point to this study as confirming our hypotheses. I say, "look, it's not statistically significantly different from zero." And you say, "yeah, but it's not statistically significantly different from .003 either. Moreover, the estimate actually IS .003! So the evidence supports .003 at least as much as zero."

That leaves me speechless (unless I want to make a Bayesian argument, which let's assume I don't). After all, it's my own fault. I didn't have enough data. My study was incapable of noticing a difference between .000 and .003.

So I go back to the drawing board, and use a lot more data. And, this time, I come up with an estimate of .001, with an SD of .002.

And we have the same conversation! I say, "look, it's not different from zero." And you say, "it's not different from .001, either. I still think clutch hitting exists at .001."

So I go and try again. And, every time, I don't have an infinite amount of data, so, every time, my point estimate is something other than zero. And every time, you point to it and say, "See? Your study is completely consistent with my hypothesis that clutch hitting exists. It's only a matter of how much."

------

What's the way out of this? The way out of this is to realize that you can't use statistics to prove a point estimate. The question, "does clutch hitting exist?" is the same as the question "is clutch hitting exactly zero?". And, no statistical technique can ever give you an exact number. There will always be a standard error, and a confidence interval, so it will always be possible that the answer is not zero.

You can never "prove" a hypothesis about a single point. You can only "disprove" it. So, you can never use statistical techniques to demonstrate that something does not exist.

What we should be talking about is not existence, but size. We can't find evidence of absence, but we can certainly find evidence of smallness. When an announcer argues for the importance of being able to step up when the game is on the line, we can't say, "we studied it and there's no such thing". But we *can* say, "we studied it, and even under the most optimistic assumptions, the best clutch hitter in the league is only going to hit maybe .020 better in the clutch ... and there's no way to tell who he is."

Or, the short form -- "we studied it, and the differences between players are so small that they're not worth worrying about."

------

But aren't there issues where it's important to actually be able to disprove a hypothesis? Take, for instance, ESP. Some people believe they can do better than chance at guessing which card is drawn from an ESP deck.

If we do a study, and the subject guesses exactly what you'd expect by chance, you'd think that would qualify as a failure to find ESP. But when you calculate the confidence interval, centered on zero, you might have to say, "our experiment suggests that if ESP exists, its maximum level is one extra correct guess in 10,000."

And, of course, the subject will hold it up, and triumphantly say, "look, the scientists say that I might have a small amount of ESP!!"

What's the solution there? It's to be common-sense Bayesian. It's to say, "going into the study, we have a great deal of "evidence of absence" that ESP doesn't exist -- not from statistical tests, but from the world's scientific knowledge and history. If you want to challenge that, you need an equal amount of evidence."

That makes sense for ESP, but not for clutch hitting. Don't we actually *know* that clutch hitting talent must exist, even at a very small level? Every human being is different in how they respond to pressure. Some batters may try to zone out, trying to forget about the situation and hit from instinct. Some may decide to concentrate more. Some may decide to watch the pitcher's face between pitches, instead of adjusting their batting glove.

Any of those things will necessarily change the results a tiny bit, in one direction or the other. Maybe concentration makes things worse, maybe it makes it better. Maybe it's even different for different hitters.

But we *know* something has to be different. It would be much, much too coincidental if every batter did something different, but the overall effect is exactly .0000000.

Clutch hitting talent *must* exist, although it might be very, very small.

So why are we so fixated on zero? It doesn't make sense. We know, by logical argument, that clutch hitting can't be exactly zero. We also know, by logical argument, that even if it *were* exactly zero, it's impossible to have enough evidence of that.

When we say "clutch hitting doesn't exist," we're using it as a short form for, "clutch hitting is so small that, for all intents and purposes, it might as well not exist."

------

When the effect is small, like clutch hitting, it's not a big deal. But when the effect might be big, it's a serious issue.

A lot of formal studies -- not just clutch hitting or baseball -- will find they can't reject the null hypothesis. They usually say, "we found no evidence," and then they go on to assume that that also means they can assume that what they're looking for doesn't exist.

They'll do a study on, I don't know, whether an announcer is right that playing a day game after a night game affects you as a hitter. And they'll get an estimate that says that batters are 40 points of OPS worse the day after. But it's not statistically significant. And they say, "See? Baseball guys don't know what they're talking about. There's no evidence of an effect!"

But that's wrong. Because, unlike clutch hitting, the confidence interval does NOT show an effect that "for all intents and purposes, might as well not exist." The confidence interval is compatible with a *large* effect, of at least 80 points. (That is, since 2 SD is enough to drop from 40 points to zero on one side, it's also enough to rise from 40 points to 80 points on the other side.)

So it's not that there's evidence of absence. There's just absence of evidence.

And that's because of the way they did their study. It was just too small to find any evidence -- just like my office is too small to find any asteroids.



Labels: ,

Wednesday, December 21, 2011

Are there "good team" goalies and "bad team" goalies?

In "The Game," Ken Dryden argues that some players are not psychologically suited to playing on good teams:

Because the demands of a goalie are mostly mental, it means that for a goalie the biggest enemy is himself. The fear of failing, the fear of being embarrassed ... The successful goalie understands these neuroses, accepts them, and puts them under control. The unsuccessful goalie is distracted by them, his mind in knots, his body quickly following.

It is why [Rogie] Vachon was superb in Los Angeles and as a high-priced free-agent messiah, poor in Detroit. It is why Dan Bouchard ... lurches annoyingly in and out of mediocrity. It is why there are good "good team" goalies and good "bad team" goalies -- Gary Smith, Doug Favell, Denis Herron. The latter are spectacular, capable of making near-impossible saves that few others can make. They are essential for bad teams, winning them games they shouldn't win, but they are goalies who need a second chance, who need the cushion of an occasional bad goal, knowing that they can seem to earn it back later with several inspired saves. On a good team, a goalie has few near-impossible saves to make, but the rest he must make, and playing in close and critical games as he does, he gets no second chance.

A good "bad team" goalie, numbed by the volume of goals he cannot prevent, can focus on brilliant saves and brilliant games, the only things that make a difference to a poor team. A good "good team" goalie cannot. Allowing few enough goals that he feels every one, he is driven instead by something else -- the penetrating hatred of letting in a goal.


Dryden seems to be saying at least three things here:

1. Some goalies, like Rogie Vachon, can't handle pressure.
2. Some goalies are better on bad teams than on good teams.
3. Those two groups are the same goalies.

I'm very skeptical about #1, especially with regards to Rogie Vachon. Yes, Vachon had a serious decline after leaving the Kings -- with Detroit, he was worse by more than a goal a game (3.90 to 2.86). But, was it really Vachon's neuroses? After all, he was 33 years old that year. Dryden may know Vachon pretty well -- they were together on the Canadiens for a few months in 1971 -- but is that enough for him to conclude that Vachon's problem is that he choked under pressure?

I'll skip over #3, also, and concentrate on #2, the part about "good team" goalies and "bad team" goalies. What Dryden seems to be saying, as an empirical hypothesis, is something like this:

There are some goalies who make brilliant saves that few others can, but also give up more weak goals. Those goalies are more valuable to bad teams, because bad teams give up more scoring chances where brilliant saves are required. It wouldn't make sense for a good team to pick up a goalie like that, because they'd get only the weak goals, but not the brilliant saves.

That's actually a pretty interesting theory! And it seems plausible. After all, what a team should care about is how many goals a guy allows, not how he looks doing it. A goalie with a 2.50 GAA is more valuable than a goalie with a 2.75 GAA, even if the first guy lets in more bad goals than the second guy.

But, is there any evidence for it?

Not in the book. Dryden gives us only those three examples of "bad team" goalies. Unfortunately, they played on bad teams for most of their careers.

Still, we have a few datapoints.

-- Gary Smith left Oakland (bad) to play two seasons for the Black Hawks (good) as Tony Esposito's backup. The first year, he was very good; the second year, he was mediocre.

-- Denis Herron moved from (bad) Pittsburgh to (good) Montreal (where he replaced Ken Dryden). Like Smith, he was great the first year, but not so great the second year.

-- Doug Favell was nothing special in his last season with Toronto (an average team). Then, he went to a below-average Colorado Rockies team, where it seems like he was pretty good. So, maybe that's a plus.

But, overall ... no real evidence either way, really. And Dryden doesn't give any examples of "good team" goalies, so there's nothing to check there.

But ... maybe here's something we can do.

Goalies play some of their games against good teams, and some against bad teams. If Dryden is correct, that Smith, Herron and Favell give up more bad goals but also make more spectacular saves, they should do better than expected against good teams, and worse than expected against bad teams. That's because they'll give up roughly the same amount of bad goals each way, but they'll make more brilliant saves against the good teams.

Does that make sense? Maybe we can find a way to check that.

Here's how that might work. In 1977-78, the Penguins, with Denis Herron as their regular goalie, gave up 321 goals, or 4.01 per game. That season, the best five teams (alphabetically) were the Bruins, Canadiens, Flyers, Islanders, and Sabres. The worst were the Barons, Blues, Canucks, Capitals, and North Stars.

From the game log, I manually calculated that against the five best teams, the Penguins gave up 5.25 goals per game. Against the five worst teams, they gave up 3.14.

For that to be evidence that the Penguins have "bad team" goalies, you'd have to show that 5.25 goals against good teams is actually better than expected for a team that gives up 4.01, and that 3.14 against bad teams is actually worse than expected for a team that gives up 4.01.

How would you do that? Well, one thing you could do is find a matching team, one that also gave up 4.01 goals per game (or close to it), but had a random goalie. If that team gave up 6.00 goals against the good teams, but 2.50 against the bad teams, that would be confirmatory evidence.

The match wouldn't be perfect, because it might have to be from another season, and the "best" and "worst" groups might not be comparable. Still, even with just those three goalies, you'd have about 33 seasons to compare (if you require a minimum 30 game season). If you found the two most comparable instead of one, that would be 66 comparisons.

That's better than nothing. But there'd still be a lot of noise. To make it workable, you'd have to limit your sample to games those goalies started (in 1977-78, Herron himself gave up only 3.57 goals per game, compared to 3.96 for the team after subtracting empty net goals). There'd be even less noise if you used save percentage instead of goals against.

The process would be a bit like searching for clutch hitting, only with a lot less data. And it would be a lot of work ... but, there's an organization, the Hockey Summary Project, that collects NHL game summaries -- like a hockey Retrosheet -- and I've asked them for access to their database. I'm hoping they have shots on goal. (Also, those summaries might help us trying to trace the games that Dryden talked about in his book, the ones that didn't match the game logs.)

Before I go any further, does this make sense as a way to test Dryden's hypothesis? Can you think of any others that might be easier?


Labels: , ,

Monday, June 06, 2011

Interpreting regression interaction terms

Last post, after talking about the results from the "choking foul shooter" study, I mentioned that there was one additional assumption I had to make. That assumption was that, in the regression, the coefficients for "last 15 seconds" and "down 1-4 points" were close to zero.

The easiest way to explain that is to go through what an interaction term means in a regression. (Warning: This is boring statistics stuff, no sports content until the end.)

------

Suppose I want to figure out if stimulants help a student do better on an exam. So I run a regression to predict the exam score. I use a bunch of variables, like age, time studying, performance on other exams, grades on assignments, number of classes missed, and so on, but I also include a dummy variable for whether the student had (both) coffee and Red Bull before the exam.

After the exam, I run the regression, and I find the coefficient for "both coffee and Red Bull" is -3, and statistically significant. I conclude that if I were a student, I might consider not taking both coffee and Red Bull.

Fair enough, so far.

But, now, suppose I do the same experiment again, but, this time, I add a couple of new dummy variables -- whether or not the student had coffee (with or without Red Bull), and whether or not the student had Red Bull (with or without coffee). I don't remove the original "had both" variable -- that stays in.

I run the regression again, and, again, the coefficient for "both coffee and Red Bull" comes out to -3 -- exactly the same as last time. What am I able to conclude this time about the desirability of drinking both coffee and Red Bull?

The answer: almost nothing. That coefficient, *on its own*, does not give much useful information at all about how performance is affected by the coffee/Red Bull combination.

Let me explain why.

-------

(First, a quick not on terminology. In a regression, the "both coffee and Red Bull" variable would be referred to as "the interaction of coffee and Red Bull". That would be written as "coffee x Red Bull," or a suitable abbreviation (In fact, I'm going to start referring to "coffee" as C, "Red Bull" as R, and "Coffee x Red Bull" as CxR). The "x" is a multiplication sign -- it's there because you can get the coefficient by multiplying together the dummy values for C and R. That is, if either C or R is zero, CxR equals zero; if both coffee and Red Bull are 1, then CxR equals one. That's exactly what we want.)

-------

In a regression result, the simplest way to interpret the coefficient of a dummy variable is, "what happens when you change the value from 0 to 1 and leave all the other variables the same." In the first regression, that works fine. But in the second regression, it can't work. Because if you change CxR and leave everything else constant, your data and regression become inconsistent. You wind up with CxR being 1 (meaning both coffee and Red Bull), but you'll have either C=0 (no coffee) or R=0 (no Red Bull). Those three variables are tied together, so you can't just change CxR and leave the other two constant.

Put another way, there are four possible combinations for C, R, and CxR:

C = 0, R = 0, CxR = 0
C = 1, R = 0, CxR = 0
C = 0, R = 1, CxR = 0
C = 1, R = 1, CxR = 1

You can't change CxR from 0 to 1, and still have a combination that's on the list. So the "change CxR but leave all other variables the same" strategy no longer works. If you change CxR from 0 to 1, you'll have to change one of the other variables, too.

-------

Which ones should you change? It depends what question you're trying to answer. For example, suppose you do the regression and you get these coefficients:

C = -5
R = -10
CxR = -3

If you're trying to ask, "what's the effect of taking coffee alone versus nothing at all," it's like asking, "what is the effect of changing (C=0, R=0, CxR=0) to (C=1, R=0, CxR = 0)?" The answer is -5.

If you're trying to ask, "what's the effect of taking both coffee and Red Bull versus nothing at all?", it's like asking, what's the effect of changing (C=0, R=0, CxR=0) to (C=1, R=1, CxR =1)?" The answer is -18.

And so on. But none of those kinds of questions lead to the answer of -3 points, because none of these questions can be answered by changing CxR alone.

So what does the -3 represent? The non-linearity of the coffee and Red Bull variables. Or, put another way, the "increasing or diminishing returns" to combining coffee and Red Bull. Or, put a third way, the effects of the *interaction* of coffee and Red Bull, independent of their individual effects. Or, put a fourth way, the amount of effects *duplicated* from both coffee and Red Bull, that you can't count twice even if you take both drinks.

The -3 is NOT any indication of whether it's a good thing to take coffee and Red Bull together. Even though the coefficient of the interaction is negative, coffee and Red Bull together might be a positive thing. Suppose the regression coefficients had looked like this:

C = +10
R = +20
CxR = -3

The CxR coefficient is still -3, but now look what happens:

Take coffee, score 10 points higher
Take Red Bull, score 20 points higher
Take both coffee and Red Bull, score 27 points higher!

In this case, you're still going to want to take both coffee and Red Bull. What the -3 is telling you is, there are diminishing returns to taking both. You might think that, since coffee improves you by 10, and Red Bull improves you by 20, that, if you take both, you'll improve by 30. That's not right. There are diminishing returns of -3, so, if you take both, you'll only improve by 27.

-------

Of course, if the coefficients of C and R are both zero, then the CxR variable is indeed the entire effect. So if coffee does nothing, and Red Bull does nothing, but, when you take them together, you lose 3 points ... in that case, the CxR variable actually IS the effect of taking both C and R.

-------

This is fairly standard stuff, I would think ... I looked for an explanation on the web, so I wouldn't have to type all this, but I couldn't find one.

Anyway, going back to the choking study ... there, I looked at a variable called "Last15 x Down1_4", which was the interaction of shots that happen in the last 15 seconds (a dummy variable called "Last15"), and with the shooting team up by 1 to 4 points (dummy variable "Down1_4").

It turned out that the coefficient for that was -0.058 (as compared to 11-point+ blowouts). I implied that meant that shooters were 5.8 percentage points worse in those clutch situations than in blowouts.

But that wasn't right, because "Down1_4" and "Last15" were also in the regression. It's like the "Coffee / Red Bull / Both" case. If I want to compare the effects of shooting in the last 15 seconds down by 1-4 points, against shooting where *neither* of those is true, I have to add up all three coefficients:

Down1_4 = A
Last15 = B
Down1_4 x Last 15 = -0.058

To get the true clutch effect, I have to compute A + B - 0.058. It could turn out that A and B are huge: maybe they're +7 points each! In that case, the effect would be 0.07 + 0.07 - .058, which would be +0.082 -- which would mean shooters were GREAT in the clutch.

The study doesn't give us A and B. However, the authors do tell us (and author Dan Stone reiterated in the comments to the previous post) that almost all the omitted coefficients are less than 0.01.

Still, suppose they are as high as exactly +0.01. That means that A + B - 0.058 would be -0.038, which would less significant a choke effect than I thought. Or, suppose they were as low as negative 0.01. In that case, players would be even chokier -- at -0.078.

That's why I added a note to the end of my post, saying I had to make one additional assumption. That assumption is that A and B were both close to zero. If they're exactly zero, the -0.058 stands.

Labels: , , , ,

Monday, May 30, 2011

A new basketball free throw choking study

"Performance Under Pressure in the NBA," by Zheng Cao, Joseph Price, and Daniel F. Stone; Journal of Sports Economics, 12(3). Ungated copy here (pdf).

------

There's a new paper demonstrating evidence that NBA players choke when shooting free throws under pressure. The link to the paper is above; here's a newspaper article discussing some of the claims.

Using play-by-play data for eight seasons -- 2002-03 to 2009-10 -- the authors find that players' percentage on foul shots goes down significantly late in the game when they're behind. They present their results through a regression, but it's more obvious just by using their summary statistics. Let me show you the trend, which I got by doing some arithmetic on the cells in the paper's Table 1:

In the last 15 seconds of a game, foul shooters hit

.726 when down 1-4 points (922 attempts)
.784 when tied or up 0-4 points (5505)
.776 when up or down 5+ points (4510)

With 16-30 seconds left, foul shooters hit

.748 when down 1-4 points (727 attempts)
.775 when tied or up 1-4 points (2652)
.779 when up or down 5+ points (6174)

With 31-60 seconds left, foul shooters hit

.752 when down 1-4 points (922 attempts)
.742 when tied or up 0-4 points (1634)
.767 when up or down 5+ points (8969)

In all other situations, foul shooters hit about .751 regardless of score differential (400,000+ attempts).

------

Take a look at the first set of numbers, the "last 15 seconds" group. When down 1 to 4 points, it appears that shooters do indeed "choke," shooting almost 2.5 percentage points (.025) worse than normal. In 5+ point "blowout" situations late in the game, they shoot more than 2.5 percentage points *better* than normal.

But neither of these numbers is statistically significantly different from the overall average (which I'm guessing is about .751). The difference of .025 is about 1.7 SDs.

The real statistical significance comes when you compare the "down by 1-4" group, not to the average, but to the "5+ points" group. In that case, the difference is double: the "down 1-4" is .025 below average, and the "5+" group is .025 *above* average. The difference of .050 is now significant at about 3 SDs.


UPDATE: The above paragraphs are incorrect in one aspect. Dan Stone, one of the paper's authors, corrected me in the comments. What I didn't notice was that in Table 1, the overall free throw percentage of each group was provided. Those percentages are .779 (down 1-4 group), .795 (up 0-4 group), and .782 (5+ group). So the average for those particular players is move like .787 than .751. All three groups shot below expected, but the "down 1-4" group shot WAY below expected.

So the "down 1-4" group is, on its own, statistically significant from expected, without regard to the other two groups. My apologies for not noticing that earlier.

So, if you look only at the last 15 seconds of games, it looks like players down by 1-4 points choke significantly compared to players who are up or down by at least five points.

There are similar (but lesser) differences in the 16-30 seconds group, and the 31-60 seconds group. I haven't done the calculation, but I'm pretty sure you also get statistical significance if you combine the three groups, and compare the "down 1-4" to the "5+" group.

-------

So that's what we're dealing with: when you compare "down 1-4 late in the game" to "up or down 5+ late in the game", the difference is big enough to constitute evidence of choking. The most obvious explanation is that the foul shooters in the two groups might be different. However, that can't be the case, because the authors controlled for who the shooter was, and the results were roughly the same. Indeed, they controlled for a lot of other stuff, too: whether the previous shot was made or not, which quarter it is, whether it's a single foul shot or multiple, and so on. But even after all those controls, the results are pretty much the way I described them above.

Again, I repeat: the authors (and the data) do NOT say that the "choke" group shoots significantly worse than average. They can only say that the "choke" group is significantly worse than one specific group of players: the "don't care" group, shooting late in the game when the result is pretty much already decided, with a gap of 5+ points.

But this fits in with the authors' thesis: that the higher the leverage of the situation, the more choking you see. They later break down the "5+" group into "5-10" and "11+", and they find that even that breakdown is consistent -- the 11+ group shoots better than the (slightly) higher leverage 5-10 group. Indeed, for most of the study, they compare to "11+" instead of "5+". For some of the regressions, they post two sets of results, one relative to the "5-10 points" group, and one relative to the "11+" group. The "11+" results are more extreme, of course.

-------

As I said, the authors don't present the results the way I did above ... they have a big regression with lots of tables and results and such. The result that comes closest to what I did is the first column of their Table 5. There, they say something like,

"In the last 15 seconds of a game, a player down 1-4 points will shoot 5.8 percentage points (.058) worse than if the game were an 11+ point blowout. The SD of that is 2.1 points, so that's statistically significant at p=.01."

-------

Oh, and I should mention that the authors did try to eliminate deliberate misses, by omitting the last foul shot of a set with 5 seconds or less to go. Also, they omitted all foul shots with less than 5 minutes to go in the game (except those in the last 15/30/60 seconds that they're dealing with). I have absolutely no idea why they did that.

-------

Although the authors do mention the "down 1-4" effect above, it's almost in passing -- they spend most of their time breaking the effect down in a bunch of different ways.

The biggest effect they find is for this situation:

-- shooting a foul shot that's not the last of the set (that is, the first of two, or the first or second of three);
-- in the last 15 seconds of the game;
-- team down exactly one point.

compared to

-- shooting a foul shot that's not the last of the set (that is, the first of two, or the first or second of three);
-- in the last 15 seconds of the game;
-- score difference of 11+ points in either direction.

For that particular situation, the difference is a huge 10.8 percentage points (.108), significant at 2.5 SDs.

Also: change "down by one point" to "down by two points", and it's a 6.0 percentage point choke. Change "not the last of the set" to "IS the last of the set," and the choke effect is 6.6 points when down by 1, and 6.0 points when down by 2.

This highly specific stuff doesn't impress me that much ... if you look at enough individual cases, you're bound to find some effects that are bigger and some that are smaller. My guess is that the differences between the individual cases and the overall "down 1-4" case are probably random. However, the authors could counter with the argument that the biggest sub-effects were the ones they predicted -- the "down by 1" and "down by 2" case. On the other hand, late performance is actually *better* than blowouts when the score is tied (by around 0.2 points), a finding the authors say they didn't expect.

So my view is that maybe the "1-4 points" result is real, but I'm wary of the individual breakdowns. Especially this one: in this situation:

-- last 15 seconds of the game
-- for a visiting team
-- where the most recent foul shot was missed
-- down by 1-4 points

the player is 9.6 percentage points (.096) less likely to make the shot than

-- last 15 seconds of the game
-- for a visting team
-- where the most recent foul shot was missed
-- score 11+ points in favor of either team.

Despite the large difference in basketball terms, this one's only significant at .05.

------

Another thing about the main finding is ... we actually already knew it. Last year, I wrote about a similar study (which the authors reference) that found roughly the same thing. Here, copied from that other post, are the numbers those researchers found, for all foul shots in the last minute of games, broken down by score differential:

-5 points: -3% [percentage points]
-4 points: -1%
-3 points: -1%
-2 points: -5% (significant at .05)
-1 points: -7% (significant at .01)
+0 points: +2%
+1 points: -5% (significant at .05)
+2 points: +0%
+3 points: -1% ("also significant")
+4 points: +1%
+5 points: -1%

There are some differences in the two studies. The older study controlled for player career percentages, instead of player season percentages. It didn't control for quarter (which is why commenters suspected it might just be late-game fatigue). It didn't control for a bunch of other stuff that this newer study does. And it used only three seasons of data instead of eight.

But the important thing is: the newer study's eight seasons *include* the older study's three seasons. And so, you'd expect the results to be somewhat similar. It's possible that the three significant seasons are enough to make all eight seasons look significant, even if the other five seasons were just average.

Let's try, in a very rough way, to see if we can tease out the new study's result for those other five seasons.

In the older study, if we average the -1, -2, -3, and -4 effects, we get -3.5. So, in the last minute, down by 1-4 points, shooters choked by 3.5 percentage points.

How do we get the same number for the newer study? Well, in the top-left cell of Table 5, we get that, in the last 30 seconds and down by 1-4 points, shooters choked by 3.8 percentage points.

That's our starting point. But the new study's selection criteria are a little different from the old study's, so we need to adjust.

First, the "-3.8" in the new study is from comparing to games in which the point differential is 11 or more. The "5-10" is probably a more reasonable comparison to the previous study. The difference between "11+" and "5-10", at 30 seconds, appears to be about one percentage point (compare the second columns of Tables 3 and 4). So we'll adjust that 3.8 down to 2.8.

Second, the new study is for the last 30 seconds, while the old study is for the last minute. From earlier in this post, we see that the 31-60 difference between the "down 1-4 group" and the "5+" group is only about -1.5 percentage points. Averaging that with the -2.8 from the above step (but giving a bit more weight to the -2.8 because there were more shots there), we get to about -2.4.

So we can estimate, very roughly, that for the same calculation,

Old study (three seasons): -3.5
New study (eight seasons): -2.4

Let's assume that if the new study had confined itself to only the same three seasons as the older study, it would have come up with the same result (-3.5). In that case, to get an overall average of -2.4, the other five seasons must have averaged -1.74. That's because, if you take five seasons of -1.74, and three seasons of -3.5, you get -2.4.

So, as a rough guess, the new study found:

-3.5 -- same three seasons as the old study;
-1.7 -- five seasons the old study didn't cover;
-------------------------------------------------
-2.4 -- all eight seasons combined.

So, in the new data, this study finds only half the choke effect that the other study did. Moreover, I estimate it's only 1 SD from zero.

-------

That's for "down 1-4 points." Here's the same calculation, broken down by individual score. Here "%" means percentage point difference:

-2 points: First three: -5%. Next five: -2.3%. All eight: -3.3%.
-1 points: First three: -7%. Next five: -1.7%. All eight: -3.7%.
+0 points: First three: +2%. Next five: -1.8%. All eight: -1.0%.
+1 points: First three: -5%. Next five: -0.5%. All eight: -2.2%.
+2 points: First three: +0%. Next five: -0.3%. All eight: -1.3%.


Generally, five new seasons are closer to zero than the three original seasons. That's what you would expect if the original numbers were mostly luck.

-------

So, in summary:

-- The study finds that in the last seconds of games, players behind in close games shoot significantly worse than in blowouts.

-- In the last 30 seconds, they're maybe about 2.8 percentage points worse. In the last 15 seconds, they're maybe about 4.8 percentage points worse.
-- The effect is biggest when down by 1 in the last fifteen seconds.

-- However, they are not statistically significantly better or worse than *average,* just statistically significantly worse than blowouts (although they certainly are "basketball significantly" worse than average).

-- The effect is mostly driven by the three seasons covered in the earlier study. If you look at the other five seasons, the effect is not statistically significant (but still has the same sign).

What do you think? I'm not absolutely convinced there's a real effect overall, but yeah, it seems like it's at least possible.

However, I do think the most extreme individual breakdowns are overstated. For instance, the newspaper article says that in the last 15 seconds, down by 1, players will shoot "5 to 10 percentage points worse than normal." (They really mean "worse than 11+ blowouts," but never mind.) Given that that's the most extreme result the study found, I think it's probably a significant overestimate. I'd absolutely be willing to bet that, over the next five seasons, that the observed effect will be less than five percentage points.

--------

P.S. One last side point: the newspaper article says,

"Shooters who average 90 percent from the line performed slightly better than that under pressure, while 60 percent shooters had a choking effect twice as great as 75 percent shooters. That suggests that a lack of confidence begets less confidence, and vice versa."

This is a correct summary of what the authors say in their discussion, but I think it's wrong. The regressions that this comes from (Table 5, columns 2 and 6) don't include an adjustment for the player. So what it really means is that the 60 percent shooter will be *twice as far below the average player* as the 75 percent shooter. That makes sense -- because he's a worse shooter to begin with, even before any choke effect.

------

UPDATE: After posting this, I realized that I may have missed one aspect of the regression ... but I think my analysis here is correct if I make one additional assumption that's probably true (or close to true). I'll clarify in a future post.


Labels: , , ,

Saturday, May 07, 2011

Clutch hitting and getting killed by a puck

There was something someone said about clutch hitting a few months ago -- I think it was Tango -- that took a while to sink in for me.

It went something like this: saying that clutch hitting ability "does not exist" is silly. Humans are different, and different people react to pressure in different ways. So, *of course* there must be differences in clutch hitting ability. The question isn't whether or not clutch hitting exists, because we know it must, but *to what extent* it exists.

It turns out that extent is small. The best we can say is that clutch hitting studies have found that the SD of individual clutch tendencies is about 3 percent of the mean. So for players who are normally .250 hitters, two out of three of them will be between .242 and .258 in the clutch, and 19 out of 20 of them will be between .235 and .265. ("The Book" study on the topic used wOBA, rather than batting average, but this is probably still roughly true.)

That's pretty minor, especially compared to other factors like platoon advantage, and so on. More importantly, there's not nearly enough data to know which are the real clutch hitters and which are the real clutch chokers. When your favorite pundit pronounces player X as "clutch," that's still completely pulled out of his butt.

So
the conclusion remains that calling specific players "clutch" is silly, but the stronger statement "clutch hitting does not exist" is unjustified.

------

As I said, that took a while to sink in for me. I'm willing to agree with it. But I still have reservations. Because, you can take this "humans are different" stuff to extremes.

Suppose you found that a certain bench player hits much better than usual on the first day of the month. Some announcer notices and says the manager should always play him on those days. The sabermetricians step in and say, "that's silly: there's no reason to believe it's a real effect, and the guy's not very good on all the other days."

But, by the same logic ... all humans are different. For me personally, every day on the first of the month, I'm amazed how fast the past month flew by. It seems to me that time is going faster as I get older, and it makes me a little sad. On the other hand, other people might be happier on the first of the month. Maybe that's when they get paid, and they're feeling rich.

So, since humans are different, and circumstances are different, why couldn't their be a *real* "first of month" effect? The same logic says there *has* to be.

But, the thing is ... any such effect is probably very, very small. Too small to measure. There's no way it'll affect player performance (in my estimation) even one one-hundredth as much as clutch.

It's real, but it's too small to matter.

In cases like that, is it OK to say that "there's no such thing" as "first of month hitting talent"? Maybe it's not technically true, but I don't think that'll always stop me from saying it anyway. But, if I remember, I'll say "if it exists, it's probably infinitesimal."

For clutch, the Andy Dolphin study mentioned in "The Book" came up with an SD of clutch talent estimated at .008 in wOBA. I'm not 100% willing to accept that, mostly because, as Guy points out in a recent clutch thread on Tango's blog, there might be a natural clutch difference due to batter handedness or batting style that doesn't really reflect "clutch talent" as it's normally understood.

But, what I might choose to say is that there's weak evidence of a small clutch effect, and add that there isn't nearly enough evidence to know who's weakly clutch and who's weakly choke.

Or, I might say that there's "no real evidence of a meaningful clutch effect," which says the same thing, but with a more understandable spin.

------

Anyway, it occurred to me, while thinking about this, that we like to think about things as "yes/no" when what's really important is "how much". By that, I don't mean the moral argument about "black and white" versus "shades of grey." What I mean is something more quantitative -- "zero or not zero" versus "how much?"

Take, for example, Omega 3 fatty acids. I hear they're good for you. You read about them in the paper, and on milk cartons, and fish products.

But, *how* good for you are they? Isn't that important, to quantify it? The media doesn't seem to.

Here, for instance, is an article from the CBC. It talks about which Omega 3s are better than others, and ends with a recommendation that people eat 150 grams of fish a week. But ... how big are the benefits? The CBC article doesn't tell us. You'd expect better from the Wikipedia entry, but even that doesn't tell us.

So, how are we supposed to decide whether it's worth it?

I mean, suppose you really, really don't like fish. Won't your decision on whether or not to eat fish anyway, depend on how much the benefit is?

And, don't other foods have benefits too? If I eat fish for lunch, that means I won't be eating oat bran instead. How can I decide which is better for me without the numbers?

Or, suppose I have a choice ... go to a fish restaurant for lunch, or go for a jog around the block. Both might help my heart. Which one will help more? Or, suppose a nice piece of salmon in my favorite restaurant is $5 more than a piece of chicken. Could I do better with the $5? Maybe I could save up and use the money to switch to a more expensive gym, where I'll go more often. Or, I could put it in a travel fund, to buy better travel health insurance next trip I go abroad. Which option for my $5 is best for my long-term health?

There are lots of people who hear "fish is good for your heart," so they start buying fish. What they're really buying is a security and good feelings. Because if you ask them to quantify the benefits are, they just look at you blankly, or they quote some authority saying it's good for you. They treat it as a yes/no question -- is it good for you? -- when the important question is "how much is it good for you"?

Maybe I'm just cynical, but when someone tells me a food is good for me, without quantifying the benefits, I just assume the benefits are tiny, like clutch hitting. After all, for a study to make a journal, all you need is 5 percent significance, which means that most of the claims are probably not true. And, even if they *are* true, the study probably found that the benefits are small (though possibly statistically significant).

In his book "Innumeracy," back in the 80s, John Allen Paulos suggested there be a logarithmic scale for risk, so when the media tell you something is risky, they could also use that number to tell you how much. I'd like to see that for everything, not just risk. And, I'd like to see the scale changed. Because, if it's just a logarithm, people will still ignore it. If someone tells me I should have had the salmon, I might say, "Why? It's only a 0.3 on the Paulos scale, which is almost nothing." And they'll say, "0.3 is better than nothing. It all adds up. You should look out for your health."

But, what if you expressed the benefits in terms of something real, like exercise? Everyone intuitively understands the benefits of exercise, so the scale would make sense. And, it would make the insignificance of small numbers harder to rationalize away. So when someone insists I eat the salmon, I can say, "look, it's $5, and it's only the equivalent of 30 meters of jogging." They can still say, "30 meters is better than nothing!" And I'll say, "look, I'll have the chicken, and when we get outside the restaurant, I'll jog the 30 meters to the car, and save the $5."

They could use that scale in the supermarket, too. I buy milk with Omega 3 (not for the health benefits, but because it stays fresh longer). Wouldn't it be great if it said on the carton, "Each glass gives you the benefits of 0.3 push-ups"? That would be awesome.

I'm not sure if exercise is the absolute best reference point. Maybe "days of life" is better. ("Each floret of broccoli adds 3.2 seconds to your lifespan!") But, still, it should be possible to come up with *something* that'll work.

------

For risk, a good unit of measure would be "miles of driving". That works well because it's widely recognized that driving is dangerous, and we all know people who died in car accidents, so we have an idea of the risk. But, there's no moral stigma associated with it (unlike, say, smoking), so we can be fairly rational about it.

In a comment thread yesterday on Tango's blog, there was a discussion about putting a protective barrier down the lines of baseball stadiums, to prevent people from getting hurt by foul balls. That would be similar to what the NHL did, when they installed a mesh partition behind each net after the death of a spectator in 2002.

Suppose the hockey netting wasn't there. What would the risk be?

From 1917 to the end of the 2002 season, there were 37,480 regular season NHL games. Assuming 15,000 fans per game, that's 562 million fans. Suppose one-quarter of those fans are sitting in high-risk seats; that's 140 million fans. Finally, suppose that players shoot a lot harder now, so today's risk is double the historical average. That means it only takes 70 million of today's fans to shoulder the same risk as throughout the NHL's history. (We could add a bit for playoffs and pre-season, but never mind.)

So, that's one death out of 70 million people, or a risk of 1 / 70,000,000 of dying at any given game. Is that big or small? It's hard to say. What is it in terms of driving?

In 2010, there were 1.09 deaths per 100 million miles travelled. Let's round that down to 1.00, just to make the calculation easier. So there's 1 / 100,000,000 of a death per mile.

That means the hockey game is the equivalent of 1.43 miles.


So that's how I'd say it: putting the mesh up at hockey games makes each fan behind the net safer by 1.43 miles of driving.

Doesn't that give you a really good intuitive idea of the risk involved?

Of course, that's death only, and not injury. But I'm sure you could find injury data for car accidents, and for puck injuries, and come up with some kind of reasonable scale. I'm guessing that if you did that, you'd probably find that it was still about the same order of magnitude of a couple of miles. But I don't know for sure.

And, hey, now that I think about it, you could treat *healthy* things as "miles of driving saved*. If something saves you one minute of lifespan, that's easily converted to driving. 100 million miles, at an average of 30 miles an hour, is about 380 years of driving. (At six hours of driving (or passenging) a week, that's about 10,000 years of life per fatal accident. That would mean that about one American in 200 eventually winds up dying in a car accident. Sound about right?)

Suppose the average driver has 40 years of life left, on average. Then every 380 years of driving wipes out 40 years of life. That works out to 9.5 years of driving per year of life. Round that to 10. That means that every hour you drive -- 30 miles -- costs you six minutes of life. So five miles of driving costs you a minute of life.

(Again, that's death only, and not injury. In terms of quality of life, you'd probably want to bump up the "5 miles" figure, because bad health usually makes you miserable before you die, but car accidents often kill you instantly. But let's stick with five miles for now.)


So if eating salmon for a week saves you one minute of lifespan, eating salmon for a week is like cutting 5 miles off your commute one day. I made that "one minute" number up; if anyone knows how to figure out what the real number is, let me know.

------

Let's do cigarettes. Actually, let's just do lung cancer, to make it easier.

According to Wikipedia, 22.1% of male smokers will die of lung cancer before age 85. Let's assume that entire amount is from cigarettes.

Since this is a back-of-the-envelope calculation, let's just make some reasonable guesses. I'll assume the average male smoker starts at age 18 and smokes a pack a day. I'll also assume that when a smoker dies of lung cancer, it's at age 70 on average, and it cuts 15 years off his lifespan.

So: 52 years times 365 days times 20 cigarettes equals ... 379,600 cigarettes. 15 years of lost life equals 7,884,000 minutes. So each cigarette equals 21 minutes.

Multiply the 21 minutes by 22.1% and you get 4.6 minutes.


Google "cigarette minutes of life" and you get figures ranging from 3 to 11 minutes ... and that's of *all causes*, not just lung cancer. So, we're in the right range.

4.6 minutes equals 23 miles.

If you're a pack-a-day smoker, your risk is the same as if you drove from New York to Los Angeles every week or so.

------

If I were made evil dictator of the world, I would insist that every media report on risks and benefits tell you *how much*. Every report of a health scare, every quote from a safety group, every recommendation from a nutritionist, would need to include a number. Because, really, when someone tells you "vegetables are healthy," that's useless. Even if it's true, how true? Is it true like "the platoon advantage exists?" Is it true like "clutch hitting exists"? Or is it true like "first of month hitting exists?"

The difference matters. We need the numbers.


Labels: , , ,