Tuesday, March 06, 2012

Are early NFL draft picks no better than late draft picks? Part IV

This is the last post about the Berri/Simmons NFL draft paper, in which they say that draft position doesn't matter much for predicting quarterback performance. Here are Parts I and II and III.

------

In his Freakonomics post, Dave Berri argues, reasonably, that quarterbacks are harder to predict from season to season than basketball players.


When he runs a regression to predict NFL quarterbacks' completion percentage this season, based on only the stat from last season, he gets an r-squared of .311. On the other hand, if he does the same thing for "many statistics in the NBA," his r-squared "exceeds 70 percent."

According to Berri,
Link

"This is not surprising since much of what a quarterback does depends upon his offensive line, running backs, wide receivers, tight ends, play calling, opposing defenses, etc. Given that many of these factors change from season to season, we should not be surprised that predicting the performance of veteran quarterbacks is difficult."

But ... aren't basketball players also subject to changes in the quality of their teammates? Why should teammates be so much more important for football than basketball?

Well, they're not. Almost the entire difference is just sample size. Let me show you.

The r-squared from season to season depends on the variances of what kinds of things are constant between seasons, and what kinds of things are not. For the most part, we can call these "talent" (t) and "noise" (n), respectively.

If the r-squared for QBs between seasons is .31, that means

(t/(t+n)) * (t/(t+n)) = .31

Taking the square root of both sides gives

t / (t+n) = .56

And from there, you can multiply both sides by (t+n), and discover that

n = .79 * t

So, for a single NFL season, the variance due to noise is 79% of the variance due to talent.

Now, in the NFL, a quarterback will get maybe 450 passing attempts per season. In the NBA, a full-time player might get three times as many (FG attempts, FT attempts (even if you take those at half weight), and 3P attempts). So, the noise should be only 1/3 as large. Instead of noise being 79% of talent, it will be only maybe 26%. Call the new value of noise n'. Then,

n' = .26 * t

If you sub that back into the first equation, you get

(t/(t+n')) * (t/(t+n')) = .63

See? Just considering opportunities raises the r-squared of .31 all the way up to .63. Berri says it should "exceed 70%", and we probably could get that to happen if we included rebounds, or used a more sophisticated stat than just shooting percentage.

So, if quarterbacks are harder to predict than basketball players, it's simply because they don't play enough for their stats to be as reliable.

(UPDATE: As Alex alludes to in the first comment to this post, my logic assumes that "t" -- the variance of talent -- is roughly the same for QBs and NBA shooters. It might not be. But the point is, assuming they're the same is a reasonable first approximation, and that leads to the conclusion that sample size is the biggest difference.

So, maybe I should have been more conservative and said that it *could be* that they don't play enough for their stats to be as reliable.)

------

Which brings me back to Berri's (and Simmons's) academic study. There, they write,

"[Our] results suggest that NFL scouts are more influenced by what they see when they meet the players at the combine than what the players actually did playing the game of football."


Well, yes -- and perhaps the scouts SHOULD be more influenced by the combine. There's lots of noise in only one season of performance, and a rational scout won't weight it too heavily. What if the scout only saw one play? Then, it's obvious that he should be more influenced by the combine than the results. The less data you have, the more you have to weight the combine results.

Look at it this way. You have two pitchers. One throws 100 mph and had an ERA of 3.50 in 50 innings. The other throws 80 mph and had an ERA of 3.20. Which do you draft? Well, *of course* you draft the 100 mph guy. It's only after 200, 300, 400, 1000 innings that you might have enough evidence to change your mind.

------

The idea of random noise and sample size never figures into this paper at all. I don't think the authors even think about it. When they see unexplained variance, they always argue that it's something like the effect of teammates, instead of looking at binomial randomness. In fact, you get the impression they think there's no randomness at all, and the scouts could be perfect if only they were smarter.

For the record, the paper has no occurrences of the words "luck," "random," "binomial," or "sample size."



Labels: , , , ,

Saturday, March 03, 2012

Are early NFL draft picks no better than late draft picks? Part III

This is about the Berri/Simmons NFL draft paper, in which they say that draft position doesn't matter much for predicting quarterback performance. Here are Parts I and II.

-----

One of the paper's most important claims is that scouts are looking at the wrong things -- specifically, the results of the NFL combine.

At the combine, prospects are tested on a bunch of objective measurements. How fast they run the 40-yard dash. Their BMI (a measure of weight to height ratio). Their intelligence, as measured by the Wonderlic test. And, of course, their height.

But, the authors argue, those things don't matter, and scouts are completely misguided in looking at what happens at the combine. They say that a QB prospect's height, BMI, 40-yard-dash time, and Wonderlic score have almost no effect on performance.

And that one point is key. Because, most of the authors' argument goes (in my words):

Premise 1: Scouts care about combine stats.
Premise 2: Combine stats affect draft position.
Premise 3: Combine stats don't predict performance.

Conclusion: Scouts don't know what they're doing.

So, premise 3 is key. How do the authors prove it?

Here's what they did. They took the 121 QBs drafted from 1999 to 2008, for which they had full combine data. Then they ran a regression to predict the QB's senior year performance based on those factors.

They found no statistical significance for any of them. And they conclude:

"Such results indicate that the combine measures are not able to capture key attributes of the quarterback."


And, in a related footnote,

"Such results indicate that there is little relationship between the combine statistics and per play performance."


Again, I think the problem is that the effect is there, but there just isn't enough data for significance. Indeed, it seems to me that they almost COULDN'T find significance in a study of that type.

Look, how much of QB performance is affected by height? Probably not much, right? There are so many other things involved. I mean, this isn't basketball: you don't see a lot of quarterbacks who are 6-foot-9, which suggests that height can't be that big a deal.

If the effect is that small, how are you going to find statistical significance with only 121 datapoints? Especially when you're trying to predict ONE SINGLE YEAR of college performance, which is very noisy (made even more so because, first, the authors chose to predict a measure that's dependent on playing time)?

You can't, and the authors didn't.

But ... that just means your study isn't precise enough. It doesn't show the effect isn't there. You can't look for a needle in a haystack, from fifty feet away, looking through the wrong end of a pair of binoculars, then say, "we didn't see a needle, so it doesn't exist."

Berri and Simmons didn't even show the results of that regression, even though it's key to their story. They just mention "not significant therefore zero" and move on. But if they HAD given their results, I bet you'd see the standard error is wide enough to encompass not just zero, but also many possible values that are perfectly reasonable and perfectly in line with what scouts think height is worth.

The same thing for the other factors -- 40 yard dash speed, Wonderlic, and BMI. It was almost guaranteed that the regression wouldn't find small effects in that sample.

What about the overall results for the four factors? You might get none of the individual combine stats being significant, but the overall correlation might be. Was it?

We really need to see the estimates for the coefficients. How many of them are reasonable individually? If you add them all up, are they also reasonable? If they are, that's all the more reason to point out that the lack of significance doesn't prove anything.

Again, the authors don't show the results ... but they do give a little hint. They run a second regression, this time using rate statistics instead of playing time stats. In a footnote to that, the authors say,

"The adjusted R-squared from these regressions, though, is in the negative range and the F-statistic is statistically insignificant."


A negative adjusted R-squared ... at first glance, that seems to say no relationship.

Except ... I looked up "adjusted R-squared". And, it turns out, for a regression with 5 variables and 121 rows, you can have a negative adjusted R-squared even if the "real" R-squared is as high as .042. That's not as small as it looks. An r-squared of .042 is an r of around 0.2, which is nothing to sneeze at.

(That makes sense. According to this calculator, a single-variable regression on 121 rows needs to find a correlation of 0.178 to find statistically significance, and I think the "adjusted" is meant to make the 5-variable case comparable to the 1-variable case.)

But 0.2 is probably higher than the effect we're looking for. Or at least, on par with it.

Suppose you ranked all the QBs on their combine stats. And then you took a QB who was +1 in SD in combine stats, and compared him to one that was -1 SD in combine stats. What kind of difference would you expect in on-field performance between the two?

Well, to get a correlation of 0.2, you'd have to expect a difference of about 3 points of NFL Quarterback Rating, or 3 or 4 positions in the performance rankings. (To estimate that, I looked here, and added 3 points to the rating of a typical QB.)

Now, remember, QB performance is very noisy. 3-4 positions in the performance rankings probably means 5-6 positions in *talent* rankings.

That seems to me like it's too much. There's no way height, BMI, Wonderlic, and 40-yard-dash speed could be *that* important, could it, that it's 5 or 6 rankings?

If not, then we're looking for an effect that's too small to find with only 121 datapoints to look at.

So, I think Berri and Simmons' regression was doomed from the start. They were guaranteed not to find significance, even if the scouts were right.




Labels: , , , ,

Wednesday, February 29, 2012

Are early NFL draft picks no better than late draft picks? Part II

This is about the Dave Berri/Rob Simmons paper that concludes that QBs who are high draft choices aren't much better than QBs who are low draft choices. You probably want to read Part I, if you haven't already.

--------

If you want to look for a connection between draft choice and performance, wouldn't you just run a regression to predict performance from draft choice? The Berri/Simmons paper doesn't. The closest they come is their last analysis of many, the one that starts at the end of page 47.

Here's what the authors do. First, they calculate an "expected" draft position, based on a QB's college stats, "combine stats" (height, body mass index, 40 yard dash time, Wonderlic score), and whether he went to a Division I-A school. That's based on another regression earlier in the paper. I'm not sure why they use that estimate -- it seems like it would make more sense to use the real draft position, since they actually have it, instead of their weaker (r-squared = 0.2) estimate.

In any case, the use that expected draft position, and run a regression to predict performance for every NFL season in which a QB had at least 100 plays. They also include terms for experience (a quadratic, to get a curve that rises, then falls).

It turns out, in that regression, that the coefficient for draft position is not statistically significant.

And, so, Berri and Simmons conclude,


"Draft pick is not a significant predictor of NFL performance. ... Quarterbacks taken higher do not appear to perform any better."


I disagree. Two reasons.


1. Significance

As the authors point out, the coefficient for draft position wasn't nearly significant -- it was only 0.52 SD from zero.

But, it's of a reasonable size, and it goes in the right direction. If it turns out to be non-significant, isn't that just that the authors didn't use enough data?

Suppose someone tells me that candy bars cost $1 at the local Kwik-E-Mart. I don't believe him. I hang out at the store for a couple of hours, and, for every sale, I mark down the number of candy bars bought, and the total sale.

I do a regression. The coefficient comes out to $0.856 more per bar, but it's only 0.5 SD.

"Ha!" I tell my friend. "Look, that's not significantly different from zero! Therefore, you're wrong! Candy bars are free!"

That would be silly, wouldn't it? But that's what Berri and Simmons are doing.

Imagine two quarterbacks with five years' NFL experience. One was drafted 50th. The other was drafted 150th. How much different would you expect them to be in QB rating? If you don't know QB rating, thing about it in terms of rankings. How much higher on the list would you expect the 50th choice to be, compared to the 150th? Remember, they both have 5 years' experience and they both had at least 100 plays that year.

Well, the coefficient would say the early one should be 1.9 points better. I calculate that to be about 15 percent of the standard deviation for full-time quarterbacks. It'll move you up in the rankings two or three positions.

Is that about what you thought? It's around what I would have thought. Actually, to be honest, maybe a bit lower. But well within the bounds of conventional wisdom.

So, if you do a study to disprove conventional wisdom, and your point estimate is actually close to conventional wisdom ... how can you say you've disproven it?

That's especially true because the confidence interval is so wide. If we add 2 SD to the point estimate, we find that the effect of draft choice could be as high as -0.091. That means that 100 draft positions is worth 9.1 points. That's a huge difference between quarterbacks. Nine points would move you up at least 6 or 7 positions -- just because you were drafted earlier. It's almost 75 percent of a standard deviation.

Basically, the confidence interval is so wide that it includes any plausible value ... and many implausible values too!

That regression doesn't disprove anything at all. It's a clear case of "absence of evidence is not evidence of absence."


2. Attrition

In Part I, I promised an argument that doesn't require the assumption that QBs who never play are worse than QBs who do. However, we can all agree, can't we, that if a QB plays, but then he doesn't play any more because it's obvious he's not good enough ... in *that* case, we can say he's worse than the others, right? I can't see Berri and Simmons claiming that Ryan Leaf would have been a star if only his coaches gave him more playing time.

If we agree on that, then I can show you that the regression doesn't work -- that the coefficient for draft choice doesn't accurately measure the differences.

Why not? Again, because of attrition. The worse players tend to drop out of the NFL earlier. That means they'll be underweighted in the regression (which has one row for each season). So, if those worse players tend to be later draft choices, as you'd expect, the regression would underestimate how bad those later choices are.

Here, let me give you a simple example.

Suppose you rate QBs from 1 to 5. And suppose the rating also happens to be the number of seasons the QB plays.

Let's say the first round gives you QBs of talent 5, 4, 4, and 3, which is an average of 4. The second round gives you 4, 2, 1 and 1, which averages 2.

Therefore, what we want is for the regression to give us a coefficient of 4 minus 2, which is 2. That would confirm that the first round is 2 better than the second round.

But it won't. Why not? Because of attrition.

-- The first year, everyone's playing. So, those years do in fact give us a difference of 2.

-- The second year, the two "1" guys are gone. The second round survivors are the "4" and "2", so their average is now 3. That means the difference between the rounds has dropped down down to 1.

-- The third year, the "2" guy is gone, leaving the second round with only a "4". Both rounds now average 4, so they look equal!

-- The fourth year, the "3" drops out of the first round pool, so the difference becomes 0.33 in favor of the first round.

-- The fifth year ... there's nothing. Even though the first round still has a guy playing, the second round doesn't, so the results aren't affected.

So, see what happens? Only the first year difference is correct and unbiased. Then, because of attrition, the observed difference starts dropping.

Because of that, if you actually do the regression, you'll find that the coefficient comes up 1.15, instead of 2.00. It's understated by almost half!

This will almost always happen. Try it with some other numbers and assumptions if you like, but I think you'll find that the result will almost never be right. The exact error depends on the distribution and attrition rate.

Want a more extreme case? Suppose the first round is four 4s (average 4), and the second round is a 7 and three 1s (average 2.5). The first round "wins" the first year, but then the "1"s disappear, and the second round starts "winning" by a score of 7-4.

In truth, the first round players are 1.25 better than the second round. But if you do the Berri/Simmons regression, the coefficient comes out negative, saying that the first round is actually 0.861 *worse*!

So, basically, this regression doesn't really measure what we're trying to measure. The number that comes out isn't very meaningful.

------

Choose whichever of these two arguments you like ... or both.

I'll revisit some of the paper's other analyses in a future post, if anyone's still interested.


------

UPDATE: Part III is here.



Labels: , , , ,

Monday, February 27, 2012

Are early NFL draft picks no better than late draft picks? Part I

Dave Berri thinks that NFL teams are inexplicably useless in how they evaluate quarterback draft choices. He believes this is true because of data he presents in a 2009 study, co-written with Rob Simmons, in the Journal of Productivity Analysis. The study is called "Catching the Draft: on the process of selecting quarterbacks in the National Football League amateur draft."

The study was in the news a couple of years ago, gaining a little bit of fame in the mainstream media when bestselling author Malcolm Gladwell debated it with Steven Pinker, the noted author and evolutionary psychologist.

In his book "What the Dog Saw," Gladwell wrote,

"... Berri and Simmons found no connection between where a quarterback was taken in the draft -- that is, how highly he was rated on the basis of his college performance -- and how well he played in the pros."


Pinker, reviewing the Gladwell book in the New York Times, flatly disagreed.

"It is simply not true that a quarter­back’s rank in the draft is uncorrelated with his success in the pros."

Gladwell wrote to Pinker, asking for evidence that would contradict Berri and Simmons' peer-reviewed published study. Pinker referred Gladwell to some internet analyses, one of which was from Steve Sailer. Gladwell was not convinced, but responded mostly with ad hominem attacks and deferrals to the credentialized:

"Sailer, for the uninitiated, is a California blogger with a marketing background who is best known for his belief that black people are intellectually inferior to white people. Sailer’s “proof” of the connection between draft position and performance is, I’m sure Pinker would agree, crude: his key variable is how many times a player has been named to the Pro Bowl. Pinker’s second source was a blog post, based on four years of data, written by someone who runs a pre-employment testing company, who also failed to appreciate—as far as I can tell (the key part of the blog post is only a paragraph long)—the distinction between aggregate and per-play performance. Pinker’s third source was an article in the Columbia Journalism Review, prompted by my essay, that made an argument partly based on a link to a blog called “Niners Nation." I have enormous respect for Professor Pinker, and his description of me as “minor genius” made even my mother blush. But maybe on the question of subjects like quarterbacks, we should agree that our differences owe less to what can be found in the scientific literature than they do to what can be found on Google."

Pinker replied:

"Gladwell is right, of course, to privilege peer-reviewed articles over blogs. But sports is a topic in which any academic must answer to an army of statistics-savvy amateurs, and in this instance, I judged, the bloggers were correct."


And, yes, the bloggers *were* correct. They pointed out a huge, huge problem with the Berri/Simmons study. It ignored QBs who didn't play.

----

As you'd expect, the early draft choices got a lot more playing time than the later ones. Even disregarding seasons where they didn't play at all, and even *games* where they didn't play at all, the late choices were only involved in 1/4 as many plays as the early choices. Berri and Simmons don't think that's a problem. They argue -- as does Gladwell -- that we should just assume the guys who played less, or didn't play at all, are just as good as the guys who did play. We should just disregard the opinions of the coaches, who decided they weren't good enough.

That's silly, isn' t it? I mean, it's not logically impossible, but it defies common sense. At least you should need some evidence for it, instead of just blithely accepting it as a given.

And, in any case, there's an obvious, reasonable alternative model that doesn't force you to second-guess the professionals quite as much. That is: maybe early draft choices aren't taken because they're expected to be *better* superstars, but because they're expected to be *more likely* to be superstars.

Suppose there is a two-round draft, and a bunch of lottery tickets. Half the tickets have a 20% chance of winning $10, and the other half have a 5% chance of winning $10. If the scouts are good at identifying the better tickets, everyone will get a 20% ticket in the first round, and a 5% ticket in the second round.

Obviously, the first round is better than the second round. It has four times as many winners. But, just as obviously, if you look at only the tickets that win, they look equal -- they were worth $10 each.

Similarly for quarterbacks. Suppose, in the first round, you get 5 superstar quarterbacks and 5 good ones. In the last round, you get only one of each. By Berri's logic, the first round is no better than the last round! Because, the 10 guys from the first round had exactly the same aggregate statistics, per play, as the 2 guys from the last round.

I don't see why Gladwell doesn't get it, that the results are tainted by the selective sampling.

Anyway, others have written about this better than I have. Brian Burke, for instance, has a nice summary.

Also,
Google "Berri Gladwell" for more of the debate.

-----

The reason I bring this up now is that, a couple of days ago, Berri reiterated his findings on "Freakonomics":

"We should certainly expect that if [Andrew] Luck and [Robert] Griffin III are taken in the first few picks of the draft, they will get to play more than those taken later. But when we consider per-play performance (or when we control for the added playing time top picks receive), where a quarterback is drafted doesn’t seem to predict future performance."


What he's saying is that Andrew Luck, who is widely considered to be the best QB prospect in the world, is not likely to perform much better than a last-round QB pick, if only you gave that last pick some playing time.

Presumably, Berri would jump at the chance to trade Luck for two last-round picks. That's the logical consequence of what he's arguing.

-----

Anyway, I actually hadn't looked at Berri's paper (.PDF) until a couple of days ago, when that Freakonomics post came out. Now that I've looked at the data, I see there are other arguments to be made. That is: even if, against your better judgment, you accept that the unknowns who never got to play are just as good as the ones who did ... well, even then, Berri and Simmons's data STILL don't show that late picks are as good as early picks.

I'll get into the details next post.

-----

UPDATE: That next post, Part II, is here. Part III is here.



Labels: , , , ,

Tuesday, May 19, 2009

Don't always blindly insist on statistical significance

Suppose you run a regression, and it turns out that the input you're investigating turns out to appear to have a real-life relationship to the output. But it also turns out that the despite being significant in the real-life sense, the relationship is not statistically significant. What do you do?

David Berri argues (scroll down to the second half of his post) that once you realize the variable is statistically insignificant, you stop dead:

We do not say (and this point should be emphasized) the “coefficient is insignificant” and then proceed to tell additional stories about the link between these two variables.

One of my co-authors puts it this way to her students.

“When I teach econometrics I tell my students that a sentence that begins by stating a coefficient is statistically insignificant ends with a period.” She tells her students that she never wants to see “The coefficient was insignificant, but…”


Well, I don't think that's always right. I explained why in a post two weeks ago, called "Low statistical significance doesn't necessarily mean no effect." My argument was that, if you already have some reason to believe there is a correlation between your input and your output, the result of your regression can help confirm your belief, even if it doesn't rise to statistical significance.

Here's an example with real data. I took all 30 major league teams for 2007, and I ran a regression to see if there was a relationship between the team's triples and its runs scored. It turned out that there was no statistically-significant relationship: the p-value was 0.23, far above the 0.05 that's normally regarded as the threshold.

Berri would now say that we should stop. As he writes,

"Even though we have questions, at this point it would be inappropriate to talk about the coefficient we have estimated ... as being anything else than statistically insignificant."


And maybe that would be the case if we didn't know anything about baseball. But, as baseball fans, we know that triples are good things, and we know that a triple does help teams score runs. That's why we cheer our team's players when they hit them. There is strong reason to believe there's a connection between triples and runs.

So I don't think it's inappropriate at all to look at our coefficient. It turns out that the coefficient is 1.88. On average, every additional triple a team hit was associated with an increase of 1.88 runs scored.

Of course, there's a large variance associated with that 1.88 estimate -- as you'd expect, since it wasn't statistically significant from zero. The standard deviation of the estimate was 1.53. That means a 95% confidence interval is approximately (-1.18, 4.94). Not only is the 1.88 not significantly different from zero, it's also not significantly different from -1, or from almost +5!

But why can't we say that? Why shouldn't we write that we found a coefficient of 1.88 with a standard deviation of 1.53? Why can't we discuss these numbers and the size of the real effect, if any?

Berri and his co-author would argue that it's because we have no good evidence that the effect is different from zero. But what makes zero special? We also have no good evidence that the effect is different from 1.88, or 4.1, or -0.6. Why is it necessary to proceed as if the "real" value of the coefficient is zero, when zero is just one special case?

As I argued before, zero is considered special because, most of the time, there's no reason to believe there's any connection between the input and the output. Do you think rubbing chocolate on your leg can cure cancer? Do you think red cars go faster than black cars just by virtue of their color? Do you think standing on your head makes you smarter?

In all three of these examples, I'd recommend following Berri's advice, because there's overwhelming logic that says the relationship "should" be zero. There's no scientific reason that red makes cars go faster. If you took a thousand similarly absurd hypotheses, you'd expect at least 999 of them to be zero. So if you get something positive but not statistically significant, the odds are overwhelming that the non-zero point estimate got that way just because of random luck.

But, for triples vs. runs, that's not the case. Our prior expectation should be that the result will turn out positive. How positive? Well, suppose we had never studied the issue, or read Bill James or Pete Palmer. Then, we might naively figure, the average triple scores a runner and a half on base, and there's a 70% chance of scoring the batter eventually. That's 2.2 runs. Maybe half the runners on base would score eventually even without the triple, so subtract off .75, to give us that the triple is worth 1.45 runs. (I know these numbers are wrong, but they're reasonable for what I might have guessed pre-Bill James.)

If our best estimate going in was that a triple should be worth 1.45 runs, and the regression gave us something close to that (and not statistically significantly different), then why should we be using zero as a basis for our decision for whether to consider this valid evidence?

Rather than end the discussion with a period, as Berri's colleague would have us do, I would suggest we do this:

-- give the regression's estimate of 1.88, along with the standard error of 1.53 and the confidence interval (-1.18, 4.94).
-- state that the estimate of 1.88 is significant in the baseball sense.
-- admit that it's not significantly different from zero.
-- BUT: argue that there's reason to think that the 1.88 is in the neighborhood of what theory predicts.

If I were writing a paper, that's exactly what I'd say. And I'd also admit that the confidence interval is huge, and we really should repeat this analysis with more years' worth of data, to reduce the standard error. But I'd argue that, even without statistical significance, the results actually SUPPORT the hypothesis that triples are associated with runs scored.

You've got to use common sense. If you got these results for a relationship between rubbing chocolate on your leg and cancer, it would be perfectly appropriate to assume that the relationship is zero. But if you get these results for a relationship between height and weight, zero is not a good option.

And, in any case: if you get results that are significant in the real world, but not statistically significant, it's a sign that your dataset is too small. Just get some more data, and run your regression again.

------

Here's another example of how you have to contort your logic if you want to blindly assume that statistical insignificance equals no effect.

I'm going to run the same regression, on the 2007 MLB teams, but I'm going to use doubles instead of triples. This time, the results are indeed statistically significant:

-- p=.0012 (signficant at 99.88%)
-- each double is associated with an additional 1.50 runs scored
-- the standard error is 0.417, so a 95% confidence interval is (0.67, 2.33)

Everyone would agree that there is a connection between hitting doubles and scoring runs.

But now, Berri and his colleague are in a strange situation. They have to argue that:

-- there is a connection between doubles and runs, but
-- there is NO connection between triples and runs!

If that's your position, and you have traditional beliefs about how doubles lead to more runs (by scoring baserunners and putting the batter on second base), those two statements are mutually contradictory. It's obvious to any baseball fan that, on the margin, a triple will lead to at least as many runs scoring as a double. It's just not possible that a double is worth 1.5 runs, but the act of stretching it into a triple makes it worth 0.0 runs instead. But if you follow Berri's rule, that's what you have to do! Your paper can't even argue against it, because "the coefficient was insignificant, but ..." is not allowed!

Now, in fairness, it's not logically impossible for doubles to be worth 1.5 runs in a regression but triples 0.0 runs. Maybe doubles are worth only 0.1 runs in current run value, but they come in at 1.5 because they're associated with power-hitting teams. Triples, on the other hand, might be associated with fast singles-hitting teams who are always below average.

In the absence of other evidence, that would be a valid possibility. But, unlike the chocolate-cures-cancer case, I don't think it's a very likely possibility. If you do think it's likely, then you still have to make the argument using other evidence. You can't just fall back on the "not significantly different from zero."

Using zero as your baseline for significance is not a law in the field of statistical analysis. It's a consequence of how things work in your actual field of study, an implementation of Carl Sagan's rule that "extraordinary claims require extraordinary evidence." For silly cancer cures, for red cars going faster than black cars, saying there's a non-zero effect is an extraordinary claim. And so you need statistical significance. (Indeed, silly cancer cures are so unlikely that you could argue that 95% significance is not enough, because that would allow too many false cures (2.5%) to get through.)

But for triples being worth about the same as doubles ... well, that's not extraordinary. Actually, it's the reverse that's extraordinary. Triples being worth zero while doubles are worth 1.5 runs? Are you kidding? I'd argue that if you want to say triples are worth less than doubles, the burden is reversed. It's not enough to show that the confidence interval includes zero. You have to show that the confidence interval does NOT include anything higher than the value of the double.


According to David Berri, the rule of thumb in econometrics is, "if you don't have signficance, ignore any effect you found." But that rule of thumb has certain hidden assumptions. One of those assumptions is that on your prior beliefs, the effect is likely to be zero. That's true for a lot of things in econometrics -- but not for doubles creating runs.

-----

This doubles/triples comparison is one I just made up. But there's a real life example, one I talked about a couple of years ago.

In that one, Cade Massey and Richard Thaler did a study (.pdf) of the NFL draft. As you would expect, they found that the earlier the draft pick, the more likely the player was to make an NFL roster. Earlier choices were also more likely to play more games, and more likely to make the Pro-Bowl. Draft choice was statistically significant for all three factors.

Then, the authors attempted to predict salary. Again as you'd expect, the more games you played, and the more you were selected to the Pro Bowl, the higher your salary. And, again, all these were statistically significant.

Finally, the authors held all these constant, and looked at whether draft position influenced salary over and above these factors. It did, but this factor did not reach statistical significance. Higher picks earned more money, but by somewhere between 1 and 2 SDs.

From the lack of significance, the authors wrote:

" ... we find that draft order is not a significant explanatory variable after controlling for [certain aspects of] prior performance."

I disagree. Because for that to be true, you have to argue that

-- higher draft choices are more likely to make the team
-- higher draft choices are more likely to play more games
-- higher draft choices are more likely to make the Pro-Bowl

but that

-- higher draft choices are NOT more likely to be better players in other ways than that.

That makes no sense. You have two offensive linemen on two different teams -- good enough to play every game for five years, but not good enough for the Pro Bowl. One was drafted in the first round; one was drafted in the third round. What Massey and Thaler are saying is that, despite the fact that the first round guy makes, on average, more money than the third round guy, that's likely to be random coincidence. That flies in the face of the evidence. Not statistically significant evidence, but good evidence nonetheless -- a coefficient that goes in the right direction, is signficant in the football sense, and is actually not that far below the 2 SD cutoff.

That isn't logical. You've shown, with statistical significance, that higher picks perform better than lower picks in terms of playing time and stardom. The obvious explanation, which you accept, is that the higher picks are just better players. So why would you conclude that higher picks are exactly the same quality as lower picks in the aspects of the game that you chose not to measure, when the data don't actually show that?

In this case, it's not only acceptable, but required, to say "the coefficient was insignificant, but ..."



Labels: , ,

Friday, May 08, 2009

Why r-squared doesn't tell you much, revisited

In a blog post I wrote about yesterday, "Wages of Wins" author Stacey Brook ran a regression to try to figure out what kind of relationship there is between an NBA team's payroll and its success on the court.

The regression gives you several pieces of information. Which ones should you use to best explain the relationship?

Brook says it's the r-squared. He writes,

"We use R2 since we are interested in the proportion of variance that is in common between NBA team payroll and NBA team performance."


But is that truly what we're interested in? I don't think so.

I do agree with Brook when he says that R-squared gives you "the proportion of variance that is in common between NBA team payroll and NBA team performance." But what does that mean? Almost nothing, unless you're a statistician.

When you do research like this, there's a question that you want to answer. In this case, if your question is "what proportion of variance is in common between NBA team payroll and NBA team performance?," well, then, there's your answer. But that's not the question. It's not even Brook's real question. His real question is implied by the first paragraph of his post:

"I have to disagree that NBA (or for that matter NHL, MLB or NFL) teams that have high payrolls result in higher winning percentages; nor am I the first to say this."


The question is: do teams with higher payrolls do better on the court? And that question is different from "what proportion of variance is in common between NBA team payroll and NBA team performance?"

If you want to see what payroll does to performance, what you want to see is the regression equation. The way regression works, of course, is to plot all the datapoints on a graph, then draw the best fit straight line among those points. That line represents the best-fit relationship between payroll and wins.

If you do that for the 2008-09 NBA teams, you get

Wins = 0.61 (millions of $ spent) - 0.76

This, basically, answers your question, in several ways

-- every extra million dollars you spend on salaries gives you three-fifths of a win.
-- every extra $1.64 million you spend gives you an extra win.
-- if you spend $100 million, like the Knicks, you should win about 60 games.
-- if you spend only $45 million, like the Grizzlies, you should win only about 27 games.

Not that complicated, right? If you want to know about the direct relationship between salary and wins, the regression equation does it.

Of course, you want to check the statistical significance; it's possible that while the best-fit straight line says $1.64 million per win, that might not be significantly different from zero. (As it turns out, it IS significant, at the 99.5% level. In fairness to Brook, it appears his data source had incorrect information, and because of that, his results were not, in fact, significant.)

I think we can all agree, from these results, that it certainly does appear that spending leads to winning. When the highest-spending team is expected to go 60-22, and the lowest-spending team is expected to go 27-55, you can't really claim that payroll is irrelevant. (Again, in fairness to Brook, he didn't get results this extreme. With the incorrect data, the regression suggests the highest-spending team should only be 45-37.)

So if the regression equation is the gold standard for making these kinds of calculations, what's with the r-squared? Well, the r-squared answers a different question.

Let's suppose that you had no idea what makes teams win basketball games. You see the Cavs go 66-16, and you see the Clippers go 19-63, and you think, what causes the difference?

What you could do is list as many plausible things as you could think of. Payroll would be one of them. Maybe average days of rest. Maybe whether they're an offensive or defensive team. Maybe average age. Maybe pace of play. Just list them all, as many as you want. Then, run a regression, and look at the r-squared.

What the r-squared will do is tell you, in a certain mathematical sense, after correcting for all those variables, what percentage of all the variation in wins have you explained? What you're trying to do is get as close to 100% as you can. The closer you get, the more you've explained what makes teams win and what makes teams lose. Maybe, if you actually ran this regression, you'd get to something like 40%. If you adjusted team wins for all those variables, as best you could, your variance would decrease by 40%.

In this particular case, our regression didn't include all that other stuff, like pace of play or average age. We only had one variable, payroll. And it turned out that the r-squared was .256, which means that 25.6% of the variation is "explained" by payroll.

It doesn't sound like a lot. In "The Wages of Wins," Brook (and co-authors David Berri and Martin Schmidt) did that for MLB, and came up with only 18%. That doesn't sound like a very big number either, and those authors decide that means that payroll isn't very important.

But that doesn't follow.

The r-squared, the seemingly-low 25.6% number, does NOT tell you about the relationship between payroll and wins. It just tells you that payroll is 25.6% of the total variance, and other factors are 74.4%. But, if the total variance is large, 25.6% of it would be substantial.

When you go into the car dealership and ask for a price, you want the amount in dollars. If you ask "how much for that Camry," and the salesman says, "it's 700% of your monthly pay," it may sound like a lot. If he says, "it's 9.5% of your net worth," it may sound cheaper. And if he says, "it's less than 0.01% of Bill Gates' disposable income for the week," it may sound cheaper still. But those all represent the same number of dollars. The fact that one percentage is a large number, and one percentage is a small number, doesn't change that fact.

It's the same thing for r-squared. The size of the percentage number depends what it's a percentage of -- which happens to be the total variance of wins in the league. Do you know, intuitively, what that variance is? I don't. But I know that a lot of it is random chance. And random variation depends on sample size. You could have exactly the same relationship between salary and wins, but, in one case, the r-squared is .25, and in another case, it's .04, and in another case, it's .5.

I wrote before about one example of how that can happen. But I can do another right now.


Want to see how you can use the same data to get a larger r-squared? Easy. I'm going to take the actual data for the 30 teams, but group them into threes according to payroll. So instead of the three data points "$100 million, 32 wins" (Knicks), "$90.1 million, 66 wins" (Cavs), and "$86 million, 50 wins" (Mavericks), I'm going to add them all up into the one data point "$276.1 million, 148 wins". Then I'm going to repeat for the other 27 teams, until I have 10 sums of three teams. Then, I'm going to run a regression on those 10 data points.

What happens? The r-squared now goes up to .497 -- almost double what it was!

But while I was able to arbitrarily double the r-squared, the regression line stayed almost the same -- which makes sense, since the actual relationship between salary and wins shouldn't change just because we arranged the data differently. Using all 30 teams, we got 0.61 wins per million dollars. Using the 10 groups of three teams, we get 0.68 wins per million dollars. Pretty close.

Here, let me give you everything in one place:

30 teams.... r-squared = 0.256
10 groups... r-squared = 0.497

30 teams.... Wins = 0.61 ($millions) - 0.76
10 groups... Wins = 0.68 ($millions) - 5.5


If Stacey Brook did the analysis his way, using all 30 teams, he'd say "salary explains 25.6% of the variance in wins." If I do the analysis my way, using groups of three teams, I'd say "salary explains 49.7% of the variance in wins." Which one of us would be right? Both of us! Because we are using different denominators, different variances. The same Toyota Camry can be a smaller percentage of Brook's salary than of my salary, because our salaries are different.

And so saying "payroll explains 25.6% of the variance of wins" is like saying "a Camry costs 35% of salary." Whose salary, and how much does he earn? Unless you know that, the "35%" figure is useless.

But, again, despite the fact that Brook and I did our regression differently, the equation should come out very similar. It won't come out exactly the same, because of random fluctuation, but you should *expect* it to come out the same, in the same sense as you expect a coin to come up heads 50% of the time. 0.61 wins per $million and 0.68 wins per $million are pretty close.

The regression equation is meaningful, it requires less information to interpret, and its expected value is the same regardless of your sample size. Most importantly, it answers the exact question that you want to know.

The r-squared, on the other hand, is unintuitive, can be made to come out to almost anything you like by tweaking the sample size to get a different total variance, and requires you to know how the study was done in order to interpret what it means. In terms of answering real-life questions, it's not very useful at all.


Labels: , , , , ,

Wednesday, May 06, 2009

Low statistical significance doesn't necessarily mean no effect

The "Wages of Wins" blog is written mostly by David Berri, but, as it turns out, co-author Stacey Brook also blogs. Recently, Brook had a post on the relationship between salary and wins.

He says there is none. Seriously. Not that the relationship is weak, not that money doesn't help much. Brook seems to honestly believe that salary doesn't buy wins at all. Read the full post to see if I'm interpreting him correctly, but here's a quote:

"So not only the proportion of variance that is common between the two tiny, but here I am able to show that the correlation coefficient between the two populations (NBA payroll and NBA performance) for the 2008-2009 season is statistically zero."


I have several problems with this analysis. The first one is not unique to Brook, and it drives me nuts. It's the idea that if you do a regression, and the significance level is less than 95%, it's OK to claim that there is no relationship between the variables.

That's not always right. It's often right; I suppose you could even say it's *usually* right. But this is one of those exceptions where it's not right at all.

Let's suppose that somehow you get it into your head that rubbing chocolate on your legs can help cure cancer. So you set up a double-blind experiment, where one set of patients gets the chocolate rub, and the other set gets a rub with fake chocolate. It turns out that the first group actually improves more than the second group -- by a small amount, maybe 1%. But the result is not statistically significant. Maybe, instead of the 95% you were looking for, you only have 80% significance.

In this case, I agree with Brook -- it would be wrong to argue that the 1% improvement you saw was real. It's probably just random chance, and you'd be justified in saying that there's no reason to believe that a chocolate rub has any therapeutic value at all.

But, now, let's turn to salary and wins. Suppose you study actual NBA payrolls and records, and you find a similar small effect: every $1 million gives you 0.1 extra wins. Again, suppose that's significant at only the 80% level.

In this case, can you draw the same conclusion, that money has no effect on wins at all? No, you can't. In this case, it's likely that the effect is real, despite the low significance level.

Why the difference? Because in the first case, there was absolutely no reason to believe that chocolate can have any effect on cancer. There's no previous scientific evidence for it, and there isn't a plausible mechanism for how the effect might work.

Suppose that, going in to the study, you (generously) thought there couldn't be more than a one in a million chance that chocolate helps treat cancer. So imagine a million different universes where you run the experiment. One time, you'll get a real effect. 200,000 times, you'll get 80% significance just by chance. So the chance that the chocolate actually works in this universe is roughly 1 in 200,001. That's still no reason to believe.

But the salary case is very different. There's no basis to believe that chocolate can cure cancer, but there's very good reason to believe that spending money buys better players and leads to more wins. In fact, every serious basketball fan in the world (except maybe Stacey Brook) believes that you can buy wins. When the Celtics pay Kevin Garnett some $25 million, does anyone really believe that the signing won't help the team? That if the Celtics instead paid $500,000 for some mediocre guy, they'd be doing just as well?

In the salary case, when you run regressions and get only 80% significance, the calculation works out differently. Suppose that going into the study, you figured there was a 99% chance that money helped buy performance (which is again conservative). Then, in a million different universes, you'd get 2,000 where the 80% signfiicance came up just by chance; and you'd get 990,000 universes where the effect is real. The chance, then, that salary actually does buy wins in this particular universe is 99.8% (990,000 divided by 992,000). The effect that Brook found is probably a real one.

(The above argument can be put into more formal mathematics using Bayesian probability, but I won't bother -- first, because it makes more sense to explain it in plain English, and, second, because I don't remember all the terminology and notation from the one Bayesian course I took in 1996.)

-----

Here's another way to look at it, if you don't like the "multiple universes" approach.

There are two possible reasons you might get a non-significant correlation between two variables:

1. There really is no relationship between the variables; or

2. There *is* a relationship, but you haven't looked at enough data to get a high enough significance level.

Almost any relationship, no matter how strong, will give you low significance if your sample size is too small. If you look at one random Ted Williams game, and one random Mario Mendoza game, what kind of significance level will you get? Pretty low. Even if Ted goes 2-for-5, and Mario goes 1-for-5 -- both of which are more extreme than their career averages -- you won't find the difference to be significant at the 95% level. One game is just not enough.

That doesn't mean this particular experiment is useless. You can still show the effect that you found, and invite further investigation. In this case, the difference between Williams and Mendoza is huge in the baseball sense -- .400 vs. .200. As a general rule, when you find an effect that's significant in the real-life sense, but not in the statistical sense, that's an indication that you might need more data. If the observed effect does have real-life importance, you are NOT entitled to conclude that there is no relationship between the variables. You are only entitled to conclude that you need more data.

And, in my opinion, you MUST show the size of the effect you found, not just the signficance level. Brook doesn't do that in his blog post. He gives us significance levels, and r, and r-squared, but the purpose of the study was to estimate the relationship between payroll and wins. Is it $5 million per win? $10 million per win? $15 million per win? Because, regardless of the significance level, the slope of the best-fit line is still the best estimate of that relationship. And I suspect that the results are reasonable, very close to what other analysts have estimated as the rate at which you can buy wins.

I suspect if we were able to look more closely at Brook's study, we'll find that:

-- he got an estimate of wins per dollar that's close to conventional wisdom;
-- but he didn't have enough data to get statistical significance;
-- so he claims that the proper estimate of wins per dollar is zero.

That ain't right.

-----

P.S. Probably more on this topic in the next post -- for a preview, this is why I think Brook got such low correlation.

UPDATE: Actually, I think Brook got a low correlation because the data was flawed. Details in my next post here.





Labels: , , , , ,

Wednesday, June 20, 2007

Bill James' job

Here, from the Wall Street Journal, is another one of those profiles of Bill James, and how he now works for the Red Sox after a quarter-century of writing books. The article's first line:


"After 25 years on the outside, Bill James was invited to take a seat at the center of the baseball universe."

When I read that line, it struck me as somehow wrong. Bill James, it seems to me, was at the "center of the baseball universe" when he was discovering important things about baseball, writing about them brilliantly, and sharing them with hundreds of thousands of rabid fans. Now, he discovers smaller, less-important things, doesn't write about them, and shares them only with the Red Sox front office. He's obviously less influential now than he was then.

Everyone's personality and taste is different, but, speaking for myself, I'd rather have Bill James' old job than his new one. What fun is discovering new things if you can't tell anyone?

To me, the best job in sports analysis is something like Dave Berri's. They pay you to do sabermetric research, you get to teach it, you get to write books, and you know that front offices are reading what you have to say. Can't get much better than that.

Labels: ,

Wednesday, November 29, 2006

Does "Win Score" overvalue rebounds?

In the past little while, there's been a debate about a basketball statistic from "The Wages of Wins" called "Win Score." The statistic, invented by authors Berri, Schmidt, and Brook, attempts to calculate how many wins each player contributed to the team. One of its forms is


Win Score = Points + Rebounds + Steals + ½*Assists + ½*Blocked Shots - Field Goal Attempts – ½*Free Throw Attempts – Turnovers – ½*Personal Fouls


The debate, for which details can be found at TWOW posts here and here, is this: does this statistic overrate rebounds?

King Kaufman believes it does. John Hollinger believes it does. I also believe it does.

First, the data shows that not every player has the same opportunity to try for a rebound. After a missed shot, only about 30% of rebounds are secured by the offense; the other 70% by the defense. (I got that 30% figure from
this comment.)

Obviously, the circumstances of where players find themselves has a bearing on who gets the rebound. Otherwise, the breakdown would be 50-50, not 70-30.

So, for some reason, players have different chances of rebounding that are related to positioning, rather than raw skills. Crediting a player for plays he makes only because of his position tends to overrate the value of those skills. I don't know enough about basketball to know if or how certain players are somehow set up for more rebounds – but to the extent to which that happens, if any, rebounds will tend to be overvalued in players' accounts. Just like cleanup hitters have more RBI opportunities just due to circumstances, some players may have more rebounding opportunities due to circumstances. And the 70-30 split shows there is certainly some of that going on. And the more it's circumstances, the less it's skill on the part of the player.

To see why, consider a more extreme example. Imagine that the NBA institutes a new rule: the offense is prohibited from touching a rebound until it has bounced three times on the floor.

That rule change will do nothing to affect TWOW's regression or logic. A defensive rebound still constitutes a change of possession, and is therefore still worth exactly the same number of wins as it was before. But, now, instead of 70% of rebounds going to the defense, the number is now 99%. Dennis Rodman might still snag a large proportion of rebounds, but now, instead of having to run and jump and position himself and maybe fight off an opposing player, he can just jog to where the ball is and pick it up.

Given that there is now no skill at all, doesn't it overrate Rodman to give him credit for those rebounds? Obviously, any excess rebounds picked up by Rodman, instead of his teammates, are positioning, luck, or opportunities given him by his coach and team. Even a caveman could get them.

The argument for 99% also applies to 70%, but to a lesser extent. Some, but not all, of Rodman's rebounds are, in effect, his team "letting him" have the ball more. Those are perhaps better classified as team rebounds, rather than individual rebounds. Since they aren't, Rodman winds up overrated.

That's opportunities. But there's a second reason rebounds are overrated, a much more important reason, and it has to do with the construction of Win Score itself.

It's the reason John Hollinger gives, the one TWOW disputes in the above links. That argument is that part (or even most) of the credit for a rebound should go to the other members of the team, for making the rebound possible. As Hollinger writes
here, "missed shots can be rebounded while turnovers can't, and ... a defensive rebound is merely the completing piece of a sequence that began by forcing a missed shot."

Suppose the NFL makes a rule change. Starting immediately, a touchdown is worth zero points instead of six – but, to compensate, the convert [extra point] is now worth seven points instead of one. A touchdown and convert is still worth seven points total. And since almost all converts are good, this doesn't change scoring in the NFL very much.

But now, running a regression assigns the entire seven points to the kicker. So suddenly, kickers are overrated, because they get credit for seven points instead of one! There's a 90-yard drive... the quarterback takes the team down the field, the receivers make some great catches, the running back drags two defenders three yards down the field for a third-down conversion, and they finally get the ball into the end zone. But, if you do a regression, it's the kicker that gets all the points ... the rest of the players come out at zero!

And the regression is absolutely correct – all things being equal, only the kicker matters. It's the interpretation of the regression that's questionable.

Really, the touchdown drive and the kick are one unit. No matter how good the kicker is, the only way he can get an opportunity to try for seven points is to have the rest of the team score a touchdown first. We know in our gut that it's really the touchdown that's worth the seven points, not the kick, because that's where the important skills came out. But the regression has no idea where the skills lie. It has no idea about what really caused the points, in the human sense. It sees when a kick is good, that's seven points. When a kick is bad, it's zero points. And everything else is irrelevant.

A similar situation happens for rebounds in basketball. To get the opportunity for a defensive rebound (convert), the defense must first force the opposition to miss (touchdown). The defensive rebound is a combination of the two acts: good defense for up to 24 seconds, and one grab of the ball. Crediting the rebounder with the full value of the defensive play is like crediting the kicker with all seven points of the touchdown.

And to get the opportunity for an offensive rebound, the shooter must have missed a field goal attempt. Win Score sees the two events superficially – the missed field goal is a turnover, and gets scored as such, and the offensive rebound is treated like a steal back from the defense. The shooter is charged with minus one possession, and the rebounder is credited with plus one possession.

But that's the wrong weighting. Any field goal attempt has, intrinsically, built into it, the embedded feature that a missed shot results in a 30% chance of getting the ball back. The miss includes a consolation prize, a lottery ticket with a 30% chance of winning back the possession. The shooter figured that into his decision about whether to make the shot. That 30% chance belongs to the shooter. In effect, he hasn't wasted a whole possession with his miss, he's only wasted 70% of a possession. Remember Hollinger's point – a missed shot gives the team a chance to recover, but a turnover doesn't. Obviously, the shooter should be debited less for getting a shot away than for letting the shot clock expire.

I think the correct way to handle rebounds in a stat like Win Score is to start by ignoring them. Take the league average rebounding stats, and give the entire contribution to the shooter and defense.

For offensive rebounds, note that on average, a missed field goal causes no damage 30% of the time. And so give the shooter back his 30% and charge him with only 70% of a turnover.

For defensive rebounds, note that they are the statistically average outcome of a defense good enough to force a missed shot. And so give all the credit for defensive rebounds – 70% of opposition missed shots -- to the defense, and ignore the rebounder.

(Remember that assigning values this way is completely compatible with the empirical data. If you were to run a regression that leaves out rebounds entirely, those are the weights you'd get – 70% of a turnover for a missed shot by either team.)

After all that, if the team turns out to be different from average, we can figure out how much different, and assign the credit or debit it to the players in proportion to what we think their contribution is. The hard part is figuring that out. Is Dennis Rodman a great rebounder with average opportunities, or an average rebounder with lots of opportunities? That's something you have to analyze properly, or you'll get bad results.

How much can the TWOW method overrate a rebounder? Let's take Kevin Garnett as an example. In 2005, the
Timberwolves had 947 offensive rebounds and 3527 defensive rebounds. Garnett was responsible for about 16% of the team's playing time. If rebounding were exactly proportional to playing time, Garnett would have come in at 150 offensive rebounds and 559 defensive. His actual numbers were 247 and 861. Garnett got to 399 more rebounds than average, or about 56% more than expected.

Is that difference a matter of skill, or opportunity? It's hard to argue that it's completely a matter of skill. The average team gets 70% of defensive rebounds. If Garnett is 56% better, a team of five Garnetts would get 109% of defensive rebounds! Now, you could argue that the five Garnetts would get in each other's way and take rebounds away from each other – there's only one ball, after all. But if you argue that five Garnetts would take rebounds away from each other, then you have to admit that there are times when two players both have a chance to make the play. And, therefore, there must be cases where Garnett takes rebounds away from his existing teammates! And so we have deduced that not all of that 56% can be simply Garnett's exceptional skill, because some of his rebounds would be snagged by a teammate if he weren't there. There must be at least some effect of opportunity there, and possibly a lot.

Now take the other extreme -- suppose Garnett is just an average rebounder, and his numbers are completely the result of opportunity. Then Garnett is being credited with wins that should really be going to the defense (for defensive rebounds) and the shooters who missed (for offensive rebounds). 401 rebounds is worth 14 wins. When we reallocate the defensive-rebound wins among all the players, about two will come back to Garnett. If we reallocate the offensive-rebound wins among shooters who missed, maybe half a win will come back to Garnett. Call it three wins total.

So if Garnett's rebounding is simply a matter of other players deferring to him, Garnett would be overrated by 11 wins. That's huge. Instead of being responsible for 30 wins out of his team's 45, he'd be responsible for only 19.


The correct number is somewhere between 19 and 30. Logic and evidence suggest that it has to be at least somewhat lower than 30. And so I think Hollinger and Kaufman are right -- and that TWOW's Win Points do indeed seriously overvalue rebounders.

Labels: , , , ,

Monday, November 20, 2006

Increasing NBA competitive balance

Today, on the "Wages of Wins" blog, David Berri reprises an argument from the book, that NBA competitive balance is low because of "the short supply of tall people." He writes,

"Given the supply of talent the NBA employs, there is very little the league can do to achieve the levels of competitive balance we see in soccer or American football."

I disagree. There are many ways the NBA could substantially increase competitive balance. Here are a few. Some are more realistic than others, of course:

-- add a 4-point, 5-point, and 6-point line behind the 3-point line.


-- make the 3-point shot worth 5 points.


-- make a "nothing but net" shot worth an extra point.


-- make the hoop a few inches smaller.


-- make the shot clock 12 seconds instead of 24.


-- overinflate the ball, like they do in carnival games.


-- make games 20 minutes in length instead of 48.


-- adjust the draft rules so that the worst teams get even more draft choices and the best teams get even fewer, thus more quickly evening out team talent over time.


-- institute a "talent cap" instead of a salary cap, so that teams with too many good players have to get rid of some.


-- count young players at estimated free agent value towards the salary cap, so that teams can't dominate just because of drafting ability.


-- divide the game into seven "quarters" instead of four. Whoever wins 4 out of 7 quarters is declared the winner of the game.


-- like in baseball, allow each player to take only about 1/9 of his team's offensive opportunities.


-- like in soccer and hockey, allow goaltending.


-- like in football, give teams points only when they complete a long sequence of successful plays – for instance, give them a seven-point "touchdown" when they score on six consecutive posessions, or a three-point "field goal" when they hit two three-pointers in a row.

(Just for one example, making the hoop smaller would help substantially. I simulated a simplified 100-possession game between a team that shoots field goals at 50% vs. a team that scores at 48%. The first team's record was .620. Then, I changed the probabilities from 50%/48% to 40%/38.4%. The first team's record dropped to .590.)

The point, of course, is that the rules of the game are at least as important as the supply of talent. For a full exposition of this argument, see Roland Beech's review
here. TWOW's recent rebuttal to Beech is here. (My own argument is on page 3 here.)





Labels: , , , ,