Tuesday, March 06, 2012

Are early NFL draft picks no better than late draft picks? Part IV

This is the last post about the Berri/Simmons NFL draft paper, in which they say that draft position doesn't matter much for predicting quarterback performance. Here are Parts I and II and III.

------

In his Freakonomics post, Dave Berri argues, reasonably, that quarterbacks are harder to predict from season to season than basketball players.


When he runs a regression to predict NFL quarterbacks' completion percentage this season, based on only the stat from last season, he gets an r-squared of .311. On the other hand, if he does the same thing for "many statistics in the NBA," his r-squared "exceeds 70 percent."

According to Berri,
Link

"This is not surprising since much of what a quarterback does depends upon his offensive line, running backs, wide receivers, tight ends, play calling, opposing defenses, etc. Given that many of these factors change from season to season, we should not be surprised that predicting the performance of veteran quarterbacks is difficult."

But ... aren't basketball players also subject to changes in the quality of their teammates? Why should teammates be so much more important for football than basketball?

Well, they're not. Almost the entire difference is just sample size. Let me show you.

The r-squared from season to season depends on the variances of what kinds of things are constant between seasons, and what kinds of things are not. For the most part, we can call these "talent" (t) and "noise" (n), respectively.

If the r-squared for QBs between seasons is .31, that means

(t/(t+n)) * (t/(t+n)) = .31

Taking the square root of both sides gives

t / (t+n) = .56

And from there, you can multiply both sides by (t+n), and discover that

n = .79 * t

So, for a single NFL season, the variance due to noise is 79% of the variance due to talent.

Now, in the NFL, a quarterback will get maybe 450 passing attempts per season. In the NBA, a full-time player might get three times as many (FG attempts, FT attempts (even if you take those at half weight), and 3P attempts). So, the noise should be only 1/3 as large. Instead of noise being 79% of talent, it will be only maybe 26%. Call the new value of noise n'. Then,

n' = .26 * t

If you sub that back into the first equation, you get

(t/(t+n')) * (t/(t+n')) = .63

See? Just considering opportunities raises the r-squared of .31 all the way up to .63. Berri says it should "exceed 70%", and we probably could get that to happen if we included rebounds, or used a more sophisticated stat than just shooting percentage.

So, if quarterbacks are harder to predict than basketball players, it's simply because they don't play enough for their stats to be as reliable.

(UPDATE: As Alex alludes to in the first comment to this post, my logic assumes that "t" -- the variance of talent -- is roughly the same for QBs and NBA shooters. It might not be. But the point is, assuming they're the same is a reasonable first approximation, and that leads to the conclusion that sample size is the biggest difference.

So, maybe I should have been more conservative and said that it *could be* that they don't play enough for their stats to be as reliable.)

------

Which brings me back to Berri's (and Simmons's) academic study. There, they write,

"[Our] results suggest that NFL scouts are more influenced by what they see when they meet the players at the combine than what the players actually did playing the game of football."


Well, yes -- and perhaps the scouts SHOULD be more influenced by the combine. There's lots of noise in only one season of performance, and a rational scout won't weight it too heavily. What if the scout only saw one play? Then, it's obvious that he should be more influenced by the combine than the results. The less data you have, the more you have to weight the combine results.

Look at it this way. You have two pitchers. One throws 100 mph and had an ERA of 3.50 in 50 innings. The other throws 80 mph and had an ERA of 3.20. Which do you draft? Well, *of course* you draft the 100 mph guy. It's only after 200, 300, 400, 1000 innings that you might have enough evidence to change your mind.

------

The idea of random noise and sample size never figures into this paper at all. I don't think the authors even think about it. When they see unexplained variance, they always argue that it's something like the effect of teammates, instead of looking at binomial randomness. In fact, you get the impression they think there's no randomness at all, and the scouts could be perfect if only they were smarter.

For the record, the paper has no occurrences of the words "luck," "random," "binomial," or "sample size."



Labels: , , , ,

Saturday, March 03, 2012

Are early NFL draft picks no better than late draft picks? Part III

This is about the Berri/Simmons NFL draft paper, in which they say that draft position doesn't matter much for predicting quarterback performance. Here are Parts I and II.

-----

One of the paper's most important claims is that scouts are looking at the wrong things -- specifically, the results of the NFL combine.

At the combine, prospects are tested on a bunch of objective measurements. How fast they run the 40-yard dash. Their BMI (a measure of weight to height ratio). Their intelligence, as measured by the Wonderlic test. And, of course, their height.

But, the authors argue, those things don't matter, and scouts are completely misguided in looking at what happens at the combine. They say that a QB prospect's height, BMI, 40-yard-dash time, and Wonderlic score have almost no effect on performance.

And that one point is key. Because, most of the authors' argument goes (in my words):

Premise 1: Scouts care about combine stats.
Premise 2: Combine stats affect draft position.
Premise 3: Combine stats don't predict performance.

Conclusion: Scouts don't know what they're doing.

So, premise 3 is key. How do the authors prove it?

Here's what they did. They took the 121 QBs drafted from 1999 to 2008, for which they had full combine data. Then they ran a regression to predict the QB's senior year performance based on those factors.

They found no statistical significance for any of them. And they conclude:

"Such results indicate that the combine measures are not able to capture key attributes of the quarterback."


And, in a related footnote,

"Such results indicate that there is little relationship between the combine statistics and per play performance."


Again, I think the problem is that the effect is there, but there just isn't enough data for significance. Indeed, it seems to me that they almost COULDN'T find significance in a study of that type.

Look, how much of QB performance is affected by height? Probably not much, right? There are so many other things involved. I mean, this isn't basketball: you don't see a lot of quarterbacks who are 6-foot-9, which suggests that height can't be that big a deal.

If the effect is that small, how are you going to find statistical significance with only 121 datapoints? Especially when you're trying to predict ONE SINGLE YEAR of college performance, which is very noisy (made even more so because, first, the authors chose to predict a measure that's dependent on playing time)?

You can't, and the authors didn't.

But ... that just means your study isn't precise enough. It doesn't show the effect isn't there. You can't look for a needle in a haystack, from fifty feet away, looking through the wrong end of a pair of binoculars, then say, "we didn't see a needle, so it doesn't exist."

Berri and Simmons didn't even show the results of that regression, even though it's key to their story. They just mention "not significant therefore zero" and move on. But if they HAD given their results, I bet you'd see the standard error is wide enough to encompass not just zero, but also many possible values that are perfectly reasonable and perfectly in line with what scouts think height is worth.

The same thing for the other factors -- 40 yard dash speed, Wonderlic, and BMI. It was almost guaranteed that the regression wouldn't find small effects in that sample.

What about the overall results for the four factors? You might get none of the individual combine stats being significant, but the overall correlation might be. Was it?

We really need to see the estimates for the coefficients. How many of them are reasonable individually? If you add them all up, are they also reasonable? If they are, that's all the more reason to point out that the lack of significance doesn't prove anything.

Again, the authors don't show the results ... but they do give a little hint. They run a second regression, this time using rate statistics instead of playing time stats. In a footnote to that, the authors say,

"The adjusted R-squared from these regressions, though, is in the negative range and the F-statistic is statistically insignificant."


A negative adjusted R-squared ... at first glance, that seems to say no relationship.

Except ... I looked up "adjusted R-squared". And, it turns out, for a regression with 5 variables and 121 rows, you can have a negative adjusted R-squared even if the "real" R-squared is as high as .042. That's not as small as it looks. An r-squared of .042 is an r of around 0.2, which is nothing to sneeze at.

(That makes sense. According to this calculator, a single-variable regression on 121 rows needs to find a correlation of 0.178 to find statistically significance, and I think the "adjusted" is meant to make the 5-variable case comparable to the 1-variable case.)

But 0.2 is probably higher than the effect we're looking for. Or at least, on par with it.

Suppose you ranked all the QBs on their combine stats. And then you took a QB who was +1 in SD in combine stats, and compared him to one that was -1 SD in combine stats. What kind of difference would you expect in on-field performance between the two?

Well, to get a correlation of 0.2, you'd have to expect a difference of about 3 points of NFL Quarterback Rating, or 3 or 4 positions in the performance rankings. (To estimate that, I looked here, and added 3 points to the rating of a typical QB.)

Now, remember, QB performance is very noisy. 3-4 positions in the performance rankings probably means 5-6 positions in *talent* rankings.

That seems to me like it's too much. There's no way height, BMI, Wonderlic, and 40-yard-dash speed could be *that* important, could it, that it's 5 or 6 rankings?

If not, then we're looking for an effect that's too small to find with only 121 datapoints to look at.

So, I think Berri and Simmons' regression was doomed from the start. They were guaranteed not to find significance, even if the scouts were right.




Labels: , , , ,

Wednesday, February 29, 2012

Are early NFL draft picks no better than late draft picks? Part II

This is about the Dave Berri/Rob Simmons paper that concludes that QBs who are high draft choices aren't much better than QBs who are low draft choices. You probably want to read Part I, if you haven't already.

--------

If you want to look for a connection between draft choice and performance, wouldn't you just run a regression to predict performance from draft choice? The Berri/Simmons paper doesn't. The closest they come is their last analysis of many, the one that starts at the end of page 47.

Here's what the authors do. First, they calculate an "expected" draft position, based on a QB's college stats, "combine stats" (height, body mass index, 40 yard dash time, Wonderlic score), and whether he went to a Division I-A school. That's based on another regression earlier in the paper. I'm not sure why they use that estimate -- it seems like it would make more sense to use the real draft position, since they actually have it, instead of their weaker (r-squared = 0.2) estimate.

In any case, the use that expected draft position, and run a regression to predict performance for every NFL season in which a QB had at least 100 plays. They also include terms for experience (a quadratic, to get a curve that rises, then falls).

It turns out, in that regression, that the coefficient for draft position is not statistically significant.

And, so, Berri and Simmons conclude,


"Draft pick is not a significant predictor of NFL performance. ... Quarterbacks taken higher do not appear to perform any better."


I disagree. Two reasons.


1. Significance

As the authors point out, the coefficient for draft position wasn't nearly significant -- it was only 0.52 SD from zero.

But, it's of a reasonable size, and it goes in the right direction. If it turns out to be non-significant, isn't that just that the authors didn't use enough data?

Suppose someone tells me that candy bars cost $1 at the local Kwik-E-Mart. I don't believe him. I hang out at the store for a couple of hours, and, for every sale, I mark down the number of candy bars bought, and the total sale.

I do a regression. The coefficient comes out to $0.856 more per bar, but it's only 0.5 SD.

"Ha!" I tell my friend. "Look, that's not significantly different from zero! Therefore, you're wrong! Candy bars are free!"

That would be silly, wouldn't it? But that's what Berri and Simmons are doing.

Imagine two quarterbacks with five years' NFL experience. One was drafted 50th. The other was drafted 150th. How much different would you expect them to be in QB rating? If you don't know QB rating, thing about it in terms of rankings. How much higher on the list would you expect the 50th choice to be, compared to the 150th? Remember, they both have 5 years' experience and they both had at least 100 plays that year.

Well, the coefficient would say the early one should be 1.9 points better. I calculate that to be about 15 percent of the standard deviation for full-time quarterbacks. It'll move you up in the rankings two or three positions.

Is that about what you thought? It's around what I would have thought. Actually, to be honest, maybe a bit lower. But well within the bounds of conventional wisdom.

So, if you do a study to disprove conventional wisdom, and your point estimate is actually close to conventional wisdom ... how can you say you've disproven it?

That's especially true because the confidence interval is so wide. If we add 2 SD to the point estimate, we find that the effect of draft choice could be as high as -0.091. That means that 100 draft positions is worth 9.1 points. That's a huge difference between quarterbacks. Nine points would move you up at least 6 or 7 positions -- just because you were drafted earlier. It's almost 75 percent of a standard deviation.

Basically, the confidence interval is so wide that it includes any plausible value ... and many implausible values too!

That regression doesn't disprove anything at all. It's a clear case of "absence of evidence is not evidence of absence."


2. Attrition

In Part I, I promised an argument that doesn't require the assumption that QBs who never play are worse than QBs who do. However, we can all agree, can't we, that if a QB plays, but then he doesn't play any more because it's obvious he's not good enough ... in *that* case, we can say he's worse than the others, right? I can't see Berri and Simmons claiming that Ryan Leaf would have been a star if only his coaches gave him more playing time.

If we agree on that, then I can show you that the regression doesn't work -- that the coefficient for draft choice doesn't accurately measure the differences.

Why not? Again, because of attrition. The worse players tend to drop out of the NFL earlier. That means they'll be underweighted in the regression (which has one row for each season). So, if those worse players tend to be later draft choices, as you'd expect, the regression would underestimate how bad those later choices are.

Here, let me give you a simple example.

Suppose you rate QBs from 1 to 5. And suppose the rating also happens to be the number of seasons the QB plays.

Let's say the first round gives you QBs of talent 5, 4, 4, and 3, which is an average of 4. The second round gives you 4, 2, 1 and 1, which averages 2.

Therefore, what we want is for the regression to give us a coefficient of 4 minus 2, which is 2. That would confirm that the first round is 2 better than the second round.

But it won't. Why not? Because of attrition.

-- The first year, everyone's playing. So, those years do in fact give us a difference of 2.

-- The second year, the two "1" guys are gone. The second round survivors are the "4" and "2", so their average is now 3. That means the difference between the rounds has dropped down down to 1.

-- The third year, the "2" guy is gone, leaving the second round with only a "4". Both rounds now average 4, so they look equal!

-- The fourth year, the "3" drops out of the first round pool, so the difference becomes 0.33 in favor of the first round.

-- The fifth year ... there's nothing. Even though the first round still has a guy playing, the second round doesn't, so the results aren't affected.

So, see what happens? Only the first year difference is correct and unbiased. Then, because of attrition, the observed difference starts dropping.

Because of that, if you actually do the regression, you'll find that the coefficient comes up 1.15, instead of 2.00. It's understated by almost half!

This will almost always happen. Try it with some other numbers and assumptions if you like, but I think you'll find that the result will almost never be right. The exact error depends on the distribution and attrition rate.

Want a more extreme case? Suppose the first round is four 4s (average 4), and the second round is a 7 and three 1s (average 2.5). The first round "wins" the first year, but then the "1"s disappear, and the second round starts "winning" by a score of 7-4.

In truth, the first round players are 1.25 better than the second round. But if you do the Berri/Simmons regression, the coefficient comes out negative, saying that the first round is actually 0.861 *worse*!

So, basically, this regression doesn't really measure what we're trying to measure. The number that comes out isn't very meaningful.

------

Choose whichever of these two arguments you like ... or both.

I'll revisit some of the paper's other analyses in a future post, if anyone's still interested.


------

UPDATE: Part III is here.



Labels: , , , ,

Monday, February 27, 2012

Are early NFL draft picks no better than late draft picks? Part I

Dave Berri thinks that NFL teams are inexplicably useless in how they evaluate quarterback draft choices. He believes this is true because of data he presents in a 2009 study, co-written with Rob Simmons, in the Journal of Productivity Analysis. The study is called "Catching the Draft: on the process of selecting quarterbacks in the National Football League amateur draft."

The study was in the news a couple of years ago, gaining a little bit of fame in the mainstream media when bestselling author Malcolm Gladwell debated it with Steven Pinker, the noted author and evolutionary psychologist.

In his book "What the Dog Saw," Gladwell wrote,

"... Berri and Simmons found no connection between where a quarterback was taken in the draft -- that is, how highly he was rated on the basis of his college performance -- and how well he played in the pros."


Pinker, reviewing the Gladwell book in the New York Times, flatly disagreed.

"It is simply not true that a quarter­back’s rank in the draft is uncorrelated with his success in the pros."

Gladwell wrote to Pinker, asking for evidence that would contradict Berri and Simmons' peer-reviewed published study. Pinker referred Gladwell to some internet analyses, one of which was from Steve Sailer. Gladwell was not convinced, but responded mostly with ad hominem attacks and deferrals to the credentialized:

"Sailer, for the uninitiated, is a California blogger with a marketing background who is best known for his belief that black people are intellectually inferior to white people. Sailer’s “proof” of the connection between draft position and performance is, I’m sure Pinker would agree, crude: his key variable is how many times a player has been named to the Pro Bowl. Pinker’s second source was a blog post, based on four years of data, written by someone who runs a pre-employment testing company, who also failed to appreciate—as far as I can tell (the key part of the blog post is only a paragraph long)—the distinction between aggregate and per-play performance. Pinker’s third source was an article in the Columbia Journalism Review, prompted by my essay, that made an argument partly based on a link to a blog called “Niners Nation." I have enormous respect for Professor Pinker, and his description of me as “minor genius” made even my mother blush. But maybe on the question of subjects like quarterbacks, we should agree that our differences owe less to what can be found in the scientific literature than they do to what can be found on Google."

Pinker replied:

"Gladwell is right, of course, to privilege peer-reviewed articles over blogs. But sports is a topic in which any academic must answer to an army of statistics-savvy amateurs, and in this instance, I judged, the bloggers were correct."


And, yes, the bloggers *were* correct. They pointed out a huge, huge problem with the Berri/Simmons study. It ignored QBs who didn't play.

----

As you'd expect, the early draft choices got a lot more playing time than the later ones. Even disregarding seasons where they didn't play at all, and even *games* where they didn't play at all, the late choices were only involved in 1/4 as many plays as the early choices. Berri and Simmons don't think that's a problem. They argue -- as does Gladwell -- that we should just assume the guys who played less, or didn't play at all, are just as good as the guys who did play. We should just disregard the opinions of the coaches, who decided they weren't good enough.

That's silly, isn' t it? I mean, it's not logically impossible, but it defies common sense. At least you should need some evidence for it, instead of just blithely accepting it as a given.

And, in any case, there's an obvious, reasonable alternative model that doesn't force you to second-guess the professionals quite as much. That is: maybe early draft choices aren't taken because they're expected to be *better* superstars, but because they're expected to be *more likely* to be superstars.

Suppose there is a two-round draft, and a bunch of lottery tickets. Half the tickets have a 20% chance of winning $10, and the other half have a 5% chance of winning $10. If the scouts are good at identifying the better tickets, everyone will get a 20% ticket in the first round, and a 5% ticket in the second round.

Obviously, the first round is better than the second round. It has four times as many winners. But, just as obviously, if you look at only the tickets that win, they look equal -- they were worth $10 each.

Similarly for quarterbacks. Suppose, in the first round, you get 5 superstar quarterbacks and 5 good ones. In the last round, you get only one of each. By Berri's logic, the first round is no better than the last round! Because, the 10 guys from the first round had exactly the same aggregate statistics, per play, as the 2 guys from the last round.

I don't see why Gladwell doesn't get it, that the results are tainted by the selective sampling.

Anyway, others have written about this better than I have. Brian Burke, for instance, has a nice summary.

Also,
Google "Berri Gladwell" for more of the debate.

-----

The reason I bring this up now is that, a couple of days ago, Berri reiterated his findings on "Freakonomics":

"We should certainly expect that if [Andrew] Luck and [Robert] Griffin III are taken in the first few picks of the draft, they will get to play more than those taken later. But when we consider per-play performance (or when we control for the added playing time top picks receive), where a quarterback is drafted doesn’t seem to predict future performance."


What he's saying is that Andrew Luck, who is widely considered to be the best QB prospect in the world, is not likely to perform much better than a last-round QB pick, if only you gave that last pick some playing time.

Presumably, Berri would jump at the chance to trade Luck for two last-round picks. That's the logical consequence of what he's arguing.

-----

Anyway, I actually hadn't looked at Berri's paper (.PDF) until a couple of days ago, when that Freakonomics post came out. Now that I've looked at the data, I see there are other arguments to be made. That is: even if, against your better judgment, you accept that the unknowns who never got to play are just as good as the ones who did ... well, even then, Berri and Simmons's data STILL don't show that late picks are as good as early picks.

I'll get into the details next post.

-----

UPDATE: That next post, Part II, is here. Part III is here.



Labels: , , , ,