Thursday, January 20, 2011

Sabermetric basketball statistics are too flawed to work

You know all those player evaluation statistics in basketball, like "Wins Produced," "Player Evaluation Rating," and so forth? I don't think they work. I've been thinking about it, and I don't think I trust any of them enough put much faith in their results.

That's the opposite of how I feel about baseball. For baseball, if the sportswriter consensus is that player A is an excellent offensive player, but it turns out his OPS is a mediocre .700, I'm going to trust OPS. But, for basketball, if the sportswriters say a guy's good, but his "Wins Produced" is just average, I might be inclined to trust the sportswriters.

I don't think the stats work well enough to be useful.

I'm willing to be proven wrong. A lot of basketball analysts, all of whom know a lot more about basketball than I do (and many of whom are a lot smarter than I am), will disagree. I know they'll disagree because they do, in fact, use the stats. So, there are probably arguments I haven't considered. Let me know what those are, and let me know if you think my own logic is flawed.

------

The most obvious problem is rebounds, which I've posted about many times (including these posts over the last couple of weeks). The problem is that a large proportion of rebounds are "taken" from teammates, in the sense that if the player credited with the rebound hadn't got it, another teammate would have.

We don't know the exact numbers, but maybe 70% of defensive and 50% of offensive rebounds are taken from a teammates' total.

More importantly, it's not random, and it's not the same for all players. Some rebounders will cover much more of other players' territory than others. So when player X had a huge rebounding total, we don't know whether he's just good at rebounds, whether he's just taking them from teammates, or whether it's some combination of the two.

So, even if we decide to take 70% of every defensive rebound, and assign it to teammates, we don't know that's the right number for the particular team and rebounder. This would lead to potentially large errors in player evaluations.

The bottom line: we know exactly what a rebound is worth for a team, but we don't know which players are responsible, in what proportion, for the team's overall performance.

------

Now, that's just rebounds. If that were all there were, we could just leave that out of the statistic, and go with what we have. But there's a similar problem with shooting accuracy.

I ran the same test for shooting that I ran for rebounds. For the 2008-09 season, I ran regression for each of the five positions. Each row of the regression was a single team for that year, and I checked how each position's shooting (measured by eFG%) affected the average of the other four positions (the simple average, not weighted by attempts).

It turns out that there is a strong positive correlation in shooting percentage among teammates. If one teammate shoots accurately, the rest of the team gets carried along.

Here are the numbers (updated, see end of post):

PG: slope 0.30, correlation 0.63
SG: slope 0.40, correlation 0.62
SF: slope 0.26, correlation 0.27
PF: slope 0.28, correlation 0.27
-C: slope 0.27, correlation 0.43

To read one line off the chart: for every one percentage point increase in shooting percentage by the SF (say, from 47% to 48%), you saw an increase of 0.26% in each of his teammates (say, from 47% to 47.26%).

The coefficients are a lot more important than they look at first glance, because they represent a change in the average of all four teammates. Suppose all five teammates took the same number of shots (which they don't, but never mind right now). That means that when the SF makes one extra field goal, each teammate also makes an extra 0.26, for a team team total of 1.04 extra field goals.

That's a huge effect.

And, it makes sense, if my logic is right (correct me if I'm wrong). Suppose you have a team where everyone has a talent of .450, but then you get a new guy on the team (player X) with a talent of .550. You're going to want him to shoot more often than the other players. For instance, if X and another guy are equally open for a roughly equal shot, you're going to want to give the ball to X. Even if Y is a little more open than X, you'll figure that X will still outshoot Y -- maybe not .550 to .450, but, in this situation, maybe .500 to .450. So X gets the ball more often.

But, then, the defense will concentrate a little more on X, and a little less on the .450 guys. That means X might see his percentage drop from .550 to .500, say. But the extra attention to X creates more open shots for the .450 guys, and they improve to (say) .480 each.

Most of the new statistics simply treat FG% as if it's solely the achievement of the player taking the shot, when, it seems, it is very significantly influenced by his teammates.

------

Some of that, of course, might be that teams with good players tend to have other good players; that is, it's all correlation, and not causation. But there's evidence that's not the case, as illustrated by a recent debate on the value of Carmelo Anthony.

Last week, Nate Silver showed that if you looked at Carmelo Anthony's teammates' performance, and then looked at that performance when Anthony wasn't on their team, you see a difference of .038 in shooting percentage. That's huge -- about 15 wins a season.

Dave Berri responded with three criticisms. First, that Silver weighted by player instead of by game; second, that Silver hadn't considered the age of the teammates (since very young players improve anyway as they get older); and, third, that if you control for age and a bunch of other things, the results aren't statistically significant from zero. (However, Berri didn't post the full regression results, and did not claim that his estimate was different from .038.)

Finally, over at Basketball Prospectus, Kevin Pelton ran a similar analysis, but within games instead of between seasons (which eliminates the age problem, and a bunch of other possible confounding variables). He found a difference of .028. Not quite as high as Silver, but still pretty impressive. Furthermore, a similar analysis of all of Anthony's career shows similar improvements in team performance, which suggests the effect is real.

To be clear, this kind of analysis is the kind that, I'd argue, works great -- comparing the team's performance with the player and without him. What I think *doesn't* work is just using the raw shooting percentages. Because how do you know what those percentages mean? Suppose one team is all at .460, and another team is all at .490. The .490 means that you have more players on the team above average than below average. But, the above average players are lifting the percentages of the below average players, and the below-average players are reducing the percentages of the above-average players. But which are which? We have no way of telling.

Here's a hockey example. Of Luc Robitaille's eight highest-scoring NHL seasons, six of them came while he was a teammate of Wayne Gretzky. In 1990-91, Robitaille finished with 101 points. How much of the credit for those points do you give to Robitaille, and how much of the credit do you give to Gretzky? There's no way to tell from the single season raw totals, is there? You have to know something about Robitaille, and Gretzky, and the rest of their careers, before you can give a decent estimate. And your estimate will be that Gretzky that should get some of the credit for some of Robitaille's performance.

Similarly, when Carmelo Anthony increases all his teammates' shooting percentages by 30 points, *and it's the teammates that get most of that credit* ... that's a serious problem with the stat, isn't it?

------

So far, we've only found problems with two components of player performance -- rebounds and shooting percentage. However, those are the two biggest factors that go into a player's evaluation. And, additionally, you could argue that the same thing applies to some of the other stats.

For instance, blocked shots: those are primarily a function of opportunity, aren't they? Some players take a lot more shots than others, so the guy who defends against Allen Iverson is going to block a lot more shots than his teammates, all else being equal.

------

Still, it could be possible that the problems aren't that big, and that, while the new statistics aren't perfect, they're still better than existing statistics. That's quite reasonable. However, I think that, given the obvious problems, the burden of proof shifts to those who maintain the stats still work.

The one piece of evidence that I know of, with regard to that issue, is the famous study from David Lewin and Dan Rosenbaum. It's called "The Pot Calling the Kettle Black – Are NBA Statistical Models More Irrational than 'Irrational' Decision Makers?" (I wrote about it here; you can find it online here; and you can read a David Berri critique of it here.)

What Lewin and Rosenbaum did was try to predict how teams would perform last year, based on their previous year's statistics. If the new sabermetric statistics were better evaluators of talent than, say, just points per game, they should predict better.

They didn't. Here are the authors' correlations:

0.823 -- Minutes per game
0.817 -- Points per game
0.820 -- NBA Efficiency
0.805 -- Player Efficiency Rating
0.803 -- Wins Produced
0.829 -- Alternate Win Score

As you can see, "minutes per game" -- which is probably the closest representation you can get to what the coach thinks of a player's skill -- was the second highest of all the measures. And the new stats were nothing special, although "Alternate Win Score" did come out on top. Notably, even "points per game," widely derided by most analysts, finished better than PER and Berri's "Wins Produced."

When this study came out, I thought part of the problem was that the new statistics don't measure defense, but "minutes per game" does, in a roundabout way (good defensive players will be given more minutes by their coach). I still think that. But, now, I think part of the problem is that the new statistics don't properly measure offense, either. They just aren't able to do a good job of judging how much of the team's offensive performance to allocate to the individual players.

Now that I think I understand why Lewin and Rosenbaum got the results they did, I have come to agree with their conclusions. Correct me if I'm wrong, but logic and evidence seem to say that sabermetric basketball statistics simply do not work very well for players.

-----

UPDATE: some commenters in the blogosphere are assuming that I mean that basketball sabermetric research can't work for basketball. That's not what I mean. I'm referring here only to the "formula" type stats.

I think the "plus-minus"-type approaches, like those in the Carmelo Anthony section of the post above, are quite valid, if you have a big enough sample to be meaningful.

But, just picking up a box score or looking up standard player stats online, and trying from that which players are how much better than others (the approach that "Wins Produced" and other stats take) ... well, I don't think you're ever going to be able to make that work.


UPDATE: I found a slight problem with the data: one team was missing and one team I entered twice. I've updated the post. The conclusions don't change.

For the record, the wrong slopes were .30/.39/.31/.25/.24. The corrected slopes, as above, are .30/.40/.26/.28/.27.

The wrong correlations were .59/.58/.37/.26/.40. The corrected correlations are .63/.62/.27/.27/.43
.






Labels: , , , ,

Thursday, January 13, 2011

2010-11 NBA rebounding correlations

My last post showed that, in the NBA, players generally "steal" 2/3 of their rebounds away from their teammates. That is, when a player grabs a rebound, there's almost a 70% chance that someone else on their own team would have got it if they hadn't.

That study was based on the 2008-09 season, and I didn't have the data broken down into offensive rebounds and defensive rebounds. In addition, I used raw numbers of rebounds, which means my analysis would underestimate the amount of stealing (for reasons given in the "gravity" argument in the previous post).

But, now, commenter DSMok1 solved both problems. He's provided data for the current 2010-11 season (up to last night, I assume?), broken down, and his numbers are in rebound percentage, rather than raw numbers.

And the results are even stronger than before.

For defensive rebounds:

-- PG: every extra DREB% reduces his teammates' DREB% by 0.79.
-- SG: every extra DREB% reduces his teammates' DREB% by 0.92.
-- SF: every extra DREB% reduces his teammates' DREB% by 0.52.
-- PF: every extra DREB% reduces his teammates' DREB% by 0.94.
--- C: every extra DREB% reduces his teammates' DREB% by 0.77.

It looks like the "stealing" rate is about 0.8, on average, for defensive rebounds.

Having said that, however, I should say that the standard errors of these estimates are pretty high. Here are those estimates again, along with the SE:

PG -- 0.79 +/- 0.35
SG -- 0.92 +/- 0.21
SF -- 0.52 +/- 0.20
PF -- 0.94 +/- 0.13
C --- 0.77 +/- 0.14

Here's the same chart for offensive rebounds. This time, I'll put the SEs in brackets at the end.

-- PG: every extra OREB% reduces his teammates' OREB% by 1.18 (+/- 0.55).
-- SG: every extra OREB% reduces his teammates' OREB% by 1.06 (+/- 0.46).
-- SF: every extra OREB% reduces his teammates' OREB% by 0.15 (+/- 0.40).
-- PF: every extra OREB% reduces his teammates' OREB% by 0.13 (+/- 0.16).
--- C: every extra OREB% reduces his teammates' OREB% by 0.38 (+/- 0.27).

A lot different. As expert basketball sabermetricians have said, diminishing returns are at a much lower level for offensive rebounds. The PG and SG numbers are probably enhanced by random errors -- it would be hard to come up with an explanation of why a shooting guard's OREB would cost his teammates *more than one* OREB in exchange. [UPDATE: not true! See comments. I now agree a coefficient more extreme than -1 could actually be correct.]

Even though you could argue that the bottom three estimates are not statistically significantly different from zero (they're all less than 2 SDs), I think our best guess is still the point estimates we have here (since there's no prior reason to believe the diminishing returns rate should be zero).

Overall, other analysts have suggested an average of somewhere between .2 and .3, which seems very reasonable looking at the above table.


Finally, overall rebounds:

-- PG: every extra REB% reduces his teammates' REB% by 0.95 (+/- 0.15).
-- SG: every extra REB% reduces his teammates' REB% by 1.05 (+/- 0.21).
-- SF: every extra REB% reduces his teammates' REB% by 0.99 (+/- 0.26).
-- PF: every extra REB% reduces his teammates' REB% by 0.75 (+/- 0.12).
--- C: every extra REB% reduces his teammates' REB% by 0.71 (+/- 0.14).

All five of these seem to be about the average of OREB and DREB, except for the SF. Not sure why the SF is so different.

------

So, I think this roughly confirms what certain basketball analysts are doing: assigning only about 0.3 of a defensive rebound to the player, and about 0.7 of an offensive rebound. [UPDATE: Guy makes an excellent argument for why my estimate of 0.3 could be too low. See comments.]

Thanks again to DSMok1 for the data.

------

P.S. one other finding: overall, there is a negative correlation between a team's *overall* OREB% and its overall DREB%.

For every one percentage point extra on a team's OREB%, its DREB goes down by 0.3 percentage points. It's not statistically significant -- about 1.5 SDs -- but I thought I'd mention it anyway.


Labels: , , ,

Wednesday, January 12, 2011

Do players "steal" rebounding opportunities from teammates?

In my previous post on rebounding, I promised to review some of the evidence supporting the "diminishing returns" hypothesis. That's the theory that, in general, a player's individual rebound totals are not mostly a product of his own ability to snare rebounds, but, rather, predominantly because of his positioning on the court, or the role he is assigned on his team.

That is: if a player gets a lot of rebounds, even adjusted for what position he plays, a large part of his total is rebounds that other players would have gotten to regardless.

Actually, instead of reviewing the evidence, I thought I'd just create my own. However, my numbers here follow on arguments made by others, in other places. (Here, for instance, is a comment by Guy listing four reasons why we should believe that "diminishing returns" is a real phenomenon.)

The easiest test, perhaps, is simply whether, when a player gets more rebounds, it means that his teammates get fewer. If they do, that's evidence that player is simply grabbing balls his teammates would have gotten anyway. If not, that's support for David Berri's hypothesis, that differences in rebounds are the result of the talent of the particular player credited with the rebound.

So, here's what I did. For every team in the 2008-2009 season, I ran a regression comparing the total rebounds by the team's centers (combined) to the total for the rest of the team (combined). I typed in the data manually from some team pages at the 82games.com website, because that's where I happened to find it broken down by position. (Strangely, basketball-reference.com doesn't do that.)

It's only one season's worth of data, but, still, the results are striking.

Every extra rebound by the center results in 0.61 fewer rebounds by teammates.

That means that only 39 percent of center rebounds are "real", in the sense that the other team would have got them if not for the center's efforts. If a center is (say) 20 rebounds above average for the season, then, on average, you should estimate that 12 of those rebounds would have been grabbed by a teammate anyway.


(For those of you who (unlike me) prefer correlation coefficients, the r was -.63, and the r-squared was .395.)

There's a possible alternative explanation for this: it might be that teams just don't want to acquire too much rebounding talent. So, if they get a center who's good at rebounding, they'll get below-average rebounders at the other positions. Still, even that possibility supports the "diminishing returns" hypothesis. Because, why else would teams limit their rebounders? Baseball teams don't limit their home run hitters ... if baseball teams try to cap the amount of rebounding talent on their team, it must be because they realize that the additional rebounders would be somewhat wasted.

And, in any event, I suspect the effect is much, much too strong to be the result of general manager choice. Here are some other ways of looking at the strength of this finding:

1. Let's convert to a won-loss record.

The center is either above or below average for the league. And the rest of the positions, combined, are either above or below average for the league. Suppose you call it a "win" (for the David Berri theory) when both tendencies are the same (both are above or below average). And suppose you call it a "loss" when the tendencies are different (centers above average and the rest of the team below average, or vice versa).

In that case, the Berri hypothesis goes 7-23.

In extreme cases, it does even worse. Breaking the 7-23 record down into how far centers were from the average:

4-9: Teams' centers within 1 Reb/G of average
3-7: Teams' centers between 1 and 2 Reb/G from average
0-7: Teams' centers more than 2 Reb/G from average

This exactly supports the hypothesis. The farther the center is above (below) average, the less likely his teammates are to also be above (below) average.

2. As others have pointed out before, teams are surprisingly close to each other in total rebounds.

So close, in fact, that it leads to this shocking result:

The variance in team rebounds per game is LESS than the variance in rebounds per game by the center alone!

Think about that, how unusual that is. It doesn't work for many other player skill stats in the world of team sports, does it?

Suppose I ask you to predict how many home runs a random full-time player hit last year. You might guess, say, 18. You'd be off by 18 if the guy hits zero, or you may be off by 36 if the guy turns out to be Jose Bautista. But 36 is your worst-case scenario.

Now, suppose I ask you to predict how many home runs a random TEAM will hit next year. You guess 154, which was the average in 2010. Now, you're in worse shape. The Blue Jays hit 257, and the Mariners hit only 101. The worst case is now 103 for a team, almost triple the worst case of 36 for a player. Instead of -18 to 36, the range of error is now -56 to 103. The variance, and the margin of error, is much higher for a team.

But for NBA rebounds, it's the other way around -- it's actually easier to predict rebounds for a random team than for a random team's centers!

-- For centers, the SD was 1.63, and the range from worst to best was 7.4 (10.3 to 17.7).
-- For teams, the SD was only 1.43. And the range was only 4.0 (38.8 to 43.8).

This piece of evidence is so surprising, and so strong, that it should be enough, by itself, to convince you that something's going on. The SD of rebounds by centers is higher than the SD of rebounds by their teams. That's it in one sentence.


3. There are five positions on a basketball team. In terms of rebounds, every one of those five positions had a negative correlation with the rest of their team. That's 5-0 for the diminishing returns hypothesis.

There are also 10 different pairs of two positions. If you run the correlation between rebounds for those 10 pairs, eight of them come out negative. So that's 8-2.

(The two positive ones were (1) shooting guard vs. small forward (r = .312; every additional rebound by the small forward leads to .31 more rebounds by the shooting guard), and (2) shooting guard vs. power forward (r = .024, every additional rebound by the power forward leads to .01 more rebounds by the shooting guard).


4. There was very, very interesting result for point guard vs. the rest of the team.

Every additional rebound taken by the point guard reduces the rest of the team's rebound total by ... almost exactly one rebound: .96, to be exact.

Taken at face value, that means that 24 out of 25 times, when a point guard gets a rebound, it's because the rest of the team deliberately leaves it for him!

DSMok1 posted this comment last week:

"For a significant number of defensive rebounds, there are multiple defensive players present for the rebound (could get the rebound), while the offense has already cleared out to cut off the fast break. These rebounds do not show value or skill to the player who gets them, but are rather a random/confounding variable. For some teams, their center will grab such "garbage" rebounds. For other teams, maybe the PG will grab them himself (I see OKC and Russel Westbrook do this)."


That's consistent with the data. If the PG only gets the rebound when the entire offense has cleared out, it means that there can't be any value added.

I don't know if this is true or not -- I really don't watch a lot of basketball -- but I find it interesting that the regression and DSMok1 are saying exactly the same thing, only several days apart. Admittedly, it could be just random error ... you'd want to check this for other years to make sure.


5. As I said, every position had a negative correlation with the other positions on the team. Here they are. (UPDATE, 1/23: I realized I accidentally left out one team, and entered one team twice. Corrections have been made below.)

-- PG: every extra rebound reduces his teammates' rebounds by 0.96 0.87.
-- SG: every extra rebound reduces his teammates' rebounds by 0.65 0.64.
-- SF: every extra rebound reduces his teammates' rebounds by 0.73 0.73.
-- PF: every extra rebound reduces his teammates' rebounds by 0.63 0.68.
--- C: every extra rebound reduces his teammates' rebounds by 0.65 0.69.

------

So, there you go: about 2/3 of marginal rebounds are taken away from a teammate.

Actually, it's worse than that! I'm not sure how much worse, but it's worse.

Why? Because the number of rebounds is mostly a function of your opponent's missed shots. The more missed shots, the more defensive rebounds are available. So your defensive rebounds depend, in part, on how good your team defense is.

If you have a good defense, and you get more rebounds than expected, you'd expect all players' rebounds should go up above average together. Same, in reverse, if you have a bad defense.

Again, the same is true for offensive rebounds. The worse your shooting, the more defensive rebounds come available, and vice-versa. But you'd expect all five of your players to get more rebounds, to some extent. So, again, rebounding moves together.

And, finally, there's another factor that should cause a certain amount of positive correlation -- pace. Teams that play faster or slower will see their rebound totals rise or fall in unison.

What all that means is that if the "diminishing returns" factor were zero, you'd have a *positive* correlation between teammates' rebounds. The fact that the actual correlation is negative, means that it has to be more negative than it looks -- it has to first overcome the positive correlation caused by team defense and shooting.

Does that make sense?

Look at it this way. The defense, shooting, and pace correlations are like gravity, pushing rebounds towards a positive relationship. For the "stealing rebounds" factor to turn that negative *against gravity* means that the negative is a bigger factor than it looked.

-----

Since I have no idea how to compensate for the "gravity" effect, let's ignore it for now, and stick with the 2/3 estimate. Assume that a typical player rebound takes 2/3 of a rebound away from a teammate, which means that only 1/3 of all individual rebounds are really individual. What does that mean for player evaluation?

The first instinct is to take every rebound, give 1/3 of the credit to the player who grabbed it, and spread the remaining 2/3 among the rest of the players.

Would that work? Well, it's better than giving the entire rebound to one player. But it still has a big problem.

And that problem is: the 2/3 figure is not random, nor is it evenly distributed. It varies by player and team.

There are some centers who deliberately let their teammates have more rebounds. There are others who deliberately take rebounds away from their teammates (not necessarily due to selfishness, by the way -- it could be the coach's strategy). If you just give centers 1/3 of their marginal rebounds, some centers will be regularly overestimated, and some will be regularly underestimated. Because it's not random, it won't even out over a career.

In his "Win Shares" baseball book, Bill James talked about assists by first basemen. When fielding a ground ball, some 1Bs would almost always take it themselves (Steve Garvey), and some would always toss it to the pitcher (Bill Buckner). As a result, Buckner would always wind up with more assists than Garvey. But it doesn't mean he was a better fielder.

Suppose that exactly half of first basemen tossed to the pitcher, and half of them didn't. If you just regress all of them by 50%, you may be closer, but you're still wrong. You want to regress Buckner 100%, and Garvey 0%, in order to get an accurate rating.

The same applies here. When you have two players who are above average in rebounding, how do you know which one is above average because he's really, really good at getting to the ball before the opposition, and which one is above average because he's just taking the easy ones away from his teammate?

One suggestion is, don't even try -- just give all the credit for rebounds to the defense. The problem with that, of course, is that the legitimately great rebounders no longer get credit.

Another suggestion is: use subjective evaluations. What do NBA observers think about who the best rebounders are? Combine that information with the numbers, and try to work out "custom" estimates for each player that still add up to the team's overall performance. The problem with that is that even the "experts" are often wrong about these things, both because they're not perfect, and because they almost certainly let the raw numbers influence their evaluations.

Still, there's probably *some* value there ... everyone was able to tell that Brooks Robinson and Ozzie Smith were great baseball fielders, even without statistics. So it should also be possible to tell, just from observation, who the best rebounders are. But, then, there's also the case of Derek Jeter, who won a lot of Gold Glove awards despite the objective statistics pointing to him as among the worst-fielding shortstops in baseball. So, you have to be careful.

-----

I don't have a good solution here. Obviously, it would be best if there were a way to figure the best rebounders objectively, using the evidence of the statistical record. It would be great if we could just take David Berri's stat, and adjust it a little, and wind up with the right answer.

But I don't see how it's possible, given the limitations of the data, to tell the talented from the "stealers."


In that light, we should maybe just admit our ignorance, for now, and treat rebounding as something that has to be evaluated subjectively. Lump it in with defense, as a team measure, while keeping in mind that individual rebounding skill does exist and needs to be considered for adjusting individual players when appropriate.

It's not a perfect solution, but it's obviously much, much better than giving the entire statistical credit to one player.

---

UPDATE: Eli Witus' excellent 2008 post (and others) point out that diminishing returns are much lower for offensive rebounds than for defensive. Does anyone know of a good source for ORB and DRB breakdowns by team by position, so I can rerun the analysis for them separately?



Labels: , , ,

Wednesday, January 05, 2011

David Berri's FAQ and rebounding

A few years ago, "The Wages of Wins" introduced a basketball rating statistic called "Wins Produced" (WP). Since the book came out, there's been some debate about whether WP has a problem with rebounds when evaluating players. I have argued that it does; two of my posts on it are here and here, but you can probably find others elsewhere.

Dave Berri has recently updated a FAQ that tries to take on us doubters. I'll get to that, but first I guess I should summarize the disagreement, since it's been a few years.

-----

WP values a rebound at +1 point. That's because a rebound takes a possession that would go to the other team, and effectively eliminates it. Since a possession is worth about one point, on average, so is the rebound.

There's no argument about that. The argument is about *who should get credit* for the +1 point. "The Wages of Wins" (or, more specifically, Berri, who is co-author and main blogger and spokesperson for the book), gives the entire +1 to the player who snagged the ball. Others argue that this is wrong.

The opposing argument goes something like this:

When a shot is missed, the ball is likelier to go some places than others. Whoever is at that position is more likely to be in a position to pick up the rebound. When you award the entire value of the rebound to that player, you are mostly rewarding him for being in that spot. As Guy pointed out in a comment way back, it's like putouts in baseball. Getting an out is quite valuable, and many putouts are made at first base. But that doesn't mean that your 1B is five times more valuable than your CF just because he makes five times the putouts. His high total is because of where he plays.

It's obvious that this is also true in basketball. Here is one sample of overall rebounding percentage based on position played:

15.0% Center
13.8% Power Forward
8.9% SF/SG
5.9% Point Guard

Obviously, it's not the case that centers are 250% as good at rebounding as point guards -- it's just that because of the way offenses and defenses work, they happen to be in position for a rebound much more often. That's why Berri adjusts WP scores for position. Otherwise, the numbers wouldn't make sense, and point guards as a group would look like they're horrible basketball players.

A slight variation on this is "diminishing returns". This is an argument that, when you have one player who snags a lot of rebounds, it's not just that he's good at rebounds -- it's that he's being given more opportunities that would otherwise go to his teammates. Perhaps he's also going into other players' "territory" to get them. Or, perhaps, the team has assigned him the role of primary defensive rebounder, reducing other players' rebounding responsibilities to allow them to better transition to offense.

If that's the case, a player shouldn't necessarily get credit for every extra rebound, because, if he didn't get it, one of his teammates would have. That is, there are diminishing opportunities available to the other four players on the team.

So, while there's no question that a rebound is worth +1 to the team, it certainly doesn't seem that the full value should be credited to the skill of the individual player.

Berri, however, is not as convinced. He acknowledges that there is *some* diminishing returns happening, but he still gives the entire +1 to the player who picked up the rebound.

------

OK, now to Berri's FAQ. He makes four separate arguments for why his WP stat doesn't overvalue or misappropriate credit for rebounds. All four of those arguments, I think, are easily rebutted.

I'll go through them one at a time, using Berri's own numbering and titles. Keep in mind that I am *not* trying to provide evidence here for the other side of the debate -- I'm just trying to show why Berri's arguments do not prove his position.


Response #1 -- The Consistency of Rebounds

Rebounds per minute, for individual players, are consistent from year to year. Berri reports a correlation coefficient of over .9, and 0.83 even after adjusting for position played. This is higher than similar correlations in other sports. For instance (examples are Berri's):

0.65 -- Baseball OPS
0.47 -- Baseball batting average
0.37 -- Baseball ERA
0.36 -- NFL rushing yards per attempt (for running backs)
0.24 -- NHL goalie save percentage
0.07 -- NFL QB interceptions per attempt

First, you can't really compare the numbers that way. The actual correlations depend on all kinds of things other than skill -- mostly, the number of opportunities and the variance of the circumstances in which those opportunities happen. The fact that one correlation coefficient is higher than another doesn't necessarily mean that the underlying cause is more consistent.

But, suppose we let that go, and assume, with Berri, that rebounding is more consistent than (say) batting average.

So what? Even if you show that rebounding is consistent, that doesn't prove that rebounding is a skill. To go back to Guy's analogy, the consistency of putouts in baseball would be just as high: Albert Pujols had a lot of putouts in 2009 and 2010, and Alex Rodriguez had a lot fewer putouts in both 2009 and 2010. That doesn't prove that Pujols is a much better "putouter" than A-Rod ... it just proves that Pujols plays first base and Rodriguez plays third base.

To that, you could argue that the analogy isn't perfect, because Berri did adjust by position. Still, there are other reasons you could get a high baseball correlation, other than skill. Maybe some 3B play for teams with pitchers who give up a lot of ground balls, and others don't. That would create a higher correlation, while having nothing to do with skill. Maybe some LF play for teams with lots of RH pitching, so they get fewer fly balls hit to them. And so on.

Or, consider saves. There is a very high correlation between saves one year and saves the next year. That doesn't mean that David Aardsma has a talent for saves, but Felix Hernandez doesn't. It just means that, even though they play the same position on paper -- pitcher -- they are used in very different ways. In this case, the consistency isn't of talent, but of managerial decision-making. (I previously expanded on this thought here.)

So, when Berri says,

"When we look at rebounds, we see a higher correlation than all of these [other sports'] statistics. This leads one to conclude that rebounding is a skill that is primarily about the player credited with the rebound."


... it's obvious that doesn't follow. A high degree of consistency in rebounding rate could mean a consistency of talent, or it could mean a consistency of covering more of the other players' territory.

Consistency just means you're measuring something real. It doesn't mean that the "something real" is necessarily talent.


Response #2 -- Rebounds Are Not the Same For All Teams

Berri writes,

"If a player's rebounds are all "stolen" from his teammates, then teams would have to be getting the same number of rebounds. So do all teams end up with the same number of rebounds?"


As written, this is an egregious straw man. Nobody is saying that rebounds are *all* "stolen" from teammates -- just enough to make the raw statistic unreliable. And nobody is saying teams are exactly the same -- we're saying that teams show more similarity than you'd expect by just adding up the individuals. But I'll assume that Berri knows that, and is just exaggerating for effect.

To show how rebounds differ highly across teams, Berri goes on to compare various statistics by "coefficient of variation" (the SD divided by the mean). Again, as I have written before, that number is not meaningful in the way Berri thinks it is.

For offensive rebounding percentage, Berri gets a figure of .106, which is probably something like .027/.265. The .027 is the SD of OR%, and the .265 is the overall average.

But, what if you changed "offensive rebounding percentage" to "offensive rebounding missed percentage"? That is, suppose you start counting missed rebounds instead of made rebounds. In that case, the SD stays the same, but the mean reverses, from .265 to .735 (26.5% made is 73.5% missed). Now, you now get a "coefficient of variation" of .027./.735, which is .036. That now almost exactly matches the other stats Berri cites (which range from .035 to .043). Still, that doesn't matter, because, as just a raw number, "coefficient of variation" has little do to with the subject at hand.

Intuitively, it may *look* like it does, at least to Berri. But it doesn't.

More generally, I don't understand Berri's argument that the more variation there is among teams, the more skill there is in the statistic. There's a lot more variation in sacrifice bunts than there is in batting average, isn't there? But bunting numbers vary mostly because of managerial decisions, not because of talent. The same is true for intentional walks by pitchers. And, to a lesser extent, it's also true for stolen bases.


Response #3 -- Do We Overvalue Rebounds?

Berri makes an argument that goes something like this: suppose rebounds were overvalued, the way his opponents think they are. Then, if we credit a player for only half the rebounds he makes (and spread the other half around to his teammates), that should change things a lot. But, when you look at the top 20 players in the league, the ranking doesn't change that much. (Chart in FAQ, or alone here.) And the new and old statistic correlate with each other at 0.95.

To which the response is:

First: It DOES make a significant difference in the rankings. Some of those top-20 players drop significantly. Carlos Boozer, for instance, goes from 16.2 wins to 12.5 wins. More importantly, it's the evaluations of Boozer's teammates that would change a lot. Since the Jazz player's stats still have to sum to Utah's total wins, Boozer's teammates will get quite a boost. The standings of the top 20 players may not change a whole lot, but, in the middle, where players are very close together, there will be a wholesale re-evaluation, with non-rebounders moving up and rebounders moving down.

Second: Of the top 20 players last year, 19 of them drop in total wins produced when you credit them only half their rebounds (the 20th one stays the same). That means that every one of the top 20 players was at or above his team's average in rebounds (otherwise, replacing half his rebounds with half his teammates' rebounds would make him look better). It looks like the average drop among the top 20 players is a win or two.

That means that it makes a big difference to whether you get it right. If you're an NBA general manager, whether a player is worth 12 wins or 14 is very significant at contract negotiation time.

Third: A correlation coefficient of .95 does not imply that there's not much difference. It's true that .95 seems like a "big number," but you have to evaluate it in context. I feel pretty certain that I could take the established, proven values for baseball events, screw them all up to make them significantly wrong, and still come up with a .95 correlation to the original. I mean, think about it: any not-too-far wrong stat will put Babe Ruth at the top and Mario Mendoza at the bottom. In that light, mismeasuring some of the components will still leave the correlation pretty high.


Response #4 -- WP isn't just about rebounds

This argument of Berri's says that rebounds aren't such a big deal in the entire context of the WP calculation. They're just one small part. Even if rebounds *were* misallocated, it doesn't matter all that much in context, not nearly enough to invalidate WP.

What's the evidence? Well, Berri shows how much a 1 percent change in various statistics changes the final value of WP:

+5.2% -- points per FG attempt
+3.2% -- rebounds
+1.2% -- free throw percentage
-1.1% -- personal fouls
-0.9% -- turnovers
+0.7% -- steals
+0.2% -- blocked shots


Berri concluded,

"Rebounding certainly matters. ... But WP is more "responsive" to shooting efficiency from the field."


Yes, except: a 1% change in Points Per FG Attempt is much less common in the NBA than a 1% change in Rebounding.

For an analogy, consider baseball. The average player might hit .260 with 12 home runs. Now, a 100% change will increase home runs from 12 to 24 -- a significant increase, but not out of this world. On the other hand, a 100% change in batting average will have the player go from .260 to .520 -- which is pretty much impossible.

So the extent to which a statistic is influenced by one of its components is the product of two factors: "elasticity" (responsiveness to change), as Berri calculated, and the extent to which players actually differ in real life (that is, the variance). Berri has only considered the first.

What he could have done, instead, is something that's commonly done in other studies: show the response, not to a 1% change in the value, but to a 1 *standard deviation* change in the value. If Berri had done that, he would have noticed that the SD of rebounds is (I think) approximately 45% of average, while the SD of shooting percentage is only about 11% of average.

So, a 1 SD change in shooting percentage increases value by 5.2 times 11 -- 57.2%. And a 1 SD change in rebounds increases value by 3.2 times 45 -- 144%. So rebounds are indeed much more influential than shooting.

Now, in fairness to Berri, the real-life results won't be that extreme. Berri adjusted all players' stats by position, and, as we saw above, some positions rebound a lot more than others. The adjustment, therefore, will pull the SD of rebounding down. (Having said that, field goal percentage was also adjusted by position, and some positions probably shoot better than others too, so the SD of shooting percentage will drop too. But probably not as much.)

But my point is not to come up with a definitive answer to the question -- it's to argue that Berri's elasticity calculation doesn't mean what Berri thinks it does.

The strange thing is, that, for this particular narrow question, it would actually make sense to compare correlation coefficients. You could look at the r (or even r-squared) for player rebounds vs. WP, and compare it to the one for player Points Per FG Attempt vs. WP. That would give you an intuitive idea of which season stat affects WP the most. But, in this case, Berri chose not to run a regression.

(And, while I know I promised not to argue for the facts either way, one note. Commenter Guy, in an e-mail, told me that last year, WP had a .75 correlation with rebounds, but only a .5 correlation with shooting percentage.)

------

So, that's why I think that Berri's four counterarguments are not relevant to the question of whether rebounds are misallocated to players. As for actual evidence and argument one way or another, there have been some posts lately at various basketball sabermetrics sites, that perhaps I will comment on in future. Here, for instance, is one of them -- both Guy and Berri make appearances.



Labels: , , ,

Friday, June 04, 2010

Payroll and wins and correlation

Yesterday, Stacey Brook posted about payroll and wins. Brook ran a regression and found that, as of June 3, about a third of the way through the 2010 baseball season, the correlation between team salaries and team wins was .224. He writes,

" ...that the two variables move together just over 22%; or there is about 78% not moving together. That does not seem to me a great deal of support for the hypothesis that as team's spend more on payroll, it results in higher team winning percent (or better quality teams)."


What he's saying, paraphrased, is: ".22 is a low number. Therefore, the relationship between payroll and wins is low. QED."

But that's simplistic and wrong. You know what's an even lower number? .00142. Much lower, right? More than 100 times lower. Really small number. Turns out that's Yao Ming's height, in miles. Boy, Yao must be a pretty short little man!

On the other hand, here's a big number: 2,290,000,000. That's Yao's height in nanometers. Huge number! Yao must be really tall!

Well, which is it? Is Yao short or tall?

Obviously, the number alone isn't enough -- you need the units. Brook is simply saying "0.22 is low" without figuring the units, and that's where the problem is. (If anyone reading this thinks otherwise, I invite you to offer me 0.22 Megadollars for my car.)

What are the units of the correlation coefficient of .22? Well, Brook is right when he says it measures how the two variables move together. It means that for every 1 SD that salary moves, you'll see winning percentage moving by 0.22 SDs. Just like "one knot" is a rate of distance divided by time, the correlation coefficient 0.22 is a measure of the SD of winning percentage divided by the SD of money. So we should be able to convert that to wins per dollar. There are 5280 feet in a mile. How many wins per dollar are there in one "correlation coefficient"?

One SD of salary this year is about $38 million, or about $12 million in the 55 games so far. One SD of winning percentage so far this season is .095, or about 5 wins in the 55 games so far.

So one correlation coefficient = (5 wins /12 million dollars) = .42 wins / million dollars = 4.2 * 10^-7 win/dollar.

So 0.22 correlation coefficients equals 0.22 times that, which is 9.2 * 10^-8 win/dollar.

THAT is the number that Brook should be checking to see if it's big or small. Which is it? Well, if you take the inverse, it's about $11 million dollars per win. $11 million is the number Brook should be looking at, not .22.

At the margin, an extra $11 million in spending buys you an extra win. That's the number that the regression is telling us. That's exactly what the 0.22 means, when you figure out what units it's denominated in.

----

Over at "The Book" blog, Tom Tango criticizes Brook on the same grounds: that the correlation coefficient can be made as high or as low as you like just by using a larger or smaller sample. Again, it's a matter of units. If you use (say) only ten games, you get a very small correlation number, but a large unit of variance. If you use many seasons' worth of games, you get a higher correlation coefficient, but a small unit of variance. It's .001 kilometers vs. .999 meters. The numbers are extreme, but the units make up for it, and the end result is almost exactly the same.

It makes sense that the result should be the same -- after all, if one thing causes another thing at a certain rate, it should cause it at a certain rate no matter the size of the sample. If smoking causes cancer, smoking causes cancer. If payroll causes wins over 10 seasons, then payroll causes wins over 55 games. It doesn't matter that the correlation coefficient over 10 seasons is big, and the correlation coefficient over 55 games is small. They are not comparable without computing the units. Once you put in the units, they'll tell you exactly the same story, subject, of course, to the fact that the 55-game sample will have more random variation.

----

Another problem is that Brook dismisses the idea that money has been buying wins because the results of his regression are statistically insignificant:

"In other words in statistical terms payroll has zero effect on winning percentage at this point in the season."


That's just not right, for two reasons.

First: Suppose I claim that I can have an ability called "sensory perception," which other people call "eyesight". You toss a coin, and I will be able to tell you whether it landed heads or tails -- just by looking at it! You don't believe me. So you toss a coin. I look at it and tell you, "heads." You do it again, and I look and say "tails". Then you do it a third time, and after looking I say "tails" again.

I've called it correctly three times in a row. You run a statistical test on it, and find that the chance of me getting three in a row is 0.125 -- a lot higher than the threshold of 0.05 that you need for statistical significance.

And so you say, "in statistical terms you looking at the coin has no effect on whether you are able to guess it right."

Well, that's not a fair argument about me being wrong about having eyesight. Because, after all, I did exactly what I said I could do. If that's not enough evidence for you, that's a fault of your own experiment. You could have made me call ten tosses, or 20 tosses, or 100 tosses, and then you certainly would have had enough evidence! The fact that three tosses isn't enough to convince you is an issue with your experiment, not with real life.

It's true that your weak experiment doesn't show statistical significance for my ability to call coins. But it doesn't show statistical significance against my hypothesis of *always* being able to call coins. So the results are as consistent with my hypothesis as they are with your hypothesis -- even more so, in fact. So why are you rejecting my hypothesis but not yours?

If the results of your experiment are consistent with both your hypothesis (money doesn't buy wins) and your critics' hypothesis (money *does* by wins, at the rate of several million dollars each), you haven't proven anything.

Second, Brook contradicts himself. First, he claims that " in statistical terms payroll has zero effect on winning percentage." But then, he claims that's false:

"Over time (in other words adding more seasons) we do find a statistically significant relationship ... "


So what's the point of trumpeting the new experiment? It doesn't contradict what we already know -- it actually confirms it. If a big experiment finds a statistically significant relationship between salary and wins, and a small experiment finds approximately the same relationship, but without enough data to be statistically significantly different from zero ... then why would you argue that the small one contradicts the big one? It doesn't -- it's exactly what you would expect as confirmation!

In fairness, I think what Brook is doing is again just looking at a single number and giving a gut reaction. For this regression, he looks at the significance level, sees it's not significant, and realizes that, if not for the other study, he would be allowed to conclude that the relationship between salary and wins was zero. If he can "almost" conclude that the relationship is zero, then at the very least it must be small, right?

Well, no, not right. There's a big difference between the size of the effect, and the size of the evidence. Suppose I claim there's an elephant in the room, and show you a picture, but you choose to dismiss my claim on grounds of insufficient evidence. That doesn't give you the right to conclude, "therefore, if there IS an elephant in the room, he's probably a very small one."


Labels: , , , ,

Sunday, March 28, 2010

"Stumbling on Wins:" is there really little difference between goalies?

My copy of "Stumbling on Wins," the new book by David Berri and Martin Schmidt, arrived on Friday. It's a quicker read than their first book, "The Wages of Wins"; for one thing, it's shorter, at 140 pages (before appendices and endnotes). For another thing, the writing style is a bit breezier and less technical, more suited to the non-academic (but serious) sports fan.

The theme of this book is how decision-makers in sports make bad decisions because they don't know how to properly evaluate the information they have. Irrationality in decision making is a subject that's been popularized quite a bit lately. In the last few years, you've got "Predictably Irrational" by Dan Ariely, "Nudge" by Richard Thaler and Cass Sunstein, "Sway" by Ori and Rom Brafman, "Priceless" by William Poundstone, and others. The authors of this book acknowledge the trend, and that they chose their title in tribute to Daniel Gilbert's "Stumbling on Happiness."

I disagree with many (but not all) of the conclusions the authors reach ... it seems like, too often, the authors will do a quick study, look at the results superficially, jump to conclusions that I don't think are justified, and argue from those conclusions that decision-makers are doing it wrong.

For now, I'll just give you one example. In Chapter 3, they argue that NHL goalies are overpaid. Why? Because

"... there simply is little difference in the performance of most NHL goalies."


Why evidence to they give for this?

First, they ran a correlation between a goalie's save percentage (SV%) in consecutive seasons. They got an r-squared of .06, or 6%. That's a small number. So goalies are inconsistent, and what is being observed is not really the goalie's talent.

That's not correct at all.

As I wrote before, and Tango has repeatedly said on his own blog, you can't just observe that because the r-squared is a small number, that the relationship between two variables is weak. Indeed, the same relationship can give you very different r-squareds, depending on other factors in your data, the most obvious of which, here, is sample size.

The r-squared, by definition, is the variance of talent as a percentage of total variance. But the smaller your sample, the more total variance you have just because of luck. And so, the smaller the sample, the lower the r-squared, regardless of whether the talent is low or high.

A low r-squared might mean a small needle -- or it might mean a large haystack.

Unless you take a few seconds to figure out which it is, your r-squared doesn't tell you much of anything about the relationship between the two variables.

What *does* that .06 mean? Well, if the r-squared is .06, then the r is about .25. Roughly speaking, that means you can expect 25% of a goalie's difference from the mean to be repeated next year. Put another way, you have to regress the goalie 75% towards the mean.

Yes, that's not as much as you'd expect. By that calculation, if the average save percentage is .904, and goalie X comes in one season at .924, you'd expect next year he'd be at .909 -- one quarter of the distance between .904 and .924. That's still something: it's .005 above average, which is one goal every 200 shots, or about 10 goals a season.

What do you think -- the idea that a .924 goalie is really .909, does that mean "there's little difference between goalies?" That's more a matter of opinion ... but at least now you have the numbers you need to get a grip on what's going on. The "r-squared equals only .06" doesn't really help you decide.

----

Anyway, that's one problem, that the .06 isn't as small as it looks. A bigger problem is that I don't think the .06 is accurate.

I repeated the same correlation for two sets of two consecutive years, 2005-06 to 2006-07, and 2007-08 to 2008-09. I looked at only the 20 goalies with the most minutes played. I got r-squareds of .30 and .25, respectively, both much higher than the authors' .06.

Why? I think it's because the authors included goalies with many fewer shots against. They don't say exactly what their criteria were, except that they "adjusted for time on the ice" (whatever that means: SV% doesn't depend on time played). In other studies in the same chapter, they used 1000 minutes as a criterion, so maybe that's what they did here.

Now, to simplify, suppose the variance of SV% consists of only talent and luck. A full-time goalie plays about 3,500 minutes. In my regression, it turns out that you get 1 part talent to three parts luck (that's where the .25 comes from: 25% of the total is talent). Now, suppose Berri and Schmidt's average goalie played only half that, or 1,750 minutes. Then the luck variance would be twice as high, and they'd get one part talent to *six* parts luck. That would drop the r-squared down from .25 to .14.

I don't know how the authors got .06 when my analysis shows .14 ... maybe their cutoff was lower than 1,000 minutes. Maybe there's some selection bias in my sample of top goalies only. Maybe my four seasons just happened to be not quite representative. Regardless, the fact that the r-squared varies so much with your selection criterion shows that you can't take it at face value without doing a bit of work to interpret it.

In any case, going back to my r-squared of .25 ... the square root of .25 is .50. That means that exactly half a full-time goalie's observed difference from the mean is real, and will be repeated next season; if a goalie is .020 better than average this year, expect him to be .010 better than average next year. That's pretty reasonable. In that light, I don't think you can say "there's little difference between goalies" at all.

----

And, in fact, we should be able to figure out the spread in goalie talent directly, by a method I learned from Tango a few years ago.

Suppose a goalie faces 1,700 shots, and is expected to save 90% of them. By random chance, he'll sometimes save more than 90%, and sometimes less. By the binomial approximation to the normal distribution, the standard deviation of his save percentage due to luck will be .0073.

Now, for the five seasons I checked, the top 20 goalies that year had an actual SD between .007 and .013 ... let's call it about .011.

That's higher than .0073, as you'd expect. The .0073 is what you'd get if all goalies were identical. But there's also extra variance from the fact that some goalies are better than others. Since

(Observed SD)^2 = (Non-luck SD)^2 + (Luck SD)^2

we can say

.011 ^2 = (Non-luck SD) ^2 + .0073 ^2

So the non-luck SD should be about .0082. If we consider everything that's not binomial luck to be talent, then we can say that the SD of top-20 goalie talent is .008. (I dropped the last decimal because our numbers are very rough here.)

If everything that's "non-luck" should repeat next year, we should get an r-squared of about (.008/.011)^2, which is .53. I only got .25 or .30. Why? Well, there could be more luck involved than just binomial. Not all shots are created equal; maybe some goalies got easier shots, and some harder (search for "Shot Quality" here). Maybe there's some variation in talent because of injury or age. There's definitely the quality of the goalie's defense, and that varies a bit from year to year.

Still, there's quite a bit of evidence of talent here. The theoretical value for r-squared was .53, which means the theoretical value for r is .73. That means that if a goaltender was absolutely perfectly consistent, and every shot gave him the same chance of stopping it, each and every year ... then, 73% of his observed talent would be real. That's what it means to be absolutely consistent.

I didn't find .73, but I found about .50. That's a pretty good proportion of the theoretical maximum. I think we can say that a good part of what we see of a goalie's performance is real.

But, does all this mean that "there's little difference between goalies?" Well, let's check. We got an r-squared of .25, which means that 25% of the variance is talent. The variance observed is .011^2, so the variance due to talent is a quarter of that, which is .0055^2.

That means that a goalie who's one SD above average will have a save percentage .0055 better than average. A goalie who's two SDs above average will be .011 better than the mean.

In the context of 1700 shots, one SD is about 9 goals. Two SDs is about 18 goals. And that's from only the 20 goalies with the most playing time. You'd imagine that if you included backup goalies, the variance would be larger. But, to be conservative, I'll leave the SD at 9 goals for now.

Berri and Schmidt looked at Martin Brodeur's career and found he saved an average of 13.6 goals per year, compared to an average goalie. That's consistent with a 9 goal SD; it implies that Brodeur is about one and a half SDs above average, which seems very reasonable. The authors also point out that, in terms of wins, an advantage of 13.6 goals a year is very small compared to what an NBA superstar can provide. That's true, but it doesn't mean that goalies don't matter in the context of hockey. To address that point, you need to look at the 9 goal SD. Is that a lot?

Well ... I'm not sure. I think it's more than it looks. Let's compare goalies to skaters.

Looking at the plus-minus statistics from 2008-09, a bunch of Bruins come up near the top, with numbers scattered around +30. That means that, when those players were on the ice in non-power-play situations, the Bruins scored 30 more goals than they gave up. Along with Detroit, that seems to be the highest bunch in the league.

Since five players are on the ice, you could give each of them credit for 6 extra goals. But they're not all equal -- some are better than others. Let's say that instead of 6/6/6/6/6, they might be 10/8/6/4/2.

That means that the best player on the Bruins might be worth 10 goals. Regressing that to the mean, let's call it 8 goals. Adding power plays, which weren't included in plus/minus, let's move it back to 10 goals.

That's the best player on the best team. But maybe the best player in the league wasn't on the Bruins -- he might have been on a mediocre team, and his teammates caused his plus/minus to drop. How do we adjust for that? I don't know, but let's bump it up 4 goals, and estimate that the best player in the NHL was worth 14 goals last year.

Now, figure the best goalie is about 2 SD above average, for 18 goals. So, the best goalie in the league is better than the best skater! That doesn't suggest, at all, that there's little difference between goalies.

Except ... last year's top plus/minus figure of +37 (David Krejci) is low by historical standards. In 1981-82, the top five players had plus-minuses above 66, almost twice what the Bruins had last year (although in a higher-scoring offensive environment). And, in 1970-71, Bobby Orr had a plus-minus of +124. Back then, you could certainly argue that goalies were more homogeneous than skaters, and the best skater (Gretzky, Orr, or Lemieux) was easily better than the best goalie. And I think that coincides with the intuition that people had back then, that a good goalie could help, but would never be a factor like a Gretzky would.

Still, maybe we should bump the 14 goal estimate for the best skater up a little bit, closer to the 18 goals we found for the best goalie.

I may be wrong in my logic somewhere, but, if I've done everything right, it seems that top goalies in this era are very similar in importance to top skaters. So when Berri and Schmidt accuse GMs of signing goalies to big contracts because "the people that write the checks" don't "understand [the] story" that goalies don't matter much ... well, I think they underestimate the capabilities of those hockey executives. Their judgment might not be perfect, but I think they understand the variation of talent at least as well as Berri and Schmidt seem to.

-----

So I think Berri and Schmidt got into trouble by just looking at the number .06 without thinking about what it meant. They do this again, a bit later, when they run a correlation between SV% in the regular season, and SV% in the playoffs. That's just doomed to fail, because the playoff sample is so small. That makes the variance due to luck very large, which, in turn, brings the r-squared very close to zero.

Actually, they find an r-squared of .07, which is actually larger than the .06 they found over two consecutive regular seasons. You'd think it would be smaller, since playoff samples are so much smaller. I wonder if the .06 was maybe they used very small samples over the regular season, including goalies with only a couple of games played?

Anyway, after that, they try the correlation between two consecutive playoff appearances. They found "none" of the performance was predictable, which suggests an r-squared of .00 (or maybe they assume it's .00 because it wasn't statistically significant). But that's probably just a sample size issue. If their intention was to show that playoff performance by goalies has a lot of random luck in it, well, yes, of course it does. But if their intent is to conclude that goalie performance is completely unpredictable, that one r-squared isn't enough evidence of that. And I'd bet that if they looked a little closer, they'd find that goalies perform in the playoffs exactly as you'd expect them to, subject to a substantial amount of binomial random luck. Or maybe not -- maybe playoff hockey is so different from regular season that different goalies excel at it. But if you want to check that, you have to do more than just run a single regression and look at a single r-squared.

----

Finally, another non sequitur arises where they write,

"Looking at ... goalies ... one sees an average save percentage of [.895]. The standard deviation of that percentage, though, is only .018. Hence the coefficient of variation of save percentage [the SD divided by the mean] is only 0.02. Hence, there simply is very little difference in the performance of most NHL goalies."


Now, I don't get this at all. How does the coefficient of variation tell you whether or not there's a qualitative difference in performance? It just doesn't. The fact that the SD is a small fraction of the mean doesn't have anything to do with how important the statistic is.

Inutitively, I can see how you might jump to that conclusion, if you don't think about it much. But if you do, it makes no sense. The proportion doesn't matter. When it comes to goals, it's the absolute number that matters. If you let in 10 more goals than average over a season, you cost your team 10 goals. It doesn't matter if you and the other goalies get 100 shots, 1000 shots, 10,000 shots, or 100,000 shots -- ten goals in a season is ten goals in a season.

Another way to look at it is that the SV% statistic is arbitrary, which means the coefficient of variation is arbitary. Suppose the NHL had decided to use "goal percentage" instead of "save percentage", counting up the percentage of shots that went in, instead of the percentage that did not. In that case, the SD would be exactly the same, .018. But the average is now the opposite of what it was -- if 89.5% of shots are stopped, then 10.5% of shots are NOT stopped. And so now your coefficient of variation is .17.

One way, you get .02. Another way, you get .17. So how can the size of the arbitrary coefficient of variation possibly have anything to do with how important goaltending is?

I'm sure the coefficient of variation has its uses, but this isn't one of them.

-----

In summary: as I read it, Berri and Schmidt's argument goes something like this:


-- The r-squared of SV% in consecutive seasons is .06.
-- The r-squared of SV% between a season and the playoffs is .07.
-- The r-squared of SV% between two consecutive playoffs is .00.
-- The coefficient of variation for SV% is .02.

--> These are all small numbers. Therefore, goalies' performances aren't consistent. That means there's not much difference between them, and GMs don't seem to realize this.


As I wrote, I don't think that logic makes sense. I think the evidence shows that, in the current era, good goalies are about as valuable as good skaters. I haven't looked, but I bet that salary data would show that to be roughly consistent with what GMs think.



Labels: , , , ,

Friday, May 08, 2009

The regression equation versus r-squared

OK, I hope I'm not beating a dead horse here, but here's another way to think of the difference between r-squared and the regression equation.

The r-squared comes from the standpoint of stepping back and looking at the distribution of wins among teams in your dataset. Some teams have over 60 wins, some teams have under 20 wins, and some teams are in the middle. If you look at the standings, and ask yourself, "how important are differences in salary to how we got this way?", then you're asking about r-squared.

The regression equation matters more if you're interested in the future, if you care about how much you can influence wins by increasing payroll. If you ask yourself, "how much do I have to spend to get a few extra wins?", then you want the regression equation.

The r-squared looks at the past, and asks, "was salary important to how we got to this variance in wins?". The regression equation looks to the future, and says, "can we use salary to influence wins?"

It's very possible, and very easy, to have two different answers to these two questions. Here's an example.

Suppose you're trying to see what activities 25-year-olds partake in that affect their life expectancy. You might discover that the average 25-year-old lives to 80, but you want to try to figure out what factors influence that. You run a multiple regression, and you figure out that if the person smokes at 25, it appears to cut five years off his life expectancy. If he eats healthy, it adds four years. If he commits suicide at 25, it cuts off 55 years (since he dies at 25 instead of 80).

Your regression equation would look something like:

life expectancy = 80 - (5 * smoker) + (4 * eats healthy) - (55 * commits suicide).

We should all agree that committing suicide has a big effect on life expectancy, right?

Now, let's look at the r-squared. To do that, look at all the 25-year-olds in the sample (which might be several thousand). You'll see a few that live to 25, some that live to 45, a bunch that live to 65, a larger bunch that live to 80, and some that live to 100. The distribution is probably bell-shaped.

For the r-squared, ask yourself: how much did suicide contribute to the curve looking like this? The answer: very little. There are probably very few suicides at 25, and even if you adjusted for those, by taking those points out of the left side of the curve and moving them to the peak, the curve would still look roughly the same. Suicide is not a very big factor in making the curve look like it does.

And so, you get a very low r-squared for suicide. Maybe it would be .01, or even less.

See the apparent contradiction?

-- suicide has a HUGE effect on lifespan.
-- r-squared for suicide vs. lifespan is very low

And, again, that's because:

-- the regression equation tells you what effect the input has on the output;
-- the r-squared tells you how important that input was in creating the distribution you see.

The regression equations tell you that having a piano drop on your head is very dangerous. The low r-squared tells you that pianos haven't historically been a major source of death.

----

Here's a different way to explain this, which might make more sense to gamblers:

Suppose that you had to predict the lifespan of a random 25-year-old. Obviously, the more information you have, the more accurate your estimate will be. And, imagine the amount you lose is the square of the error in your guess. So if you guess 80, and the random person dies at 60, you lose $400 (the square of 80 minus 60).

Without any information, your best strategy is to guess the average, which we said was 80. Your average loss will be the variance, which is the square of the SD. Suppose that SD is 15. Then, your average loss would be $225.

Now, how valuable is knowing the value of whether or not the guy committed suicide? It's probably not that valuable. Most of the time, the answer will be "no", and you're only slightly better off than when you started (maybe you guess 80.05 now instead of 80). A tiny, tiny proportion of the time, the answer will be "yes," and you can safely guess 25 and be right on. On balance, you're a little better off, but not much.

On average, how much less will you lose given the extra information? The answer is given by the r-squared. If the r-squared of the suicide vs. lifespan regression is .01, as estimated above, then your loss will be reduced by 1%. Instead of losing $225, on average, you'll lose only about $222.75.

Again: the r-squared doesn't tell you that suicide is dangerous. It just tells you that, because of *some combination of dangerousness of suicide and historical frequency of suicide*, you can shave 1% off your error by taking it into account.

----

Now, let's reapply this to basketball. The r-squared for salary vs. wins was .2561. The SD of wins was 14.1, so the variance was the square of that, or 199.

If you took a bet where you had to guess a random team's wins, and had to pay the square of the difference, you'd pick "41" and, on average, owe $199. But let's suppose someone tells you the team's payroll. Now, you can adjust your guess, to predict higher if the team has a high payroll, or lower if the team has a low payroll. If you adjust your guess optimally -- by using the results of the regression equation -- you'll cut your average loss by 25.61%. So, on average, you'd lose only 74.39% as much as before. That works out to $148.11.

What Berri, Brook and Schmidt are saying, in "The Wages of Wins," is, "look, if you can only cut your losses by 25.61% by knowing salary, then money can't be that important in buying wins." But that's wrong. What they should conclude is that "how important money is, combined with how often it's been used to buy wins," isn't that important.

And, really, if you look at the full results of the regression, it turns out that money IS important in buying wins, but that not too many teams took advantage of that fact in 2008-09.

The equation shows that every $1.6 million dollars in additional salary will buy you a win -- so if you want to go 61-21, it should only cost you $32 million more than the league-average payroll of $68.5 million.

That's pretty important, and so the low r-squared must be that not a lot of teams varied much in salary. If you look at the salary chart, there's a huge group bunched near the average: there are 18 teams between $62mm and $75mm, within $6.5 million of the average. Those teams are so close together that there's not much difference in their expected wins.

If you have to bet, and the random team you pick turns out to be the lowest-spending in the league, you'll reduce your estimate. You would have lost a lot of money guessing 41, so the information that you picked a low-spending team will cut your losses a lot. If it turns out be be one of the highest-spending in the league, same thing. But if it turns out to be one of the 18 teams in the mdidle, the salary information won't help you much. And why the r-squared is only about 25% -- for many of the teams in the sample, knowing the salary doesn't help you cut your losses much.


What if we take out those 18 teams, and regress only on the remaining 12? Well, the regression equation stays almost the same -- $1.5 million per win instead of $1.6. But the r-squared increases to .4586. Why does the r-squared increase? Because salary is much more significant a factor for those 12 teams than for the ones in the middle. Before, knowing the salary might not do you much good for your estimate if it's one of the teams bunched in the middle. But, now, those teams are gone. Your random team is much more likely to be the Cavaliers or the Clippers, so knowing the salary is a much bigger help, and it lets you cut your betting losses by almost half.

----

One last summary:

1. The regression equation tells you how powerful the input is in affecting output -- is it a nuclear weapon, or a pea-shooter?

2. The r-squared tells you how powerful the input is, "multiplied by" how extensively the input was historically used. That is: a nuclear weapon used once might give you the same r-squared as a pea-shooter used a billion times.

So a low r-squared might mean

-- an input that doesn't have much effect on the output (e.g., shoe size probably doesn't affect lifespan much);

-- an input that has a big effect on output but doesn't happen much (e.g., suicide curtails 100% of lifespan but happens rarely); or

-- an input that doesn't affect output and also doesn't happen much. (e.g., fluorescent purple shoes' effect on lifespan).

In the case of the 2008-09 NBA, the regression equation shows that salary is a fairly powerful bomb. And the moderate r-squared shows that not every team uses it to its full potential.

Bottom line: salary can indeed very effectively buy wins. The r-squared is as small as it is because, in 2008-09, NBA teams differed only moderately in how they chose to vary their spending.


Labels: , , , , ,