Tuesday, October 16, 2012

Can money buy meat?

If you want to have meat in your diet, you have to spend money in the grocery store.  At least, that's the conventional wisdom.

But is that really true, or is it just a myth?  Let's look at the evidence.

I took a (made up) random sample of 30 shoppers in my local supermarket earlier this year.  I ran a regression to predict the total amount of meat they had, from the total amount of money they spent.  It did turn out that the York family, the one who spent the most money by far, did get the most meat.  And, that there was a positive slope, meaning that spending more money leads to more meat.

However, there was one very important issue: the link between meat and money was not statistically significant.  In other words, we can't argue that money spent and meat obtained are actually related to each other in 2012.

It's easy to understand why we got this result.  Some of the lowest-spending families wound up with a lot of meat -- one was stocking up for a BBQ, and one owned a cattle ranch.  And a few rich-spending families barely had any meat in their houses at all -- they paid a lot for only a few ounces of filet mignon.

But 2012 isn't typical.  When I (pretended that I) did the same experiment for other years, I got statistically significant results.  But even for those years, explanatory power is quite low.  Only 17 percent of the variation in meat over the last 25 years is explained by variation in spending.  So much of the variation in meat obtained is not explained by how much money was spent.

And if you look at each year individually -- as the following table (with made-up numbers) illustrates -- the power of money buying meat seems to vary quite a bit:

2012: not significant
2011: r-squared = .17, p = .01
2010: r-squared = .13, p = .04
2009: r-squared = .21, p = .02
2008: r-squared = .10, p = .06
2007: r-squared = .25, p = .00
2006: r-squared = .29, p = .00
2005: r-squared = .24, p = .00
2004: r-squared = .29, p = .00
2003: r-squared = .18, p = .02
2002: r-squared = .20, p = .01
2001: r-squared = .10, p = .04
2000: r-squared = .10, p = .04
1999: r-squared = .50, p = .00
1998: r-squared = .47, p = .00
1997: r-squared = .22, p = .01
1996: r-squared = .34, p = .00
1995: not significant
1994: r-squared = .16, p = .07
1993: r-squared = .09, p = .09
1992: not significant
1991: not significant
1990: not significant
1989: not significant
1988: r-squared = .18, p = .00


From 1996 to 2001, supermarket spending and meat were statistically linked each and every year.  However, explanatory power varied.  If we look at shoppers before 1993, we see four years where where the relationship was again not significant.

So here is the big question: Why is the relationship not stronger?  One would think that as shoppers spend more, they would wind up with more meat.  But, often, that's not what we see in the data.

One issue is that you can get meat at other places than the supermarket -- butchers, say, or gifts, or the slaughter of animals you own yourself.  Another issue is that it's hard to predict what shoppers will buy any given week. 

But, does the result from 2012 show that spending and meat will not be statistically related in future?  We don't know.  But what we *do* know is that spending does not guarantee a shopper more meat.

That's the nature of shopping.  Sometimes you don't get enough meat, and, it seems, no amount of spending can change that reality.

--------

So: do you believe me?  Do you believe that how much meat you have in 2012 doesn't depend on how much money you spend?  I hope not.

What, specifically, is wrong with the logic?  Lots of things, many of which I've written about before.

------

1.  Even if you don't get a statistically-significant relationship between spending and meat, that does NOT mean that you "can't argue that money spent and meat obtained are actually related to each other".  Of course you can!  Lack of significance just means that, in one specific, narrow, sense, you don't have enough grounds to assert a relationship *on this evidence alone*. 

But, of course, there's LOTS of other evidence that meat and spending are related.  For one thing, there's a big sign in front of the steaks, that says, "$7.99 per pound."  For another thing, millions of people will tell you that they have successfully exchanged money for meat.

You can only argue that there's no relationship if you choose to ignore all those things.  Which, I hope, you wouldn't.

2.  The implicit assumption in the argument is that every year is different.  That is: money bought meat in 1993 and 1994, but not in 1992 or 1995.  Why would you assume that, that the nature of shopping changes so often and so much that you can buy meat in 1994, but not 1995?  If we're going to assert that, we need some kind of explanation of how that could be plausible.

3.  Also, if you're interested in statistical significance, shouldn't you care about checking if 1994 and 1995 are actually significantly different?  What do you do if there's not statistically significant difference between them, as there probably isn't?  How can you say money bought meat one year, but not the other, when the p value of the difference is very high? 

You have a contradiction:

1994 is significantly different from zero
1995 is NOT significantly different from zero
1994 is NOT significantly different from 1995.


Isn't it just as reasonable to say there's no difference, than to say that money bought meat in 1994 but not 1995?  Even if you're depending on statistical significance, you still have to make an argument.

4.  Why use "different from zero" as your significance criterion anyway?  In this particular situation, there is no real reason to think that zero is more likely than any other value -- and, in fact, there's very, very good reason to believe it's different, unless you have good reason to believe that big spenders don't buy more meat than the guy in the express lane with one item. 

In some cases, like whether prayer cures cancer, a default of zero makes sense.  But not here.  Saying, "we'll assume money can't buy meat until we see strong evidence otherwise" ... well, that's just privileging your hypothesis.

5.  If you get a value that's significant in the real world sense, but isn't statistically significant, you need more data.  You can say, "I don't have enough evidence."  You can say, "there isn't enough evidence HERE."  But you can't just assume that there's no relationship.  Otherwise, it would be easy to argue that smoking is harmless.  You just do a double-blind study that's really small.  And then you say, "even though 40 percent of the five smokers got lung cancer, and only 20 percent of the five non-smokers got cancer, we got an r-squared of only .1, and that's not statistically significant.  So, there's no evidence that smoking causes cancer."

Yes, the evidence of THAT study is weak.  But that's because the study is too small.  Twice the risk of cancer is plausible, and important, and you can't just dismiss it because you deliberately designed your study the way you did.  And there are lots and lots of other studies showing a link, and a biological mechanism by which it happens.

If you did that study, and you deliberately ignore all the other evidence, than it's fair to say that YOU can't conclude that smoking causes cancer.  But WE can certainly conclude it. 

Similarly, if all you know is that within the dataset of your 30 individuals, the correlation between meat and spending is low ... YOU can conclude you don't have evidence that meat can be bought.  But WE cannot, because WE have other evidence: we've been to a supermarket.  We know something about how the market for meat works.

6.  Even noting that the r-squareds jump around a bit -- and that the jumping around is statistically significant -- that doesn't necessarily mean that the relationship between money and meat has changed.  The r-squared depends not just on the relationship, but on the scattering of the values in the actual dataset. 

So an increasing r-squared could simply indicate a larger variation in overall spending.  Think about it ... if some families spend $1, and some spend $1000, it should be easier to notice the relationship between spending and meat, which means a large r-squared.  But if everyone spends exactly $100, it's going to be harder -- a lower r-squared -- even if money buys the same amount of meat as always.

So when you see a changing r-squared, you can't really be sure what's going on.  It would be better to look at the coefficient estimate of the regression equation, rather than the r-squared.

In fact, for any arbitrarily low r-squared, I can construct a dataset where the coefficient is as statistically significant as you like, and meat costs any amount per pound you like.  (I thought I wrote about this fact before, but I can't find it.)

7.  Even though an r-squared less than .10 may look small intuitively, it probably isn't.  A low-looking r-squared can be very important in real life.  You can't just say ".10 is small".  You have to *argue* that, in context, it's small.

If you did a regression of suicide vs. life expectancy, the r-squared would be at least as small as the ones here.  But suicide and life expectancy are most definitely linked. 

You have to interpret the r-squared for what it is.  It's not really an indicator of how easily money buys meat.  It's a measure of how well you can predict meat from money, *relative to all the other things* that help you predict meat*. 

If cancer kills a million people, and suicide kills 10, the r-squared between suicide and life expectancy will be low, because suicide is being compared to cancer.  That's true even though a single suicide has a bigger effect on life expectancy than a single case of cancer.

8.  You'll notice how large the relationship really is if you look at the r, instead of the r-squared.  The square root of .17 is .41.  That means that for every standard deviation difference in money spent, you get 41% of a standard deviation in meat obtained.  That's a pretty strong association: if you move two inches to the right on the supermarket-spending bell curve, you move 2/5 of an inch to the right on the meat curve.

9.  The r-squared doesn't really tell you whether meat CAN be increased by increasing supermarket spending.  It tells you how much meat WAS increased with supermarket spending.  Obviously, you'd expect an imperfect correlation.  People use money on all kinds of things -- TVs, cars, tofu, vegetables.  They get meat from sources -- their own animals, gifts, butchers -- other than supermarkets.  And, they buy different kinds and forms of meat, at various prices: steaks, hamburger, spam, TV dinners, dog food, and so on.

Given all that variation, *of course* you're going to find a less-than-100% correlation between supermarket spending and meat purchased.  That doesn't mean that there's no cause-and-effect relationship of deliberately spending more money and getting more meat, at the margin.

This is easier to understand if we look at something other than meat -- say, hair. 

Hair CAN be bought for money.  If you're bald, and you want to have hair, you can write a check to Hair Club For Men, and they'll add hair to your head.  But if you look at whether hair HAS BEEN bought for money, very little of it has -- most of it we got free, from God.  The r-squared between "hairs on head" and "money spent" is low, because most hair is not bought for money, and most money is not spent on hair. 

But if you have money, and you choose to buy hair, you'll get it.  


Same for meat.

------

And so, the botttom line is: even if we get a legitimately small correlation, you CANNOT say that "no amount of spending can buy meat."  That's exactly like noting that the correlation between shooting yourself in the head and lifespan is small, and saying, "no amount of shooting yourself in the head can change your life expectancy." 

That's just not true, because it's just not what r-squared means.

-------







(Inspiration: this Freakonomics post.)



Labels: , , ,

Tuesday, August 12, 2014

More r-squared analogies

OK, so I've come up with yet another analogy for the difference between the regression equation coefficient and the r-squared.

The coefficient is the *actual signal* -- the answer to the question you're asking. The r-squared is the *strength of the signal* relative to the noise for an individual datapoint.

Suppose you want to find the relationship between how many five-dollar bills someone has, and how much money those bills are worth. If you do the regression, you'll find:

Coefficient = 5.00 (signal)
r-squared = 1.00 (strength of signal)
1 minus r-squared = 0.00 (strength of noise)
Signal-to-noise ratio = infinite (1.00 / 0.00)

The signal is: a five-dollar bill is worth $5.00. How strong is the signal?  Perfectly strong --  the r-squared is 1.00, the highest it can be.  (In fact, the signal to noise ratio is infinite, because there's no noise at all.)

Now, change the example a little bit. Suppose a lottery ticket gives you a one-in-a-million chance of winning five million dollars. Then, the expected value of each ticket is $5.  (Of course, most tickets win nothing, but the *average* is $5.)

You want to find out the relationship between how many tickets someone has, and how much money those tickets will win. With a sufficiently large sample size, the regression will give you something like:

Coefficient = 5.00 (signal)
r-squared = 0.0001 (strength of signal)
1 minus r-squared = 0.9999 (strength of noise)
Signal-to-noise ratio = 0.0001 (0.0001 / 0.9999)

The average value of a ticket is the same as a five-dollar bill: $5.00. But the *noise* around $5.00 is very, very large, so the r-squared is small. For any given ticketholder, the distribution of his winnings is going to be pretty wide.

In this case, the signal-to-noise ratio is something like 0.0001 divided by 0.9999, or 1:10,000. There's a lot of noise in with the signal.  If you hold 10 lottery tickets, your expected winnings are $50. But, there's so much noise, that you shouldn't count on the result necessarily being close to $50. The noise could turn it into $0, or $5,000,000.

On the other hand, if you own 10 five-dollar bills, then you *should* count on the $50, because it's all signal and no noise.

It's not a perfect analogy, but it's a good way to get a gut feel. In fact, you can simplify it a bit and make it even easier:

-- the coefficient is the signal.
-- the r-squared is the signal-to-noise ratio.

You can even think of it this way, maybe:

-- the coefficient is the "mean" effect.
-- the (1 - r-squared) is the "variance" (or SD) of the effect.

Five-dollar bills have a mean value of $5, and variance of zero. Five-dollar lottery tickets have a mean value of $5, but a very large variance.  

------

So, keeping in mind these analogies, you can see that this is wrong: 

"The r-squared between lottery tickets and winnings is very close to zero, which means that lottery tickets have very little value."

It's wrong because the r-squared doesn't tell you the actual value of a ticket (mean). It just tells you the noise (variance) around the realized value for an individual ticket-holder. To really see the value of a ticket, you have to look at the coefficient.  

From the r-squared alone, however, you *can* say this:

"The r-squared between lottery tickets and winnings is very close to zero, which means that it's hard to predict what your lottery tickets are going to be worth just based on how many you have."

You can conclude "hard to predict" based on the r-squared. But if you want to conclude "little value on average," you have to look at the coefficient.  

------

In the last post, I linked to a Business Week study that found an r-squared of 0.01 between CEO pay and performance. Because the 0.01 is a small number, the authors concluded that there's no connection, and CEOs aren't paid by performance.

That's the same problem as the lottery tickets.

If you want to see if CEOs who get paid more do better, you need to know the size of the effect. That is: you want to know the signal, not the *strength* of the signal, and not the signal-to-noise ratio. You want the coefficient, not the r-squared.

And, in that study, the signal was surprisingly high -- around 4, by my lower estimate. That is: for every $1 in additional salary, the CEO created an extra $4 for the shareholders. That's the number the magazine needs in order to answer its question.

The low r-squared just shows that the noise is high. The *expected value* is $4, but, for a particular case, it could be far from $4, in either direction.  I haven't checked, but I bet that some companies with relatively low-paid executives might create $100 per dollar, and some companies who pay their CEOs double or triple the average might nonetheless wind up losing value, or even going bankrupt.

------

Now that I think about it, maybe a "lottery ticket" analogy would be good too: 


Think of every effect as a combination of lottery tickets and cash money.

-- The regression coefficient tells you the total value of the tickets and money combined.

-- The r-squared tells you what proportion of that total value is in money.  

That one works well for me.

------

Anyway, the idea is not that these analogies are completely correct, but that they make it easier to interpret the results, and to spot errors of interpretation. When Business Week says, "the r-squared is 0.01, so there is no relationship," you can instantly respond:

"... All that r-squared tells you is, whatever the relationship actually turns out to be, the signal-to-noise ratio is 1:99. But, so what? Maybe it's still an important signal, even if it's drowned out by noise. Tell us what the coefficient is, so we can evaluate the signal on its own!"

Or, when someone says, "the r-squared between team payroll and wins is only .18, which means that money doesn't buy wins," you can respond:

"... All that r-squared tells you is, whatever the relationship actually turns out to be, 82 percent of it comes in the form of lottery tickets, and only 18 percent comes in cash. But those tickets might still be valuable! Tell us what the coefficient is, so we can see that value, and we can figure out if spending money on better players is actually worth it."

------

Does either one of those work for you?  




(You can find more of my old stuff on r-squared by clicking here.)


Labels: , ,

Wednesday, August 22, 2012

R-squared and "twenty questions"

(This is a follow-up to two previous posts on r-squared.)

------
 
You've probably played the game "Twenty Questions." Here's how it works.  I choose a subject, which can be anything I want -- "baseball glove," or "Hillary Clinton".  Then, you have twenty "yes/no" questions to try to figure out what it is.

To win the game, you try to narrow it down as fast as you can, and as much as possible.  To start, you might ask the traditional question: "is it bigger than a bread box?"  That one question won't tell you what it is immediately, but it starts narrowing it down.  According to Wikipedia, other good questions are, "can I put it in my mouth?" and "does it involve technology for communications, entertainment, or work?"

Some questions are obviously bad.  Starting off by asking, "is it a DVD of a Sylvester Stallone movie?" is a waste of a question.  The answer is probably "no," in which case you're left pretty much where you started.  Of course, if it's a "yes," you're almost certain to get it, but the chances of that "yes" are pretty slim.

-----

So, now, here's a variation of the game.  This time, I'm going to pick a random American person.  Your job is to guess his or her 2011 income, and come as close as you can. 

Instead of twenty questions, I'm only going to give you one, for now.  However, it doesn't have to be a "yes/no" question -- if you want, it can be any question that can be answered by a number.  (You can't ask specifically about the salary, though.)  Once you get the answer, you take your guess at the income. 

What kind of question do you ask? 

Well, one good question might be, "how many years of education does the person have?"  You can then go by the general rule that the more education, the more likely they were to have a higher salary.

Or, you can ask, "what's the person's IQ?"  Again, you can assume that the higher a person's people's intelligence, the more likely they are to have a higher income.

But if you ask, "did they win a lottery jackpot last year?", that's a waste.  It's the like the Stallone DVD question.  Most of the time, the answer is no, and you get barely any useful information.  It's just not worth asking, just on the off-chance that you get a yes.

-----

That all makes sense, right?  Well, if you understand the strategy of the game of "Twenty Questions," you understand r-squared.  Because, r-squared is really just a measure of how good your question is.  Seriously -- the correspondence between the two is almost perfect.  The better the question, the higher the r-squared; and, the higher the r-squared, the better the question.

If you were to run a regression of IQ against income, you'd probably wind up with a decent r-squared -- maybe, I don't know, .15 or something.  That means that if you know a random person's IQ, you can knock 15 percent off your average error squared.  Maybe if you were completely ignorant, you would guess $30,000 for everyone.  But if you know IQ, you can guess $20,000 for low, $30,000 for average, and $40,000 for high, and your guesses would be closer. 

But, the bad question, the lottery question: the r-squared of that might only be .001.  Originally, you guessed $30,000.  Now, if you find out they didn't win the lottery, you guess $29,999, and are a tiny bit closer, on average.  If you find out they *did* win the lottery, then you guess, say, $5 million, and you come a lot closer that you would have before.  But that happens very infrequently -- so infrequently that you're square error is still going to be well over 99.99 percent of what it was without the question.

-----

I said the analogy between the game and r-squared was *almost* perfect.  If you care, here's how to make it exact:

After you ask the IQ question, you're given a table of all 300,000,000 people in the US, with their income and answer to your question (IQ).  Then, before I tell you the random person's IQ, you have to decide in advance what you're going to answer for each possible IQ, and your decision has to have each point of IQ worth the same amount of income (that is, it has to be linear, since it's a linear regression). 

Once you've decided, I give you the IQ, and we figure your answer, and your negative score is the square of how much you missed it by.

Under those rules, the analogy is exact: the r-squared exactly corresponds to how good a question you asked.

(Oh, and if you want to actually ask twenty questions instead of one ... that's just a multiple regression with 20 variables.)


-----

In the past, I've been critical of analyses that find a low r-squared, and assume that, therefore, there's only a weak relationship.  For instance, I've written about the study that found, in MLB, an r-squared of .18 for team payroll vs. team wins.  The authors of the study then said something like, "The r-squared is low.  Therefore, there's not much of a relationship.  Therefore, salary doesn't lead to wins."

Well, that's not right.  It's like saying, for the lottery example, "The r-squared is low.  Therefore, winning the lottery doesn't lead to more money."

That's obviously incorrect. 

The r-squared does NOT measure the direct relationship between the variables.  It just measures how good a question it is to ask about the one variable.

But, the thing is, if what you really want is the relationship between winning the lottery and getting rich ... well, that's easy.  Just look at the regression equation!

If you do that regression, the one on lottery winnings that gives you an r-squared of .001, you'll wind up with an equation like

Expected salary = $30,000 + $5,000,000 if they won the lottery.


It gives you exactly what you want -- winning the lottery is worth $5 million.  Why would you focus on the r-squared, when the exact answer is right there?  In fact, the r-squared is completely irrelevant!

I think the reason we sometimes focus on the r-squared, though, is that we make a false assumption.  It is true that (a) if you have a high r-squared, you have a strong relationship.  But it is NOT necessarily true that (b) if you *don't* have high r-squared, you *don't* have a strong relationship.  I think that maybe we just assume because (a) is true, (b) is also true.  But it's not.

------

So, in summary, three different ways to think about it:

--- One

The regression equation answers, "how much does winning the lottery affect income?"  [lots.]

The r-squared answers, "is asking about the lottery a good "twenty questions" way to help estimate income?"  [not very.]

--- Two

The regression equation answers, "how much does winning the lottery affect income?"  [lots.]

The r-squared answers, "when people differ in income, how much of that is because some of them won the lottery?" [not much.]

--- Three

The regression equation answers, "if you change the value of the lottery variable from "no" to "yes," how much does income change?  [lots.]

The r-squared answers, "if you change the value of the lottery value from one random person's to another random person's, how much does income change?"  [not much -- two random people are probably both "no", so the change is usually zero.]

------

If you've got any more good ones, let me know, and I'll add them in.





Labels: , ,

Monday, December 03, 2007

Does pace impact defensive efficiency? Don't use r-squared, use the regression equation

When someone runs a regression, they will wind up reporting a value for r or r-squared. If the value is small, they'll argue that the two variables don't really have much of a relationship. But that's not necessarily true.

Before I talk about why, I should say that if the result goes the other way – the r or r-squared is high, and statistically significant -- that *does* mean there's a strong relationship. If the correlation between cigarettes smoked and lung cancer is, say, 0.7, that's a pretty big number, and we can conclude that lung cancer and smoking are strongly related.

But a low value doesn't necessarily mean the opposite.

For instance: inspired by the smoking example, I looked at another lifestyle choice. Then, I ran a regression on expected remaining lifespan, based on that lifestyle choice. The results:

r = -.17, r-squared = .03

What is the effect of that lifestyle choice on lifespan? It looks like it should be small. After all, it "explains" only 3% of the variance in years lived.

But that wouldn't be correct. The lifestyle choice really does have a large effect on lifespan. In fact, the lifestyle choice I used in the equation is (literally) suicide.

Here's what I did. I took 999 random 40-year-olds, and assumed their expected remaining lifespan was 40 years, with an SD of about 8. Then, I assumed the 1,000th person jumped in front of a moving subway train, with an expected remaining lifespan of zero. (These numbers are made up, by the way.)

The results were what I showed above: an r of only –.17.

Why does this happen? It happens because the r, and the r-squared, do NOT measure whether suicide and lifespan are related. Rather, they measure something subtly different: whether suicide is a big factor in how long people live.

And suicide is NOT that big a factor in how long people live. Most people don't commit suicide; in my model, only 1 in 1000. The r-squared shows how much effect suicide has *as a percentage of all other factors*. Because there are so many other factors – in real life, heart disease and cancer are about 40 times as common as suicide -- the r-squared comes out small.

If you want to know the strength of the relationship between A and B, don't look at the r or the r-squared. Instead, look at the regression equation. In my suicide experiment, the equation turned out to be

Lifespan = 40.0 – 40.0 (suicide)

That is, exactly what you would expect: the lifespan is 40 years, but subtract 40 from that (giving zero) if you commit suicide.

And, even though the r-squared was only 0.03, that r-squared is statistically significant, at greater than 99.99%.

Again: the r-squared is heavily dependent on how "frequent" the lifestyle choice is in the population. But the significance level, and the regression equation, is not.

To prove it, let me rerun my experiment a few times, with different percentages of suicide in the population:

1 in 10,000:
r-squared = .003
lifespan = 40.1 – 40.1 (suicide); p > 99.99%

1 in 1,000: r-squared = .03

lifespan = 40.0 – 40.0 (suicide); p > 99.99%

1 in 100: r-squared = .26

lifespan = 40.3 – 40.3 (suicide); p > 99.99%

1 in 10: r-squared = .71

lifespan = 37.3 – 37.3 (suicide); p > 99.7%

The r-squared varies a lot – but all the experiments tell you that suicide costs you 40 years of life, and that the result is statistically signficant.


The moral of the story:

The r-squared (or r) does NOT tell you the extent to which A causes B, or even the strength of the relationship between A and B. It tells you the extent to which A explains B relative to all the other explanations of B.

If you want to quantify the effect a change in A has on B, do not look at the r or r-squared. Instead, look at the regression equation.

------

Which brings us to today's post at "The Wages of Wins." There, David Berri checks whether teams who play a fast-paced brand of basketball (as measured by possessions per game) wind up playing worse defense (as measured by points allowed per possession) because of it. Berri quotes Matthew Yglesias:

"For example, there’s a popular conception of a link between pace and defensive orientation — specifically the idea that teams that choose to play at a fast pace are sacrificing something in the defense department. On the most naive level, that’s simply because a high pace leads to more points being given up. But I think it’s generally assumed that it holds up in efficiency terms as well. The 2006-2007 Phoenix Suns, for example, were first in offensive efficiency, third in pace, and fourteenth in defense. But is this really true? If you look at the data season-by-season is there a correlation between pace and defense?"

Berri runs a regression for 34 years of team data. So, is there a relationship? He writes,





"The correlation coefficient between relative possessions and defensive efficiency is 0.17. Regressing defensive efficiency on relative possession reveals that there is a statistically significant relationship. The more possessions a team has per game - again, relative to the league average - the more points the team’s opponents will score per possession. But relative possessions only explains 2.8% of defensive efficiency. In sum, pace doesn’t tell us much about defensive efficiency ... " [emphasis mine]


But I don't think that's right. As we saw, the r-squared of 2.8% (or the r of 0.17) means only that, historically, pace is small *compared to other explanations of defensive efficiency.* And that makes sense. Even if pace has a significant impact on defense, we'd expect other factors to be even more important. The players on the team, for instance, are a big factor. Luck is also a big factor. The coach's strategy probably has a large impact on defensive efficiency. Compared to all those things, pace is pretty minor. And we probably knew that before we started, that personnel matters more than pace.

And so I would guess that's not really what Yglesias wants to know. What I bet he's interested in, and what teams would be interested in, and what I'd be interested in, is this: if a team speeds up the pace by (say) 2 possessions per team per game, how much will its defense suffer? That's an important question: if you're evaluating how good a team (or player) is on defense, you want to know if you can take the stats at face value, or if you have to bump them up to compensate for fast play, or if you have to discount them for teams who play a little slower. It's like a park factor, but for defensive efficiency. The regression should be able to tell you just how big that park factor is.

That's the real question, and the r-squared doesn't answer it at all. Given the data Berri gives us, the effect of pace on defensive efficiency could be small, or it could be large. After all, the effect of suicide on lifespan was huge, even though the r-squared was small. And just like in the suicide case, if a lot more teams suddenly decide to start playing at a different pace, the r and r-squared will go up – but the relationship between pace and defense will likely not change.

To really understand what effect pace has on defense, we need the regression equation. Berri doesn't give it to us. He does tell us the result is statistically significant, so we do know there *is* some kind of non-zero effect. But without the equation, we don’t know how big it is (or even whether it's positive or negative). All we know is that pace *does* signficantly impact on a team's defensive stats, and that the effect (as judged by statistical signficance) appears to be real.




Labels: ,

Friday, May 08, 2009

The regression equation versus r-squared

OK, I hope I'm not beating a dead horse here, but here's another way to think of the difference between r-squared and the regression equation.

The r-squared comes from the standpoint of stepping back and looking at the distribution of wins among teams in your dataset. Some teams have over 60 wins, some teams have under 20 wins, and some teams are in the middle. If you look at the standings, and ask yourself, "how important are differences in salary to how we got this way?", then you're asking about r-squared.

The regression equation matters more if you're interested in the future, if you care about how much you can influence wins by increasing payroll. If you ask yourself, "how much do I have to spend to get a few extra wins?", then you want the regression equation.

The r-squared looks at the past, and asks, "was salary important to how we got to this variance in wins?". The regression equation looks to the future, and says, "can we use salary to influence wins?"

It's very possible, and very easy, to have two different answers to these two questions. Here's an example.

Suppose you're trying to see what activities 25-year-olds partake in that affect their life expectancy. You might discover that the average 25-year-old lives to 80, but you want to try to figure out what factors influence that. You run a multiple regression, and you figure out that if the person smokes at 25, it appears to cut five years off his life expectancy. If he eats healthy, it adds four years. If he commits suicide at 25, it cuts off 55 years (since he dies at 25 instead of 80).

Your regression equation would look something like:

life expectancy = 80 - (5 * smoker) + (4 * eats healthy) - (55 * commits suicide).

We should all agree that committing suicide has a big effect on life expectancy, right?

Now, let's look at the r-squared. To do that, look at all the 25-year-olds in the sample (which might be several thousand). You'll see a few that live to 25, some that live to 45, a bunch that live to 65, a larger bunch that live to 80, and some that live to 100. The distribution is probably bell-shaped.

For the r-squared, ask yourself: how much did suicide contribute to the curve looking like this? The answer: very little. There are probably very few suicides at 25, and even if you adjusted for those, by taking those points out of the left side of the curve and moving them to the peak, the curve would still look roughly the same. Suicide is not a very big factor in making the curve look like it does.

And so, you get a very low r-squared for suicide. Maybe it would be .01, or even less.

See the apparent contradiction?

-- suicide has a HUGE effect on lifespan.
-- r-squared for suicide vs. lifespan is very low

And, again, that's because:

-- the regression equation tells you what effect the input has on the output;
-- the r-squared tells you how important that input was in creating the distribution you see.

The regression equations tell you that having a piano drop on your head is very dangerous. The low r-squared tells you that pianos haven't historically been a major source of death.

----

Here's a different way to explain this, which might make more sense to gamblers:

Suppose that you had to predict the lifespan of a random 25-year-old. Obviously, the more information you have, the more accurate your estimate will be. And, imagine the amount you lose is the square of the error in your guess. So if you guess 80, and the random person dies at 60, you lose $400 (the square of 80 minus 60).

Without any information, your best strategy is to guess the average, which we said was 80. Your average loss will be the variance, which is the square of the SD. Suppose that SD is 15. Then, your average loss would be $225.

Now, how valuable is knowing the value of whether or not the guy committed suicide? It's probably not that valuable. Most of the time, the answer will be "no", and you're only slightly better off than when you started (maybe you guess 80.05 now instead of 80). A tiny, tiny proportion of the time, the answer will be "yes," and you can safely guess 25 and be right on. On balance, you're a little better off, but not much.

On average, how much less will you lose given the extra information? The answer is given by the r-squared. If the r-squared of the suicide vs. lifespan regression is .01, as estimated above, then your loss will be reduced by 1%. Instead of losing $225, on average, you'll lose only about $222.75.

Again: the r-squared doesn't tell you that suicide is dangerous. It just tells you that, because of *some combination of dangerousness of suicide and historical frequency of suicide*, you can shave 1% off your error by taking it into account.

----

Now, let's reapply this to basketball. The r-squared for salary vs. wins was .2561. The SD of wins was 14.1, so the variance was the square of that, or 199.

If you took a bet where you had to guess a random team's wins, and had to pay the square of the difference, you'd pick "41" and, on average, owe $199. But let's suppose someone tells you the team's payroll. Now, you can adjust your guess, to predict higher if the team has a high payroll, or lower if the team has a low payroll. If you adjust your guess optimally -- by using the results of the regression equation -- you'll cut your average loss by 25.61%. So, on average, you'd lose only 74.39% as much as before. That works out to $148.11.

What Berri, Brook and Schmidt are saying, in "The Wages of Wins," is, "look, if you can only cut your losses by 25.61% by knowing salary, then money can't be that important in buying wins." But that's wrong. What they should conclude is that "how important money is, combined with how often it's been used to buy wins," isn't that important.

And, really, if you look at the full results of the regression, it turns out that money IS important in buying wins, but that not too many teams took advantage of that fact in 2008-09.

The equation shows that every $1.6 million dollars in additional salary will buy you a win -- so if you want to go 61-21, it should only cost you $32 million more than the league-average payroll of $68.5 million.

That's pretty important, and so the low r-squared must be that not a lot of teams varied much in salary. If you look at the salary chart, there's a huge group bunched near the average: there are 18 teams between $62mm and $75mm, within $6.5 million of the average. Those teams are so close together that there's not much difference in their expected wins.

If you have to bet, and the random team you pick turns out to be the lowest-spending in the league, you'll reduce your estimate. You would have lost a lot of money guessing 41, so the information that you picked a low-spending team will cut your losses a lot. If it turns out be be one of the highest-spending in the league, same thing. But if it turns out to be one of the 18 teams in the mdidle, the salary information won't help you much. And why the r-squared is only about 25% -- for many of the teams in the sample, knowing the salary doesn't help you cut your losses much.


What if we take out those 18 teams, and regress only on the remaining 12? Well, the regression equation stays almost the same -- $1.5 million per win instead of $1.6. But the r-squared increases to .4586. Why does the r-squared increase? Because salary is much more significant a factor for those 12 teams than for the ones in the middle. Before, knowing the salary might not do you much good for your estimate if it's one of the teams bunched in the middle. But, now, those teams are gone. Your random team is much more likely to be the Cavaliers or the Clippers, so knowing the salary is a much bigger help, and it lets you cut your betting losses by almost half.

----

One last summary:

1. The regression equation tells you how powerful the input is in affecting output -- is it a nuclear weapon, or a pea-shooter?

2. The r-squared tells you how powerful the input is, "multiplied by" how extensively the input was historically used. That is: a nuclear weapon used once might give you the same r-squared as a pea-shooter used a billion times.

So a low r-squared might mean

-- an input that doesn't have much effect on the output (e.g., shoe size probably doesn't affect lifespan much);

-- an input that has a big effect on output but doesn't happen much (e.g., suicide curtails 100% of lifespan but happens rarely); or

-- an input that doesn't affect output and also doesn't happen much. (e.g., fluorescent purple shoes' effect on lifespan).

In the case of the 2008-09 NBA, the regression equation shows that salary is a fairly powerful bomb. And the moderate r-squared shows that not every team uses it to its full potential.

Bottom line: salary can indeed very effectively buy wins. The r-squared is as small as it is because, in 2008-09, NBA teams differed only moderately in how they chose to vary their spending.


Labels: , , , , ,

Friday, August 17, 2012

A benefit of r-squared

(This is a sequel to a previous post on r-squared.)

-----

Bob and Sam are arguing about how raffles are won.  Bob says, "People tend to win because they buy lots of tickets.  That's the biggest factor."  And Sam replies, "Well, you can buy a lot of tickets, but you still have to get lucky.  People win more because they get lucky, rather than because of how many tickets they buy."

You do a regression to predict prizes won, based on how many tickets were bought.  And the r-squared winds up at .52. 

Well, Bob has won the argument, hasn't he?  The number of tickets explains 52 percent of the variance in prizes won.  If Sam were right, then luck would have to explain more than 52 percent.  Since tickets and luck are presumably independent, you can just add the variances (by the Pythagorean theorem of statistics).  That adds up to 104 percent, which is impossible.

Anything above 50 percent must be the biggest factor -- at least, of all other factors independent of that one.

----

And that's where it gets tricky.  Because, it's hard to find factors that are legitimately independent.

Suppose you're looking to "explain" differences in salary.  Why does Chris make $30,000, while Pat makes $50,000?  What factors contribute to a higher or lower salary?

Someone does a regression, to predict salary based on intelligence (as imperfectly measured by an IQ test, say).  He winds up with an r-squared of, say, .41.  That's pretty impressive!  He concludes that how smart you are is a big predictor of how much money you'll make.

But, then, a colleague comes along.  She thinks it's education that leads to higher salaries.  She does her own regression, to predict salary based number of years of schooling.  The r-squared is .36.  Again, impressive!  She writes a study claiming that schooling increases salaries.

Finally, a third colleague thinks it's family culture.  He does a third regression, this time using parental income.  There's again a high r-squared, this time .43.  The conclusion, this time, is that high salary is something that parents influence you to achieve.

(These r-squareds are all made up; they're probably way too high to be realistic.)

Now, if you just add up the three r-squareds, you might conclude that those three factors, taken together, explain 120% of the variation in salary!  Obviously, that can't be right.

And it's not right -- because you can only add variances when the variables are independent. 

These aren't.  They're highly correlated.  If you have a high IQ, you're more likely to stay in school longer.  If your parents are academics, they probably had a high IQ, and therefore, probably, so would you.

So you can't just add up the r-squareds, like you could for the lottery example, or the dice example from the last post. 

All you can do, if you choose, is perhaps to say that the r-squared for IQ is the highest, so that's the most plausible theory right now.  But, really, it's probably some combination of all three factors, which overlap quite a bit.  The r-squareds don't help a whole lot, here, in supporting one hypothesis over another.

What we *can* do, to help a little bit, is run a multiple regression, using all three variables.  Multiple regression is smart enough to adjust for the fact that the variables aren't independent, by taking them all at once.

Let's suppose we did that, and we got an r-squared of 0.6. That's half of the total sum, which says that exactly half the variance is shared by the three factors.  That means that in the aggregate, our three researchers are "half right".  It doesn't mean they're all half right; one of them might be 100% right, one might be 40%, and one might be 10%.  But, on average, they're half.  That doesn't really do us a lot of good.  We still don't know which of the three factors are important in what proportion. 

Or, even, if there are other factors that just happen to correlate with these. 

------

It's obvious when I spell it out like this, with all three studies.  But, suppose only the first one was done.  You might just read that schooling explains 36% of the variation in salary, and, without thinking, conclude that getting more formal education makes you richer.  But, you'd probably be wrong.  It might be IQ.  It might be culture.  It might be other things, that correlate with schooling, that we haven't thought of yet (how much you study, for instance, or what kind of degree you have).

This is why, I say, you always have to make an argument.  Sure, you got a high r-squared when you looked at schooling.  But why do you assume cause and effect?  How do you know it's not something else instead, something that correlates with schooling?  No statistical test can tell you.  You have to argue for it.

-------

Let me give you a baseball example. 

I took every Major League Baseball team since 1970 (excluding 1981 and 1994), and ran a regression to predict their runs scored from their doubles hit.

The r-squared came out to .462.

I was surprised how high that was.  Knowing 2B, you can reduce the variance by almost half. 

Can we conclude that 46 percent of baseball is doubles?  No, of course not.  It's not the doubles -- it's a confounding factor, something that correlates with doubles.  What I think is actually happening is that teams that hit a lot of doubles also hit a lot of home runs.  (The correlation between HR and 2B was .559.)

Let's get rid of doubles and substitute home runs.  Now, the r-squared is .594.

Can we conclude that 59 percent of baseball is HR?  Again, of course not.  Again, what's likely happening is that teams that hit a lot of home runs also hit a lot of doubles (and probably other things).

All we can say is that knowing HR lets us reduce our mean squared error by 59 percent.  If we want to argue *why*, that's fine, but the r-squared alone doesn't tell us.

--------

Another thing we can do is a multiple regression: predict runs based on both HR and 2B.  I did that, and I got an r-squared of .684.

So, if you start with HR (r-squared .594), and then add doubles, you gain an extra .084.  If you start with doubles (r-squared .462), and then add home runs, you gain an extra .222. 

So, HR and 2B have .378 in common.  HR adds .222 to that.  2B adds .084 to that.  You could draw a Venn diagram to illustrate the overlap, if you wanted to.

But ... those numbers don't mean a whole lot to me, in real life terms. 

-------

So what good is the r-squared, then?

For me, there's one particular task that r-squared is great for: helping figure out how much luck is embedded in performance.  For instance, how much of the variation in team W-L records is based on clutch performance, as opposed to just scoring and preventing runs?

Most sabermetricians say, not much.  My impression is that most mainstream baseball people would say, quite a bit.

Well, here's what I did.  I ran a regression to predict team wins based on runs scored and runs allowed, for 1973-2011 (omitting 1981 and 1994).  The r-squared came out to .87.

That leaves only .13 remaining for clutch performance.  That is, assuming clutch is uncorrelated with RS and RA.  It isn't, quite, but it's probably close.  In any case, you can argue that the *independent* portion has only .13 remaining to claim.  That should satisfy the clutch advocates, since they usually argue that clutch matters in a way *not measured* by raw runs.  (As in, "sure, team A and B scored 750 runs each, but team B scored them when they counted most.")

That's what I like r-squared for.  It lets you estimate the "explanation space" available for the unusual theories, the ones that don't correlate with other variables.  Effectively, it takes some of the air out of the weird hypotheses. 

Suppose I have a theory that having a job interview on a good biorhythm day is an important explanation for differences in salary.  If I run a regression to account for the other, mainstream variables -- IQ, education, study habits, height, sex, race, and so on -- and I get an r-squared of .90, that means there's only .10 left for unrelated factors.  So, I have to be aware that those factors are *at least* nine times more important than my biorhythm theory.

So that's the thing I like about r-squared: it gives you a mental pie chart of the strength of competing explanations -- or, at least, competing *independent* explanations.






Labels: ,

Tuesday, August 22, 2006

On correlation, r, and r-squared

The ballpark is ten miles away, but a friend gives you a ride for the first five miles. You’re halfway there, right? Nope, you’re actually only one quarter of the way there.

That’s according to traditional regression analysis, which bases some of its conclusions on the square of the distance, not the distance itself. You had ten times ten, or 100 miles squared to go – your buddy gave you a ride of five times five, or 25 miles squared. So you’re really only 25% of the way there.

This makes no sense in real life, but, if this were a regression, the "r-squared" (which is sometimes called the "coefficient of determination") would indeed be 0.25, and statisticians would say the ride "explains 25% of the variance." There are good mathematical reasons why they say this, but they mean "explains" in the mathematical sense, not in the real-life sense.

For real-life, you can also use "r". That’s the correlation coefficient, which is the square root of 0.25, or 0.5. In this example, obviously the r=0.5 is the value which makes the most sense in the context of getting to the ballpark. Because you really are, in the real life sense, halfway there.

r is usually the value you use to draw real life conclusions from a regression. According to "The Hidden Game of Baseball," if you regress Runs Scored against Winning Percentage, you get an r of .737, which is an r-squared of .543. A statistician might use the r-squared to say that runs "explains 54.3% of the variation in winning percentage." Which is true if you are concerned with the sums of the squares of the differences – and only a statistician cares about those.

What real people are concerned about is what conclusions we can draw about baseball. And those conclusions are based on the "r", the 0.737. What that tells us is that (a) if a team ranks one standard deviation above average in runs scored, then (b) on average, it will rank 0.737 standard deviations above average in winning percentage.


The 73.7% is useful information about the value of runs to winning ballgames. But the 54.3% figure doesn’t tell you anything you need to know.

I made this point in my review of "The Wages of Wins," where the authors found that payroll "explains only 18%" of wins. They were using r-squared. The r is the square root of .18, which is about .42. Every SD of increased salary leads to an increase of 0.42 SD in wins. In real life, salary explains 42% of wins – although a statistician would probably never put it that way.

Sometimes, the correlation coefficient is used not to predict anything, but just to give you an idea of the relationship between variables. Everyone knows that +1 is a perfect positive relationship, -1 is a perfect negative relationship, and 0 is no relationship at all. And the higher the absolute value of the number, the stronger the relationship. So an r of 0.1 is a weak relationship, but -0.9 is a very strong relationship.

But a "very strong relationship" depends on the context. Sean Forman reports that the correlation between year-to-year players’ batting average is 0.45. That’s pretty high. But if the game-to-game correlation was 0.45, that would be enormous! It would indicate a huge "hot hand" effect. It would mean that if a player was two hits above average one night – say, he went 3-for-4 instead of 1-for-4 – he would be 0.9 hits above average the next night. That would mean that a .250 hitter turns into a .475 hitter after a 3-for-4 game!

Obviously, if you really did the experiment of computing game-to-game correlations, you’d get a very small number. I’m guessing, but, for the sake of argument, let’s say it might be 0.04.

Now, these two correlations are measuring the same ability – hitting for average. But because of context, an 0.45 can be pretty high in the season case, but earth-shattering in the game case. Conversely, 0.04 is meaningful in the game case, but, in the season case, it would show that batting average is barely a repeatable skill at all.

It all depends on context.

I mention this because of a blog entry on the "Wages of Wins" website. There, David J. Berri compares his book’s quarterback ranking to various versions of more sophisticated stats from Football Outsiders. He finds that the correlations are 90%, 92%, and 95% respectively.

And so he writes, "this exercise reveals that there is a great deal of consistency between the Football Outsiders metrics and the metrics we report in The Wages of Wins."

With which I disagree. The interpretation of correlation coefficient depends, again, upon the context. If you were completely ignorant about football statistics, then, yes, a 90% correlation would indicate that you’re measuring roughly the same thing. But given the vast amount of sabermetric knowledge we have about football, 90% could mean the statistics are very different at the margins of knowledge.

For instance, I’d bet that, in baseball, Total Average and Runs Created might correlate on the order of 90%. But, given our knowledge of baseball, we know that Total Average is unsatisfactory in many ways, and the differences are significant at the level of detail that we need for future research. 90% is enough to put Babe Ruth on the top and Mario Mendoza on the bottom. But it’s not good enough to tell the productive base stealers from the unproductive, or give us reliable information about the relative value of hits, or even to distinguish the 55th percentile player from the 45th.

To sum up: in one example, a 0.45 correlation was huge; in another example, a 0.9 correlation was mediocre. If your analysis starts and stops with the correlation coefficient, you really haven’t proven anything at all.


----

More posts on r and/or r-squared:

The regression equation vs. r-squared

Still more on r-squared

Why r-squared doesn't tell you much, revisited

R-squared abuse

"The Wages of Wins" on r and r-squared







Labels: ,

Friday, May 08, 2009

Why r-squared doesn't tell you much, revisited

In a blog post I wrote about yesterday, "Wages of Wins" author Stacey Brook ran a regression to try to figure out what kind of relationship there is between an NBA team's payroll and its success on the court.

The regression gives you several pieces of information. Which ones should you use to best explain the relationship?

Brook says it's the r-squared. He writes,

"We use R2 since we are interested in the proportion of variance that is in common between NBA team payroll and NBA team performance."


But is that truly what we're interested in? I don't think so.

I do agree with Brook when he says that R-squared gives you "the proportion of variance that is in common between NBA team payroll and NBA team performance." But what does that mean? Almost nothing, unless you're a statistician.

When you do research like this, there's a question that you want to answer. In this case, if your question is "what proportion of variance is in common between NBA team payroll and NBA team performance?," well, then, there's your answer. But that's not the question. It's not even Brook's real question. His real question is implied by the first paragraph of his post:

"I have to disagree that NBA (or for that matter NHL, MLB or NFL) teams that have high payrolls result in higher winning percentages; nor am I the first to say this."


The question is: do teams with higher payrolls do better on the court? And that question is different from "what proportion of variance is in common between NBA team payroll and NBA team performance?"

If you want to see what payroll does to performance, what you want to see is the regression equation. The way regression works, of course, is to plot all the datapoints on a graph, then draw the best fit straight line among those points. That line represents the best-fit relationship between payroll and wins.

If you do that for the 2008-09 NBA teams, you get

Wins = 0.61 (millions of $ spent) - 0.76

This, basically, answers your question, in several ways

-- every extra million dollars you spend on salaries gives you three-fifths of a win.
-- every extra $1.64 million you spend gives you an extra win.
-- if you spend $100 million, like the Knicks, you should win about 60 games.
-- if you spend only $45 million, like the Grizzlies, you should win only about 27 games.

Not that complicated, right? If you want to know about the direct relationship between salary and wins, the regression equation does it.

Of course, you want to check the statistical significance; it's possible that while the best-fit straight line says $1.64 million per win, that might not be significantly different from zero. (As it turns out, it IS significant, at the 99.5% level. In fairness to Brook, it appears his data source had incorrect information, and because of that, his results were not, in fact, significant.)

I think we can all agree, from these results, that it certainly does appear that spending leads to winning. When the highest-spending team is expected to go 60-22, and the lowest-spending team is expected to go 27-55, you can't really claim that payroll is irrelevant. (Again, in fairness to Brook, he didn't get results this extreme. With the incorrect data, the regression suggests the highest-spending team should only be 45-37.)

So if the regression equation is the gold standard for making these kinds of calculations, what's with the r-squared? Well, the r-squared answers a different question.

Let's suppose that you had no idea what makes teams win basketball games. You see the Cavs go 66-16, and you see the Clippers go 19-63, and you think, what causes the difference?

What you could do is list as many plausible things as you could think of. Payroll would be one of them. Maybe average days of rest. Maybe whether they're an offensive or defensive team. Maybe average age. Maybe pace of play. Just list them all, as many as you want. Then, run a regression, and look at the r-squared.

What the r-squared will do is tell you, in a certain mathematical sense, after correcting for all those variables, what percentage of all the variation in wins have you explained? What you're trying to do is get as close to 100% as you can. The closer you get, the more you've explained what makes teams win and what makes teams lose. Maybe, if you actually ran this regression, you'd get to something like 40%. If you adjusted team wins for all those variables, as best you could, your variance would decrease by 40%.

In this particular case, our regression didn't include all that other stuff, like pace of play or average age. We only had one variable, payroll. And it turned out that the r-squared was .256, which means that 25.6% of the variation is "explained" by payroll.

It doesn't sound like a lot. In "The Wages of Wins," Brook (and co-authors David Berri and Martin Schmidt) did that for MLB, and came up with only 18%. That doesn't sound like a very big number either, and those authors decide that means that payroll isn't very important.

But that doesn't follow.

The r-squared, the seemingly-low 25.6% number, does NOT tell you about the relationship between payroll and wins. It just tells you that payroll is 25.6% of the total variance, and other factors are 74.4%. But, if the total variance is large, 25.6% of it would be substantial.

When you go into the car dealership and ask for a price, you want the amount in dollars. If you ask "how much for that Camry," and the salesman says, "it's 700% of your monthly pay," it may sound like a lot. If he says, "it's 9.5% of your net worth," it may sound cheaper. And if he says, "it's less than 0.01% of Bill Gates' disposable income for the week," it may sound cheaper still. But those all represent the same number of dollars. The fact that one percentage is a large number, and one percentage is a small number, doesn't change that fact.

It's the same thing for r-squared. The size of the percentage number depends what it's a percentage of -- which happens to be the total variance of wins in the league. Do you know, intuitively, what that variance is? I don't. But I know that a lot of it is random chance. And random variation depends on sample size. You could have exactly the same relationship between salary and wins, but, in one case, the r-squared is .25, and in another case, it's .04, and in another case, it's .5.

I wrote before about one example of how that can happen. But I can do another right now.


Want to see how you can use the same data to get a larger r-squared? Easy. I'm going to take the actual data for the 30 teams, but group them into threes according to payroll. So instead of the three data points "$100 million, 32 wins" (Knicks), "$90.1 million, 66 wins" (Cavs), and "$86 million, 50 wins" (Mavericks), I'm going to add them all up into the one data point "$276.1 million, 148 wins". Then I'm going to repeat for the other 27 teams, until I have 10 sums of three teams. Then, I'm going to run a regression on those 10 data points.

What happens? The r-squared now goes up to .497 -- almost double what it was!

But while I was able to arbitrarily double the r-squared, the regression line stayed almost the same -- which makes sense, since the actual relationship between salary and wins shouldn't change just because we arranged the data differently. Using all 30 teams, we got 0.61 wins per million dollars. Using the 10 groups of three teams, we get 0.68 wins per million dollars. Pretty close.

Here, let me give you everything in one place:

30 teams.... r-squared = 0.256
10 groups... r-squared = 0.497

30 teams.... Wins = 0.61 ($millions) - 0.76
10 groups... Wins = 0.68 ($millions) - 5.5


If Stacey Brook did the analysis his way, using all 30 teams, he'd say "salary explains 25.6% of the variance in wins." If I do the analysis my way, using groups of three teams, I'd say "salary explains 49.7% of the variance in wins." Which one of us would be right? Both of us! Because we are using different denominators, different variances. The same Toyota Camry can be a smaller percentage of Brook's salary than of my salary, because our salaries are different.

And so saying "payroll explains 25.6% of the variance of wins" is like saying "a Camry costs 35% of salary." Whose salary, and how much does he earn? Unless you know that, the "35%" figure is useless.

But, again, despite the fact that Brook and I did our regression differently, the equation should come out very similar. It won't come out exactly the same, because of random fluctuation, but you should *expect* it to come out the same, in the same sense as you expect a coin to come up heads 50% of the time. 0.61 wins per $million and 0.68 wins per $million are pretty close.

The regression equation is meaningful, it requires less information to interpret, and its expected value is the same regardless of your sample size. Most importantly, it answers the exact question that you want to know.

The r-squared, on the other hand, is unintuitive, can be made to come out to almost anything you like by tweaking the sample size to get a different total variance, and requires you to know how the study was done in order to interpret what it means. In terms of answering real-life questions, it's not very useful at all.


Labels: , , , , ,

Tuesday, October 23, 2012

Yet another r vs. r-squared explanation

There's an election with one million voters, who randomly choose between Party A and Party B.  The margin of victory is important, not just who got the most votes.

If you run a regression to predict the margin of victory, based on a single vote, what will the r-squared be?  It will be 1/1,000,000.  That's because, if you knew all the votes, you'd know the margin of victory perfectly and the r-squared would be 1.  Since no vote is more important than any other, each must be equal in r-squared.  And the r-squareds have to add up, since the voters are independent.  So, 1/1,000,000 is the answer.

But, here's a different question: what is the impact of one vote on the margin of victory?  Well, that margin will be fairly small -- if you do the calculation, the SD of the difference between A and B will be 1,000 votes.  So, a single vote will be 1/1000 of the margin.  Not one in a million, but one in a thousand.

It's not that hard to see why.  When a million people vote, their votes will mostly cancel out, since they're all choosing randomly.  We know that they cancel out by the square root of the number of votes, since SD goes down by the square root of sample size.  So they'll cancel out to a difference of 1,000 votes.  John Smith becomes one vote in a margin of 1,000, instead of 1 vote in 1,000,000. 

That's the r: 1/1,000.  It's the square root of the r-squared of 1/1,000,000.

Effectively, a single vote sticks out higher because everyone else cancels out.

-----

This, I think, is a good analogy to visualize the difference between r-squared and r:

-- r-squared tells you how important the factor is relative to all the other factors.


-- r tells you how important the factor is relative *to the size of the final outcome*.

The size of the outcome -- in the sense of the difference from the mean -- is the square root of the size of the number of independent "others", which is why this works out.

-----

This appears to lead to a contradiction: if there are a million voters, and they're all equal, how can they ALL be 1/1000 of the outcome?  That would add up to one thousand outcomes!

But that's OK.  Remember, we're talking about the *size* of the outcome, not the responsibility for it.  If all million voters voted for A, the outcome would have been a 1,000,000 vote margin, instead of just 1,000.  So, all the voters combined DO add up to 1,000 outcomes -- by size.

If you don't like that, here's a non-statistical analogy.  Suppose there's an election where party A wins by one vote.  Whose vote tipped the balance?  Everyone's!  That is, everyone who voted for party A.  There might have been 500,001 votes for A, and 500,000 votes for B.  If *any* of the A voters had voted the other way instead, B would have won.

That is: 500,001 voters can say that *they* made 100% of the difference in the election -- and they'd all be right!  That is, it's perfectly OK that the sum of the voters' effects add up to a huge number.  It's an illusion that it seems they shouldn't be able to.

-----

Moving to a sports example ... let's go back to payroll vs. wins in baseball.  Suppose you do a regression, like the ones in this Freakonomics post, and you find that the r-squared equals .1, like it was in 2008. 

What that means is: payroll explained about 1/10 of the variance of wins.  That means that, in a sense, there could be 9 other factors that are just as important as payroll.  (Or, one factor that's "nine times" as important.  Or one factor "five times" as important, and two other factors "two times" as important, or some combination like that.) 

That is: payroll gets "one vote out of 10" in determining wins.

OK, fair enough.

But: those 9 other factors can be treated as independent and random (as an assumption of the regression -- and, also, if they were correlated to payroll, the regression would lump that in with payroll).  Therefore, they mostly cancel each other out, down to their square root.  So the SD of the other 9 factors is only 3 times the SD of payroll.

If you add payroll back in as the 10th factor, you get that the SD of the total is the square root of 10 times the SD of payroll (the square root of 3 squared from the other factors, plus 1 squared from payroll).  That's around 3.1.

If payroll represents a single vote, the margin of victory is 3.1 votes.  So, salary influences wins not by 1/10 (0.1), but by 1/3.1 (0.32), which is the square root.

Which is why we say, if you increase your payroll by 1 SD, you increase your wins by 0.32 SD.  If you move one inch up or down the normal curve for payroll, you'll move 0.32 inches up or down the normal curve for wins.

That's fairly large.

-----

The moral of the story is:

1.  The r-squared tells you what percentage of the *votes* you got.

2.  The r tells you what percentage of the *result* was because of you.

For a cause-and-effect relationship, like payroll and wins, you almost always want number 2.


------


(I've written about r and r-squared numerous times in the past, such as here.)



Labels: , ,