Tuesday, November 23, 2010

Explaining the sibling study

I promised to revisit Frank Sulloway and Richie Zweigenhaft's sibling study in light of their response to me last week. And I thought it might be best if I start over from scratch.

Today, I'm going to try to just *describe* the study and the results. In a future post, I'll add my own opinions. For now, my presumption is that on everything in this post, the authors and I would agree.

In the original study, the authors' primary finding was that younger baseball players attempt more stolen bases (per opportunity) than their older brothers, with an odds ratio of 10.58. If I explain this right, by the end of this post, it should be clear what the authors did and what the 10.58 figure actually means. I've been a bit unclear on it in the past, but I think I've got it now.

------

The purpose of the study was to check whether younger brothers attempt more stolen bases than older brothers. The authors hypothesized that would be the case. That's because the psychology literature on siblings finds that younger siblings tend to develop a more risk-taking personality than their older siblings, and stolen base attempts are the baseball manifestation of taking risks.

The authors found approximately 95 pairs of siblings for their study. When they compared the younger player to the older in each pair, they found support for their hypothesis. It turned out that the younger brother "beat" the older (in the sense of trying more SB attempts per time on base) 58 times, and the older "beat" the younger only 37 times. That's a .610 winning percentage for the younger brother.

Younger brother: 58-37 (.610)

By my calculation, that's statistically significant at the 5% level (2.15 SDs from .500).

However, for reasons to become clearer later, the authors preferred to give the results in terms of odds. Stated that way, the younger brothers had odds of 58:37, which is odds of 1.57:1.

Younger brother: 1.57:1

However, that's still not quite the way the authors chose to describe results like this one. In both the paper and the response, the authors quote the "odds ratio" of dividing the younger player's odds by the older player's odds. The younger brother is 1.57:1, and the older brother is obviously the reverse, 1:1.57. If you divide the first odds (1.57/1) by the second (1/1.57), you get 1.57 squared, or 2.46. That means:

The "odds ratio" of younger brother to older brother is 2.46.

This 2.46 does not appear in the actual paper; I'm doing it here for comparison later.

------

However, the authors argue, the numbers above are misleading. They don't adequately represent the true effects of being an older or younger brother. That's because, they argue, there is a confounding effect not controlled for -- which player got called up first.

The authors and I disagree on whether this is an appropriate control, but both of us agree that controls are often useful and revealing.

Take an unrelated example. Suppose you did a study, and you found that there was an equal proportion of Canadians and Brazilians who were world-class hockey players. Would you be able to conclude that Canadians and Bermudans are generally equal in hockey talent?

No, you wouldn't, because you'd see that Canadians tend to get a lot more practice time than Brazilians. Canadians live in a cold climate, with lots of frozen ponds. And, Canada has a lot more ice rinks per capita than Brazil does. That means Canadians should get a lot more practice than Brazilians.

Given those facts, you'd expect a lot more hockey players from Canada than Brazil. If you find equal numbers, that's evidence that Brazil must have a higher aptitude for hockey than Canadians, to be able to succeed equally despite fewer opportunities to practice.


More formally, you might control for the number of ice rinks that each group has access to. You'd find that *holding ice rinks constant*, Brazilians are a lot better at hockey than Canadians are. I won't do a numerical example, but you can probably see how this would work.

So, going back to baseball siblings: the authors argue that "getting called up first" is this study's equivalent to "having lots of ice rinks". If you control for that, the odds ratio gets a lot bigger.

To prove that, they took 80 pairs of siblings where one got called up before the other, and they split them up according to whether it was the older or the younger who got called up first. The results were:

-- When the younger brother got called up first, the brother called up first went 5:1.

-- When the *older* brother got called up first, the brother called up first went 32:42.

The authors now calculate the odds ratio as it stands with this control. They divide (5/1) by (32/42), and get 6.56.

What they're saying is something like this: "The raw data makes it look like the odds ratio is 2.46. However, that's because we didn't control for callup order, which is important, just as important as the "access to hockey rinks" control. If we do that, we see that the raw data understates the true odds ratio, which is 6.56."

------

Putting in the control for callup order is like adding a variable to a regression. Whenever you do that, you check the apparent size of the effect, and also the significance level.

As we just saw, the effect size is fairly large: we went from an odds ratio of 2.46, without the control, to 6.56 with the control.

But what about the significance level? As it turns out, the control variable turns out to be not significant.

My simple argument goes like this: the six pairs in the "younger called up first" control went 5-1, or .833. If the control was not significant, we'd expect them to go .588, like in the sample as a whole. The chance of a .588 team going 5-1 or better is more than 21%, far higher than the 5% required for significance.

The evidence from the authors' paper goes something like this: for their entire sample (which we'll talk about a bit later), the authors wound up with an odds ratio of 10.58 (instead of 6.56). On page 13 of their response, they report a 95% confidence interval of (2.21, 50.73). That is easily wide enough to include the odds ratio that would have resulted if the control had no effect, which is 2.46 for our sample, and a bit higher for the authors' larger sample.

The authors never explicitly say that the "called up first" variable is not statistically significant, but it seems clear that that's the case.

-----

Now: what does an odds ratio actually *mean*? In our example above, where we came up with an odds ratio of 6.56, what does that 6.56 mean? How do we use it in an English sentence?

The answer, I think, is this:

The Vegas odds if you bet on the older brother are 6.56 times higher than the Vegas odds if you bet on the younger brother.

If you bet on the older brother attempting more steals than the younger brother when the older brother is called up first, your odds are 42:32, which is 1.3125 to 1. If you bet on the younger brother attempting more steals than the older brother when the younger brother is called up first, you get 1:5, which is 0.2 to 1. Divide 1.3125 by 0.2, and you get 6.56.

The numbers in the odds are 6.56 times higher, but maybe it's easier to understand that your *winnings* are also 6.56 times higher. Let's check:

If you bet on the older brother, you'll get odds of 42:32. If you bet $32, and the older brother steals more, you'll make a profit of $42.

If you bet the same $32 on the younger brother, you'll get odds of 1:5. That means if the younger brother steals more, you'll make a profit of $6.40.

Divide $42 by $6.40, and you get ... 6.56, as expected.

------

So that's what odds ratio means in terms of betting. Are there other intuitive ways of explaining it?

The authors don't really give any, but they do agree that it's difficult to interpret. They write,


"Ironically, although odds ratios are often used in an attempt to clarify complex statistical findings, people who are not familiar with them sometimes misinterpret what odds ratios do and do not mean.... [A]n odds ratio of 10.58 ... does not mean that younger brothers attempted 10.58 times the number of steals as did their older brothers. Similarly, this statistic does not mean, as Schwarz mistakenly reported in the New York Times, that more than 90 percent of younger brothers attempted more steals per opportunity than their own older brothers."


The authors write what it *doesn't* mean, but not what it *does* mean. Let me try to tackle that now, in a couple of different ways.

------

The most important thing is that the odds ratio of 6.56 doesn't tell you anything about how much the younger brothers outsteal the older brothers, or vice versa. It only tells you how that ratio *changes* when you swap "younger" for "older".

Again, going back to the actual numbers, which I'll repeat from a few paragraphs ago:

-- When the younger brother got called up first, the brother called up first went 5:1.

-- When the *older* brother got called up first, the brother called up first went 32:42.

What are the odds the younger brother beats the older brother? Well, if the younger brother got called up first, the odds are 5:1. If the older brother got called up first, the odds are 42:32 (1.31:1). So the younger brother occasionally wins 5:1 (6 out of 80 times), and frequently wins 42:32 (74 out of 80 times). Neither of those numbers, alone, has anything to do with 6.56 to 1.

What are the odds the brother called up first beats the other brother? Well, if the younger brother got called up first, the odds are 5:1. If the older brother got called up first, the odds are 32:42 (0.76:1). So the first-callup brother occasionally wins 5:1 (6 out of 80 times), and frequently wins only 32:42 (74 out of 80 times). Again, neither of those numbers, alone, has anything to do with 6.56:1.

The 6.56 only comes in when you divide the 5:1 by the 0.76:1. It's the *difference* between betting on the younger brother and betting on the older brother.

Look at this sentence:

"The _________ brother was called up first. The odds that he beats his sibling are _____:1."

There are two ways to fill in this sentence:

"The younger brother was called up first. The odds that he beats his sibling are 5:1."

"The older brother was called up first. The odds that he beats his sibling are 0.76:1."

What the 6.56 is saying is that, if you switch the word "older" and "younger", you have to divide or multiply the odds number by 6.56 for the sentence to still be true.

That's regardless of what the actual odds are. If, instead of 5:1 and 0.76:1, the odds turned out to be 10:1 and 65.6:1, the odds ratio would *still* be 6.56. Again, the 6.56 represents the *difference* in how the odds change, not what the odds actually are.

------

Odds and odds ratios are used a lot in regressions that try to predict probabilities. If you're trying to figure out how various factors affect cancer survival, for instance, you'll probably use odds. Why? Because, if you use probabilities, you'll usually wind up with something greater than 1, which doesn't make sense.

Suppose you test a new treatment for cancer. If you figure that it doubles your chances of a cure, you're in trouble. Because, suppose a patient presents with a 60% chance of surviving without the new treatment. Do you really want to say that, with the treatment, his chances go up to 120%?

To solve that problem, statisticians use odds, instead. The 60% patient is actually 3:2. Suppose the treatment triples the odds. Now, he's 9:2 to survive, instead of 3:2. That's perfectly OK. 3:2 is 60%, and 9:2 is 91%. Nothing over 100%.

When there are two factors, the usual assumption is that you can just multiply out the odds. If chemotherapy doubles the odds, and surgery triples the odds, then the model will say that if you get both treatments, you multiply your odds by 6. (This is the idea behind logit regression -- you just take the log of the odds so you can add linearly, like a regular regression.)

This is done all the time in statistics, but whether it's actually how nature works, I'm not sure. I can't think of an intuitive reason why it *would* work. (And I could easily make up an example where it doesn't.)

Generally, I believe it works well when the probabilities are really small, like when the odds double from 2,000,000:1 to 1,000,000:1 -- but not when they're bigger, like when they go from 1:4 to 2:4. In any case, the technique is used all the time, so there may be something to it that I don't see, and so I certainly won't argue against it.

------

So, in the baseball context, I think the implication of quoting an odds ratio of 6.56 goes something like this:

Suppose you look at two siblings, and evaluate their baseball skills and psychology from head to toe.

Brother A is a little fatter and slower than brother B. And you figure that, if Brother A is older than brother B, and called up first, he has only a 20% chance of attempting more steals than brother B.

But then, you ask, what if Brother A is *younger* than brother B, and still called up first? Well, you start with the original 20% probability, and convert it to odds, giving you 1:4. Then, you multiply by the odds ratio of 6.56. That gives you new odds of 6.56:4, which works out to 62%.

The conclusion: A may have a 20% chance to beat a *younger* brother called up first, but he'd have a 62% chance to beat an *older* brother called up first.

If you want, you can start with something other than 20%. Suppose if A is older, he's 1:1 to beat his brother. Then, if A is younger, he's 6.56:1. The probability goes from 50% to 87%.

Or, suppose A is so much faster than even when he's older, he's 9:1. Then, if he's younger, he'd be 59:1. So he'd go from 90% to 98.3%.

Or, just use the example from the actual data. If A is older, he's 32:42. If he's younger, multiply that by 6.56, and you get that he's now 5:1.

That's what the 6.56 implies, in brothers-stealing terms.

As an aside, I know I promised no criticism, but here's just one quick thing. In their second paper, the authors write,

"For brothers called up to the major leagues first, a younger brother was 6.56 times more likely than an older brother to have a higher rate of attempted steals."


I think they didn't mean to say that. While the younger player has 6.56 times the odds, he's not 6.56 times more likely. He's only 1.93 times more likely -- from 83% (5:1) to 43% (32:42). The authors do argue elsewhere that odds ratios have to be interpreted carefully, so I think this was just a slip of the tongue.

------

There's nothing in the method that tells you *how* the 6.56 change happens. Maybe if A is the younger brother, he keeps himself in better shape, so he's not as fat as if he were the older brother. Maybe first siblings tend to be fatter and slower than second siblings for genetic reasons, and so it's because when A is older than B, he's more likely to have been born fatter, and B more likely to have been born skinnier.

Or, maybe the authors' hypothesis is correct: If A is the younger brother, his personality develops him into a bigger risk-taker. If the authors are correct that the personality issue is the primary driver of the 6.56, the implications are that even if A is fatter than B, the personality factor alone is enough to more than compensate, since it can bring him from 20% to 62%, or from 43% to 83%, for psychological reasons alone.

Also, the study's raw data came up with the effect that being the younger brother instead of the older brother increases your chances from 43% to 83%. It does NOT speculate on what happens if the other player you're being compared to is not related to you at all. The study compares "has an older brother" to "has a younger brother", but ignores the "doesn't have a brother at all" case.

Furthermore, the study doesn't consider a pair of siblings unless both of them make the major leagues. But, the authors hypothesize that the risk-taking strategy is adopted in childhood. If that's the case, you'd think that it doesn't matter if one of the brothers doesn't make it past the low minor leagues; the effect should still be there. Probably, even if one brother never made it out of little league, the effect should still exist.

That would make an interesting follow-up study.

------

I've been talking about an odds ratio of 6.56, but the authors' conclusion is an even higher odds ratio, 10.58. How do they get the 10.58?

It's a combination of the 6.56, and two other cases. One other case is where the brothers got called up the same year (giving an odds ratio of 7.00). The third case is almost exactly the same as the first case, but slightly different because of the way the authors dealt with families of more than two siblings. For that third case, they wound up with a "younger brother called up first" ratio of 5:0, instead of 5:1. That works out to "infinity".

So the authors combined the three odds ratios:

6.56
7.00
infinity

How do you combine an "infinity?" Well, they used the "Mantel-Haenzel common odds ratio" technique, which is able to handle ratios with a zero denominator. They wound up with an overall odds ratio of 10.58.

So that's where the "10 times" in the original NYT article comes from, and that's where the "10.58" in the paper and the response comes from.

------

There you go. As I said, I think it's right this time. Let me know if you find anything that needs correcting. If all is OK, I'll eventually prepare a response to the authors' response, outlining more clearly why I disagree with some of their conclusions.



Labels: , ,

Tuesday, November 16, 2010

Do younger brothers steal more bases than older brothers? Part VI

Note: I'm not happy with this post. I've already revised it twice because I found things that were wrong ... as it is now, I think it's right, but it's not very focused and I think some of the emphasis might be on the wrong issues.

So, just a warning that I plan to redo it soon.

-----------

Although this is "Part 6" of the discussion of the sibling study, I'm going to try to make it stand alone, so if you're coming to this thread for the first time, read on.

A few months ago, Frank Sulloway and Richie Zweigenhaft published a study on brothers (siblings) in baseball. They came to the conclusion that a younger brother is about 10 times as likely to attempt more stolen bases in his career (adjusted for opportunities) than his older brother. Ten times is a LOT.

After I read the paper, I believed the result was incorrect. My previous posts on the subject explained why. Following that, I had an e-mail conversation with one of the authors. I don't believe either of us was able to convince the other of the rightness of our respective positions.

Two weeks ago, the authors released a second paper, which attempted to clarify the arguments and address some of my points. I remain unconvinced. And, indeed, I think I've been able to come up with a better, more easily understood argument that explains why.

-----

The authors' study comprised approximately 95 sets of siblings. In those 95 pairs,

58 times the younger brother attempted more steals
37 times the older brother attempted more steals.

I think the authors and I would agree on this (although my numbers might be off by one or two because of the way the authors handled cases where there were more than two brothers, like the Alous).

So the younger brothers had a 58-37 record against their younger siblings, which works out to .610. The SD of winning percentage in 95 games is about .051. So .610 is a little more than 2 standard deviations above .500. That's statistically significant at the 5% level.

I believe this is a legitimate finding, and if the authors had left it at that, we'd have no disagreement.

Another thing they did is to express the result as odds instead of a winning percentage. If you divide 58 by 37, you get 1.57. So you can say something like,

"Younger brothers had odds of 1.57 to 1 of beating their older brothers in steal attempts."

That's perfectly accurate. Again, we have no disagreement.

-----

What the authors did next is where we start to disagree. They took the 95 cases, and split them up into groups, based on the order in which the brothers were called up to the major leagues. To keep things simple, I'm going to leave out the case where the brothers were called up the same season, and concentrate on the case where the brothers were called up in different seasons.

That leaves 80 cases, in which the younger brother's record was 47-33. That's an odds ratio of 1.42 (47 divided by 33). But that's not what the authors come up with.

Because, look what happens when you split them up based on who was called up first:

-- When the older brother was called up first, the brother called up first went 32-42 against the other brother.

-- When the younger brother was called up first, the brother called up first went 5-1 against the other brother.

Converting that to odds:

-- When the older brother was called up first, the brother called up first had odds of .76:1 (32:42).

-- When the younger brother was called up first, the brother called up first had odds of 5.00:1 (5:1).

Now, to compare the younger to the older, the authors divide 5.00 by .76 and get an "odds ratio" of 6.56.

As it turns out, the authors an even more extreme result, 10.58 instead of 6.56. Why? Mainly (and I'm simplifying here) because the 5-1 can also be interpreted as 5-0, depending on how you handle families with more than two siblings. 5-0 is an odds ratio of infinity. The authors use a mathematical technique to average the "infinity," the "6.58", and the result for when the siblings get called up the same year ("7.0"), and it works out to 10.58.

Which is where the authors get their statement,

"It may be seen that the common odds ratio is 10.58, as previously reported [in our original paper]."


That sentence, actually, is pretty much correct.

-----

So if the sentence is correct, what's the problem? The problem is the authors' interpretation of what an odds ratio means. Remember that odds ratio of 6.56 above? That's the correct number. But the authors write,

"For brothers called up to the major leagues first, a younger brother was 6.56 times more likely than an older brother to have a higher rate of attempted steals."


That is not true. That is not what odds ratio means.

Let's suppose you have 100 younger brothers called up first, and 100 older brothers called up first. 6.56 times more likely implies that you'll have 6.56 times as many "wins" in the first group as in the second group. But that's not the case. The younger brothers would have a ratio of 5:1, which, for 100 trials, is 83-17. The older brothers would have a ratio of 32:42, which, for 100 trials, is 43-57.

That means that the likelihood of winning rises from 43 (out of 100) to 83. That's 1.93 times more likely, not 6.56 times more likely.

The "1.93" figure is called the "relative risk". Relative risk is not the same thing as odds ratio.

So if "6.56 times more likely" is not the correct interpretation of an odds ratio of 6.56, what IS the correct interpretation? It's this:

An odds ratio of 6.56 means that if you place a $100 bet on the less likely outcome, your potential winnings will be 6.56 times as high as if you bet $100 on the more likely outcome.

Specifically for this case: If you bet $100 on the 5:1 favorite, you'll win $20. If you bet $100 on the 32:42 underdog, you'll win $131.25. And, $131.25 divided by $20 is 6.56.

*That* is what the odds ratio really means. You can decide how intuitively meaningful it is. It probably means more to you if you're a sports bettor than if you're not.

-----

So why is that a problem? Isn't it a real sense of what the 6.56 (or 10.58) figure actually means? Why, then, do I say it's misleading?

Because it exaggerates the scale of the effect. Roughly, it squares it.

Suppose home field advantage is 2:1 -- the home team has twice the chance of winning as the visiting team. That means that, in turn, the visiting team has half the chance of winning, which is an odds ratio of 0.5:1.

If I do what the authors did, and divide the home team odds by the visiting team odds, I get 2 divided by 0.5, which is 4. But I cannot say, "a home team is 4 times more likely to win than a visiting team." That would be wrong: the correct odds are obviously 2:1. What I'm actually saying is, "if I bet $100 on the visiting team, I'll win 4 times as much money as if I bet on the home team."

Now, that's all well and good, but I would argue that the important measure is the 2:1, not the 4:1. We get the "4" by comparing the 2:1 favorite to the 2:1 underdog. In effect, the odds ratio is roughly "squaring the odds". Which makes sense: if you divide X by the reciprocal of X, you get X squared.

If you take the 6.56 odds ratio, and figure the square root, you get 2.56. That, I think is a reasonable guess at what the effect actually is.

Put another way: the 6.56 occurs when you switch the status of *two* players -- you make the young one get called up first, and you make the old one get called up last. How do you split up the effect between the two players? The most obvious way is to "give" them 2.56 each.

-----

Anyway, that's mostly semantics, and it's about the odds ratio, which is not the interesting question.

The interesting question, to me, is, how often will a younger brother have more steal attempts than his older brother, even controlling for callup order? The answer is nothing near 10.58.

Look at it this way:

-- if the younger player gets called up first, the odds are 5:1.
-- if the younger player gets called up last, the odds are 1.42:1.

Doesn't it follow that the younger player's odds have to be somewhere between 1.42 and 5? After all, the younger player is either called up first, or he's called up last. The best case is when he's called up first, and the odds are 5:1 that he'll beat his brother. So the *overall* odds of beating his brother can't be *more* than 5:1, right?

-----

Back to the odds ratio: if the authors agreed with me, and reverted to 3.25 instead of 10.58, would I believe it? Well, no, because of the confidence interval issue.

There were only 5 or 6 pairs of brothers in the "younger player gets called up first" group. They either went 5-0 or 5-1. The "3.25" figure is almost entirely based on that fact. If they had gone, say, 3-3, instead, the odds would work out to something like the 1.57 we got just by counting. (If you recall, the younger brothers went 58-37 overall, without dividing the sample into "first callup" and "last callup".)

Suppose the actual odds for the "younger called up first" were really the same as the "older called up first". Then, we'd have expected a .610 winning percentage.

The chance of a .610 team going 5-0 is 8.4%. The chance of a .610 team going 5-1 is 20%.

So the observed p-value is somewhere between .084 and .2 -- both higher than the .05 required for significance.

-----

The authors don't do any explicit significance testing, but they say their confidence interval for the 10.58 odds ratio is (2.21, 50.73).

Again, suppose the odds for both callup groups were actually 1.57 in favor of the younger brother. Then the odds ratio we'd observe would be 1.57 squared, which is 2.46.

The authors actually found a confidence interval of (2.21, 50.73). They did things a little differently, and used more data, but, overall, I'd say their confidence interval is pretty consistent with what we found above. We found "almost significant but not really", and the authors are close. I'm actually not sure if we did exactly what they did, if our null hypothesis would be inside their confidence interval or not, but it would probably be close either way.

-----

So my conclusions are:

1. A basic look at the overall data show younger players with odds of 1.57:1 to beat their older brothers in career steal attempts.

2. Dividing the data into "called up first" and "called up last" appears to increase the odds to somewhere between 1.42 and 5.00.

3. The authors' odds ratio of 10.58 does not easily translate into anything intuitive about the odds of one brother beating another, except for the difference in the amount you'd win if you bet.

4. The authors' odds ratio of 10.58 is not how most sabermetricians would express the effect. Going by the example of home field advantage, we'd be more likely to go with an odds ratio of 3.25.

5. In any case, the difference between the 3.25 and the 1.57 we would obtain (if there were no "callup first" effect) is not statistically significant.

6. As I have argued in previous posts, the "called up first / called up last" split is not an appropriate control, because it reverses cause and effect. (You can disagree with this point, if you choose, and the overall argument still holds.)

----

Bottom line: the data show that younger brothers attempt more steals than their older brothers at a statistically significant rate, with odds of 1.57 to 1. Isn't that interesting enough on its own?


-----

Note: this post is substantially revised. First post was 11/16 am. Took that down, reposted 11/16 pm. Revised again 11/17 am.

Vigorous hug (think Cournoyer/Henderson) to Tango for pointing out that the reported odds ratio is actually approximately the square of the true odds ratio. I hadn't realized that was what's going on until he pointed it out.










Labels: , ,

Wednesday, November 10, 2010

Do younger brothers steal more bases than older brothers? Part V -- the authors' response

This past summer, I posted a few times about Frank Sulloway and Richie Zweigenhaft's "sibling" study. The authors have now responded online.

Their response is here (.pdf).

I'm away on vacation, and don't have the original paper, or copies of the discussion I had with the authors via e-mail, so I'll just link the paper for now, in order to get it out there as soon as possible.

Thanks to Dr. Sulloway and Dr. Zweigenhaft for the discussion, and I'll probably comment further within a couple of weeks.



Labels: , ,

Friday, July 30, 2010

Do younger brothers steal more bases than older brothers? Part IV -- Age vs. SB

A few weeks ago, I wrote a series of posts on an academic paper about siblings and stolen bases. That study claimed that when brothers play in the major leagues, the younger brothers are much more likely -- by an odds ratio of 10 -- to attempt steals more often than their older brothers.

Since then, the authors, Frank J. Sulloway and Richard L. Zweigenhaft, were kind enough to write me to clarify the parts of their methodology I didn't fully understand. They also disagreed with me on a few points, and, on one of those, they're absolutely right.

Previously, I had written,

"It's obvious: if the a player gets called up before his brother *even though he's younger*, he's probably a much better player. In addition, speed is a talent more suited to younger players. So when it comes to attempted steals, you'd expect younger brothers called up early to decimate their older brothers in the steals department."


Seems logical, right? It turns out, however, that it's not true.

For every retired (didn't play in 2009) batter in my database for which I have career SB and CS numbers, I calculated his age as of December 31 of the year he was called up. I expected that players called up very young, like 20 or 21, would have a much higher career steal attempt rate than players who were called up older, like 25 or 26.

Not so. Here are career steal rates for various debut ages, weighted by career length, expressed in (SB+CS) per 200 (H+BB):

17 -- 5.1
18 -- 6.9
19 - 11.3
20 - 14.4
21 - 13.6
22 - 14.0
23 - 14.7
24 - 15.5
25 - 13.4
26 - 14.0
27 - 16.5
28 - 10.2
29 - 11.0
30 - 13.5
31 - 10.3
32 - 17.4
33 - 13.0
34 -- 3.9
35 - 10.5
36 -- 6.1
37 -- 7.4
38 -- 5.3
39 -- 8.4
40 -- 0.0
41 - 14.0

It's pretty flat from 20 to 27 ... there is indeed a dropoff at 28, but few players make their debuts at age 28 or later.

Why does this happen? Isn't it true that young players are faster than old players? Perhaps what's happening is that players who arrive in the major leagues earlier also play longer, which means their extra early high-steal years are balanced out by their extra later low-steal years. I'm not sure that's right, but it's a strong possibility. In any case, my assumption was off the mark, applying as it does only to age 28 and up.

I could have figured out that was the case had I looked at Bill James' rookie study from the 1987 Baseball Abstract. Near the end of page 58, Bill gave a similar chart for hits and stolen bases (but on the total number, not the rate). And it looks like SBs decay not much more than hits or games played.

For instance, consider a 22 year old player compared to one who's 25. The 22-year-old, according to Bill, will wind up with 88 percent more base hits than the 25-year-old (623 divided by 331, on Bill's scale). For stolen bases, the corresponding increase is 84 percent (613 to 334). The two numbers are pretty much the same -- which means, since 1987, we've known that career SB rates don't have a lot to do with callup age.

----

Anyway, the "odds ratio of 10" finding in the sibling study was based on individual player-to-player comparisons. So, I decided to test those. Suppose you have two players, but one breaks in to the majors at a younger age than the other. What is the chance that the younger callup attempts steals at a higher rate for his career?

To figure that out, I took the 5,742 batters in the study, and compared each one of them to each of the others. I ignored pairs where both players were called up at the same age, and I ignored pairs with the same attempt rate (usually zero).

The results: younger players "won" the steal competition at a 52.9% rate, with a "W-L" record of 7,387,525 wins and 6,569,412 losses.

Young: 7387525-6569412 .529

However: that includes a lot of "cup of coffee" players. If I limit the comparisons to where both players got at least 500 AB for their careers, then, unexpectedly, the older guys actually win:

Young: 1950250-1965161 .498

The difference between those two lines comprises cases when one or both players had a very short career. When that happened, the young guys kicked butt, relatively speaking:

Young: 5437275-4604251 .541

These numbers are important because they represent exactly what the authors of the sibling study did -- compare players directly. My argument was that I believed the younger player would be the "winner" a lot higher than 52.9% of the time. That's not correct. So that part of my argument is wrong, and I appreciate Frank Sulloway and Richie Zweigenhaft pointing that out to me.

Does that mean I now agree with the study's finding that the odds of a player having a higher attempt rate than his brother are 10 times as large when he's a younger sibling? No, I don't. But it does mean that I need to refine my argument, which I will do in a future post.



Labels: , , , ,

Friday, June 18, 2010

Do younger brothers steal more bases than older brothers? Part III

(Note: please read Part I and Part II to understand what I'm talking about here.)

OK, after rereading the Sulloway/Zweigenhaft paper, and digesting some comments from the previous posts, I think I now have a pretty good idea of what the authors did.

In the previous post, I compared all brothers to their siblings. I found that when the older brother was called up first, which was most of the time, the younger brother attempted more steals 54% of the time. When the *younger* brother was called up first, which was seldom, he attempted more steals 87% of the time.

That's because when the younger brother was called up first, he must have been quite a bit better than his older sibling, in order to make the major leagues at a much younger age. So 87% is reasonable for younger brothers who are so good that they get called up even before their older brothers. On the other hand, 54% is normal for the more common case where the older brother gets called up first.

Now, since the authors say explicitly say they ran a regression and controlled for who got called up first, let's do that ourselves.

Simplifying the numbers, let's say

60% outsteal their brothers when called up last and younger.
40% outsteal their brothers when called up first and older.
90% outsteal their brothers when called up first and younger.
10% outsteal their brothers when called up last and older.

Now, let's set that up as a regression. We'll create dummy variables for "called up first" and "older". And so each line of our regression, indpendent variable followed by dependent variables, is:

0.6 ... 0 0
0.4 ... 1 1
0.9 ... 1 0
0.1 ... 0 1

However, the paper used odds ratios instead of probabilities, so let's do that.

A probability of 0.6 corresponds to 0.6 successes to 0.4 failures. That's an odds ratio of (0.6 divided by 0.4), which is 1.5.

Similarly, the odds ratio of 0.4 is 0.667. The odds ratio of 0.9 is 9 (9 successes to 1 failure). And the odds ratio of 0.1 is 0.111.

Now we have

1.50 ... 0 0
0.67 ... 1 1
9.00 ... 1 0
0.11 ... 0 1

We'll do one more thing: The authors say they did a logarithmic tranformation, so we'll take the logarithm of the odds ratios. That makes sense: you'd expect odds to be multiplicative, not additive, and the log of the odds ratio is standard in this kind of regression. So let's do that. Now we have:

+0.405 ... 0 0
-0.405 ... 1 1
+2.197 ... 1 0
-2.197 ... 0 1

Because of the symmetry, we have to tweak the numbers just a little bit to avoid singularity. (I think I changed one of the "405s" to a "407", or something, and a "187" to a "195" or something.)

If we now run that regression, what happens? We get

log of the odds ratio = (1.79 * dummy for called up first) - (2.60 * dummy for being older) + 0.404

And that means that being older subtracts 2.60 from the log of the odds ratio. The antilog of 2.60 equals 13.5. And so, being younger means your odds ratio goes up by 13.5 times!

The authors of the study came up with 10.6 times, but their data was different from mine. I think that what I've done and what the study did are pretty much the same thing.

----

(One quick note: the odds ratio of 13.5 does NOT mean that your chance is 13.5 times higher. It means that your ODDS are 13.5 times higher. Suppose your original odds were 2:1 in your favor, which is a probability of .667. Now, your odds multiply by 13.5 times, which is 27:1. That's a probability of .964. Not 13.5 times the chance -- 13.5 times the odds. Obviously, a .964 shot is not 13.5 times the chance of a .667 shot.

Where the original study and NYT article say "10 times as likely", they really mean "10 times the odds" in this narrow sense. The authors use "X times more likely" throughout the paper, when I think they mean "has X times the odds ratio".)

----

Going back to the 13.5 odds ratio: why is it wrong? Well, it isn't. It's right!

Suppose that you're an older player who got called up first. Your odds of attempting more steals are 0.667 to 1 (40%). Now, suppose you're magically turned into the younger player, holding everything else equal (so you still got called up first). Now your odds of attempting more steals are multiplied by 13.5. 0.667 times 13.5 equals 9. Your odds are 9 to 1, or 90%, just as we assumed at the beginning!

And, suppose that you're the older player who got called up last. Your odds of attempting more steals are 0.111 to 1 (10%). Now, suppose you're magically turned into the younger player, holding everything else equal (so you still got called up last). Now your odds of attempting more steals are multiplied by 13.5. 0.111 times 13.5 equals 1.5. Your odds are 1.5 to 1, or 60%, just as we assumed at the beginning!

It actually works out!

So does that really mean that younger players DO have 13.5 times (or 10.6 times, as the real study found) the odds of outstealing their older brothers? Yes, in the regression, but not in real life.

In real life, whether you get called up first doesn't really affect your steals directly. Being called up first matters because it's a proxy for ability. If you're younger and called up first, you're likely to be a great ballplayer. Medium-skill slow ballplayers might get called up at 23 or 24, but not at 20. Early callups are reserved for the most excellent players, who are also likely to be fast.

So, if you change old to young, and "keep everything constant," you're really NOT keeping everything constant. By keeping "called up first" constant, but changing older to younger, you're actually CHANGING "of roughly average ability" (old player called up first) to "of much better than average ability" (young player called up first). It's a little sleight of hand, using a proxy variable that means different things depending on the other variable. Younger brothers steal at 13.5 times the odds of old players only if, at the same time, they are much, much better players than the old players.

Analogy again: suppose I run this new regression

Amount of money = number of $5 bills in wallet + dummy for whether number of $1000 bills equals number of $5 bills

Suppose you have no $5 bills and the dummy is "yes", meaning you have no $1000 bills either. Your estimated wealth is therefore $0.

But now, if your inventory of $5 bills increases by 1, *holding everything else constant*, the regression will tell you that $5 bill is worth $1005! Because "holding everything else constant" means you "still" have the same number of thousands as of fives, which means you now have another $1000 bill! The "constant" refers only to holding the regression variable constant. It isn't holding the *real-life* variable constant, which is the number of $1000 bills. What's hidden is that in order to hold the regression variable constant, I have to give you $1000.

Same thing happening here. If you hold "called up first" constant, while changing "older brother" to "younger brother", you're not really holding things constant in the real life sense. To hold the regression variable "called up first" constant, you have to give the younger brother a hidden $1000 worth of talent. And that's why the odds ratio turns out so high.

Want another analogy? Suppose you predict whether a person is likely to be diagnosed a dwarf based on their height and age. Suppose you find that if you hold height constant at 4 feet tall, but increase the person's age from 8 to 28, he's now much more likely to be diagnosed a dwarf. Does that mean age causes dwarfism?

----

So, anyway, that's what's going on. I believe their regression does indeed find a coefficient equal to 10.6, as they report, but when you use a bit of common logic, you see that what they found is absolutely consistent with about 50 to 60% of younger brothers outstealing their older brothers -- not 90%.


Labels: , , , ,

Wednesday, June 16, 2010

Do younger brothers steal more bases than older brothers? Part II

A couple of weeks ago, I wrote about a study that purported to show huge differences in steal attempts between brothers in major league baseball. According to the New York Times article describing the study,

"For more than 90 percent of sibling pairs who had played in the major leagues throughout baseball’s long recorded history, including Joe and Dom DiMaggio and Cal and Billy Ripken, the younger brother (regardless of overall talent) tried to steal more often than his older brother."


If that's true, that would be huge. It would, I think, be one of the most surprising findings ever, that nine out of ten times the younger brother steals more than the older brother. But I checked, and found nothing close to that. I was left wondering how the authors, Frank J. Sulloway and Richard L. Zweigenhaft, got the results they did.

Since then, I managed to get hold of the actual study. And I still don't know where the 90% figure comes from.

Our raw numbers, it appears, are almost the same. The authors' results are different from mine, but only by a bit; they might have better data, or used different criteria for eliminating pairs from the study.

(Update: strikeouts below are where I misinterpreted some of the numbers in the authors' tables. The authors didn't actually give the numbers I thought they did.)

Let me start with steals. My numbers showed 56% of younger brothers outstealing their older brothers (adjusted for times on base). Their numbers actually showed only 48%. Pretty close. However, that's not the right comparison, because the authors are talking about steal *attempts* -- that is, SB+CS, not just SB. (I missed that the first time.)

Rerunning the numbers for attempts instead of steals, I get that 57 out of 98 younger brothers out-attempted their older brother, or 58%. What did the authors find? 97 out of 185, or 52%.

Putting that in table form:

58% me
52% them

So we're on the same page, right? I think so. It sure seems like their comparisons line up with mine.


So, then, why does the article say 90%? Because the authors, after showing 52% in their chart, change their estimate to 90% later.

Why? I can't figure it out for sure; I don't completely understand where they're coming from. The logic is flawed somewhere, but the authors don't explain themselves fully, so I can't actually spot where the flaw happens.

Let me show you what the authors give as the broad reason for bumping the 52% up to 90%: they say the sample is biased. Why?

"Call-up sequence and its relationship to athletic talent introduces a potential bias in athletic performance by birth order, as older brothers were more likely to be called up first owing to the difference in age. The extent of this effect turns out to be pronounced, with older brothers being 6.6 times more likely than their younger brothers to receive a call first to the major leagues. We are therefore comparing somewhat more talented athletes who were called up first (and who, by virtue of their relative age, tend to be older brothers) with somewhat less talented athletes who were called up second (and who tend to be older brothers). To correct for this bias, we have controlled performance data for call-up sequence in all analyses involving birth order."

(UPDATE: I had wondered why the authors assumed older brothers were better. I now see they ran a regression and found the older brothers had longer careers.)

But why is this bias?

Suppose you wanted to know if younger brothers grew up to be taller than older brothers. You measure a couple of hundred sets of siblings, and you find they grew to exactly the same height, on average. But, wait! At the time the older brother started high school, he's almost always taller than the younger brother! See? We have bias!

Well, of course, we *don't* have bias. Height (and talent) is related to age of the person, not calendar year. You'd expect brothers to be of equal height at equal ages, and of equal major-league experience at equal ages, NOT at equal calendar years. The "entering high school" and "getting called up" are just red herrings.

Regardless, the authors proceed to divide the sample of players into three groups:

-- younger brother called up first
-- younger and older brother called up the same year
-- older player called up first

If you do that, you'll find that the younger brother has a huge advantage in the first two groups. Why? It's obvious: if the a player gets called up before his brother *even though he's younger*, he's probably a much better player. In addition, speed is a talent more suited to younger players. So when it comes to attempted steals, you'd expect younger brothers called up early to decimate their older brothers in the steals department.

I ran my numbers for all three groups, and got:

-- 87% of younger brothers called up first attempted more steals (7/8)

-- 71% of younger brothers called up the same year attempted more steals (5/7)


-- 54% of younger brothers called up last attemped more steals (45/83)

The overall average, of course, is still 58% (57/98). But by splitting the data this way, it makes the effect seem larger. But that's not really telling you anything -- you're selectively sampling based on the results, because the younger players who come up first are precisely those players who you'd expect to steal only for reasons of being called up young.

Another analogy: home teams win 54% of games. But home teams that only use up one out in the ninth inning win 100% of games! That doesn't mean the 54% is biased, nor does it mean that you learn any more about home field advantage by breaking out walk-off wins. It's just another way of looking at the same result.

Anyway, having claimed to correct for a bias that I don't think really exists, they go on to compute odds ratios for the three cases separately, and use something called the "Mantel-Haenszel" technique to combine them. The statistical techniques look OK, and it seems to me that should have given them the same 52% that they found for the one group. But they somehow come up with 90%.

So, again: what's going on?

My only guess is that they somehow decided that the "younger brother called up first group" is the important one, and they just quoted that number. I got 87% for that, and their numbers are a little different, so maybe they got 90%. Another possibility: it might have something to do with the regression they used to correct for a bunch of things, including which "called up first" group the batter belonged to. They don't give the regression equation or the details, but depending on the way they chose dummy variables, I can see the regression looking something like:

Percentage of younger brothers outstealing older brothers equals:
-- 90%
-- minus 20% if the brothers were called up at the same time
-- minus 20% if the older brother was called up first.

So you can see "90%" coming up as a estimate in a regression equation. But still, the authors would certainly have noticed that you can't just quote the 90% figure without adjusting for the other dummies.

----

Here's another way they might have got it wrong. As you see in the quote above, the authors note that the older players turned out to play more games, and they were also likely to get called up first. So maybe they adjusted the numbers in the service of "controlling athletic performance" (the actual words they use).

Suppose that in 1950, older brother A stole 10 bases and younger brother B didn't steal any because he wasn't in the majors yet. If you "adjust" for that by including it in the regression, effectively subtracting 10 for every line of A's career, I can see how the "adjustment" might now show B having a 90% chance to steal more "adjusted" bases than A.

Again, I'm speculating, which I shouldn't. I don't think the authors did this, but I wouldn't be surprised if it's an adjustment something along those lines, just more subtle.

----

Another related stat the study comes up with is that younger brothers are 10.58 times more likely to steal a base than their older brother. I don't know where that one comes from either. Again, the authors' own raw data comes up with something much more realistic. As a ratio of times on base, the authors' SB+CS rates were

5.6% for older brothers
9.3% for younger brothers

That's only 1.66 times more likely.

So, obviously, the authors feel there's some huge bias here that, when you correct for it, brings the 1.66 up to 10.58. I can't imagine what that might be, and I can't really duplicate the authors' regression because they don't even give the equation.

Here are my calculated batting lines for the two groups. Is there really a factor of ten difference here?

------ AB -R --H 2B 3B HR RBI BB SB CS -avg RC/G
Young 553 70 146 26 04 11 062 50 12 05 .265 4.39
-Old- 546 75 148 26 03 17 074 54 08 04 .271 4.98


You'd think that they'd give a simple English explanation of how, when you look at the batting lines of the two groups, you see young players attempting steals at maybe 50% more than the older ones, but how when *they* look at the two batting lines, they see one line stealing bases at almost 1000% more than the older ones. But I couldn't find any such explanation in the paper.

So when they say,

"... younger brothers were 10.6 times more likely than their older brothers to attempt steals"

... well, I don't see any way that can possibly be true.

-----

UPDATE: batting lines updated, they were slightly off ... also updated to reflect the authors' explanation of something I had missed.

-----

ADDENDUM: as you can see, my data show the younger players indeed attempting more steals than the older players. In my sample, it's 46% more; in the authors' sample, it's 66% more.

My 46% seems like a lot ... I wondered there's perhaps a real effect there. So I ran a simulation to check for statistical significance.

For each of my 98 pairs of brothers, I *randomly* decided which one to consider as older, instead of looking at their real ages. Then I computed the relative attempt rate between the two random groups. I repeated that 60,000 times.

The results? 4,126 of the 60,000 random sets had one group attempting at least 46% more steals than the other group. That's a p-value of 0.069 ... not significant at 5%, but close.

I had expected more significance than that, but the players are "lumpy," in that some players steal a lot, and some don't. A few faster players landing in the same group can make a big difference.

-----

ADDENDUM 2: In the comments, I think Guy came up with the answer. Younger brothers have shorter careers. Shorter careers tend to be centered on lower ages (there are 40 year old players, but no 10 year old players). Young players are faster.

That's why younger brothers have more attempted steals -- they play more years during peak stealing age.

I bet if you did it season-by-season and controlled for age, most of the effect would disappear.

Good catch by Guy.





Labels: , ,

Wednesday, May 26, 2010

Do younger brothers steal more bases than older brothers?

Alan Schwarz's latest "Keeping Score" column, which appeared in last Sunday's New York Times, quotes an academic study that found startling sibling effects among baseball players.

A hypothesis in psychology is that younger siblings exhibit riskier behavior than older ones, "perhaps originally to fight for food, now for parental attention." If that's the case, you'd expect younger brothers to attempt more stolen bases (baseball's equivalent of risky behavior) than older ones.

In an just-published academic study, psychologists Frank J. Sulloway and Richard L. Zweigenhaft checked that, and found that evidence supporting their hypothesis to a very significant degree: a full 90 percent of younger brothers outstole their older siblings!

That's astonishing, to find an effect that large. Since the study isn't available online, I tried to reproduce the study. I didn't get 90% -- I got 56%. Which makes a lot more intuitive sense.

Here's what I did. I went to this Baseball Almanac page, which lists all brother combinations in history. I downloaded their list and fixed the spellings as best I could. Then, I eliminated

(a) all sets of twins;
(b) all sets of brothers where either or both was born before 1895 (Babe Ruth's birth year);
(c) all sets of brothers where one or both was primarily a pitcher;
(d) all sets of brothers who had identical SB rates (always both zero, I think).

That left 114 sets of batting brothers. I then computed their rate of SB per (1B + BB), to see which brother tended to steal more bases. (The original study used H+BB+HBP instead of 1B+BB, but I don't think that would affect the results much.)

64 of the 114 younger brothers outstole their older siblings. Since random chance would be 57, I don't think there's an effect there. It's 1.3 SD above expected.

I have no idea if I did something wrong, or if the authors of the study did something wrong. I'm betting it's them, just because 90% is kind of outrageous.

My data, in not particularly easy to read format, is here.

----

UPDATE: The authors' study appears to have been a regression where:

" ... several other factors were considered, like age differences, body size and even the order in which the players were promoted to the majors."

Still, it seems unlikely that those factors would raise the rate from 56% to 90%.

----

UPDATE: Here are the career batting lines for both groups, divided by 1000:

------ AB -R --H 2B 3B HR RBI BB SB -avg RC/G
Young 444 58 118 20 05 06 051 36 10 .267 4.22
-Old- 538 75 145 24 05 11 068 48 12 .270 4.55

And per 600 PA (differences in rate stats due to rounding):

------ AB -R --H 2B 3B HR RBI BB SB -avg RC/G
Young 554 73 148 25 06 08 064 46 13 .267 4.24
-Old- 550 76 149 25 06 11 070 50 12 .271 4.62

So the older brothers were bit better than the younger brothers, although the younger ones stole bases at a slightly higher rate.





Labels: , ,

Wednesday, May 13, 2009

How many runs are created by good baserunning?

There's a nice paper on baserunning in the latest issue of JQAS, "Using Simulation to Estimate the Impact of Baserunning Ability in Baseball." It's by Ben Baumer, the New York Mets' stats guy.

Baumer set out to quantify baserunning skill, in terms of runs. Specifically, he considered these seven skills:

-- advancing first to third on a single
-- advancing first to home on a double
-- advancing second to home on a single
-- beating out a DP attempt on a ground out
-- stealing second
-- stealing third
-- tagging up on a fly ball when on second or third

He created a (Markov) simulation using 2005-2007 league-average results for each of the seven skills, and proved that his model came close to actual league runs scored.

Then, he substituted actual team lineups, and, for every player, used their actual baserunning percentages for each of the seven situations. There were two probabilities for each situation: the probability of trying for an extra base (for double plays, this is the probability of there being a force play on the runner on first with less than two outs), and the probability of success given that an attempt was made.

For each team, he then ran the same simulation, but using league-average baserunning. The difference is an estimate of how many runs the team's players gained (or lost) with their baserunning.

The top three and bottom three:

+21.1 Mets
+18.0 Yankees
+14.7 Rockies

-12.3 Marlins
-12.7 Red Sox
-20.8 White Sox

Baumer concludes that most teams should be within 25 runs of average baserunning.

But now he wants to figure out, in theory, how many runs a really great baserunning team would gain, and how much a really bad team would lose. He tries a bunch of different selection criteria for "best" and "worst." The results, simplified a bit:

+41.0 runs, -54.6 runs -- high/low attempt rates
+39.4 runs, -35.1 runs -- high/low success rates
+68.4 runs, -42.5 runs -- high/low combination of attempts/successes


The +68 lineup consisted of: Joey Gathright, Willy Taveras, Jose Reyes, Willie Harris, Chone Figgins, Nook Logan, Josh Barfield, Pablo Ozuna, and Juan Pierre. The -54 lineup was Bengie Molina, Mike Piazza, Josh Bard, Bill Mueller, Frank Thomas, Olmedo Saenz, Ryan Garko, Jay Gibbons, and Toby Hall.

As I said, I really like this paper; it asks an interesting and well-defined question and answers it well. Moreover, it's written for readers who know baseball a bit. It does use more mathematical notation than is necessary for sabermetricians, but given that it's an academic paper, and given that the notation is not overdone and clearly explained, I'd have to say that it's very well done.

The one criticism I have is that, as far as I can tell, Baumer used actual raw success rates and didn't regress to the mean at all. That means that while the results wind up accurate in terms of what the actual run contribution was, they are exaggerated estimates of the actual skill of the players involved. If you're thinking about 2010, there's probably no way to estimate, in advance, what any given set of baserunners will do. While the "combination" group added 68.4 runs a season from 2005 to 2007, they'd regress to the mean in 2009-2001 by some amount. What's that amount? We don't really know.

Oh, and one useful point that I'll use in future: for leagues that score 0.531 runs per inning, the variance of runs per inning is 1.125. I've always used 1.000 as an estimate, based on some research I did on the 1988 AL a long time ago, but I think that league scored only .5 runs/inning. Also of note: a simulation that assumes average pitching and an average lineup has a variance a bit smaller: around 1.1 runs instead of 1.125. That's obviously because the pitching doesn't vary in the simulation, only the hitting.


Labels: ,