Friday, February 04, 2011

"Scorecasting" on players gunning for .300

A few months ago, I wrote about a study by two psychology researchers, Devin Pope and Uri Simonsohn. The study found that, for players hitting .299 in their last at-bat of the season, they wound up hitting well over .400 in that last at-bat. The authors concluded that it's because .299 hitters really want to get to .300, and, therefore, they try extra hard (and succeed).

But, really, that isn't the case. It's really just an illusion caused by selective sampling. When a player hitting .299 gets a hit to push him over .300, he is much more likely to be taken out (or held out) of the lineup, to preserve the .300. Therefore, it's not that they're more likely to get a hit in their last at-bat -- it's that their last at-bat is more likely to be one that results in a hit.

(For an analogy: when a game ends with less than 3 outs, the last batter probably hits well over .500 (since the winning run must have scored on the play). But that's not because the player rises to the situation; it's because, as it were, the situation rises to the player. When he gets a hit, he's the last batter because the game ends. When he doesn't, he's not the last batter.)

Since the original study and article, the authors have modified their paper a bit, saying that the batting average effect is "likely to be at least partially explained" by selective sampling. However, the data given in the previous posts does suggest that almost the *entire* effect is explained by selective sampling. (PDFs: Old paper; new paper.)

There is one part of the study's findings that's probably partially real, and that's the issue of walks. None of the .299 hitters walked in their last at-bat. That's partially selective sampling -- if they walked, they're still at .299, and stayed in the game, so it's not their last at-bat -- but probably partially real, in that .299 hitters were more likely to swing away.

(My results are in previous posts here and here.)

------

The study is given featured status in "Scorecasting," in the chapter on round numbers. However, while the authors of the original paper mention the selective sampling issue, the authors of "Scorecasting" do not:

"What's more surprising is that when these .299 hitters swing away, they are remarkably successful. According to Pope and Simonsohn, in that final at-bat of the season, .299 hitters have hit almost .430. ... (Why, you might ask, don't *all* batters employ the same strategy of swinging wildly? ... if every batter swung away liberally throughout the season, pitchers would probably adjust accordingly and change their strategy to throw nothing but unhittable junk.) ...

"Another way to achieve a season-ending average of .300 is to hit the goal and then preserve it. Sure enough, players hitting .300 on the season's last day are much more likely to take the day off than are players hitting .299."


"Scorecasting" treats these two paragraphs as two separate effects. In reality, the second causes the first.

You can read an excerpt -- almost the entire thing, actually -- at Deadspin, here.

------

One thing that interested me in the chapter was this:

"But no benchmark is more sacred than hitting .300 in a season. It's the line of demarcation between All-Stars and also-rans. It's often the first statistic cited when making a case for or against a position player in arbitration. Not surprisingly, it carries huge financial value. By our calculations, the difference between two otherwise comparable players, one hitting .299 and the other .300, can be as high as two percent of salary, or, given the average major league salary, $130,000."


The authors don't say how they calculated that, but it seems reasonable. A free-agent win is worth $4.5 million, according to Tom Tango and others. That means a run is worth $450,000. One point of batting average, in 500 AB, is turning half an out into half a hit. Assuming the average hit is worth about 0.6 runs and an out is worth negative 0.25 runs, that means the single point of batting average is worth a bit over 0.4 runs. That's close to $200,000.

That figure is higher than the authors' figure of $130,000. The difference is probably just that the authors used the average MLB salary, which includes players not yet free agents (arbs and slaves). However, they imply that the difference between .299 and .300 is worth more than other one-point differences. That might be true, but it would be nice to know how they figured it out and what they found.

------

Finally, two bloggers weigh in. Tom Scocca, at Slate, criticizes the original study. Then, Christopher Shea, at the Wall Street Journal, criticizes Scocca.



Labels: , , ,

Tuesday, October 19, 2010

How should the mainstream media report on sabermetric research?

A lot of ideas for posts to this blog come from mainstream media reports of academic sports studies. Often, those turn out to be flawed. Take, for instance, the recent article claiming that .300 hitters hit well over .400 in their last AB because they are highly motivated to succeed. It turns out that the result is caused by selective sampling, and not actually by batters "hunkering down" (as the New York Times put it, in its headline).

However, the media don't normally follow up after a study turns out to be flawed. As a result, there are probably a lot of people running around today believing that batters do actually exhibit clutch behavior when hitting .299. After all, two Ph.D.s said so, it passed peer review, and the New York Times reported it as fact.

That can't be a good thing, to wind up with people believing something that turns out to be false -- not just from the standpoint of sabermetrics, but also from the standpoint of the press.

Is there a better way to report these things?

My first reaction is that the press should report scientific findings the same way they report claims from political think tanks or interest groups -- with a skeptical tone, and with a response from those who might disagree. But that won't happen, for several reasons.

1. The paper is often not yet published and not yet available, so who's going to be able to say what's wrong with it when they can't even see it?

2. It takes a considerable amount time for anyone else to read and digest the paper, especially if it uses complex methodology. Reporters don't have that kind of time to wait.

3. If the paper has already been peer reviewed and accepted for publication, there is a presumption that the paper is correct. And it's been thoroughly reviewed by experts with doctorates. What could an amateur skeptic bring to the story?

4. It doesn't even matter if the paper is wrong. The story is not "players hit .463 in their final at-bat because they're motivated." The story is "Academics say that players hit .463 in their final at-bat because they're motivated." That's true, and newsworthy, even if the embedded claim turns out to be false, because, at the time of publication, there's a strong possibility it might be true.

5. Reporters rely on friendly sources. They don't want to get a reputation for being hostile to academics who come to them with newsworthy ideas. Would the two authors of the .300 paper have talked to the reporter had they expected to have their paper challenged? I doubt it.

So what's the solution?

One thing I'd like to see is for the press to insist that, if they're going to publish a story, the study has to be publicly available at the time the story comes out. To their credit, the authors of this particular study had a working version of the paper on the web. But, sometimes, the paper won't come out for days or weeks. To me, when someone says, "I've discovered X is true but I won't allow you to see the evidence until next month," that shouldn't be a story. Further, it should be something that *academia* frowns upon. Science, after all, is supposed to be open and free, not something you exploit so that your institution looks good. If you're not going to allow the world to see the evidence until November 22, you shouldn't promote it until November 22.

But, having said that ... I have to admit that if academics *do* promote a finding before the paper comes out, it's still a story -- if the academics are credible, knowledgeable experts.

And, not to sound like I'm bashing academia, but ... when it comes to sabermetrics, the Ph.D. economist is usually *not* the expert -- the sabermetric community is the expert. And that, I think, is how the press needs to see it.

In my experience, the way the mainstream media works is that, when they quote a lower-credentialed party, they will almost always go to the higher-credentialed party for a counterpoint. But when they quote the higher-credentialed party, that's often enough.

And that kind of makes sense. When some amateur insists that Saturn's rings are made of beer, you publish it as a novelty story if it's interesting, but you make sure you quote a real astronomer saying the guy is nuts. On the other hand, when you're writing a story about Saturn on the science page, you quote the astronomer, but, obviously, you don't need to go to the amateur for the "beer" counterpoint.

That's a reasonable way of doing it. But what the press doesn't understand yet is that when you evaluate credentials, you have to give the nod to *subject matter expert* (as Tango calls it). In this case, that's the sabermetricians. In this case, it's the Ph.D. who's the amateur, because the subject matter isn't established economic knowledge -- it's established *baseball* knowledge. Any established sabermetrician would instantly realize that a .463 average for a .300 hitter isn't plausible at all, given what's known about clutch hitting. They'd have been able to provide a decent opposing point of view for the article, and perhaps even convince the reporter to write about the issue with a skeptical eye.

Perhaps, for that to happen, sabermetrics needs to lose a little of its "nerd" image. But, even so, it's a reporter's responsibility, when writing about scientific research, to be aware of what's mainstream expertise and what's not. Academics who don't specialize in sabermetrics, and are making surprising or outlandish claims, definitely fall into the "not" category.


Labels: , ,

Thursday, October 07, 2010

Do batters really hit .463 when gunning for a .300 average? Part III

Last post, I forgot the most important findings. Here they are now.

I searched for all PAs from 1975-2008 in the last two games of the season, where the player was hitting less than .300, but, if he got a hit that PA, he'd be at .300 or over. In those plate appearances, the players hit .300.

Then, I looked for players who were at .300 or above, but, if they made an out that PA, they would drop below .300. In those plate appearances, the players hit .297.



Labels: , , ,

Wednesday, October 06, 2010

Do batters really hit .463 when gunning for a .300 average? Part II

Monday, I reviewed a study that found that players hitting .299 late in the season managed to get to .300 suspiciously often, hitting over .400 in their last plate appearance of the season. The authors argued that this is a result of the players expending extra effort to improve performance with .300 on the line. However, I argued that it was just a case of selective sampling: many players, once they hit .300, are benched for the remainder of the season, so cause and effect are reversed: getting a hit causes a certain at-bat to be the last.

After I wrote that, I found that the authors say their results account for substitutions. In footnote 4, they say that they didn't actually use the last plate appearance, but the last *scheduled* PA. So if they were pinch hit for after getting a hit to get to .300, that player would go into the study as "substituted for" instead of "base hit". However, the authors mentioned only pinch hits, and it's possible that if the player was taken out for defensive substitute, or a pinch runner, his hit would still have stayed in the study.

Because the paper is so unclear, I decided to try to reproduce the results. Using Retrosheet data like the authors did, I found every player from 1975 to 2008 who was hitting .299 before his last plate appearance of the season. (However, unlike the authors, I used only last PAs after September 25.) My results were similar to the study's.

According to the New York Times article describing the study, there were 62 players who went into their last PA at .299, and recorded an at-bat. My attempt, however, found 68 (listed here). Those "last PAs" break down as follows:

-- 33 played the entire final game and hit .242 in their last PA

-- 13 were replaced during the final game and hit .692 in their last PA before being replaced

-- 22 didn't play the final game and had hit .636 in their last PA before being sat for the rest of the season.


I have six players the study didn't. My best guess is that 6 of the 13 who were replaced during the final game were pinch hit for, and those were the ones the original study omitted. I'm pretty convinced that's what's going on and that I reproduced the results fairly. That's because the batting averages seem to be about right, but mostly because both the original study and my replication show that these players had zero walks -- and zero walks in 60+ PA is very unusual.

If that's the case, and the study only controlled for pinch hits, there's still a large amount of selective sampling in the 22 players that didn't play the final game. 14 of those 22 guys got a hit. How did that happen? Normally, it would take about 47 AB to get to 14 hits.

What happened is that there probably *were* 47 of those guys. The other 25, who made an out, dropped to .297 or .298, and therefore weren't benched -- they played again to try to get past .300. Therefore, they weren't included in the study, because that hitless AB wasn't their last.

So, the conclusion remains: the effect the authors found is simply the result of players quitting immediately after they get the hit that pushes them over .300.

-----

Just in case you're not convinced, I decided to check how well players hit late in the season when they're shooting for .300, in a method that avoids the sampling bias. Instead of checking just their last plate appearance, I checked a group of plate appearances chosen in advance.

I looked at every team's last two games of the season, and took every PA by a player who went up to the plate hitting .299. That is: not just their last PA, but *any* such PA in the last two games. That way, there's no bias: that PA winds up in the sample if he gets a hit then quits for the year, but, unlike the original, it also winds up in the sample if he makes an out, and bats again later to try to make up the ground he lost.

The results: not much different from other at-bats. Those players hit .313, going 101 for 323 with 23 walks.

What about batters hitting .300, who, if they made an out, would drop to .299? Those guys hit only .288.

Batters hitting .298, still close enough to three hundred to have a decent shot, hit .310. And batters currently at .302 subsequently hit only .289. In chart form:

.298: .310 in 365 AB, 45 walks
.299: .313 in 323 AB, 23 walks
.300: .288 in 233 AB, 30 walks
.301: .289 in 277 AB, 37 walks

Not a whole lot different than you'd expect. Combined, the .299/.300 group hit .302. That's far from the .463 cited by the Times, for the biased sample.

I then tried something slightly different. I stopped the clock before the team's last two games, found all players who were between .297 and .302 at that point in time, and then checked their performance afterwards, regardless of how their batting average may have moved up or down during those two days. That's under the assumption that if the batter is that close that late in the season, every at-bat is critical in his quest to finish at .300 or above.

The results: those players hit only .302.

However, they did hit higher than other nearby groups, as seen below. Plus, you'd expect the players in the .297-.302 group to regress to the mean a bit, and maybe hit .290 or something. So not only did they beat the surrounding groups, but they also beat their expectation.

.291 - .296 — 692/2434 (.284), 229 BB
.297 - .302 — 593/1962 (.302), 205 BB
.303 - .308 — 421/1587 (.265), 167 BB

This is very slight support for the hypothesis that players on the cusp do succeed in pushing themselves a little harder towards .300. I say "very slight" because the standard deviation for these batting averages is about 10 points, so the results certainly aren't statistically significant. Personally, I expect they're mostly random, and that if you did this same study for other seasons, you'd find the effect goes away. But there still might be something there.

-----

So I think the case is pretty well proven:

1. The factoid, "players hitting .299 or .300 batting a whopping .463 in their final at-bat" is true -- but it's the result of cherry-picking the AB in the sample. If the player got a hit to pass .300, it was likely to *become* his last at-bat, as he tended to sit out the rest of the season. But if he made an out, the AB wouldn't be his last.

2. If you look at at *all* AB, not just the cherry-picked "last" ones, players around .300 hit only slightly better than expected, not statistically significant at all.

So it's fair to say that, while there does seem to be motivation to hit .300, that doesn't seem to translate into higher performance when it matters. However, there is strong evidence that players who have achieved the .300 mark late in the season are motivated to stay out of games in order to avoid dropping back to .299 or below. It's the result of that motivation that biases the sample, creating the false impression that batters demonstrate superstar talent in those situations.

----

P.S. Lots of discussion on this issue at "The Book" blog. Thanks to commenters there, especially Guy and MGL, for their assistance.


Labels: , , ,

Monday, October 04, 2010

Do batters really hit .463 when gunning for a .300 average?

A player enters his final plate appearance of the season batting .299 or .300. Presumably, he wants to finish at .300. What happens?

He hits very, very well. According to the New York Times, describing a soon-to-be-published academic study,


"[The academic authors] found that the 127 hitters at .299 or .300 batted a whopping .463 in that final [plate appearance], demonstrating a motivation to succeed well beyond normal (and in what was usually an otherwise meaningless game)."


.463! Holy crap!

But ... do you see what might be going on? Selective sampling. The "final plate appearance of the season" situation is not known beforehand. It could just be that if that player gets a hit, and passes .300, he's removed from the game. So the "final appearance" sample will be biased in favor of players who got a hit.

Suppose you play a game with dice. Every roll is an at-bat. Every 1 or 2 is a hit. The expected batting average is .333. You plan to simulate a player's season of 500 AB. However, as soon as his batting average passes .300, you stop dead, and that's the end of his season.

What will you find? Over 99 percent of the time, the last AB will be a hit. (The only time it won't is if you go all 500 AB without ever passing .300.) Those dice "players" will hit over .990 in their last AB, the one where they first achieved .300. It's obviously not because the die has a motivation to succeed.

That is: it's not that the situation of being able to pass .300 makes their last AB productive. It's probably that being productive while passing .300 *makes the AB their last*. At least that's what I think is going on.

In fairness, there is a bit of evidence that there may be a motivational component too. First, when hitting .299, none of the 61 players involved walked in their final plate appearance. 0 walks for 61 does suggest those players were specifically gunning for .300.

Also, the study included batters hitting both .299 and .300. Those hitters already at .300 obviously weren't sat out, at least not right away. There were 66 players already at .300 ... those guys must have hit pretty well in order to keep the overall average at .463. (If they hadn't, presumably the study would have talked only about the .299 hitters.)

If it were indeed all a result of benchings, how many benchings would it take? Regressing to the mean a bit, suppose the 127 hitters in the sample had talent of about .290. The difference between .290 and .463 in 127 AB is the difference between 37 hits and 59 hits. That means 22 hits need to be explained by benchings. How many benchings would that take? More than 22 (because some of the benched guys might have passed .300 anyway in later AB.) Maybe we can take 31 as an estimate ... if those 31 weren't benched, 9 would have got a hit next time up, staying at .300, and 22 would make an out, dropping back below .300.

Or, what's the farthest we can go without statistical significance? Two SDs of batting average in 127 AB is about .080, or 10 hits. If, by random chance, the hitters were 2 SD better than normal in those situations, then there's only 12 hits left to explain, and only about 17 benchings are required.

(You can probably go even a bit lower if you take into account that some of the 127 PA were walks, meaning the .463 is based on less than 127 AB. UPDATE: David Pinto finds that it could have been 57-for-123.)

Is 31 benchings out of 127 batters reasonable? Is 17 benchings reasonable?

I don't know, but benchings can't be that rare. In 1980, I remember Bobby Mattick sitting Alvis Woods after he got to .300 in the last game (I can't find a reference, but the box score confirms my memory, for what that's worth). And, yesterday's NYT article actually mentions a more recent case (but doesn't seem to realize the implication for the study's findings):

"Five years ago, in a meaningless 162nd game against the Yankees, [David] Ortiz entered batting .299 for the season; he struck out in the first inning to drop to .298 and walked in the third, knowing he still had a few more chances to swing for .300.

One inning later, Ortiz singled to reach .300. He batted one more time in the sixth — he walked, refusing to swing at anything that might result in an out — and was, because of the statistical awareness of Manager Terry Francona, replaced on the bases to make sure that .300 season average would last forever."


So I'm skeptical. I guess we have to wait for the study, by Devin Pope and Uri Simonsohn, to come out to find out what's really going on. The Times says it's been accepted for publication in "Psychological Science", but not yet available.

Or, if anyone wants to reproduce the study, and check to see how many .300 hitters ended their seasons a couple of AB earlier than expected ...


UPDATE: commenter David N. kindly posted a link to the actual study.

If you look at the study, the authors actually show evidence that pinch-hitting is the cause! However, they didn't get the significance of that data.

In their "last scheduled plate appearance of the season", the average batter was pitch hit for 7 percent of the time.

But batters with a .298 or .299 average were pinch hit for only 4.1 percent of the time. Batters with .300 or .301 were pinch hit for 19.7 percent of the time.

And, most importantly, batters hitting exactly .300 were pinch hit for 34.3% of the time!

That basically confirms that the authors' results are likely to be the result of cherry-picking. If you're hitting .299, you get a chance to get a hit in your last AB to jump to .300. But if you're already hitting .300, you often don't get a chance to drop back to .299, getting an out in your last AB.

You know how when you looking for something you lost, you always find it in the last place you look? Well, the same thing applies here. When you're looking for .300, you find it with a hit in the last AB you take.


UPDATE: Over at "The Book," commenter Guy points out that I'm overstating the case a little bit. In computing the batting average, the study did ignore players who were pinch-hit for in their last game. However, it did *not* seem to ignore players who were pinch-run for, or replaced defensively before their next plate appearance. So the selective sampling issue remains.


Labels: , , ,