Thursday, October 07, 2010

Do batters really hit .463 when gunning for a .300 average? Part III

Last post, I forgot the most important findings. Here they are now.

I searched for all PAs from 1975-2008 in the last two games of the season, where the player was hitting less than .300, but, if he got a hit that PA, he'd be at .300 or over. In those plate appearances, the players hit .300.

Then, I looked for players who were at .300 or above, but, if they made an out that PA, they would drop below .300. In those plate appearances, the players hit .297.



Labels: , , ,

Wednesday, October 06, 2010

Do batters really hit .463 when gunning for a .300 average? Part II

Monday, I reviewed a study that found that players hitting .299 late in the season managed to get to .300 suspiciously often, hitting over .400 in their last plate appearance of the season. The authors argued that this is a result of the players expending extra effort to improve performance with .300 on the line. However, I argued that it was just a case of selective sampling: many players, once they hit .300, are benched for the remainder of the season, so cause and effect are reversed: getting a hit causes a certain at-bat to be the last.

After I wrote that, I found that the authors say their results account for substitutions. In footnote 4, they say that they didn't actually use the last plate appearance, but the last *scheduled* PA. So if they were pinch hit for after getting a hit to get to .300, that player would go into the study as "substituted for" instead of "base hit". However, the authors mentioned only pinch hits, and it's possible that if the player was taken out for defensive substitute, or a pinch runner, his hit would still have stayed in the study.

Because the paper is so unclear, I decided to try to reproduce the results. Using Retrosheet data like the authors did, I found every player from 1975 to 2008 who was hitting .299 before his last plate appearance of the season. (However, unlike the authors, I used only last PAs after September 25.) My results were similar to the study's.

According to the New York Times article describing the study, there were 62 players who went into their last PA at .299, and recorded an at-bat. My attempt, however, found 68 (listed here). Those "last PAs" break down as follows:

-- 33 played the entire final game and hit .242 in their last PA

-- 13 were replaced during the final game and hit .692 in their last PA before being replaced

-- 22 didn't play the final game and had hit .636 in their last PA before being sat for the rest of the season.


I have six players the study didn't. My best guess is that 6 of the 13 who were replaced during the final game were pinch hit for, and those were the ones the original study omitted. I'm pretty convinced that's what's going on and that I reproduced the results fairly. That's because the batting averages seem to be about right, but mostly because both the original study and my replication show that these players had zero walks -- and zero walks in 60+ PA is very unusual.

If that's the case, and the study only controlled for pinch hits, there's still a large amount of selective sampling in the 22 players that didn't play the final game. 14 of those 22 guys got a hit. How did that happen? Normally, it would take about 47 AB to get to 14 hits.

What happened is that there probably *were* 47 of those guys. The other 25, who made an out, dropped to .297 or .298, and therefore weren't benched -- they played again to try to get past .300. Therefore, they weren't included in the study, because that hitless AB wasn't their last.

So, the conclusion remains: the effect the authors found is simply the result of players quitting immediately after they get the hit that pushes them over .300.

-----

Just in case you're not convinced, I decided to check how well players hit late in the season when they're shooting for .300, in a method that avoids the sampling bias. Instead of checking just their last plate appearance, I checked a group of plate appearances chosen in advance.

I looked at every team's last two games of the season, and took every PA by a player who went up to the plate hitting .299. That is: not just their last PA, but *any* such PA in the last two games. That way, there's no bias: that PA winds up in the sample if he gets a hit then quits for the year, but, unlike the original, it also winds up in the sample if he makes an out, and bats again later to try to make up the ground he lost.

The results: not much different from other at-bats. Those players hit .313, going 101 for 323 with 23 walks.

What about batters hitting .300, who, if they made an out, would drop to .299? Those guys hit only .288.

Batters hitting .298, still close enough to three hundred to have a decent shot, hit .310. And batters currently at .302 subsequently hit only .289. In chart form:

.298: .310 in 365 AB, 45 walks
.299: .313 in 323 AB, 23 walks
.300: .288 in 233 AB, 30 walks
.301: .289 in 277 AB, 37 walks

Not a whole lot different than you'd expect. Combined, the .299/.300 group hit .302. That's far from the .463 cited by the Times, for the biased sample.

I then tried something slightly different. I stopped the clock before the team's last two games, found all players who were between .297 and .302 at that point in time, and then checked their performance afterwards, regardless of how their batting average may have moved up or down during those two days. That's under the assumption that if the batter is that close that late in the season, every at-bat is critical in his quest to finish at .300 or above.

The results: those players hit only .302.

However, they did hit higher than other nearby groups, as seen below. Plus, you'd expect the players in the .297-.302 group to regress to the mean a bit, and maybe hit .290 or something. So not only did they beat the surrounding groups, but they also beat their expectation.

.291 - .296 — 692/2434 (.284), 229 BB
.297 - .302 — 593/1962 (.302), 205 BB
.303 - .308 — 421/1587 (.265), 167 BB

This is very slight support for the hypothesis that players on the cusp do succeed in pushing themselves a little harder towards .300. I say "very slight" because the standard deviation for these batting averages is about 10 points, so the results certainly aren't statistically significant. Personally, I expect they're mostly random, and that if you did this same study for other seasons, you'd find the effect goes away. But there still might be something there.

-----

So I think the case is pretty well proven:

1. The factoid, "players hitting .299 or .300 batting a whopping .463 in their final at-bat" is true -- but it's the result of cherry-picking the AB in the sample. If the player got a hit to pass .300, it was likely to *become* his last at-bat, as he tended to sit out the rest of the season. But if he made an out, the AB wouldn't be his last.

2. If you look at at *all* AB, not just the cherry-picked "last" ones, players around .300 hit only slightly better than expected, not statistically significant at all.

So it's fair to say that, while there does seem to be motivation to hit .300, that doesn't seem to translate into higher performance when it matters. However, there is strong evidence that players who have achieved the .300 mark late in the season are motivated to stay out of games in order to avoid dropping back to .299 or below. It's the result of that motivation that biases the sample, creating the false impression that batters demonstrate superstar talent in those situations.

----

P.S. Lots of discussion on this issue at "The Book" blog. Thanks to commenters there, especially Guy and MGL, for their assistance.


Labels: , , ,

Monday, October 04, 2010

Do batters really hit .463 when gunning for a .300 average?

A player enters his final plate appearance of the season batting .299 or .300. Presumably, he wants to finish at .300. What happens?

He hits very, very well. According to the New York Times, describing a soon-to-be-published academic study,


"[The academic authors] found that the 127 hitters at .299 or .300 batted a whopping .463 in that final [plate appearance], demonstrating a motivation to succeed well beyond normal (and in what was usually an otherwise meaningless game)."


.463! Holy crap!

But ... do you see what might be going on? Selective sampling. The "final plate appearance of the season" situation is not known beforehand. It could just be that if that player gets a hit, and passes .300, he's removed from the game. So the "final appearance" sample will be biased in favor of players who got a hit.

Suppose you play a game with dice. Every roll is an at-bat. Every 1 or 2 is a hit. The expected batting average is .333. You plan to simulate a player's season of 500 AB. However, as soon as his batting average passes .300, you stop dead, and that's the end of his season.

What will you find? Over 99 percent of the time, the last AB will be a hit. (The only time it won't is if you go all 500 AB without ever passing .300.) Those dice "players" will hit over .990 in their last AB, the one where they first achieved .300. It's obviously not because the die has a motivation to succeed.

That is: it's not that the situation of being able to pass .300 makes their last AB productive. It's probably that being productive while passing .300 *makes the AB their last*. At least that's what I think is going on.

In fairness, there is a bit of evidence that there may be a motivational component too. First, when hitting .299, none of the 61 players involved walked in their final plate appearance. 0 walks for 61 does suggest those players were specifically gunning for .300.

Also, the study included batters hitting both .299 and .300. Those hitters already at .300 obviously weren't sat out, at least not right away. There were 66 players already at .300 ... those guys must have hit pretty well in order to keep the overall average at .463. (If they hadn't, presumably the study would have talked only about the .299 hitters.)

If it were indeed all a result of benchings, how many benchings would it take? Regressing to the mean a bit, suppose the 127 hitters in the sample had talent of about .290. The difference between .290 and .463 in 127 AB is the difference between 37 hits and 59 hits. That means 22 hits need to be explained by benchings. How many benchings would that take? More than 22 (because some of the benched guys might have passed .300 anyway in later AB.) Maybe we can take 31 as an estimate ... if those 31 weren't benched, 9 would have got a hit next time up, staying at .300, and 22 would make an out, dropping back below .300.

Or, what's the farthest we can go without statistical significance? Two SDs of batting average in 127 AB is about .080, or 10 hits. If, by random chance, the hitters were 2 SD better than normal in those situations, then there's only 12 hits left to explain, and only about 17 benchings are required.

(You can probably go even a bit lower if you take into account that some of the 127 PA were walks, meaning the .463 is based on less than 127 AB. UPDATE: David Pinto finds that it could have been 57-for-123.)

Is 31 benchings out of 127 batters reasonable? Is 17 benchings reasonable?

I don't know, but benchings can't be that rare. In 1980, I remember Bobby Mattick sitting Alvis Woods after he got to .300 in the last game (I can't find a reference, but the box score confirms my memory, for what that's worth). And, yesterday's NYT article actually mentions a more recent case (but doesn't seem to realize the implication for the study's findings):

"Five years ago, in a meaningless 162nd game against the Yankees, [David] Ortiz entered batting .299 for the season; he struck out in the first inning to drop to .298 and walked in the third, knowing he still had a few more chances to swing for .300.

One inning later, Ortiz singled to reach .300. He batted one more time in the sixth — he walked, refusing to swing at anything that might result in an out — and was, because of the statistical awareness of Manager Terry Francona, replaced on the bases to make sure that .300 season average would last forever."


So I'm skeptical. I guess we have to wait for the study, by Devin Pope and Uri Simonsohn, to come out to find out what's really going on. The Times says it's been accepted for publication in "Psychological Science", but not yet available.

Or, if anyone wants to reproduce the study, and check to see how many .300 hitters ended their seasons a couple of AB earlier than expected ...


UPDATE: commenter David N. kindly posted a link to the actual study.

If you look at the study, the authors actually show evidence that pinch-hitting is the cause! However, they didn't get the significance of that data.

In their "last scheduled plate appearance of the season", the average batter was pitch hit for 7 percent of the time.

But batters with a .298 or .299 average were pinch hit for only 4.1 percent of the time. Batters with .300 or .301 were pinch hit for 19.7 percent of the time.

And, most importantly, batters hitting exactly .300 were pinch hit for 34.3% of the time!

That basically confirms that the authors' results are likely to be the result of cherry-picking. If you're hitting .299, you get a chance to get a hit in your last AB to jump to .300. But if you're already hitting .300, you often don't get a chance to drop back to .299, getting an out in your last AB.

You know how when you looking for something you lost, you always find it in the last place you look? Well, the same thing applies here. When you're looking for .300, you find it with a hit in the last AB you take.


UPDATE: Over at "The Book," commenter Guy points out that I'm overstating the case a little bit. In computing the batting average, the study did ignore players who were pinch-hit for in their last game. However, it did *not* seem to ignore players who were pinch-run for, or replaced defensively before their next plate appearance. So the selective sampling issue remains.


Labels: , , ,

Monday, April 20, 2009

A Diamond Mind simulation as baseball strategy research

A science column from Alan Schwarz a couple of weeks ago investigates the effects of various baseball strategies, using a simulation.

To check out batting orders, Schwarz got Luke Kraemer at Diamond Mind to simulate two sets of 100 seasons of the 2008 Yankees. In one set, A-Rod batted fourth; in the other set, he batted ninth. The difference was 42 runs; the regular Yanks scored 789 runs, while the A-Rod-at-the-bottom-of-the-order Yanks scored only 747.

Schwarz doesn't tell us how he checked intentional walks, but finds that they are a bad strategy, costing five runs per season. That's not a very useful result; there are times when the IBB makes more sense, and times when it makes less sense. Which did Diamond Mind simulate?

Stolen Bases: Diamond Mind took the 2008 Rays and the 2008 A's, and reversed their respective propensities to steal ("switched their mind-sets," is what the article says). The A's dropped by 20 runs, but the Rays *improved* by 47 runs, "suggesting that perhaps the Rays were running too often in real life."

As it turns out, the real Tampa Bay team stole 142 bases and were caught only 50 times, for a 74% success rate; that should put them well in the black, compared to the rule of thumb that you need to be successful 67% of the time to break even. So I'm at a loss to explain the 47 run difference.

The only thing I can think of is a sample size issue. I think the SD of a team's runs scored in a single game is about 3. So the SD of a season's worth of runs is 3 times the square root of 162, or about 38 runs. The SD of the average of 100 season's worth is one-tenth of that, or about 4 runs. The difference between two 100-season averages is the square root of 2 times that, or about 5.4 runs.

But 47 runs is almost 9 standard deviations. So I'm still not sure what's going on.

Finally, the sacrifice bunt. When the simulation forced the bunt-avoiding Red Sox (27 SH in 2008, compared to the league-average 34) to do it more often, they lost 19 runs. But when they got the bunt-loving Mets (73, league average 66) to do it less, the result was also a loss – 15 runs. Schwarz concludes that the Mets' real-life bunting was better than the Red Sox, that they chose to bunt in more favorable situations. But, weren't both these numbers based on the simulation? If so, the real-life situations should make no difference.

If the comparisons, however, *were* based on real life, then we have sample size issues based on the real-life sample, which is only 162 games, with an SD of about 38 runs. Maybe the 2008 Mets and Red Sox scored more or fewer than the simulation because of luck? We should be able to tell by looking at Runs Created – but, for some reason, almost all teams undershot their RC estimate in 2008 (and their Base Runs estimate too, at least for the versions I tried).

Anyway, while I like the simulation method, I wish the results had been presented more clearly. As it stands, I'll stick to "The Book"'s conclusions on these issues of baseball strategy.

P.S. Here's what Tony LaRussa thinks of these results:

“There’s way too much importance given to what you can produce from a machine,” he said. “These are human beings, and I don’t think any computer is going to model that close to what we deal with at this level.”

Hat Tip: Daniel Hamermesh at Freakonomics


Labels: , , ,