Sunday, July 03, 2011

Home field advantage on pitch calls, by count

Did a bit of last minute research before putting together my presentation on home field advantage (HFA) for the SABR convention.

In "Scorecasting," Toby Moskowitz and Jon Wertheim wrote that HFA on ball-strike calls varies with the importance of the situation. They said that in clutch plate appearances, HFA is very high -- but, when it doesn't matter much, HFA actually goes the *other* way, and visiting pitchers actually get the benefit of more called strikes than home batters. They concluded that biased umpires are favoring the home team, but trying to compensate the visiting team by calling more strikes for them when it's not as important.

A few months ago, MGL did a study, and found some confirmation for Scorecasting's results. He did find that HFA went up with leverage. However, he didn't find any situations in which the home team actually had an advantage -- just situations in which they had less of an advantage than usual.

So I tried the same thing today (but not as rigorously). I used 2000-2009 Retrosheet data, and I got similar results.

Overall, not looking at leverage yet, the home pitchers got 0.6 percentage points more strikes than the visiting pitchers. (Specifically, the visiting team had 31.2% of their called pitches ruled strikes, but the home team had 31.8% of theirs ruled strikes.)


In certain higher-leverage situations (for which I used 8th inning or later, score tied), the difference was higher -- 1.2 percentage points. In another higher-leverage situation (9th inning or later, tying run at the plate), the difference was also higher -- 1.0 percentage points.

But in lower-leverage situations (one team leads by 5 runs or more), the difference was only 0.3 points.

Summary:

0.3 -- one team leading by 5+ runs
0.6 -- all situations
1.0 -- ninth inning+, tying run at bat
1.2 -- eighth inning+, score tied

Another thing I did is, for all these situations, I computed the HFA in terms of the outcomes of the plate appearances. Here are the home team advantages by wOBA points:


.0010 -- one team leading by 5+ runs
.0013 -- all situations
.0018 -- ninth inning+, tying run at bat
.0025 -- eighth inning+, score tied

As expected, an excellent correlation between HFA on ball/strike, and HFA on eventual outcome.

But here's something interesting: the home/road difference on what percentage of pitches were swung at (including foul balls and balls in play):

0.48 -- one team leading by 5+ runs
0.54 -- all situations
0.88 -- ninth inning+, tying run at bat
0.59 -- eighth inning+, score tied

For instance, overall, home teams swung at 44.9 percent of pitches, but road teams swung at 45.4 percent of pitches.

So, not only did home teams have fewer strikes *called* against them (first table), but they also had fewer *swings* (third table). That suggests that visiting teams actually throw fewer strikes than home teams, since this results holds even on pitches where the umpire has no say.

But, you could argue otherwise. It's possible that the swing difference is because the home batters know they're going to get more marginal calls in their favor, so they don't swing on iffy pitches in order to work a walk. It's also possible that batters are worse on the road, and they can't tell a good pitch from a bad pitch quite as well as they can at home.

So I'm not sure if we can draw any conclusions from this, but I thought it was worth mentioning.

------

Anyway, that does confirm the "Scorecasting" basic findings. But it occurs to me that what might be causing this is just the different ball/strike counts.

As I said above, when an umpire called a home pitch, it was a strike 31.8 percent of the time. But, I checked, and when an umpire called a home pitch *on an 0-0 count*, it was a strike 43.0 percent of the time. That's a big difference. Maybe it extends to home/road differences too?

It does. The HFA was also much bigger on 0-0: instead of 0.6 percentage points, it was 0.9 percentage points.

So maybe HFA is lower in low-leverage situations just because when a game is a blowout, teams pitch differently and you get a different frequency of the different counts. So, what I did was break down all pitches by count, and by leverage group. The leverage groups were:

High ..... 8th inning or later, 0-1 run difference
Low ...... One team leading by 6+ runs
Average .. All other plate appearances.

Here are the results, in percentage points of HFA (called strikes as a percentage of all called pitches). Standard errors are in parentheses.

-------------------------------- Leverage -----------------
----------------- Average -------- High ----------- Low ---
-----------------------------------------------------------
0-0 count ..... 1.03 (0.09) ... 1.02 (0.20) ... 0.66 (0.28)
0-1 count ..... 0.63 (0.14) ... 0.77 (0.31) ... 0.79 (0.43)
0-2 count ..... 0.51 (0.20) ... 0.62 (0.45) .. -0.13 (0.64)
1-0 count ..... 0.92 (0.14) ... 1.54 (0.31) ... 0.38 (0.44)
1-1 count ..... 0.65 (0.15) ... 0.35 (0.34) .. -0.22 (0.48)
1-2 count ..... 0.83 (0.17) ... 0.72 (0.37) ... 0.56 (0.53)
2-0 count ..... 0.98 (0.23) ... 1.11 (0.52) ... 1.18 (0.75)
2-1 count ..... 0.80 (0.20) ... 0.75 (0.46) ... 0.22 (0.65)
2-2 count ..... 0.56 (0.18) ... 0.44 (0.40) ... 1.05 (0.58)
3-0 count ..... 1.07 (0.37) ... 2.35 (0.86) ... 1.15 (1.20)
3-1 count ..... 0.90 (0.30) ... 1.31 (0.69) ... 0.64 (0.97)
3-2 count ..... 0.56 (0.22) ... 1.25 (0.50) ... 1.77 (0.72)

Is there evidence here that HFA depends on leverage? If you compare the average leverage to the high leverage, you get that the high-leverage situations have a higher HFA in 7 out of the 12 cases -- not much more than average. Comparing average to low, you get more HFA for the average situations again 7 out of 12 times. And, comparing high to low, the "high" only win 8 out of 12.

Doesn't seem like much. But I think I've just diced up the data so finely that you can't see the real pattern any more. It looks like all three of the three lowest differences do appear in the low-leverage column (although that could be partly because the SDs are high there, so you expect extreme more values than in the other columns).

Here's the equally weighted average of all three columns, each column weighted by the smaller of the frequencies (home, road) in column 1:

0.82 (~ 0.04 SD) overall
0.92 (~ 0.10 SD) high leverage
0.58 (~ 0.14 SD) low leverage

So there is something there, although smaller than it looked before adjusting for count. But the differences are not statistically significant, although the low-leverage one is close.

Conclusion: from 2000 to 2009, home teams were somewhat more likely to get a strike call in higher-leverage situations than in lower-leverage situations. This is significant only at approximately p=0.1.






Labels: , ,

Monday, February 21, 2011

"Scorecasting" reviews

Coincidentally, Chris Jaffe and I both have reviews of "Scorecasting" out today. Here's Chris, at The Hardball Times. Here's me, at Baseball Prospectus.




Labels:

Friday, February 04, 2011

"Scorecasting" on players gunning for .300

A few months ago, I wrote about a study by two psychology researchers, Devin Pope and Uri Simonsohn. The study found that, for players hitting .299 in their last at-bat of the season, they wound up hitting well over .400 in that last at-bat. The authors concluded that it's because .299 hitters really want to get to .300, and, therefore, they try extra hard (and succeed).

But, really, that isn't the case. It's really just an illusion caused by selective sampling. When a player hitting .299 gets a hit to push him over .300, he is much more likely to be taken out (or held out) of the lineup, to preserve the .300. Therefore, it's not that they're more likely to get a hit in their last at-bat -- it's that their last at-bat is more likely to be one that results in a hit.

(For an analogy: when a game ends with less than 3 outs, the last batter probably hits well over .500 (since the winning run must have scored on the play). But that's not because the player rises to the situation; it's because, as it were, the situation rises to the player. When he gets a hit, he's the last batter because the game ends. When he doesn't, he's not the last batter.)

Since the original study and article, the authors have modified their paper a bit, saying that the batting average effect is "likely to be at least partially explained" by selective sampling. However, the data given in the previous posts does suggest that almost the *entire* effect is explained by selective sampling. (PDFs: Old paper; new paper.)

There is one part of the study's findings that's probably partially real, and that's the issue of walks. None of the .299 hitters walked in their last at-bat. That's partially selective sampling -- if they walked, they're still at .299, and stayed in the game, so it's not their last at-bat -- but probably partially real, in that .299 hitters were more likely to swing away.

(My results are in previous posts here and here.)

------

The study is given featured status in "Scorecasting," in the chapter on round numbers. However, while the authors of the original paper mention the selective sampling issue, the authors of "Scorecasting" do not:

"What's more surprising is that when these .299 hitters swing away, they are remarkably successful. According to Pope and Simonsohn, in that final at-bat of the season, .299 hitters have hit almost .430. ... (Why, you might ask, don't *all* batters employ the same strategy of swinging wildly? ... if every batter swung away liberally throughout the season, pitchers would probably adjust accordingly and change their strategy to throw nothing but unhittable junk.) ...

"Another way to achieve a season-ending average of .300 is to hit the goal and then preserve it. Sure enough, players hitting .300 on the season's last day are much more likely to take the day off than are players hitting .299."


"Scorecasting" treats these two paragraphs as two separate effects. In reality, the second causes the first.

You can read an excerpt -- almost the entire thing, actually -- at Deadspin, here.

------

One thing that interested me in the chapter was this:

"But no benchmark is more sacred than hitting .300 in a season. It's the line of demarcation between All-Stars and also-rans. It's often the first statistic cited when making a case for or against a position player in arbitration. Not surprisingly, it carries huge financial value. By our calculations, the difference between two otherwise comparable players, one hitting .299 and the other .300, can be as high as two percent of salary, or, given the average major league salary, $130,000."


The authors don't say how they calculated that, but it seems reasonable. A free-agent win is worth $4.5 million, according to Tom Tango and others. That means a run is worth $450,000. One point of batting average, in 500 AB, is turning half an out into half a hit. Assuming the average hit is worth about 0.6 runs and an out is worth negative 0.25 runs, that means the single point of batting average is worth a bit over 0.4 runs. That's close to $200,000.

That figure is higher than the authors' figure of $130,000. The difference is probably just that the authors used the average MLB salary, which includes players not yet free agents (arbs and slaves). However, they imply that the difference between .299 and .300 is worth more than other one-point differences. That might be true, but it would be nice to know how they figured it out and what they found.

------

Finally, two bloggers weigh in. Tom Scocca, at Slate, criticizes the original study. Then, Christopher Shea, at the Wall Street Journal, criticizes Scocca.



Labels: , , ,

Tuesday, February 01, 2011

Scorecasting: are the Cubs unlucky, or is it management's fault?

The Chicago Cubs, it has been noted, have not been a particularly huge success on the field in the past few decades. Is Cubs' management to blame? The last chapter of the recent book "Scorecasting" says it's true. I'm not so sure.

The authors, Tobias J. Moskowitz and L. Jon Wertheim, set out to debunk the idea that the Cubs lack of success -- they haven't won a World Series since 1908 -- is simply due to luck.

How do they check that? How do they try to estimate the effects of luck on the Cubbies? Not the way sabermetricians would. Instead, the authors ... well, I'm not really sure what they did, but I can guess. Here's how they start:

"Another way to measure luck is to see how much of a team's success or failure can't be explaiend. For example, take a look at how the team performed on the field and whether, based on its performance, it won fewer games than it should have."

So far, so good. There is an established way to look at certain aspects of luck. You can look at the team's Pythagorean projection, which estimates its won-lost record from its runs scored and runs allowed. If it beat its projection, it was probably lucky.

Also, you can also compute its Runs Created estimate. The Runs Created formula takes a team's batting line, and projects the number of runs it should have scored. If the Cubbies scored more runs than their projection, they were somewhat lucky. If they scored fewer, they were somewhat unlucky.

But that doesn't seem to be what the authors do. At least, it doesn't seem to follow from their description. They continue:

"If you were told that your team led the league in hitting, home runs, runs scored, pitching, and fielding percentage, you'd assume your team won a lot more games than it lost. If it did not, you'd be within your rights to consider it unlucky."

Well, yes and no. Those criteria are not independent. If I were told that my team scored a certain number of runs, I wouldn't care whether it also led the league in home runs, would I? A run is a run, whether it came from leading the league in home runs, or leading the league in "hitting" (by which my best guess is that the authors meant batting average).

The authors do the same thing in the very same paragraph:

"How, for instance, did the 1982 Detroit Tigers finish fourth in their division, winning only 83 games and losing 79, despite placing eighth in the Majors in runs scored that season, seventh in team batting average, fourth in home runs, tenth in runs against, ninth in ERA, fifth in hits allowed, eighth in strikeouts against, and fourth in fewest errors?"

Again, if you know runs scored and runs against, why would you need anything else? Do they really think that if your pitchers give up four runs while striking out a lot of batters, you're more likely to win than if your pitchers give up four runs while striking out fewer batters?

(As an aside, just to answer the authors' question: The 1982 Tigers underperformed their Pythagorean estimate by 3 games. They underperformed their Runs Created by 2 games. But their opponents underperformed their own Runs Created estimate by 1 game. Combining these three measures shows the '82 Tigers finished four games worse than they "should have".)

Now, we get to the point where I don't really understand their methodology:

"Historically, for the average MLB team, its on-the-field statistics would predict its winning percentage year to year with 93 percent accuracy."

What does that mean? I'm not sure. My initial impression is that they ran a regression to predict winning percentage based on that bunch of stats above (although if they included runs scored and runs allowed, the other variables in the regression should be almost completely superfluous, but never mind). My guess is that's what they did, and they got an correlation coefficient of .93 ... or perhaps an r-squared of .93. But that's not how they explain it:

"That is, if you were to look only at a team's on-the-field numbers each season and rank it based on those numbers, 93 percent of the time you would get the same ranking as if you ranked it based on wins and losses."

Huh? That can't be right. If you were to take the last 100 years of the Cubs, and run a projection for each year, the probability that you'd get *exactly the same ranking* for the projection and the actual would be almost zero. Consider, for instance, 1996, where the Cubs outscored their opponents by a run, and nonetheless wound up 76-86. And now consider 1993, when the Cubs were outscored by a run, and wound up 84-78. There's no way any projection system would "know" to rank 1993 eight games ahead of 1996, and so there's no way the rankings would be the same. The probability of getting the same ranking, then, would be zero percent, not 93 percent.

What I think is happening is that they're really talking about a correlation of .93, and this "93 percent of the time you would get the same ranking" is just an oversimplification in explaining what the correlation means. I might be wrong about that, but that's how I'm going to proceed, because that seems the most plausible explanation.

So, now, from there, how do the authors get to the conclusion that the Cubs weren't unlucky? What I think they did is to run the same regression, but for Cub seasons only. And they got 94 percent instead of 93 percent. And so, they say,

"The Cubs' record can be just as easily explained as those of the majority of teams in baseball. ... Here you could argue that the Cubs are actually less unlucky than the average team in baseball."

What they're saying is, since the regression works just as well for the Cubs as any other team, they couldn't have been unlucky.

But that just doesn't follow. At least, if my guess is correct that they used regression. I think the authors are incorrect about what it means to be lucky and how that relates to the correlation.

The correlation in the data suggests the extent to which the data linearly "explain" the year-to-year differences in winning percentage. But the regression doesn't distinguish luck from other explanations. If the Cubs are consistently lucky, or consistently unlucky, the regression will include that in the correlation.

Suppose I try to guess whether a coin will land heads or tails. And I'm right about half the time. I might run a bunch of trials, and the results might look like this:

1000 trials, 550 correct
200 trials, 90 correct
1600 trials, 790 correct
100 trials, 40 correct

If I run a regression on these numbers, I'm going to get a pretty high correlation -- .9968, to be more precise.

But now, suppose I'm really lucky. In fact, I'm consistently lucky. And, as a result, I do 10 percent better on every trial:

1000 trials, 605 correct
200 trials, 99 correct
1600 trials, 869 correct
100 trials, 44 correct

What happens now? If I run the same regression (try it, if you want), I will get *exactly the same correlation*. Why? Because it's just as easy to predict the number of successes as before. I just do what I did before, and add 10%. It's not the correlation that changes -- it's the regression equation. Instead of predicting that I get about 50% right, the equation will just predict that I get about 55% right. The fact that I was lucky, consistently lucky, doesn't change the r or the r-squared.

The same thing will happen in the Cubs case. Suppose the Cubs are lucky, on average, by 1 win per season. The regression will "see" that, and simply adjust the equation to somehow predict an extra win per season. It'll probably change all the coefficients slightly so that the end result is one extra win. Maybe if the Cubs are lucky, and a single "should be" worth 0.046 wins, the regression will come up with a value of 0.047 instead, to reflect the fact that, all other things being equal, the Cubs' run total is a little higher than for other teams. Or something like that.

Regardless, that won't affect the correlation much at all. Whether the Cubs were a bit lucky, a bit unlucky, about average in luck, or even the luckiest or unluckiest team in baseball history, the correlation might come out higher than .93, less than .93, or the same as .93.

So, what, then, does the difference between the Cubs' .94, and the rest of the league's .93, tell us? It might be telling us about the *variance* of the Cubs' luck, not the mean. If the Cubs hit the same way one year as the next, but one year they win 76 games and another they win 84 games ... THAT will reduce the correlation, because it will turn out that the same batting line isn't able to very accurate pinpoint the number of wins.

If you must draw a conclusion from the the regression in the book -- which I am reluctant to do, but if you must -- it should be only that the Cubs' luck is very slightly *more consistent* than other teams' luck. But it will *not* tell you if the Cubs' overall luck is good, bad, or indifferent.

------

So, have the Cubs been lucky, or not? The book's study doesn't tell us. But we can just look at the Cubs' Pythagorean projections, and runs created projections. Actually, a few years ago, I did that, and I also created a method to try to quantify a "career year" effect, to tell if the team's players underperformed or overachieved for that season, based on the players' surrounding seasons. (For instance, Dave Stieb's 1986 was marked as an unlucky year, and Brady Anderson's 1996 a lucky year, because both look out of place in the context of the players' careers.)

My study gave a total of a team's luck based on five factors:

-- did it win more or fewer games than expected by its runs scored and allowed?
-- did it score more or fewer runs than expected by its batting line?
-- did its opponents score more or fewer runs than expected by their batting line?
-- did its hitters have over- or underachieving years?
-- did its pitchers have over- or underachieving years?

(Here's a PowerPoint presentation explaining the method, and here's a .ZIP file with full team and player data.)

The results: from 1960 to 2001, the Cubs were indeed unlucky ... by an average of slightly over half a win. That half win was comprised of about 1.5 wins of unlucky underperformance of their players, mitigated by about one win of being lucky in turning that performance into wins.

But the Cubs never really had seasons in that timespan in which bad luck cost them a pennant or division title. The closest were 1970 and 1971, when, both years, they finished about five games unluckier than they should have (they would have challenged for the pennant in 1970 with 89 wins, but not in 1971 with 88 wins). Mostly, when they were unlucky, they were a mediocre team that bad-lucked their way into the basement. In 1962 and 1966, they lost 103 games, but, with normal luck, would have lost only 85 and 89, respectively.

However, when the Cubs had *good* luck, it was at opportune times. In 1984, they won 96 games and the NL East, despite being only an 80-82 team on paper. And they did it again in 1989, winning 93 games instead of the expected 77.

On balance, I'd say that the Cubs were lucky rather than unlucky. They won two divisions because of luck, but never really lost one because of luck. Even if you want to consider that they lost half a title in 1970, that still doesn't come close to compensating for 1984 and 1989.

------

But things change once you get past 2001. It's not in the spreadsheet I linked to, but I later ran the same analysis for 2002 to 2007, at Chris Jaffe's request for his book. And, in recent years, the Cubs have indeed been unlucky:

2002: 67-95, "should have been" 86-76 (19 games unlucky)
2003: 88-74, "should have been" 86-76 (2 games lucky)
2004: 89-73, "should have been" 90-72 (1 game unlucky)
2005: 79-83, "should have been" 86-76 (6 games unlucky)
2006: 66-96, "should have been" 82-80 (16 games unlucky)
2007: 85-77, "should have been" 88-74 (3 games unlucky)

That's 42 games of bad luck over seven seasons -- an average of 7 games per season. That's huge. Even if you don't trust my "career year" calculations, just the Pythagoras and Runs Created bad luck sum to almost 5.5 of those 7 games.

So, yes ... in the last few years, the Cubs *have* been unlucky. Very, very unlucky.

------

In summary: from 1960 to 2001, the Cubs were a bit of a below-average team, with about average luck. Then, starting in 2002, the Cubs got good -- but, by coincidence or curse, their luck turned very bad at exactly the same time.

----

But if the "Scorecasting" authors don't believe that the Cubs have been unlucky, then what do they think is the reason for the Cubs' lack of success?

Incentives. Or, more accurately, the lack thereof. The Cubs sell out almost every game, win or lose. So, the authors ask, why should Cubs management care about winning? They gain very little if they win, so they don't bother to try.

To support that hypothesis, the authors show the impact (elasticity) of wins on tickets sold. It turns out that the Cubs have the lowest elasticity in baseball, at 0.6. If the Cubs' winning percentage drops by 10 percent, ticket sales drop by only 6 percent.

On the other hand, their crosstown rivals have one of the highest elasticities in the league, at about 1.2. For every 10 percent drop in winning percentage, White Sox ticket sales drop by 12 percent -- almost twice as much.

But ... I find this unconvincing, for a couple of reasons. First, if you look at the authors' tables (p. 245), it looks like it takes a year or so after a good season for attendance to jump. That makes sense. In 2005, it probably took a month or two for White Sox fans to realize the team was genuinely good; in 2006, they all knew beforehand, at season-ticket time.

Now, if you look at the Cubs' W-L record for the past 10 years, it really jumps up and down a lot; from 1998 to 2004, the team seesawed between good and bad. For seven consecutive seasons, they either won 88 games or more (four times), or lost 88 games or more (three times). So, fan expectations were probably never in line with team performance. Because the authors predicted attendance based on current performance, rather than lagged performance, that might be why they didn't see a strong relationship (even if there is one).

But that's a minor reason. The bigger reason I disagree with the authors' conclusions is that, even when they're selling out, the Cubs still have a strong incentive to improve the team -- and that's ticket prices. Isn't it obvious that the better the team, the higher the demand, and the more you can charge? It's no coincidence that the Cubs have the highest ticket prices in the Major Leagues (.pdf) at the same time as they're selling out most games. If the team is successful, and demand rises, the team just charges more instead of selling more.

Also, what about TV revenues, and merchandise sales, which also rise when a team succeeds?

It seems a curious omission that the authors would consider only that the Cubs can't sell more tickets, and not that total revenues would significantly increase in other ways. But that's what they did. And so they argue,

"So, at least financially, the Cubs seem to have far less incentive than do other teams -- less than the Yankees and Red Sox, and certainly less than the White Sox. ... Winning or losing is often the result of a few small things that require extra effort to gain a competitive edge: going the extra step to sign the highly sought-after free agent, investing in a strong farm team with diligent scouting, monitoring talent, poring over statistics, even making players more comfortable. All can make a difference at the margin, and all are costly. When the benefits of making these investments are marginal at best, why undertake them?"

Well, the first argument is the one I just made: the benefits are *not* "marginal at best," because with a winning team, the Cubs would earn a lot more money in other ways. But there's a more persuasive and obvious argument. If the Cubs have so small an incentive to win, if they care so little that they can't even be bothered to hire "diligent" scouts ... then why do they spend so much money on players?

In 2010, the Cubs' payroll was $146 million, third in the majors. In 2009, they were also third. Since 2004, they have never been lower than ninth, in a 30-team league. Going back as far as 1991, there are only a couple of seasons that the Cubs are below average -- and in those cases, just barely. In the past 20 years, the Cubs have spent significantly more money on player salaries than the average major-league team.

It just doesn't make sense to assume that the Cubs don't care about winning, does it, when they routinely spend literally millions of dollars more than other teams, in order to try to win?


Labels: , , , , ,

Monday, January 24, 2011

"Scorecasting:" is home field advantage just biased officiating?

There's a new book that's about to come out, that you might have heard about. It's called "Scorecasting," and the publisher was kind enough to send me a review copy. It's basically a Freakonomics for sports, in intent, in tone, in writing style, and even down to its similar authorship -- one academic economist (Tobias Moskowitz, a finance professor) and one journalist (Sports Illustrated's L. Jon Wertheim).

The book's website is here; it has an excerpt from one of the chapters, and you'll find other excerpts online if you search the authors' names.


The topics will be very familiar to sabermetricians and regular readers of this blog. There are chapters on the Hot Hand, on competitive balance, on NBA refereeing, on steroids, and so on. There isn't a huge amount of breakthrough stuff there, although there are certainly a few new insights. Mostly, the authors summarize what they've learned from academic articles on sports, and they add the results of a few little studies they did themselves.

Alas, by concentrating on the journals, they've missed much of the scholarship of us amateurs. For instance, in the chapter of competitive balance, they argue that baseball is less balanced than the football because MLB teams play 162 games, while NFL teams only play 16. That, of course, is only a small part of the story. There are other parts, such as the distribution of talent in the league, and the internal details of the game itself. Tom Tango has effectively solved the problem of comparing different sports (here's just one of his many posts), but the authors seem unaware of that (although they do mention Tango in the book once, in a mention of leverage, referring to him as a "stats whiz.")

And they're occasionally completely off, as when they say the sample size of the MLB playoffs is enough that the best team ought to win the series.

Still, a lot of the material is solid; the authors are at their strongest when they're reviewing one of the more famous and established studies, like the Romer "fourth down" paper, and the Massey/Thaler NFL draft study. (I'll probably do a full review of the book later, but, for now, just picture a sports Freakonomics that's not as rigorous as most of the websites, but does mention a few things that you didn't know before.)

Anyway, I'm going through the book, and suddenly I see that the authors claim to have solved the problem of what causes home field advantage (HFA). That was sort of shocking. My personal subjective view is that HFA is the biggest unsolved problem in sabermetrics, and very little progress has been made. There's so little progress, in fact, that I've started to take seriously a hypothesis that seems way off the wall -- the theory that humans have built in evolutionary programming that makes them more physically and mentally effective when defending their own turf. (I'm not saying that it's necessarily true, just that I have a bizarre attraction to it.) In that light, finding these HFA claims was a bit like picking up a newspaper article on math, and finding that the reporter has proved Fermat's Last Theorem.

So what's the authors' solution to the long-standing HFA conundrum? Refereeing. After dismissing most of the usual suspects (fan enthusiasm, travel, tailoring the team to the park), Moskowitz and Wertheim believe that most, or all, of HFA can be explained by biased officiating.

They list a bunch of supporting evidence, which I'll summarize here. If you want to follow along, some of this stuff is also in a long excerpt from the book that appeared a couple of weeks ago in the Jan. 17 issue of Sports Illustrated (the article, unfortunately, does not appear to be online).

------

Soccer

1. In soccer, the referee controls how much extra "injury time" is added to the end of a match. It turns out that injury time is longer when the home team would benefit. When the home side was ahead by a goal, there were two minutes of injury time, on average, in a sample of Spanish league games. But when they were *behind* by a goal, it was four minutes.

2. In 1998, the point structure changed to give the winning team three standings points instead of two. Immediately, the above injury time bias increased.

3. The same bias exists in England, Italy, Germany, Scotland, and the US.

4. "... home teams receive many fewer red and yellow cards even after controlling for the number of penalties and fouls on both teams."


Baseball

5. In baseball, the authors looked at the percentage of called pitches that are strikes. In crucial situations (high leverage), home teams got a lot more favorable calls. But in low-leverage situations, it was *road* teams that got more favorable calls. "This makes sense," the authors write. "If the umpire is going to show favoritism to the home team, he or she will do it when it is most valuable -- when the outcome of the game is affected the most. You might even contend that it noncrucial situations the umpire might be biased against the home team to maintain an overall appearance of fairness."

6. " ... the success rates of home teams in scoring from second base on a single or scoring from third base on an out -- typically close plays at the plate -- are much higher than they are for their visitors in high-leverage/crucial situations. yet they are no different or even slightly less successful in noncrucial situations."

7. Over a large sample of 5.5 million pitches, "called strikes and balls went the home team's way, *but only* in stadiums without QuesTec ... Not only did umpires not favor the home team when QuesTec was watching them, they actually gave *more* strikes and *fewer* balls to the home team. In short, when umpires knew they were being monitored, home field advantage on balls and strikes didn't simply vanish; the advantage swung all the way to the visiting team."

8. In low-leverage situations, even in non-QuesTec parks, there was no bias at all.

9. The authors then analyzed pitches using Pitch f/x data, to see how many pitches were miscalled based on the recorded location. For pitches on the corner of the strike zone, there were more miscalls in the home team's favor than in the visiting team's favor. The home advantage was largest on full-count pitches, followed by other three-ball counts, other two-strike counts, and, lastly, all other counts. So, the more crucial the pitch, the greater the HFA.

10. "Over the course of the season, all of this adds up to 516 more strikeouts called on away teams, and 195 more walks awarded to home teams than there otherwise should be, thanks to the home plate umpire's bias. And that includes only terminal pitches -- where the next called pitch will result in either a strikeout or a walk. Errant calls given earlier in the pitch count could confer an even greater advantage on the home team."

11. "This adds up to an extra 7.3 runs per season given to each home team by the plate umpire alone. That might not sound significant, but cumulatively, home teams outscore their visitors by only 10.5 runs in a season." [That latter number isn't correct ... in 2010, it was 23.5 runs. 23.5 runs equals 2.35 wins out of 81, which is a .530 winning percentage. (UPDATE: Oops! I forgot to adjust for the home team not batting in the bottom of the ninth when leading. If you adjust for that, the home advantage is a lot bigger than 10.5 or 23.5 runs.)]


Football

12. In the NFL, "Home teams receive fewer penalties than away teams -- about half a penalty less per game -- and are charged with fewer yards per penalty. Of course, this does not necessarily mean officials are biased. But when we looked at more crucial situations in the NFL ... we found that the penalty bias [increases]."

13. When instant replay came to the NFL, the home winning percentage declined from 58.5 percent (1985-98) to 56 percent (1999-2008). "Before instant replay, home teams enjoyed more than an 8 percent edge in turnovers ... When instant replay came along ... the turnover advantage was cut in half." Also, "the home team does not actually fumble or drop the ball less often than the away team ... they simply lose fewer fumbles than away teams. After instant replay was installed, however, the home team advantage of *losing* fewer fumbles miraculously disappeared, whereas the frequency of fumbles remained the same. ... In close games, where referees' decisions may *really* matter ... home teams enjoyed a healthy 12 percent advantage in recovering fumbles. After instant replay was installed, that advantage simply vanished."

14. After instant replay, there was no change in the relative frequency of home and away penalties. That might be because penalties can't be challenged.

15. Away teams have their challenges upheld 37 percent of the time, versus 35 percent for home teams. But when the home team is losing, the visiting team wins 40 percent, versus only 28 percent for the home team. So it looks like the referees favor the home team more when they need it more.


Basketball

16. In the NBA, fouls and turnovers that are not subjective referee calls (like shot clock violations) are equal for home and road teams. But for subjective calls, away teams get between 1 and 1.5 more of those per game. Visiting players are 15 percent more likely to be called for traveling than home players.

17. "How much of the [HFA] in the NBA is due to referee bias? If we attribute the differences in free throw attempts to referee bias, this would account for 0.8 points per game. If we gave credit to the referees for the more ambiguous turnover differences ... this would also capture another quarter of the home team's advantage. Attributing some of the other foul differences to the referees and adding the effects of those fouls (other than free throws) ... brings the total to about three-quarters of the home team's advantage. And, remember, scheduling in the NBA [visiting teams play more back-to-back games than home teams] explained about 21 percent of [HFA]. This adds up to nearly all of the NBA home court advantage."


Hockey

18. In the NHL, home teams get 20 percent fewer penalties and receive fewer minutes per penalty. "On average, home teams get two and a half more minutes of power play opportunities ... than away teams. That is a *huge* advantage." If you multiply that by a 20 percent success rate, you get an extra 0.25 goals per game for the home team. Since the average overall differential is only 0.3 goals for the home team, "this alone accounts for more than 80 percent of the home ice advantage in hockey."

19. There is no apparent HFA in shootouts, where refereeing makes no difference. Also, in NBA foul shooting. And, even in Pitch f/x data. Visiting pitchers throw no worse, according to Pitch f/x, than home teams do. It's only the umpires' calls that are different.

-----

It's an impressive array of evidence and argument. But, at least some of it doesn't hold up.

Look at number 5: in baseball, in low leverage situations (I believe this means the bottom 50%), the authors say that umpires favor the visiting team. That would mean that, in less critical situations, we should find a "visiting field advantage." But home teams outscore visiting teams even in medium-leverage situations. For instance, here's the breakdown of home and road runs scored by inning (1954 to 2007). The last column is the percentage by which the home team outscored the visiting team:

1 61872-52071 +18%
2 46823-42539 +10%
3 53590-48188 +11%
4 53357-49593 +8%
5 53203-48448 +10%
6 54401-50603 +8%
7 52231-48641 +7%
8 50451-47781 +6%

You would think that you'd have more high-leverage events in the later innings -- but the HFA goes *down* in the last few innings, not up.

But I might be wrong about that, maybe the eighth inning has no more high-leverage situations than the first inning (after all, there are more 8-1 games in the eighth than in the first). So, let's look at innings where, at the start, one team was at least four runs ahead of the other. Those should all be low leverage, for the most part, and should show the visiting team having the advantage.

Nope:

2 2543-2139 +19%
3 4583-4176 +10%
4 8817-7801 +13%
5 10940-10057 +9%
6 14371-13279 +8%
7 15698-14583 +8%
8 16935-16180 +5%

Now, this could be just because, in a four-run game, the home teams are a lot better than the visiting teams. What if we look at situations when the *visiting* team is ahead by at least four runs? Then, we should see a huge effect in favor of the visiting team: first, they're probably a much better team, and, second, the low leverage means the umpire should still be favoring them.

But, no. Even in those situations, the home team still performs a little better, on average, having the advantage in five of the seven cases:

2 957-1022 -6%
3 1974-1799 +10%
4 3609-3355 +8%
5 4435-4645 -5%
6 6269-5705 +10%
7 6627-6562 +1%
8 7309-7179 +2%


So, I just don't see it. If umpires DO call more strikes for visiting teams in low-leverage situations, maybe that's compensated for by those pitches actually being strikes ... but being worse pitches in location and movement and velocity. That is, maybe HFA comes from pitchers throwing more accurately, but more hittably.

In any case, if my data are correct, and the authors' data are also correct, it can't be the case that the authors' findings are an explanation of HFA.

------

Now, let's look at number 18, the hockey case. The authors argue that HFA is caused almost entirely by penalties. If that's the case, then you'd expect home and visiting teams to have similar numbers at equal strength.

They do not. The NHL.com website has home/road goal breakdowns. Here they are for the 2008-09 season, averaged by team:

Even strength... 124-110 (home advantage 12.5%)
Power play...... 35-30 (home advantage 15.1%)
Shorthanded..... 4-4 (home advantage 1.0%)

There's almost as large an advantage at even strength as there is on the power play. Admittedly, the extra power play boost is probably caused by more penalties, as the authors say, but the overall contribution of the extra penalties seems to be pretty small.

Just to make sure it wasn't a fluke, I ran the same numbers for 2009-10:

Even strength... 121-106 (home advantage, 13.9%)
Power play...... 30-25 (home advantage, 21.0%)
Shorthanded..... 4-3 (home advantage, 32.9%)

A bit more extreme in favor of power play. But how do you explain the sizeable advantage for home teams at even strength? One possible explanation is that visiting teams have to play an overcautious game, to avoid being penalized by biased referees. But for a 13.9% disadvantage, that caution would have to be way out of line, wouldn't it?


------

Both of these examples -- and, by the way, they're the only two I checked -- cast doubt on the authors' hypothesis that HFA is almost all refereeing. I have never disagreed that *some* of it might be refereeing, but there's obviously a lot more going on.

And I have to say that the authors have indeed provided a blueprint for how this kind of research should go -- try to break down performance into its constituent parts, and check those.

If there's no home advantage in foul shooting, why not? If there's no HFA in hockey shootouts, why not? If we get a list of areas with high HFA, and a list of areas with low HFA, we can maybe start narrowing down what the causes might be.

But the authors have amassed a lot of evidence, and the must be something to at least some of it, no? For instance, I can't think of any explanation for the injury time phenomenon (maybe I should look up the relevant study). And it seems reasonable that referees will call more fouls on visitors, even if they're unbiased. Why? Because they might be using crowd noise as a guide ot what is and what isn't a foul. If the fans scream when a visitor trips an opponent, but not when a home player trips an opponent, that will simply make it more likely that an unbiased referee will have enough evidence to correctly "convict" the visiting player.

But the question is not just whether referee bias exists, but *how much* of it there is, and how much of HFA it's responsible for. The authors of "Scorecasting" seem more focused on "existence" evidence, and it seems to me they've made only a small dent in terms of explaining the real-life observed HFA. I wish the authors had provided more details of some of their findings, so we can figure out what's going on and maybe quantify it a bit more ... but I guess it is what it is.

I know there are a lot of working sabermetricians reading this ... if you have expertise or evidence on any of the authors' points, please weigh in.

-------


UPDATE: I have a full review of the book here.


Labels: ,