Monday, October 24, 2016

Why the 2016 AL was harder to predict than the 2016 NL

In 2016, team forecasts for the National League turned out more accurate than they had any right to be, with FiveThirtyEight's predictions coming in with a standard error (SD) of only 4.5 wins. The forecasts for the American League, however, weren't nearly as accurate ... FiveThirtyEight came in at 8.9, and Bovada at 8.8. 

That isn't all that great. You could have hit 11.1 just by predicting each team to duplicate their 2015 record. And, 11 wins is about what you'd get most years if you just forecasted every team at 81-81.

Which is kind of what the forecasters did! Well, not every team at 81-81 exactly, but every team *close* to 81-81. If you look at FiveThirtyEight's actual predictions, you'll see that they had a standard deviation of only 3.4 wins. No team was predicted to win or lose more than 87 games.

Generally, team talent has an SD of around 9 wins. If you were a perfect evaluator of talent, your forecasts would also have an SD of 9. If, however, you acknowledge that there are things that you don't know (and many that can't be known, like injuries and suspensions), you'll forecast with an SD somewhat less than 9 -- maybe 6 or 7.

But, 3.4? That seems way too narrow. 

Why so narrow? I think it was because, last year, the AL standings were themselves exceptionally narrow. In 2015, no American League team won or lost more than 95 games. Only three teams were at 89 or more. 

The SD of team wins in the 2015 AL was 7.2. That's much lower than the usual figure of around 11. In fact, 7.2 is the lowest for either league since 1961. In fact, I checked, and it's the lowest for any league in baseball history! (Second narrowest: the 1974 American League, at 7.3.)

Why were the standings so compressed? There are three possibilities:

1. The talent was compressed;

2. There was less luck than normal;

3. The bad teams had good luck and the good teams had bad luck, moving both sets closer to .500.

I don't think it was #1. In 2016, the SD of standings wins was back near normal, at 10.2. The year before, 2014, it was 9.6. It doesn't really make sense that team talent regressed so far to the mean between 2014 and 2015, and then suddenly jumped back to normal in 2016. (I could be wrong -- if you can find trades and signings those years that showed good teams got significantly worse in 2015 and then significantly better in 2016, that would change my mind.)

And I don't think it was #2, based on Pythagorean luck. The SD of the discrepancy in "first-order wins" was 4.3, which larger than the usual 4.0. 

So, that leaves #3 -- and I think that's what it was. In the 2015 AL, the correlation between first-order-wins and Pythagorean luck was -0.54 instead of the expected 0.00. So, yes, the good teams had bad luck and the bad teams had good luck. (The NL figure was -0.16.)

-------

When that happens, that luck compresses the standings, it definitely makes forecasting harder. Because, there's not as much information on how teams differ. To see that, consider the extreme case. If, by some weird fluke, every team wound up 81-81, how would you know which teams were talented but unlucky, and which were less skilled but lucky? You wouldn't, and so you wouldn't know what to expect next season.

Of course, that's only a problem if there *is* a wide spread of talent, one that got overcompressed by luck. If the spread of talent actually *is* narrow, then forecasting works OK. 

That's what many forecasting methods assume, that if the standings are narrow, the talent must be narrow. If you do the usual "just take the standings and regress to the mean" operation, you'll wind up implicitly assuming that the spread of talent shrank at the same time as the spread in the standings shrank.

Which is fine, if that's what you think happened ... but, do you really think that's plausible? The AL talent distribution was pretty close to average in 2014. It makes more sense to me to guess that the difference between 2014 and 2015 was luck, not wholesale changes in personnel that made the bad teams better and the good teams worse.

Of course, I have the benefit of hindsight, knowing that the AL standings returned to near-normal in 2016 (with an SD of 10.2). But it's happened before -- the record-low 7.3 figure for the 1974 AL jumped back to an above-average 11.9 in 1975.

I'd think when I was forecasting the 2016 standings, I might want to make an effort to figure out which teams were lucky and which ones weren't, in order to be able to forecast a more realistic talent SD than 3.5 wins.

Besides, you have more than the raw standings. If you adjust for Pythagoras, the SD jumps from 7.2 to 8.6. And, according to Baseball Prospectus, when you additionally adjust for cluster luck, the SD rises to 9.4. (As I wrote in the P.S. to the last post, I'm not confident in that number, but never mind for now.)

An SD of 9.4 is still smaller than 11, but it should be workable.

Anyway, my gut says that you should be able to differentiate the good teams from the bad with a spread higher than 3.4 games ... but I could be wrong. Especially since Bovada's spread was even smaller, at 3.3.

-------

It's a bad idea to second-guess the bookies, but let's proceed anyway.

Suppose you thought that the standings compression of 2015 was a luck anomaly, and the distribution of talent for 2016 should still be as wide as ever. So, you took FiveThirtyEight's projections, and you expanded them, by regressing them away from the mean, by a factor of 1.5. Since FiveThirtyEight predicted the Red Sox at four games above .500 (85-77), you bump that up to six games (87-75).

If you did that, the SD of your actual predictions is now a more reasonable 5.1. And those predictions, it turns out, would have been better. The accuracy of your new predictions would have been an SD of 8.4. You would have beat FiveThirtyEight and Bovada.

If that's too complicated, try this. If you had tried to take advantage of Bovada's compressed projections by betting the "over" on their top seven teams, and the "under" on their bottom seven teams, you would have gone 9-5 on those bets.

Now, I'm not going to so far as to say this is a workable strategy ... bookmakers are very, very good at what they do. Maybe that strategy just turned out to be lucky. But it's something I noticed, and something to think about.

-------

If compressed standings make predicting more difficult, then a larger spread in the standings should make it easier.

Remember how the 2016 NL predictions were much more accurate than expected, with an SD of 4.5 (FiveThirtyEight) and 5.5 (Bovada)? As it turns out, last year, the SD of the 2015 NL standings was higher than normal, at 12.65 wins. That's the highest of the past three years:

2014  AL= 9.59, NL= 9.20
2015  AL= 6.98, NL=12.65
2016  AL=10.15, NL=10.71

It's not historically high, though. I looked at 1961 to 2011 ... if the 2015 NL were included, it would be well above average, but only 70th percentile.*

(* If you care: of the 10 most extreme of the 102 league-seasons in that timespan, most were expansion years, or years following expansion. But the 2001, 2002, and 2003 AL made the list, with SDs of 15.9, 17.1, and 15.8, respectively. The 1962 National League was the most extreme, at 20.1, and the 2002 AL was second.)

A high SD won't necessarily make your predictions beat the speed of light, and a low SD won't necessarily make them awful. But both contribute. As an analogy: just because you're at home doesn't mean you're going to pitch a no-hitter. But if you *do* pitch a no-hitter, odds are, you had the help of home-field advantage.

So, given how accurate the 2016 NL forecasts were, I'm not surprised that the SD of the 2015 NL standings was higher than normal.

-------

Can we quantify how much compressed standings hurt next year's forecasts? I was curious, so I ran a little simulation. 

First, I gave every team a random 2015 talent, so that the SD of team talent came out between 8.9 and 9.1 games. Then, I ran a simulated 2015 season. (I ran each team with 162 independent games, instead of having them play each other, so the results aren't perfect.)

Then, I regressed each team's 2015 record to the mean, to get an estimate of their talent. I assumed that I "knew" that the SD of talent was around 9, so I "unregressed" each regressed estimate away from the mean by the exact amount that gets the SD of talent to exactly 9.00. That became the official forecast for 2016. 

Finally, I ran a simulation of 2016 (with team talent being the same as 2015). I compared the actual to the forecast, and calculated the SD of the forecast errors.

The results came out, I think, very reasonable.

Over 4,000 simulated seasons, the average accuracy was an SD of 7.9. But, the higher the SD of last year's standings, the better the accuracy:

SD Standings    SD next year's forecast
------------------------------------
7.0             8.48 (2015 AL)
8.0             8.31
9.0             8.14
10.0            7.98
11.0            7.81
12.0            7.64
12.6            7.54 (2015 NL)
13.0            7.47
14.0            7.31
20.1            6.29 (1962 NL)

So, by this reckoning, you'd expect the 2016 NL predictions to have been one win more accurate than than the AL predictions. 

They were "much more accurater" than that, of course, by 3.4 or 4.5. The main reason, of course, is that there's a lot of luck involved. Less importantly, this simulation is very rough. The model is oversimplified, and there's no assurance that the relationship is actually linear. (In fact, the relationship *can't* be linear, since the "speed of light" limit is 6.4, and the model says the 1974 AL would beat that, at 6.3). 

It's just a very rough regression to get a very rough estimate. 

But the results seem reasonable to me. In 2016, we had (a) the narrowest standings in baseball history in the 2015 AL, and (b) a wider-than-average, 70th percentile spread in the 2016 NL. In that light, an expected difference of 1 win, in terms of forecasting accuracy, seems very plausible. 

--------

So that's my explanation of why this year's NL forecasts were so accurate, while this year's AL forecasts were mediocre. A large dose of luck -- assisted by a small (but significant) dose of extra information content in the standings.












Labels: , , , , ,

Friday, October 21, 2016

National League forecasts were too accurate in 2016

FiveThirtyEight predicted the National League surprisingly accurately this year.

The standard error of their predictions -- that is, the SD of the difference between their team forecasts, and what actually happened -- was only 4.5 games.* (Here's the link to their forecast -- go to the bottom and choose "April 2".)

(* The SD is the square root of the average squared error. If you prefer just the average error, in this case, it was three-and-a-third games. But I'll be using just the SD in the rest of this post. In most cases, to estimate average error when you only have the SD, you can multiply by 2/pi (approximately 0.64).)

4.5 games is very, very good. In fact, it's so good it can't possibly be all skill. The "speed of light" limit on forecasting MLB is about 6.4 games. That is, even if you knew absolutely everything about the talent of a team and its opposition, every game, an SD of 6.4 is the very best you could expect to do.

Of course, you can get lucky, and beat 6.4 games. You could even get to zero, if fortune smiles on you and every team hits your projection exactly. But, 6.4 is the best you can do by skill.**  

(** Actually, it might be a bit less, 6.3 or something, because 6.4 is what you get when teams are evenly matched ... mismatches are somewhat easier to predict. But never mind.)

How unusual is an SD of 4.5? Well, not *that* unusual. By my estimate, the SD of the observed SD -- sorry if that's a little confusing -- is somewhere around 1.7, for a league of 15 teams. So, FiveThirtyEight was a little over one standard deviation lucky, which isn't really a big deal. Even taking into account that FiveThirtyEight couldn't have been perfectly accurate in their talent assessments, it's still not that big a deal. If they were off, on talent, by around 3 games per team, that would bring them to only about 1.5 SDs of luck.

Still not a huge deal, but interesting nonetheless.

------

It wasn't just FiveThirtyEight whose projections did well ... the Vegas bookmakers did OK too. Well, at least the one I looked at, Bovada. (I assume the others would be pretty close.)  They had an SD of 5.5 games, which is also better than the "speed of light."  (I can't find the page I got them from, but this one, from a month earlier, is close.)

That suggests that it probably wasn't any particular flash of brilliance from either FiveThirtyEight or Bovada ... it must have been something about the way the season unfolded. 

Maybe, in 2016, there was less random deviation than usual? One type of random variation is whether a team exceeds their Pythagorean Projection -- that is, whether they win more (or fewer) games than you'd expect from their runs scored and allowed. To check that, I used Baseball Prospectus's numbers -- specifically, the difference between actual and "first-order wins."***

(*** Why didn't I use second-order wins? See the P.S. at the bottom of the post.)

In the National League in 2016, the SD of Pythagorean error was 3.55. That is indeed a little smaller than the average of around 4.0. But that small difference isn't nearly enough to explain why the projections were so good.

Here's what I think is the bigger factor -- actually, a combination of two factors.

First, by random chance, the better teams happened to undershoot their Pythagorean expectation, and the worse teams happened to exceed it. 

The Cubs were the best team in the league, and also the team with the most bad luck, -4.8 games. The Phillies were the worst team in the league with luck removed; you'd expect them to have won only 61.4 games, but they but played +9.6 games above their Pythagorean projection to go 71-91.

Those two were the most obvious examples, but the pattern continued through the league. Overall, the correlation between first-order wins (which is an approximation of talent) and Pythagorean error was huge: +0.61. Normally, you'd expect it to be close to zero. (In the American league, it was, indeed, close to zero, at -0.06.)

Second, there was a similar, offsetting relationship in the predictions themselves. 

It turns out that the forecast errors had a strong pattern this year.  Instead of being random, they came out too "conservative" -- they underestimated the talent of the better teams, and overestimated the talent of the worse teams. Here's the distribution of FiveThirtyEight's forecast errors, with the teams sorted by their forecast:

Top 5 teams: average error -4 wins (underestimate)
Mid 5 teams: average error +4  win (overestimate)
Btm 5 teams: average error +1  win (overestimate)

So, in summary:

-- FiveThirtyEight predicted teams too close to the mean
-- Teams' Pythagorean luck moved them closer to the mean

Those two things cancelled each other out to a significant extent. And that's why FiveThirtyEight was so accurate.

-------

Next post: The American League, which is interesting for completely different reasons.

-------

P.S. Baseball Prospectus also produces "second-order wins," which attempts to remove a second kind of luck, what I call "Runs Created luck" (and others call "cluster luck"), which is teams scoring more or fewer runs than would be expected by their batting line. I started to do that, but ... I stopped, because I found something weird.

When you remove luck from the standings, you expect to make them tighter, to bring teams closer together. (To see that better, imagine removing luck from coin tosses. Every team reverts to .500.)

Removing first-order (Pythagorean) luck does seem to reduce the SD of the standings. But, removing second-order (Cluster) luck seems to do the *opposite*.

I checked four seasons of BP data, and, in every case, the SD of second-order wins (for the set of all 30 teams) was higher than the SD of first-order wins:  

         Actual  First-order  Second-order
------------------------------------------
2016      10.7        10.8        13.1
2015      10.4        10.1        11.8
2014       9.6         8.9         9.6
2013      12.2        12.2        12.8

So, either the good teams got lucky all four years, or there's something weird about how BP is computing second-order luck. 










Labels: , , , , ,

Saturday, April 11, 2015

MLB forecasters are more conservative this year

Every April, sabermetricians, bookies and sportswriters issue their won-lost predictions for each of the 30 teams in MLB. And, every year, some of them are overconfident, and essentially wind up trying to predict coin flips.

As I've written before, there's a mathematical "speed of light" limit to how much of a team's record can be predicted. That's the part that's determined by player talent. Any spread that's wider than the spread of talent must be just random luck, and, by definition, unpredictable.

Based on the historical record, we can estimate that the spread of team talent in MLB is somewhere around an SD of 9 games. Not all of that talent can be predicted beforehand, because some of it hasn't happened yet -- trades, injuries, players unexpectedly blossoming or declining, and so on. My estimate is that if you were the most omnicient baseball insider in the universe, maybe you could predict an SD of 8 games.

Last year, many pundits exceeded that "speed of light" limit anyway. I thought that there would be fewer this year, that the 2015 forecasts would project a narrower range of team outcomes. That's because last year's standings were particularly tight, and there's been talk about how we may be entering a new era of parity.

And that did happen, to some extent.

I'll show you the 2015s and the 2014s together for easy comparison. A blank space is a forecast I don't have for that year. (For 2015, I think Jonah Keri and the ESPN website predicted only the order of finish, and not the actual W-L record.) 

Like last year, I've included the "speed of light" limits, the naive "last year regressed to the mean" forecast, and the "every team will finish .500" forecast. Links are for 2015 ... for 2014, see last year's post.


 2015  2014
--------------------------------------------------
 9.32 11.50  Sports Illustrated
 9.07  8.76  Jeff Passan (Yahoo!)
 9.00  9.00  Speed of Light (theoretical est.)
 8.79        Bruce Bukiet
       8.53  Jonah Keri (Grantland)
       8.51  Sports Illustrated (runs/10)
 8.00  8.00  Speed of Light (practical est.)
       7.79  ESPN website
 7.92  7.78  Mike Oz (Yahoo!)
 6.99        Chris Cwik (Yahoo!)
 6.38  6.38  Naive previous year method (est.)
 6.34  9.23  Mark Townsend (Yahoo!)
 6.10  6.90  Tim Brown (Yahoo!)
 6.03  7.16  Vegas Betting Line (Bovada)
 5.46  5.55  Fangraphs 
 4.93  8.72  ESPN The Magazine 
 0.00  0.00  Predict 81-81 for all teams
--------------------------------------------------

Of those who predicted both seasons, only two out of eight increased the spread of their forecasts from last year. And those two, Jeff Passan and Mike Oz, increased only a little bit. 

On the other hand, some of the other forecasters see *dramatically* more equality in team talent. Yahoo's Mark Townsend dropped from 9.23 to a very reasonable 6.34. And ESPN dropped from one of the highest spreads, to the lowest -- by far -- at 4.93. 

Which is strange, because ESPN's words are so much more optimistic than their numbers. about the Washington Nationals, they write,


"It's the Nationals' time."
"They're here to stay."
"Anything less than an NL East crown will qualify as a big disappointment."

But their W-L prediction for the Nationals, whom they projected higher than any other team?  A modest 91-71, only ten games above .500.

------

In any case ... I wonder how much of the differences between 2014 and 2015 are due to (a) new methodologies, (b) the same methodologies reflecting a real change in team parity, and (c) just a gut reaction to the 2014 standings having been tighter than normal.

My guess is that it's mostly (b). I'd bet on Bovada's forecasts being the most accurate of the group. If that's true, then maybe teams really *are* tighter in talent than last year, by around 1 win of SD. Which is, roughly, in line with the rest of the forecasts.

I guess we'll know more next year. 



Labels: , , ,

Saturday, July 12, 2014

Nate Silver and the 7-1 blowout

Brazil entered last Tuesday's World Cup semifinal match missing two of their best players -- Neymar, who was out with an injury, and Silva, who was sitting out a red-card suspension. Would they still be good enough to beat Germany?

After crunching the numbers, Nate Silver, at FiveThirtyEight, forecasted that Brazil still had a 65 percent chance of winning the match -- that the depleted Brazilians were still better than the Germans. In that prediction, he was taking a stand against the betting markets, which actually had Brazil as underdog -- barely -- at 49 percent. 

Then, of course, Germany beat the living sh!t out of Brazil, by a score of 7-1. 


"Time to eat some crow," Nate wrote after Brazil had been humiliated. "That prediction stunk."

I was surprised; I had expected Nate to defend his forecast. Even in retrospect, you can't say there was necessarily anything wrong with it.

What's the argument that the prediction stunk?  Maybe it goes something like this:

-- Defying the oddsmakers, Nate picked Brazil as the favorite.
-- Brazil suffered the worst disaster in World Cup history. 
-- Nate's prediction was WAY off.
-- So that has to be a bad prediction, right?

No, it doesn't. It's impossible to know in advance what's going to happen in a soccer game, and, in fact, anything at all could happen. The best anyone can do is try to assign the best possible estimate of the probabilities. Which is what Nate did: he said that there was a 65% chance that Brazil would win, and a 35% chance they would lose. 

Nate said Brazil had about twice as much chance of winning as Germany did. He did NOT say that Brazil would play twice as well. He didn't say Brazil would score twice as many goals. He didn't say Brazil would control the ball twice as much of the time. He didn't say the game would be close, or that Brazil wouldn't get blown out. 

All he said was, Brazil has twice the probability of winning. 

The "65:35" prediction *did* imply that Nate thought Brazil was a better team than Germany. But that's not the same as implying that Brazil would play better this particular game. It happens all the time, in sports, that the better team plays like crap, and loses. That's all built in to the "35 percent". 

Here's an analogy. 

FIFA is about to pick a random number of dollars between 1 and 1,000,000. I say, there's a 65 percent chance that the number drawn will be higher than the value of a three-bedroom bungalow, which is $350,000. 

That's absolutely a true statement, right?  650,000 "winning" balls out of a million is 65 percent. I've made a perfect forecast.

After I make my prediction, FIFA reaches into the urn, pulls out out one of the million balls, and it's ... number 14. 

Was my prediction wrong?  No, it wasn't. It was exactly, perfectly correct, even in retrospect.

It might SEEM that my prediction was awful, if you don't understand how probability works, or you didn't realize how the balls were numbered, or you didn't understand the question. In that case, you might gleefully assume I'm an idiot. You might say, "Are you kidding me?  Phil predicted you could buy a house for $14! Obviously, there's something wrong with his model!"

But, there isn't. I knew all along that there was a chance of "14" coming up, and that factored into my "35 percent" prediction. "14" is, in fact, a surprisingly low outcome, but one that was fully anticipated by the model.

When Nate said that Brazil had a 35 percent chance of losing, a small portion of that 35 percent was the chance of those rare events, like a 7-1 score -- in the same way my own 35 percent chance included the rare event of a really small number getting drawn. 

As unintuitive as it sounds, you can't judge Nate's forecast by the score of the game. 

-------

Critics might dispute my analogy by arguing something like this:

"The "14" result in Phil's model doesn't show he was wrong, because, obviously, which ball comes out of the urn it just a random outcome. On the other hand, a soccer game has real people and real strategies, and a true expert would have been able to foresee how Germany would come out so dominating against Brazil."

But ... no. An expert probably *couldn't* know that. That's something that was probably unknowable. For one thing, the betting markets didn't know -- they had the two teams about even. I didn't hear any bettor, soccer expert, sportswriter, or sabermetrician say anything otherwise, like that Germany should be expected to win by multiple goals. That suggests, doesn't it, that it was legimately impossible to foresee?

I say, yes, it was definitely unknowable. You can't predict the outcome of a single game to that extent -- it's a violation of the "speed of light" limit. I would defy you to find any single instance where anyone, with money at stake, seriously predicted a single game outcome that violates conventional wisdom to anything near this extent. 

Try it for any sport. On August 22, 2007, the Rangers were 2:3 underdogs on the road against the Orioles. They won 30-3. Did anyone predict that?  Did anyone even say the Rangers should be heavy favorites?  Is there something wrong with Vegas, that they so obviously misjudged the prowess of the Texas batters?

Of course not. It was just a fluke occurrence, literally unpredictable by human minds. Like, say, 7-1 Germany.


Huh? [Nate Silver] says his prediction  “stunk,” but it was probabilistic. No way to know if it was even wrong. 

Exactly correct. 

--------

So I don't think you can fault Nate's prediction, here. Actually, that's too weak a statement. I don't mean you have to forgive him, as in, "yeah, he was wrong, but it was a tough one to predict."  I don't mean, "well, nobody's perfect."  I mean: you have no basis even for *questioning* Nate's prediction, if your only evidence is the outcome of the game. Not as in, "you shouldn't complain unless you can do better," but, as in, "his prediction may well have been right, despite the 7-1 outcome."  

But I did a quick Google search for "Brazil 7-1 Nate Silver," and every article I saw that talked about Nate's prediction treated it as certain that his forecast was wrong.

1. Here's one guy who agrees that it's very difficult to predict game results. From there, he concludes that all predictions must therefore be bullsh!t (his word). "Why did they even bother updating their odds for the last three remaining teams at numbers like 64 percent for Germany, 14 percent for the Netherlands, when we just saw how useless those numbers can be?"

Because, of course, the numbers *aren't* bullsh!t, if you correctly interpret them as probabilities and not certainties. If you truly believe that no estimate of odds is useful unless it can successfully call the winner of every game, then how about you bet me every game, taking the Vegas underdog at even money?  Then we'll see who's bullsh!tting.

2. This British columnist gets it right, but kind of hides his defense of Nate in a discussion of how sports pundits are bad at predicting. Except that he means that sabermetricians are bad at correctly guessing outcomes. Well, yes, and we know that. But we *are* fairly decent at predicting probabilities, which is all that Nate was trying to do, because he knows that's all that can realistically be done.


"I love Nate Silver and 538, but this result might be breaking his model. Haven't been super impressed with the predictions."

What, in particular, wasn't this guy impressed with?  He can't just be talking about game results, can he?  Because, in the knockout round, Nate's predicted favorites won *every game* up to Germany/Brazil. Twelve in a row. What would have "super impressed" this guy, 13 out of 12?

Here's another one: 

"To be fair to Nate Silver + 538, their model on the whole was excellent. It's how they dealt with Brazil where I (and others) had problems."

What kind of problems?  Not picking them to lose 7-1?  

In fairness, sure, there's probably some basis for critiquing Nate's model, since he's been giving Brazil siginficantly higher odds than the bookies. But, in this case, the difference was between 65% and 49%, not between 65% and "OMG, it's a history-making massacre!"  So this is not really a convincing argument against Nate's method.

It's kind of like your doctor says, "you should stop smoking, or you're going to die before you're 50!"  You refuse, and the day before your fiftieth birthday, a piano falls on your head and kills you. And the doctor says, "See? I was right!" 

4. Here's a mathematical one, from a Guardian blogger. He notes that Nate's model assumed goals were independent and Poisson, but, in real life, they're not -- especially when a team collapses and the opponent scores in rapid-fire fashion.

All very true, but that doesn't invalidate Nate's model. Nate didn't try to predict the score -- just the outcome. Whether a team collapses after going down 3-0, or not, doesn't much affect the probability of winning after that, which is why any reasonable model doesn't have to go into that level of detail.

Which is why, actually, losing 7-1 loss isn't necessarily inconsistent with being a favorite. Imagine if God had told the world, "if Brazil falls behind 2-0, they'll collapse and lose 7-1." Nate, would have figured: "Hmmm, OK, so we have to take subtract off the chance of 'Brazil gives up the first two goals, but then dramatically comes back to win the game,' since God says that can't happen."

Nate would have figured that's maybe a 1 percent of all games, and say, "OK, I'm reducing my 65% to 64%."  

So, that particular imperfection in the model isn't really a serious flaw. 

But, now that I think about it ... imagine that when Nate published his 65% estimate, he explicitly mentioned, "hey, there's still a 1-in-3 chance that Brazil could lose ... and that includes a chance that Germany will kick the crap out of them. So don't get too cocky."  That would have helped him, wouldn't it?  It might even have made him look really good!

I mean, he shouldn't need to say it to statisticians, because it's an obvious logical consequence of his 65% estimate. But maybe it needs to be said to the public.


"It's hard to imagine how Silver could have been more wrong."

No, it's not hard to imagine at all. If Nate had predicted, "Germany has a 100% chance of winning 7-1," that would have been MUCH more wrong. 

6. Finally, and worst for last ... here's a UNC sociology professor lecturing Nate on how he screwed up, without apparently really understanding what's going on at all. I could spend an entire post on this one, but I'll just give you a summary. 

First, she argues that Nate should have paid attention to sportswriters, who said Brazil would struggle without those missing players. Researchers need to know when to listen to subject-matter experts, who knew something Nate's mathematical models don't. 

Well, first, she's cherry-picking her sportswriters -- they didn't ALL say Brazil would lose badly, did they?  You can always find *someone*, after the fact, who bet the right way. So what?

As for subject-matter experts ... Nate actually *is* a subject matter expert -- not on soccer strategy, specfically, but on how sports works mathematically. 

On the other hand, a sociology professor is probably an expert in neither. And it shows. At one point, she informs Nate that since the Brazilian team has been subjected to the emotional trauma of losing two important players, Nate shouldn't just sub in the skills of the two new players and run with it as if psychology isn't an issue. He should have *known* that kind of thing makes teams, and statistical models, collapse.

Except that ... it's not true, and subject-matter experts like Nate who study these things know that. There are countless cases of teams who are said to "come together" after a setback and win one for the Gipper -- probably about as many as appear to "collapse". There's no evidence of significant differences at all -- and certainly no evidence that's obvious to a sociologist in an armchair. 

Injuries, deaths, suspensions ... those happen all the time. Do teams play worse than expected afterwards?  I doubt it. I mean, you can study it, there's no shortage of data. After the deaths of Thurman Munson, Lyman Bostock, Ray Chapman, did their teams collapse?  I doubt it. What about other teams that lost stars to red cards?  Did they all lose their next game 7-1?  Or even 6-2, or 5-3?

Anyway, that's only about one-third of the post ... I'm going to stop, here, but you should read the whole thing. I'm probably being too hard on this professor, who didn't realize that Nate is the expert and not her, and wrote like she was giving a stock lecture to a mediocre undergrad student quoting random regressions, instead of to someone who actually wrote a best-selling book on this very topic.

So, moving along. 

------

There is one argument that would legitimately provide evidence that Nate was wrong. If any of the critics had chosen to argue convincing evidence for Brazil actually having much less TALENT than Nate and others estimated, evidence that was freely available before the game ... that would certainly be legitimate.

Something like, "Brazil, as a team, is 2.5 goals above replacement with all their players in action, but I can prove that, without Neymar and Silva, they're 1.2 goals *below* replacement!"

That would work. 

And, indeed, some of the critiques seem to be actually suggesting that. They imply, *of course* Brazil wouldn't be any good without those players, and how could anyone have expected they would be?  

Fine. But, then, why did the bookmakers think they still had a 49% chance? Are you that smart that you saw something? OK, if you have a good argument that shows Brazil should have been 30%, or 20%, then, hey, I'm listening.

If the missing two players dropped Brazil from a 65% talent to a 20% talent, what is each worth individually? Silva is back for today's third-place game against Holland. What's your new estimate for Brazil ... maybe back to 40%?

Well, then, you're bucking the experts again. Brazil is actually the favorite today. The betting markets give them a 62% chance of beating the Netherlands, even though Neymar is still out. Nate has Brazil at 71%. If you think the Brazilians are really that bad, and Nate's model is a failure, I hope you'll be putting a lot of money on the Dutch today. 

Because, you can't really argue that Brazil is back to their normal selves today, right?  An awful team doesn't improve its talent that much, from 1-7 to favorite, just from the return of a single player, who's not even the team's best. No amount of coaching or psychology can do that.

If you thought Brazil's 7-1 humiliation was because of bad players, you should be interpreting today's odds as a huge mistake by the oddsmakers. I think they're confirmation that Tuesday's outcome was just a fluke. 

As I write this, the game has just started. Oh, it's 2-0 Netherlands. Perfect. You're making money, right?  Because, if you want to persuade me that you have a good argument that Nate was obviously and egregiously incorrect, now you can prove it: first, show me where you wrote he was still wrong and why; and, second, tell me how much you bet on the underdog Netherlands.

Otherwise, I'm going to assume you're just blowing hot air. Even if Brazil loses again today. 


-----

Update/clarification: I am not trying to defend Nate's methodology against others, and especially not against the Vegas line (which I trust more than Nate's, until there's evidence I shouldn't).  

I'm just saying: the 7-1 outcome is NOT, in and of itself, sufficient evidence (or even "good" evidence) that Nate's prediction was wrong.  



Labels: , , , ,

Sunday, June 08, 2014

Team fatigue and the NBA playoffs, part II

Last month, a FiveThirtyEight study by Nate Silver found that, after winning their first round series 4-0, NBA teams significantly outperformed expectations in the second round, by about 3 points per game. Conversely, teams that took seven games in the first round fell well short in the second, underperforming expectations by a hefty 5.7 points per game.

The study argued that it's fatigue. Teams that sweep their series have lots of time to rest and recover, while the 4-3 teams have to jump right back in immediately. 

I was skeptical of the fatigue explanation. Last week, I thought it might be just a mathematical anomaly. I posted about it, and then immediately realized that, while the analysis was correct, the effect wasn't nearly enough to explain what Nate got. In fact, it explained less than one-tenth.

So, I figured, if it's not that, maybe I can try to figure out what it really is. After a couple of days of working on simulations, I have an answer. Well, I think I have an answer for part of it, and an opinion for the other part.

-----

First, there's the effect I talked about last post, where the expected point differential for a favorite should be artificially high because of how games are weighted. My estimate last post was kind of back-of-the-envelope, so I decided to use a simulation to get a better handle on it.

I created a conference of 15 teams, and assigned each a random point differential "talent," with a mean of zero and an SD of 4 points. Then, I played independent 82-game seasons for each, where a game was 100 possessions, two-point field goals only. After the season, I ranked the teams by W-L record, and had the top 8 make the playoffs. I paired them off 1-8, 2-7, 3-6, and 4-5, and played a first-round best-of-seven series. I ranked the four winners, paired them up 1-4 and 2-3, and played a second-round best-of-seven series. 

After all that, I took the four second-round teams -- actually, close to 120,000, because I ran 30,000 repetitions -- and compared their simulated second-round point differential to their talent. I expected the teams that went 4-0 in the first round would appear to exceed their talent in the second. (If they did so, the only possible reason would be the weighting anomaly, because of the way the simulation was set up.)

They did appear to score higher than their talent, but not by much:

4 games: +0.17 points/game
5 games: +0.03 
6 games: -0.03
7 games: -0.11

I found a difference of 0.28 points between the 4-0 teams and the 4-3 teams. The FiveThirtyEight study found 8.7 points. So, the logic is right, but the magnitude is nowhere near enough to explain the real-life differential.

------

There's another good reason you'd expect the 4-0 teams to outperform in the second round -- their first round gives us more information about the team. Specifically, the first round sweep suggests that the team is probably better than our original talent estimate. 

The FiveThirtyEight study used season SRS as their estimate of talent. That's the regular-season point differential after adjusting for strength of schedule. Like any other observed performance, it's subject to randomness, and will vary from true talent.  

Taking the simplified "100 possessions, 2-point attempts only" model, you can calculate that the SD of single game point differential is 14.1 points (10 times the square root of 2). For a season average, you divide that by the square root of 82, which gives 1.56. That means that even in this oversimplified model, the typical team's SRS is more than 1 point different from its true talent.

Some teams' SRSses are underestimates, and some are overestimates. The teams that go 4-0 are now more likely to be underestimates. So, you'd expect them to outperform in the next round.

To check, I re-ran the simulation, but, this time, instead of checking whether 4-0 teams performed better than their talent, I checked whether they performed better than their 82-game SPS. As expected, they did. But the effect was still pretty small:

4 games: +0.22 points/game
5 games: +0.28
6 games: -0.09
7 games: -0.35

We're up to 0.58 points difference between 4-0 and 4-3, still far short of FiveThirtyEight's finding of 8.7 points.

However: the simulation is still missing some hidden talent variation. For one thing, it assumes team talent is exactly the same every game. But that's not the case. Aside from home court advantage (which I ignored, because it wouldn't change the results much), there's things like injuries, trades, changes in player talent as they learn, and so forth.

For instance: if a team acquires a 2-point star player halfway through the season, a single SRS will blend the "before" and "after."  As a result, the overall SRS for the year will be 1 point short of playoff reality. 

To simulate that, I introduced a "playoff variation" factor. For each of the eight playoff teams, I tweaked their talent by a random number of points, mean 0 and SD 1. The results:

4 games: +0.39 points/game
5 games: +0.17
6 games: -0.08
7 games: -0.38

Larger again. Now, the difference is up to 3/4 of a point. 

When I up the "playoff variation" to have an SD of 2, it gets bigger still:

4 games: +0.75 points/game
5 games: +0.17
6 games: -0.19
7 games: -0.63

Now, we're up to 1.38 points. Still well short of 8.7, but enough that we would be able to say that this is at least *part* of what the FiveThirtyEight study found.

This suggests that, if fatigue isn't the explanation, maybe it has something to do with differences between SRS and actual talent.

-----

Well, I finished writing all of that, and then I thought, hey, we have an easy way to figure out how well SRS estimates talent -- the Vegas betting line!  So I went to check on those, and now I'm writing the rest of this post two days later.

In the FiveThirtyEight study, which covered 2003-2013, there were 17 teams that went 4-0 in the first round.

In 2013, the Heat took on the Bulls in the second round after sweeping the first. SRS ratings had the Heat at 7.04 point-per-game favorites over the Bulls. But the Vegas line had them at 9.5 points better. (The betting line was 13 points, and I subtracted 3.5 for home court advantage.)

So, we can say Vegas estimated the Heat as 2.46 points better (relative to Chicago) than their SRS estimate. I'll chart that like this:

----------------------------------------
                    Vegas   SRS     Diff
----------------------------------------
2013 Heat/Bulls      +9.5  +7.04   +2.46
----------------------------------------

Notes: (a) The Vegas numbers may vary, because I used more than one site, and they might vary by a half point here and there. (b) For all Vegas lines, I looked only at the first game of the series. (c) It's conventional to write a Vegas favorite as "-9.5", but I'm going to use "+9.5" for consistency with SRS; hope that's not too confusing. (d) I used 3.5 points for home court advantage. (e)  I'm going to talk about how SRS rated a single team (the one I'm talking about at the time), even though it's actually how SRS rated that team minus how SRS rated the opposing team. That's just to make things easier to read.

Here's all seventeen of the 4-0 teams:

----------------------------------------
4-0 teams           Vegas   SRS     Diff
----------------------------------------
2013 Heat/Bulls      +9.5  +7.04   +2.46
2013 Spurs/Warriors  +6    +5.35   +0.65
2012 Spurs/Clippers  +8    +4.36   +3.64
2012 Thunder/Lakers  +4.5  +5.32   -0.82
2011 Celtics/Heat    -1.5  -1.93   +0.43
2010 Magic/Hawks     +5.5  +2.78   +2.72
2009 Cavs/Pistons    +8    +6.97   +1.03
2008 Lakers/Jazz     +4.5  +0.47   +3.13
2007 Bulls/Pistons   -1.5  +0.84   -2.34
2007 Pistons/Bulls   +1.5  -0.84   +2.34
2007 Cavs/Nets       +2.5  +4.33   -1.83
2006 Mavs/Grizzlies  -1.5  -0.73   -0.77
2005 Suns/Mavs       +3    +1.23   +1.77
2005 Heat/Wizards    +7.5  +6.48   +1.02
2004 Spurs/Lakers    +1.5  +3.16   -1.66
2004 Nets/Pistons    -2    -3.16   +1.16
2004 Pacers/Heat     +8    +5.06   +2.94
----------------------------------------
Average              +3.74 +2.81   +0.93 
----------------------------------------

For those 17 series, the bookmakers rated the average favorite 0.93 points better than their SRS differential. In other words, the 4-0 teams were a point better than the FiveThirtyEight study gave them credit for. That explains roughly one point of the three points by which FiveThirtyEight found the favorites outperforming. 

Does that mean the "fatigue" effect can now only be two points?  Not necessarily. You could still argue that the reason for the extra Vegas point is that bookmakers and bettors *knew* about the fatigue factor, and adjusted their expectations accordingly. 

But, in that case, you could also ask, why did Vegas only adjust for one point out of the three?  Actually, I don't think even the one point is fatigue adjustment. There's a better explanation, in my view.

Suppose the the +0.93 was all fatigue adjustment. In that case, if we repeat the chart for the first round, the discrepancy should be zero, right?  Because all teams had roughly equal rest before the first round.

It's not zero. I won't give you the full chart, but the first round average is still positive, at +0.49 points. 

You could still defend fatigue. You could say, sure, maybe SRS was wrong by +0.49 points all along, but the remaining +0.44 points for the second round favorites could still be Vegas acknowledging the fatigue factor.

But, I think there's a better explanation for that +0.49 points. After the first round, we have good reason to believe those teams are better than we thought before the first round. After all , they just swept. 

The 17 teams were six-point favorites, on average, in the first round. According to a simulation I did, when a +6 team wins, it wins by an average of 14 points. (When it loses, it loses by 9.5 points, but that doesn't factor in to these 4-0 series.)  

If you add four +14 games to a +6 regular season, it increases the SRS from 6.00 to 6.37 -- almost exactly the +0.44 the favorites moved.

So, I think it's not fatigue that the lines were correcting for, but just new evidence for what the team talent was all along.

-----

But we still have those remaining 2.07 points to account for. Actually, let's adjust for the mathematical anomaly effect, which is .17 points. (It's not a big adjustment, but I did all that work, dammit, and I don't want to waste it.)

That brings us down to 1.9 points. Where did those come from?

Well, it could just be luck. If the SD of a single game point differential is 14, the SD of the average of 56 independent games would be 1.87 points. In that light, the 1.9 point differential is only one SD.

Actually, it's probably a bit more than that. Two of the 17 series were actually identical, just in reverse -- when the 2007 Bulls faced the Pistons, after they both went 4-0 in the first round. Since those two teams have to cancel to zero, the effect is bigger than it looks, perhaps by a factor of the square root of 17/15. (The observed effect rises by 17/15, and the SD rises by only the square root 17/15.)

Still, it's not statistically significant by normal standards, if that's what you like to look at.

Another way of looking at the same discrepancy: in the second round, the 17 teams went a combined 50-41 against the spread. Eliminating the two duplicates, they went 44-35. That's also about one SD away from .500, which you'd expect, since the W-L and point differential are essentially two ways of looking at the same result. 

It still *could* be fatigue, but I think you need better evidence than 44-35.

------

Now let's look at the fourteen 4-3 teams, the ones that underperformed in the second round:

----------------------------------------
4-3 teams           Vegas   SRS     Diff
----------------------------------------
2013 Bulls/Heat      -9     -7.04  -1.96
2012 Clippers/Spurs  -8     -4.85  -3.15
2012 Lakers/Thunder  -4     -4.48  +0.48
2010 Hawks/Magic     -5.5   -2.68  -2.82
2009 Hawks/Cavs      -8     -6.97  -1.03
2009 Celtics/Magic   -2     +0.95  -2.95
2008 Celtics/Cavs    +6     +9.84  -3.84
2007 Jazz/Warriors   +0.5   +3.07  -2.57
2006 Suns/Clippers   +2     +3.73  -1.73
2005 Pacers/Pistons  -5     -2.82  -2.18
2005 Mavs/Suns       -3     -1.23  -1.77
2004 Heat/Pacers     -8     -4.80  -3.20
2003 Mavs/Kings      -5     +1.22  -6.22
2003 Pistons/76ers   -5.5   +1.22  -6.72
----------------------------------------
Average              -3.89  -1.06  -2.73 
----------------------------------------

These are much bigger SRS errors than for the 4-3 teams ... compared to the bookies, SRS overestimated the teams by an average 2.7 points. The FiveThirtyEight study found a 5.7 point difference, which leaves three points unexplained -- 2.8 points after adjusting for the mathematical anomaly.

SRS had also rated those teams higher than the bookies in the first round -- but only by 0.45 points. That's only 1/6 of the full effect. So, this time, if you believe Vegas is adjusting for fatigue, you have a better argument. (But, again: why did bettors only adjust by half the observed effect?)

The 2.8-point shortfall, relative to Vegas, resulted in those teams going 29-45 against the spread. That's a bit less than 2 SDs below .500. (Actually, by "30 points equals one win," 2.8 points works out to 30-44. So there's one game Pythagorean error.)

------

So we have 44-35 for the sweeping teams, and 29-45 for the seven-game teams. Even if they aren't significant individually, doesn't the *combination* of the two suggest something real is going on?

Not as much as it seems, because the 4-0 and 4-3 results aren't independent. The 4-0 teams played the 4-3 teams some of the time. So, some of the extreme results are counted in both samples. 

In 2010, the Magic went 4-0 in the first round, while the Hawks went 4-3. When they faced each other next round, Orlando absolutely crushed Atlanta, with an *average* score of 107-82. That's 22 points per game more than expected.

That shows up as a +22 for the 4-0 teams, and a -22 for the 4-3 teams. You can see those as the two most extreme dots in the FiveThirtyEight chart: 





In fact, exactly half the series are independent, and half are exact mirror images, except for which column of the chart they appear in. In the chart, for every dot, you'll find a mirror image dot in one of the four columns somewhere.

If I erased the 4-3 column from the chart, you could reproduce it perfectly. You'd just find all the dots in the first three columns that aren't offset by a mirror-reflection dot, and those must be the ones that go in the last column. 

You can still say, "Not only did the 4-0 teams outperform, but, also, the 4-3 teams underperformed!"  But that's like saying, "Not only did the Magic score 20 more points than the Hawks last night, but the Hawks also scored 20 fewer points than the Magic!"  Well, not exactly, because not every 4-0 team played a 4-3 team -- some of them played 4-1 teams and 4-2 teams. So it's only partially like that. But still enough that you have to keep it in mind.

-----

In the FiveThirtyEight study, the traditional evidence for significance comes when they do a regression, and they find a "first round games" effect that's 3 SDs from the mean. But that's an overestimate of the significance, for the reasons discussed:

1. The series aren't independent; half are duplicates of the other half.

2. The regression doesn't adjust for the mathematical weighting anomaly, which is around 0.15 points for the first and last columns. 

3. The regression doesn't adjust for SRS under/overestimating Vegas by the 0.5 points we should all be able to agree on (looking at the first round, where fatigue didn't apply).

4. The regression doesn't adjust for the fact that our estimate of team skill should change for the second round, even independent of fatigue, because of the evidence of how they played in the first round. 

After all that, what's left?

1. The unexplained observed point differentials: +1.85 points for the first group, -2.8 points for the second group. Or, equivalently converted to wins against the spread: the unexplained record of 44-35 for the sweeping teams, and 29-45 for the seven-game teams.

2. The possible argument that the difference between SRS and Vegas in the second round -- after subtracting off the difference in the first round, and the new information about team talent -- might be evidence that Vegas is adjusting for fatigue.

Even without having a formal significance test for what's left, those effects seem small enough to me like they could just be random luck. 

-------

Going from data interpretation to personal opinion, here's my argument for it being just random:

(a) It's probably not significant at the 5% level, or, at best, just barely.

(b) It's rare to find a large effect that bookies and sharp bettors haven't also found;

(c) Nate said the effect didn't repeat for other rounds;

(d) This year failed to follow the pattern (after the study appeared). The 4-3 Pacers performed against the 4-0 Wizards exactly as SRS predicted. The 4-0 Heat performed against the 4-3 Nets a tiny bit worse than predicted. And the 4-3 Spurs handily beat expectations against the 4-2 Trail Blazers. (The remaining two 4-3 teams faced each other, cancelling out.)

(e) The difference is so huge that it's just implausible on its face. Even Nate doesn't really believe it: "the effects are so pronounced I don't trust them." Taking the results at face value would have made Indiana a 35% underdog to win its series against Washington, instead of a 76% favorite. (The Pacers wound up winning, 4 games to 2.)

(f) It seems implausible that a fatigue advantage would persist throughout the entire second round. The first game or two, maybe. But, the numbers are too big for that. A 3-point-per-game advantage over a 5-game series is 15 points overall ... and there's no way one team could have a 7.5 point advantage in the first two games, or a 15-point advantage in the first game.

Feel free to disagree with me.

------

(P.S. Credit to regular commenter GuyM for suggestions in an e-mail conversation we had. Guy was big on the "teams may be different in the playoffs from the regular season" explanation, which turns out to be the important one.)










Labels: , , ,