Sunday, February 23, 2014

Did Team Canada overestimate its players?

Joel Colomby, the fantasy columnist for the Sun chain of newspapers, has an interesting finding today showing how, in 2010, Team Canada's NHL players seemed to get worse after the Olympics.


Of nineteen of Canada's "top scorers," only five had a higher NHL points-per-game rate after the Olympics than before.  The other 14 dropped.

Colomby included only players still in the NHL in 2014, and who are "likely owned in a fantasy pool" this year.  That creates a selective-sampling issue, but if you take the results at face value, a 5-14 record is about 2 SDs from 50-50.

Looking at the other countries, there didn't seem to be any real effect.  It did vary from team to team, but, if you put all the other teams together, you get 16 of 34.  So:

 5-14 Canada
16-18 Rest of World

Going by actual performance differences, the results are similar.  The average player's points-per-game (PPG) dropoffs were:

-0.086 PPG Canada
-0.011 PPG World

By my rough estimate, Canada came in at roughly 1.8 SD from zero.  

------

These averages are from calculations I did by hand; Colomby gives only the individual player numbers.  He hints that we might see repeats for certain players.  For instance, "Leafs fans, and Randy Carlyle, have to hope Phil Kessel rebounds better than he did four years ago [-0.17]."

Of course, I don't think the individual numbers mean anything at all -- small sample size, as Colomby mentions.  But ... this got me wondering about something else.

The Canadian players dropped off in the latter part of the season.  Players who dropped off are more likely to have been lucky earlier in the season, just before Team Canada would have been selected.  

So: isn't it possible that the selection process was perhaps too influenced by randomness?   Could it be that GM Steve Yzerman and his staff put a little too much weight on players' recent "hot" performances, and wound up with a team that perhaps wasn't as good as it could have been?

That theory also explains why Canada's players dropped off more than the rest of the world's.  All the players on Team Canada came from the NHL, so Canada's roster is based on whom management thought were the best Canadian NHL players.  For other teams, the list of NHL players to draw from is smaller, so those decisions are a lot easier.  In fact, several of the teams would obviously need to include *all* of their NHL players, lucky or not.

Also, in this case, the Sun's selective sampling doesn't hurt this hypothesis, and may actually help it.  The players omitted from Colomby's list are the ones who are no longer regulars in the NHL.  Those guys, you'd think, would have been *more* likely to have declined after the Olympics, not less.

Are there other explanations?  Probably.  I bet some of the dropoff has to do with injuries.  The team is selectively sampled for prior good health (injured players don't go to the Olympics), so you'd expect a certain amount of dropoff regardless when those players later get their normal share of injuries.  (I wonder if that might be the source of the slight decrease for the other countries' players.)

And it goes without saying that it could just be random.

------

My gut says ... I bet the "luck" explanation is at least part of what happened, that Team Canada management wound up slightly overestimating players who were hot the first half of 2009-10.  

I could be wrong.  Those of you who know hockey better than I do, check out the list of players, and see if there are any questionable selections that seem to have been based on the player's uncharacteristically good recent play. I'll list the players for you (and also repeat the link).

Player                  Pre   Post  Diff
----------------------------------------
Sidney Crosby PIT      1.28 / 1.55 / +27
Patrice Bergeron BOS    .68 /  .79 / +11
Brenden Morrow DAL      .59 /  .65 /  +6
Jonathan Toews CHI      .89 /  .90 /  +1
Scott Niedermayer ANA   .60 /  .61 /  +1
Mike Richards PHIL      .77 /  .73 /  -4
Drew Doughty LA         .74 /  .67 /  -7
Eric Staal CAR         1.02 /  .95 /  -7
Patrick Marleau SJ     1.03 /  .95 /  -8
Corey Perry ANA         .95 /  .85 / -10
Duncan Keith CHI        .87 /  .76 / -11
Chris Pronger PHIL      .70 /  .59 / -11
Rick Nash CLB           .90 /  .77 / -13
Dan Boyle SJ            .80 /  .65 / -15
Shea Weber NAS          .59 /  .42 / -17
Joe Thornton SJ        1.21 /  .82 / -19
Dany Heatley SJ        1.06 /  .80 / -26
Ryan Getzlaf ANA       1.09 /  .80 / -29
Jarome Iginla CGY       .92 /  .60 / -32


-----

Of course, we will eventually be able to check whether the same thing happens this year.  Maybe, for the current season, we might also see an effect for Team USA.  According to quanthockey.com, there were 136 American players with at least 30 games played at the Olympic break.  With fewer than half the (309) equivalent Canadian NHLers to choose from, the USA might have faced fewer tough decisions, which means less reliance on luck.  But, it's still worth checking.

If the result repeats for 2014, we'd have the cleanest evidence I've seen that sports GMs fail to fully consider luck when predicting future performance.  We already have a strong intuition that happens, but it's been hard to tell for sure.  

All decent players get contracts, even unlucky ones.  So, if John Doe has a career year, we need to know not *whether* he was signed, but for *how much*.  And even then, it's orders of magnitude more difficult to compare performance to salary than it is to compare first-half performance to second-half performance.  

It's a small sample, but this time we have an unambiguous "yes/no" of whether Team Canada thought this player was among the best.  And it turned out that almost three-quarters of the players chosen had, at the time, been playing over their (later-selves') heads.  

Was Team Canada fooled by randomness?






Labels: , , ,

Monday, January 20, 2014

Luck and the Olympic hockey tournament

There are twelve countries represented in Olympic men's hockey.  They play only three or four games each before getting to the quarter-finals; then it's single elimination after that.  

In the NHL, it takes a huge number of games -- 36 or 73 games per team -- to get to the point where talent is as important as luck.  But the entire Olympic tournament takes only 30 games.  Not 30 games per team, but 30 games period.

If that's the case, then how can the Olympics possibly filter out the best teams in such a short span?  

That's the subject of my article in the Ottawa Citizen today.  If you don't want to read the whole thing, the answer, basically, is:

1.  There is a much wider range of talent in the Olympics than in the NHL.  The top teams are almost as good as an NHL all-star team; the bottom teams are below minor-league.  That makes it much, much easier to separate good from bad.

2.  The IIHF (which structured the tournament) created an unbalanced schedule -- the bad teams disproportionately face the good teams, and vice-versa.  This makes it easier for the good teams to rise to the top.

3.  The IIHF noticed that, roughly speaking, there are six strong teams and six weak teams.  Therefore, luck will affect the middle of the standings more than the top or the bottom.  So, the teams in the middle get an extra game (against the bottom), in order to make it more likely that the better teams rise and the worse teams fall.

My conclusion is that the IIHF did an outstanding job in terms of squeezing the most "talent information" out of so few games.  For the full argument, check out the link.

Labels: , , ,

Tuesday, March 13, 2012

An economist predicts the Olympic medal standings: summer edition

Dan Johnson, a professor at Colorado College, did a regression to predict medal wins at the Olympics for a given country. There are articles about it in various newspapers.

I wrote about this a couple of years ago, when he did the same thing for the Winter Olympics. At the time, I was a little skeptical. Nothing much has changed.

Except ... this time we have the equation! The National Post was kind enough to provide it in the print edition.

Here it is:

Medals = 0.33 +
.00271 * total medals available +
.02 * income +
.0000549 * income squared +
.024 * population +
19.02 * population squared +
11.91 if it's the home nation +
3.85 if it'll be the next home nation +
3.35 if it was the home nation last time or the time before +
0.29 if it borders the current home nation +
coefficients for dummy variables for nation

Apparently, the dummy variables are new -- they represent a

... "'cultural specific factor' to account for things that are hard to quantify, like the prevalence of doping or the culture of competition, and also to correct the historical under-predictions for countries such as China and Australia."


The dummy variables should make the predictions even more accurate.

BTW, an interesting thing about the regression is that if you have a dummy variable for each country, your accuracy should be the same very similar even if you don't include the income and population variables! Those variables are nice to have because they illustrate the effects of income and population, but if you leave them out, the dummy variables will pick up the slack and adjust themselves for the income and population of each country.
(UPDATE: it won't be exactly the same, just close. See the bottom of the post for an explanation.)

(The Post story also says that Johnson also updated his previous regression to remove variables for climate and politics. By the same token, I think he'd get the same results if he left them in.)

So I think that if you leave out income and population, you'll get the same coefficients for everything except the dummy variables:

Medals = 0.33 +
.00271 * total medals available +
11.91 if it's the home nation +
3.85 if it'll be the next home nation +
3.35 if it was the home nation last time or the time before +
0.29 if it borders the current home nation +
new coefficients for dummy variables for nation

I don't think this is the best way to account for the variables still included. That's because (as I think I said in the other post) all the factors are linear. I'm not sure that's right: it implies that if the USA is the home team, it gains 12 medals on top of the 100 or so it usually wins, but if Canada is the home team, they also gain 12 medals (to go with their 15 or so).


The "12 extra medals" is, roughly speaking, the average. It should work for a roughly average country. Since the UK isn't too far from average, the formula should probably work pretty well this year.

----------

UPDATE: it now occurs to me that that's not exactly right. I was assuming that the population and income were constants for each country. They're not -- they vary over time. So, the regression coefficients with population/income will be *close* to the ones without, but won't be exactly equal.

(I should have also realized that the regression wouldn't work if population and income were constants, because you'd get multicollinearity.)



Labels: ,

Thursday, February 18, 2010

An economist predicts the Olympic medal standings

Daniel Johnson is an economics professor who, according to Forbes magazine, makes "remarkably accurate" predictions on how many Olympic medals each country will win. But I'm not sure, based on the description given, that the predictions are all that remarkable.

From the article, it sounds like what Johnson is doing is running some kind of regression, on "per-capita income, the nation's population, its political structure, its climate and the home-field advantage for hosting the Games or living nearby." He doesn't consider anything specific about the sports or athletes.

How accurate are Johnson's predictions? I'm not really sure. Forbes says,

"Over the past five Olympics, from the 2000 Summer Games in Sydney through the 2008 Summer Games in Beijing, Johnson's model demonstrated 94% accuracy between predicted and actual national medal counts. For gold medal wins, the correlation is 87%."


What does that mean? From the word "correlation," my guess is that those numbers are the correlation coefficient, or "r". But an r of .94 doesn't mean that the predictions are 94% accurate. It just means that the *best fit straight line* is 94% accurate. It's possible to be wrong, perhaps badly wrong, in every guess, but still have an r of 1, which means 100% correlation. For instance, if you underpredict every below-average country by 50% of the difference, and you overpredict every above-average country by 50% of the difference, you'll get a perfect correlation, but really crappy guesses. For instance, here's an example of how that might happen:

Country A: estimate 80, actual 65
Country B: estimate 60, actual 55
Country C: estimate 40, actual 45
Country D: estimate 20, actual 35

Regressing estimate on actual (or is it actual on estimate? I forget which way the word "on" implies, but never mind) gives a 100% correlation, but the actual guesses aren't spectacular in the least.

Anyway, it might be some other method that Johnson uses to compute the accuracy percentage, but it's hard to evaluate the claims without an explanation.

(UPDATE: as this blog post was going to press, I found Johnson's website, which confirms that it *is* correlation. It still could be that it's some kind of method that doesn't have the flaw of my example above. The site contains a media release, but no actual copy of the paper, which was published in "Social Science Quarterly" in December, 2004.)

More importantly, you can't tell how impressive a set of predictions is without something to compare it to. At Forbes, commenter "Doubter" points this out, and tries using the results of the previous Olympics to predict the current one. For the top five countries, he gets an 85% accuracy rating, and correctly points out that "include a bunch of countries with stable medal counts (Jamaica, Japan, Nigeria, Kenya, most European countries) and I am sure it gets much better."

I'm pretty confident that if you were to just use a weighted average of previous Olympics results and adjust for home field advantage, you'd come pretty close to what Mr. Johnson was able to do. Forbes should have realized that the results probably aren't "remarkably accurate" -- just "accurate".

Also confusing is the estimate of home field advantage. There were no results given for the Winter Olympics, but for the summer games, the host team "typically garners 25 additional medals compared with its expected performance, 12 of them gold." It doesn't really make sense that the home field advantage should be a fixed number of medals. Shouldn't it be a percentage increase? Canada won 11 medals in 1976 in Montreal, none of them gold. That was a few more medals than usual, probably because of home field. Should they really have been expected to win minus 14 medals in 1988 in Seoul, of which minus 12 would be gold? Or, on the other hand, were they just unlucky in Montreal, where they should have won about 30, when they were in the single digits in 1968 and 1972?

Or, if Pakistan were to host the Olympics, would you really expect them to jump from (say) 1 medal to 26?


Oh, and one more thing: in his 2010 predictions, Johnson has the top 13 countries winning 250 medals, but only 57 golds. Overall, gold are 33.3% of medals, but for those 13 countries, Johnson has them winning only 22.8% golds. How come? An eyeballing of the 2006 chart shows about 1/3 golds for those countries then ... I wonder why the drop?


Labels: , ,

Friday, August 15, 2008

Why is the US lagging in 3-point percentage?

Tyler Cowen reports (via ESPN) that so far, the US basketball team has the worst 3-point percentage of any team in the Olympics.

He asks: "can you build a simple model showing this is likely the case for the best team?"

The comments are pretty high quality. The best comments, IMO, are the ones that don't try to answer the question, but try to figure out if there might be other reasons. My two favorites:

-- luck

-- the shorter international 3-point line (compared to the NBA) means the US players can't use their muscle memory, and have to think about the shots.

Both of these are testable: the first by waiting a few more games; the second (as a commenter points out) by seeing if the European NBA players are hitting more threes.

But I don't know much about basketball. What do you guys think?

Labels: , ,