Friday, August 05, 2016

log5 estimates are biased when we use the wrong measure of "talent"

The "log5" method tries to predict a team's chance of winning a game based on its talent and that of its opponent. The basic formula, for teams A and B, is  

P = (A - AB)/(A+B-2AB)

A few months ago, I wrote that there's no theoretical reason for the formula to always work. In fact, there's an obvious counterexample where it doesn't work. Consider "height baseball," where the taller team always wins. Suppose team A is .700, because it's taller than 70 percent of its opponents, while team B is .400, being taller than only 40 percent of its opponents. The formula predicts team A will win 77.8 percent of games against B, but, of course, it will win 100 percent.

So why doesn't log5 work? I think I've found one reason, which I'll explain in this post. 

(There's a second reason -- which is actually a first reason, since it came back in 2011. In a blog post, Tango showed another example, using sprinter times, of how the odds ratio method (on which log5 is based) doesn't work, and Kincaid explained why in the comments. When I started writing this post, I originally thought mine was the same argument, just explained differently. But it's not. My argument actually doesn't apply to Tango's example ... I'll try to explain the Tango/Kincaid logic in a future post.)

------

Suppose you have team A, an .800 talent, playing team D, a .500 talent. What is the probability team A wins?

It seems that the answer should be ... well, .800. If team A is .800 against the league, which averages .500, then you'd think it should be .800 against a bona fide average team. And, the log5 method confirms the inutition -- plug in the numbers, and you do, indeed, get .800.

But that can't be right. I think it has to be the case that the .800 team plays *better* than .800 against the league-average team, and that it's easy to see why without any fancy math.

It actually doesn't depend on any technicality about what it actually means to be an .800 team. 

For instance, it's not because, if a team is .800 against the rest of the league, it must be *worse* than .800 in general, since it doesn't have to play itself. Even if you fix that problem, team A will have to be better than .800 against team D.

It's not because of home/road issues either, or the difference between observed .800 and talent .800 ... adjust for those, and the result still holds.

Let me restate the question in more detail, to try to eliminate some of those technicalities: 

----

In a league with no home field advantage, there are seven teams, A through G. 

If team A played a balanced schedule against all of them -- including itself (or a clone of itself) -- you would expect it to finish with an .800 record. So, in that respect, team A "has .800 talent".

By the same definition, teams B through G, respectively, have .700 talent, .600, .500 ... all the way down to .200 talent.

When team A (.800) plays team D (.500), what's the probability A wins?

----

The answer: not .800.

----

Let's create a little spreadsheet of team A's performance against all seven teams. It looks like this:

 matchup         probability
---------------------------
.800 vs .800
.800 vs .700
.800 vs .600
.800 vs .500
.800 vs .400
.800 vs .300
.800 vs .200

Now, let's fill in the log5 estimate for each one of those matchups:

 matchup           log5
---------------------------
.800 vs .800       .500
.800 vs .700       .727
.800 vs .600       .631
.800 vs .500       .800
.800 vs .400       .857
.800 vs .300       .903
.800 vs .200       .941

Those look quite reasonable, except that ... they don't average out to .800! They average out only to .766.

 matchup           log5
---------------------------
.800 vs .800       .500
.800 vs .700       .631
.800 vs .600       .727
.800 vs .500       .800
.800 vs .400       .857
.800 vs .300       .903
.800 vs .200       .941
---------------------------
 Average           .766

There's no trick here. This is a real, valid counterexample, one that shows that log5 doesn't actually work. And there's nothing special about our choice of .800. The average would always wind up too low, except for a team that's exactly .500.

Suppose we abandon the log5 estimates, then, and just try to fill in probabilities that seem reasonable. Can we do that, while insisting that the middle number stay .800? 

We have to hold the first number at .500, since, when a team plays a clone of itself, it must win 50 percent of its games, by definition. So we start with a chart that looks like this:

 matchup       probability
---------------------------
.800 vs .800        .500
.800 vs .700
.800 vs .600
.800 vs .500        .800
.800 vs .400
.800 vs .300
.800 vs .200
---------------------------
 overall avg        .800

From here, how do we fill in the second and third lines? One obvious way, that seems not too unreasonable, is just to stick in ".600" and ".700". 

 matchup       probability
---------------------------
.800 vs .800        .500
.800 vs .700        .600
.800 vs .600        .700
.800 vs .500        .800
.800 vs .400
.800 vs .300
.800 vs .200
---------------------------
 overall avg        .800

Having done that, it seems reasonable to just continue the pattern:

 matchup       probability
---------------------------
.800 vs .800        .500
.800 vs .700        .600
.800 vs .600        .700
.800 vs .500        .800
.800 vs .400        .900
.800 vs .300       1.000 
.800 vs .200       1.100
---------------------------
 overall avg        .800

That does, indeed, keep the average at .800. But it's obviously wrong -- it makes no sense to estimate that team A beats the .200 team 110% of the time.

So, is there another way we can fill this in, while keeping the .500 and .800 estimates, so that it all makes sense? No, I don't think that's possible. 

Right now, the first three lines of the chart average .600, which is .200 points below the .800 average we're shooting for. Therefore, the bottom three lines must average .200 points *above* .800. In other words, the bottom three lines have to average 1.000! Clearly, that can't be done.

So, we have to decrease the second and third lines. Maybe we change them to, say, .650 and .750. If we do that, then the first three lines average only .167 points below .800. Now, the bottom has to average "only" .967. Which, again, doesn't pass the sniff test.

Try if you want, but I'm pretty sure that you're not going to find anything that seems like a plausible breakdown. The only way to get something that looks reasonable, I think, requires the middle line to be something higher than .800.

--------

How much higher? From the original log5 chart, we see that

(a) team A was .766 overall, but
(b) team A played .800 ball against the .500 team.

If a .766 team goes .800 against an average team, maybe we can extrapolate that an .800 team would go, say, .840 against an average team.

Plugging .840 into the middle slot, and filling in the rest in some plausible fashion to average .800, maybe the chart would look something like this:

 matchup       probability
---------------------------
.800 vs .800        .500
.800 vs .700        .690
.800 vs .600        .770
.800 vs .500        .840
.800 vs .400        .900
.800 vs .300        .940
.800 vs .200        .960
---------------------------
 overall avg        .800

That's just a guess, of course. But, no matter what the true values are, the point remains: the middle entry must be significantly higher than .800.

And, another consequence: all the outcome probabilities, other than between equal teams, are *more extreme* than log5 suggests. The log5 formula is too conservative, always underestimating the favorite's chances of winning, when there is a favorite.

--------

So, log5 doesn't actually work. But, I think, there's an easy way to tweak it so that it DOES work. 

And that is: instead of using the log5 formula with the respective teams' expected talent against the league, we use their expected record against a .500 team. 

In our league, a team that finished .800 overall beats an average team 84% of the time, not 80%. Which means, for this new definition of log5, it's not an .800 talent, it's an .840 talent.

Let's reserve the word "talent" for its usual meaning, the expected record against the league, and use the made-up word "5talent" to mean talent against a .500 team. In our seven-team league, a team with a talent of .800 has a "5talent" of .840.

What's the 5talent of the rest of the teams? We can guess. If an .800 talent is an .840 5talent, maybe a .700 team is .720, a .600 team is .610, and so on. 

Repeating the log5 calculation using 5talent instead of talent, we get:

5talent matchup     log5
------------------------------
.840 vs .840        .500
.840 vs .720        .671
.840 vs .610        .771
.840 vs .500        .840
.840 vs .390        .891
.840 vs .280        .931
.840 vs .160        .965
------------------------------
 Overall avg        .796

Not bad! Under our estimates, a team that's an .840 5talent works out to a .796 talent. You could easily tweak the assumptions to get the average to .800 exactly, if you wanted to.

--------

Why does this happen, that log5 doesn't work if you use league performance? 

Because the win probabilities are based on the odds ratio. 

The log5 method works like this: suppose you have an .800 team against a .400 team. The .800 team has average 4:1 odds of winning. The .400 team has average 2:3 odds of winning. Divide 4/1 by 2/3, and you get 6/1. So the .800 team has 6:1 odds of beating the .400 team. That works out to an .857 winning percentage.

(I used odds ratios instead of the "usual" log5 formula, but it's exactly the same thing. If you do some algebra on the odds ratio calculation, you can actually derive the log5 formula at the top of this post.

When I calculate log5 probabilities, I actually use the odds ratio method, because the method is easy to remember, and I don't have to memorize the formula.)

The log5 formula is based on *multiplication* of *odds ratios*. But a team's overall average record, the one we normally talk about, is based on *addition* of *probabilities*. Those are two different sets of two different things. 

We calculated the bottom line of the chart, the overall winning percentage, as the arithmetic mean of the win probabilities. But, the odds ratio method doesn't know about arithmetic means and win probabilities. It knows only about geometric means of odds ratios. 

And, as it turns out, if we calculate the average as the geometric mean of the odds ratios ... well, then everything works! The geometric mean of the odds ratios against all teams is the same as the odds ratio against the average team.

Going back to the chart, and going back to the .800 talent, I'll convert the probabilities to odds ratios. (The odds ratio is the probability of winning divided by the probability of losing.)

talent matchup     log5     odds ratio
------------------------------------------
.800 vs .800       .500        1.00
.800 vs .700       .631        1.71
.800 vs .600       .727        2.67
.800 vs .500       .800        4.00
.800 vs .400       .857        6.00
.800 vs .300       .903        9.33
.800 vs .200       .941       16.00
------------------------------------------
 arithmetic mean   .766        
  geometric mean               4.00 (.800)

     
The .766 arithmetic mean doesn't equal the .800 talent, but the 4.00 geometric mean of the odds ratios *does* equal the 4.00 odds ratio talent. 

In other words, a team that's a 4.00 odds ratio talent against the league overall is also a 4.00 talent against the league average team. 

-------

OK, I cheated a bit. The reason the geometric mean works out perfectly is that, in our league, the team talents are symmetrically distributed around .500. 

If the talents aren't symmetrical, it doesn't work out perfectly -- I found the geometric mean comes out a little too high. However: it's a lot closer than the arithmetic mean works out to be. And, most leagues are symmetrical enough that it wouldn't be an issue in real life. There aren't many leagues with, say, 20 teams with .480 talent, but one team at .900. 

(Also... I wonder if you use the 5talents in the chart instead of the talents, if the geometric mean might not work out perfectly in that case, even for non-symmetrical leagues. That's just a gut feeling, and I need to think about it more.)

-------

I think that what makes the arithmetic mean give such a serious underestimate is that our league has large extremes of team talent, spanning a range from .200 to .800. The farther the numbers are from .500, the more the geometric mean of the odds ratio differs from the arithmetic mean of the probabilities. 

One measure of "extremes" is the standard deviation. In our hypothetical seven-team league, the SD of talent is .200. In actual Major League Baseball, the SD of talent is only around .055.

So, let's simulate the MLB spread with an example where, instead of teams varying from .800 to .200, they only vary from .565 to .435. Here's the log5 chart:

 talent matchup    log5     odds ratio
------------------------------------------
.565 vs .565       .500      1
.565 vs .500       .565      1.2989
.565 vs .435       .628      1.6870
------------------------------------------
arithmetic mean    .564
 geometric mean              1.2989 (.565)

This league has a talent SD of .053, similar to MLB. And, with that smaller SD, log5 is a very good fit. Now, a team with a 5talent of .565 has a talent of .564 -- so close that it makes no material difference. 

It looks like the log5 method is very, very sensitive to the dispersion of talent in the league -- bumpng the SD from .053 to .200 -- only a factor of 4 -- made the log5 discrepancy jump from 1 point all the way to 40 points.

So that explains why, even though log5 produces conservatively biased estimates when we use the "usual" definition of talent, it nonetheless works so well for baseball.  It's because in MLB, the spread in talent is reasonably small.

------

We can repeat this for other sports. Ten years ago, Tom Tango found the SD of talent to be .134 in the NBA, and .143 in the NFL.

 SD of talent
-------------------------------
.200 hypothetical 7-team league
.143 NFL
.134 NBA
.055 MLB
.053 hypothetical 3-team league
-------------------------------

We know that log5 is pretty far off for the 7-team league, and pretty good for MLB. The two leagues in between -- the NFL and NBA -- are right in the middle. 

If we do a five-team league, with talents ranging from .700 to .300, the SD is .141, which is pretty close to the NBA and NFL. Here's the calculation for the .700 team:

talent matchup     log5     odds ratio
------------------------------------------
.700 vs .700       .500        1.00
.700 vs .600       .609        1.56
.700 vs .500       .700        2.33
.700 vs .400       .778        3.50
.700 vs .300       .845        5.44
------------------------------------------
 arithmetic mean   .686        
  geometric mean               2.33(.700)


We can estimate that in the NFL or NBA, a .686 talent has a "5talent" of .700, meaning that it'll beat a .500 team 70% of the time. 

That means that if you're estimating a .686 talent against a .500 talent, your estimate is going to be .014 too conservative. That's approximately what it'll be for any pair of teams about that far apart. It'll be more accurate for opponents closer in talent, and less accurate for talent mismatches, at least to the point of diminishing returns (if log5 says your team is .990, it's not possible that it really should be 1.005).

Does that bother you, that when you use log5 on teams with significantly different talent, your estimate is off by as much as .014? If it doesn't, just go ahead and keep using log5. 

But for any study that's looking for small effects ... well, to me, it seems to me that .014 points is probably as big as the effect you're looking for. If you don't correct for it, you're underestimating the favorite by, what, maybe a quarter as much as home field advantage?

If you do some kind of "record after a time zone change on a hot streak" study, and you find an effect of .014 points ... since hot streaks mean good teams, could it just be that all you did was rediscover that log5 is biased too conservatively when you use the wrong definition of talent?  



UPDATE, 8/25/16: Changed title of post and reworded a few sentences to emphasize that the "bias" in log5 is not intrinsic to the formula itself, but occurs when we improperly use "talent" instead of "5talent." Hat tip: Ted Turocy. See longer explanation here.  

Labels: , , , ,

Monday, December 09, 2013

Do Western teams dominate in NFL night games?

You're getting ready for an NFL night game.  A team from the Eastern time zone is playing a team from the Pacific time zone.  How should you bet?

From 1970 to 2011, you should have bet on the West Coast team.  In games starting at 8:00 pm (Eastern) or later, they beat the spread 70 out of 106 times.  That's a record of 70-36, or .660.  The odds of something that extreme happening by chance (either way) is 1 in 806.  It's 3.3 standard deviations from the mean.

In a control sample of afternoon games, there was no such effect.  In fact, the Western teams went only 143-150 (against the spread) in those.

What's going on?  Well, the academic authors who found this result claim it's due to the circadian rhythms of the human body.  Physiologists and psychologists believe athletic performance peaks in the late afternoon.  So, for a game that starts at 8:00 pm Eastern time, the players from the West are actually playing at 5:00 pm "body time," which is why they perform better.

That result comes from a recent academic study: "The Impact of Circadian Misalignment on Athletic Performance in Professional Football Players," by Roger S. Smith, Bradley Efron, Cheri D. Mah, and Atul Malhotra.  Here's a Business Week rundown that actually just came out today.  Deadspin mentioned it here, and Brian Burke mentioned it here.

When I read the reports, I couldn't believe that the 70-36 could actually be accurate.  I downloaded the study ($8), and then went to Pro Football Reference to confirm for myself with their game finder.  And, yup, it all checks out!  

------

Wow, eh?  Could that actually be what's happening, that time-of-day effects on the human body are so big that they're almost twice the size of the home field advantage?  

Well, as you might expect, I'm skeptical.  I can think of a whole bunch of other things that might be going on.

Nothing of what I'm going to say is conclusive ... you should take this post not as a definitive rebuttal, but, perhaps, as a case for the defense, a "devil's advocate" kind of argument.

------

1.  There's no actual evidence that the West teams played better.  The only data we have is that they consistently beat the spread.  

It's just as possible, isn't it, that the bookmakers were shading the spread in favor of the east teams, at least in the night games?

I broke down the night games by day of week.  (They don't quite add up because I did everything manually, and probably screwed up somewhere.)

Sunday: 18-12
Monday: 46-19
Other:   3- 7

It turns out almost all the effect happens on Monday.

Does this support the "line shading" hypothesis?  I think it does, a little bit.  Monday night games, I'd imagine, get the most action from bettors ... if bookmakers shade the line when betting gets too heavy and unbalanced, it seems like Monday games should be the best candidates for when that happens.

---

2.  Night games are not random.  They're selected specifically because the NFL believes that they'll be the best games.  Maybe they're what the NFL thinks will be the games with the best teams, or the most exciting games, or the games with the most serious playoff implications.

For the most part, the NFL schedules the night games before the season starts (the exception: late-season Sunday games, which, in recent years, are chosen on the fly).  So, the league is, to some extent, guessing which the good teams are.  

Because of that, certain teams will appear on Monday nights more than others.  In this particular sample, the San Francisco 49ers appeared 28 different Mondays; the Seattle Seahawks, only six.

Could it be that there's something about the 49ers that caused them to beat the spread so much?  Perhaps the oddsmakers, and bettors, consistently underestimated San Francisco, those years.  For 16 consecutive years -- 1983 to 1998 -- the 49ers went 10-6 or better.  You'd think they'd have regressed to the mean at some point, but they didn't.  So, perhaps they were consistently lucky?

There's a bit of support for that -- every year from 1983 to 1989, the Niners had a winning record against the spread. Over their entire streak, they never went worse than 7-9.  Of course, a lot of that is due to their 19-9 record on Monday night ... but, still.

---

3.  Suppose that we assume that SOME of the effect is due to these kinds of factors.  Suppose that, because of shaded or incorrect spreads, the Western teams had a 53% chance of beating the spread, instead of 50%.

In that case, the odds of a 70-36 record now drop to only 1 in 225.  That gets a bit easier to accept as random chance.  

At 55 percent, you're down to 1 in 73.  

---

4.  If an effect actually exists, it's most likely a lot lower than what was actually observed.  

First, you need to regress the 70-36 to the mean, since, in nature, small effects are found much more frequently than large ones.  

Second, any effect smaller than 2 SD wouldn't have been published, which means, in general, effects that DO get published are overestimates.  

The study found the western teams beat the spread by an average of 5.26 points, with an SD of 1.33.  That means any result less than 2.66 points wouldn't have made it into print.  

Suppose that, unbeknownst to us, the actual circadian effect is 2.5 points.  If every study takes a different random sample of 106 games, fewer than half the studies will find statistical significance.  And ALL of those studies that *do* find significance will overestimate the real effect, because the minimum significant effect is 2.66.

That wouldn't be a problem if the SD was, say, 0.01, or something, because then almost ANY real effect would be found.  But, in this case, when the bar is set so high, the selectively-sampled observed effects are likely to be inflated by luck.  

---

5.  The effect for home games and road games is almost the same.  That is, no matter which of the teams is jet-lagged, the West team has the same advantage.  

That's fine, if the theory that only time of day matters.  

It does imply, though, that jet lag, or adjusting to a new time zone, doesn't matter much at all.  Which may be true, but I've seen other psychologists argue the opposite.  

---

6.  If there's such a huge effect for a three-hour difference, you'd expect to still have a substantial effect for a two-hour difference.  So I checked west-coast teams playing night games on the road in Central Time.  In those games, the effect disappeared.  The Western teams went 17-26 against the spread.

---

7.  If the effect does depend on time of day, then the effect should be similar for the fourth quarter of late-afternoon games, right?  Those games might end at 7:00 pm, while the night games start at 8:00 pm.  Not much difference there.  

Does that happen?  I haven't checked, but that would be a good test.  You could also see if the effect diminishes as the night game goes on ... by the end of the Monday night game, the West team is playing at 8:00 or 9:00 pm circadian time.  Of course, the East team is at the actual time of 11:00 pm to midnight, but, as far as I read, the paper doesn't posit that there should be a big difference between early evening and late evening.

(UPDATE: One author says the effect is based on distance from the 3:00 am physiological low point.  Going with that ... an 8:00 game is 10 hours benefit for the PST team, and only 7 hours benefit for the EST team.  But, then, an afternoon 1:00 game would be the opposite, 7 hours to 10!)

---

8.  You could also check late-afternoon games in general.  The East team is on 4:00 time, while the West team is on 1:00 time.  So East should have the advantage.  But, the authors explicitly say (in the Business Week article) that the "ramp up" effect is smaller than the "ramp down" effect, so maybe you wouldn't see anything.  

And, while we're here ... Daylight Savings Time.  The effect should be different the week the clocks change, right?  Instead of the West team playing at 5:00 and the East playing at 8:00, it's really (from a circadian standpoint) 6:00 and 9:00.  

---

9.  The author published a similar study in 1997, with similar results.  But the effect continued -- which suggests that bettors didn't react to it, and bookmakers didn't bother adjusting their lines.  

It's possible that the betting community was just shortsighted in not believing the paper's claims, but ... sharp bettors are usually quick to seize on inefficiencies like this.  To me, that's at least a bit of evidence that it might be something else.

---

10.  There are afternoon games in many other sports -- college football, NBA, NHL, major-league baseball.  You could check to see if this happens there.  

Even better, you could check sports in which *actual* performance can be measured, not just performance relative to another team.  Do golfers hit better in late-afternoon?  Do Rubik's-cube solvers have better times in events that take place later in the day?  What about bowlers, or dart-throwers?  There should be lots of ways to check.  

---

11.  In fairness, My view is that someone would have noticed if there were large time-of-day effects in other sports.  

From my own introspection, I think I'm better certain times of day than others.  I tend to get drowsy late in the afternoon, and I wouldn't be surprised at all if my (extremely minimal) athletic ability drops during that time, and other times I'm tired.  

But, I *notice* my tiredness.  Shouldn't professional athletes have noticed something, too, especially when they're so focused on their bodies and their performances?

Maybe it's possible for a team to drop from .500 to .333 without noticing the physiological changes that caused it, just like they may not notice any difference when they play worse on the road.  But ... I dunno, my gut says that's just too big an effect that nobody even *suggested* it before the academics.

---

12.  Oops!  As I was doing the final edit for this post, I remembered an NBA study that found Western teams travelling east had an advantage.  I checked, and the advantage was huge, just like this one.  And most NBA games are at night.  So ... hmmm.  

However, that study wasn't as clean as this one, with a complicated regression.  And it was denominated in actual winning percentage, rather than against the spread.  But, still ... hmmm.

In fairness, I have to say that study supports the pro-circadian argument, to some extent.


-------

My completely arbitrary, intuitive, Bayesian guess as to what's actually causing this effect?  I'd say ... 20 percent line shading, 75 percent luck, and maybe 5 percent physiology.  

I'd also guess -- again, without any justification other than my gut -- that there's a 35 percent chance that there's no real measurable effect of circadian physiology at all, and a 65 percent chance that there's a measurable, but small, effect.  (I had it at 50/50 before I recalled the study in #12.)

Regardless, I definitely don't want to imply that this study isn't important.  Any time you find a huge, 70-36 result, after a prior prediction with a plausible mechanism ... well, that's something you definitely want to put out there for serious consideration.  I'm just not as confident as the authors that what they've found is an actual thing.

Prove me wrong, somebody!








Labels: , ,

Sunday, September 29, 2013

Do NFL underdogs consistently beat the spread?

I learned two things from investment blogger Eddy Elfenbein this week.  

First, I learned that if you were invested in the S&P 500 from 1932 to 2009, you'd have made a total return of 63,000% 14,000% (not including dividends).  But if you were invested only the middle 2/3 of each month, you'd have LOST money.  Wow.

Second, I learned that, in the NFL, heavy favorites consistently fail to beat the spread.

------

Since 1978, teams favored by 12 points or more were 220-275-9 against the spread.  Ignoring the nine pushes, that's a winning percentage of only .444.  The effect was easily large enough to turn a profit, even after the bookie's vigorish.

I wondered if, maybe, this was an anomaly that existed in earlier years, but that bookies eventually caught on to and erased.  But, since 2005, those heavy favorites are 64-94-2.

Does the effect disappear for "less heavy" favorites?  Again going back to 1978, and looking at teams favored by 6 to 11.5 points ... they were 1361-1475-57 (.480 excluding pushes).  Teams favored by 0.5 to 5.5 points were 2197-2330-156 (.485).  So, yes, it seems the effect is more pronounced for the heavy favorites.

I broke it down a bit further, into "one point" buckets from 8.5 up.  Teams favored by 8.5 or 9 points were 163-187 (.466).  Teams favored by 9.5 or 10 were 186-193-12 (.491).  And so on.

For every one of those groups, except one, the favorites had a losing record.  (The exception was teams favored by 15.5 or 16, who went 19-15 against the spread.)  No favorite above 20 points has ever covered (0-7). 

------

Is this a known anomaly?  Maybe it's just my ignorance, but I've never heard that this happens.  Well, actually, I should have known ... it was obvious in the numbers for the "home underdog" effect.  

But, actually, it seems to applies equally to home/road.  Home favorites (12+) were 194-236-8, while road favorites were 26-39-1.  

I'm very, very surprised.  

Add this to the list of arguments against NCAA basketball point shaving.  If favorites failing to cover is evidence of point shaving in the NCAA, then it must also be evidence of point shaving in the NFL too, right?

But hardly anyone argues that.  I still think it's just a case of bookies shading their lines towards the underdog favorite.  



(P.S.  Good discussion of bookies' lines in some of MGL's comments here.)

Labels: , , , ,

Tuesday, March 19, 2013

NFL coaching decisions cost 0.73 wins per team


By making bad decisions on fourth down, NFL coaches are sacrificing almost three-quarters of a win per season.  That's from Matt Meiselman, who crunched some numbers with Brian Burke and posted on Brian's site.  

In 2012, the Cleveland Browns were the "worst", sacrificing a probabilistic 1.02 wins by making 42 "wrong" decisions.  The Packers were the least "worst", giving up only around half an expected win.  I would have expected New England to represent well in this measure, since Bill Belichick has often been touted as a sabermetrically-savvy coach, but the Patriots were only a bit better than average, at 0.6.

Those numbers are based on expectations for an average team.  It's quite likely that they overstate the cost, if the probabilities vary a lot based on quality of team.  My suspicion is that the quality effect is pretty small, because the spread of "wrongness" is so narrow.  In fact, the spread suggests to me that coaches are generally following the same "book" of conventional wisdom, with individual differences being pretty minor.  

The article implies that the losses are due to coaches generally being risk-averse, but doesn't give the numbers.  Is *every* bad decision caused by playing it too safe?  95 percent?  50 percent?  I don't know the answer.  My gut says ... I dunno, I'll guess 92 percent of cases are when the coach should have gone for it and didn't, instead of when he shouldn't have and did.  Matt/Brian, if you're reading this, am I close?  

-------

I'm shocked at how high the numbers are.  Losing 0.73 wins is huge, considering that the difference between a playoff team and an average team is only, what, two games out of sixteen?  

I'd bet that's by far the biggest in-game coaching factor in any major sport (leaving out the decision of who plays).  In baseball, it's the equivalent of 4.6 games per 162, which is about the same percentage of distance to the playoffs.  But I can't see that MLB managers would have anything near that much influence.

-------

At the Sloan convention, there was a lot of talk about how analytics people can increase their influence ... like, what to do or say to get coaches and management to listen to us numbers geeks.  

But, in this case, I think there's an easier path.  Any time there's a fourth-down decision, the TV broadcast could put the probabilities on the screen.  Like, for instance, "teams that go for it should be expected to win 48% of the time, while teams that punt should win only 30% of the time."  That's simple enough for viewers to understand ... which means, fans will be second-guessing the coach based on the numbers, rather than random feelings.  It would still be fun to discuss ... the ESPN guys could argue about why the percentages don't apply in this particular case, because the offense is poor, or the defense has momentum, or whatever.

In any case, it would change the nature of the second-guessing.  Right now, a coach may attract 1 pound of criticism when he plays it by the book, and 5 pounds when he goes for it.  With the probabilities on the screen, maybe the 1:5 ratio will immediately change to 1:3 or something, and then, over time, as the stats gain acceptance, all the way to 1:1.  Then, you've reached the tipping point where the coaches' incentives change.  Now, they take more flak, and sacrifice more job security, when they *don't* go by the percentages.  It wouldn't take long, I suspect, for things to change after that.



Labels: , ,

Tuesday, March 06, 2012

Are early NFL draft picks no better than late draft picks? Part IV

This is the last post about the Berri/Simmons NFL draft paper, in which they say that draft position doesn't matter much for predicting quarterback performance. Here are Parts I and II and III.

------

In his Freakonomics post, Dave Berri argues, reasonably, that quarterbacks are harder to predict from season to season than basketball players.


When he runs a regression to predict NFL quarterbacks' completion percentage this season, based on only the stat from last season, he gets an r-squared of .311. On the other hand, if he does the same thing for "many statistics in the NBA," his r-squared "exceeds 70 percent."

According to Berri,
Link

"This is not surprising since much of what a quarterback does depends upon his offensive line, running backs, wide receivers, tight ends, play calling, opposing defenses, etc. Given that many of these factors change from season to season, we should not be surprised that predicting the performance of veteran quarterbacks is difficult."

But ... aren't basketball players also subject to changes in the quality of their teammates? Why should teammates be so much more important for football than basketball?

Well, they're not. Almost the entire difference is just sample size. Let me show you.

The r-squared from season to season depends on the variances of what kinds of things are constant between seasons, and what kinds of things are not. For the most part, we can call these "talent" (t) and "noise" (n), respectively.

If the r-squared for QBs between seasons is .31, that means

(t/(t+n)) * (t/(t+n)) = .31

Taking the square root of both sides gives

t / (t+n) = .56

And from there, you can multiply both sides by (t+n), and discover that

n = .79 * t

So, for a single NFL season, the variance due to noise is 79% of the variance due to talent.

Now, in the NFL, a quarterback will get maybe 450 passing attempts per season. In the NBA, a full-time player might get three times as many (FG attempts, FT attempts (even if you take those at half weight), and 3P attempts). So, the noise should be only 1/3 as large. Instead of noise being 79% of talent, it will be only maybe 26%. Call the new value of noise n'. Then,

n' = .26 * t

If you sub that back into the first equation, you get

(t/(t+n')) * (t/(t+n')) = .63

See? Just considering opportunities raises the r-squared of .31 all the way up to .63. Berri says it should "exceed 70%", and we probably could get that to happen if we included rebounds, or used a more sophisticated stat than just shooting percentage.

So, if quarterbacks are harder to predict than basketball players, it's simply because they don't play enough for their stats to be as reliable.

(UPDATE: As Alex alludes to in the first comment to this post, my logic assumes that "t" -- the variance of talent -- is roughly the same for QBs and NBA shooters. It might not be. But the point is, assuming they're the same is a reasonable first approximation, and that leads to the conclusion that sample size is the biggest difference.

So, maybe I should have been more conservative and said that it *could be* that they don't play enough for their stats to be as reliable.)

------

Which brings me back to Berri's (and Simmons's) academic study. There, they write,

"[Our] results suggest that NFL scouts are more influenced by what they see when they meet the players at the combine than what the players actually did playing the game of football."


Well, yes -- and perhaps the scouts SHOULD be more influenced by the combine. There's lots of noise in only one season of performance, and a rational scout won't weight it too heavily. What if the scout only saw one play? Then, it's obvious that he should be more influenced by the combine than the results. The less data you have, the more you have to weight the combine results.

Look at it this way. You have two pitchers. One throws 100 mph and had an ERA of 3.50 in 50 innings. The other throws 80 mph and had an ERA of 3.20. Which do you draft? Well, *of course* you draft the 100 mph guy. It's only after 200, 300, 400, 1000 innings that you might have enough evidence to change your mind.

------

The idea of random noise and sample size never figures into this paper at all. I don't think the authors even think about it. When they see unexplained variance, they always argue that it's something like the effect of teammates, instead of looking at binomial randomness. In fact, you get the impression they think there's no randomness at all, and the scouts could be perfect if only they were smarter.

For the record, the paper has no occurrences of the words "luck," "random," "binomial," or "sample size."



Labels: , , , ,

Saturday, March 03, 2012

Are early NFL draft picks no better than late draft picks? Part III

This is about the Berri/Simmons NFL draft paper, in which they say that draft position doesn't matter much for predicting quarterback performance. Here are Parts I and II.

-----

One of the paper's most important claims is that scouts are looking at the wrong things -- specifically, the results of the NFL combine.

At the combine, prospects are tested on a bunch of objective measurements. How fast they run the 40-yard dash. Their BMI (a measure of weight to height ratio). Their intelligence, as measured by the Wonderlic test. And, of course, their height.

But, the authors argue, those things don't matter, and scouts are completely misguided in looking at what happens at the combine. They say that a QB prospect's height, BMI, 40-yard-dash time, and Wonderlic score have almost no effect on performance.

And that one point is key. Because, most of the authors' argument goes (in my words):

Premise 1: Scouts care about combine stats.
Premise 2: Combine stats affect draft position.
Premise 3: Combine stats don't predict performance.

Conclusion: Scouts don't know what they're doing.

So, premise 3 is key. How do the authors prove it?

Here's what they did. They took the 121 QBs drafted from 1999 to 2008, for which they had full combine data. Then they ran a regression to predict the QB's senior year performance based on those factors.

They found no statistical significance for any of them. And they conclude:

"Such results indicate that the combine measures are not able to capture key attributes of the quarterback."


And, in a related footnote,

"Such results indicate that there is little relationship between the combine statistics and per play performance."


Again, I think the problem is that the effect is there, but there just isn't enough data for significance. Indeed, it seems to me that they almost COULDN'T find significance in a study of that type.

Look, how much of QB performance is affected by height? Probably not much, right? There are so many other things involved. I mean, this isn't basketball: you don't see a lot of quarterbacks who are 6-foot-9, which suggests that height can't be that big a deal.

If the effect is that small, how are you going to find statistical significance with only 121 datapoints? Especially when you're trying to predict ONE SINGLE YEAR of college performance, which is very noisy (made even more so because, first, the authors chose to predict a measure that's dependent on playing time)?

You can't, and the authors didn't.

But ... that just means your study isn't precise enough. It doesn't show the effect isn't there. You can't look for a needle in a haystack, from fifty feet away, looking through the wrong end of a pair of binoculars, then say, "we didn't see a needle, so it doesn't exist."

Berri and Simmons didn't even show the results of that regression, even though it's key to their story. They just mention "not significant therefore zero" and move on. But if they HAD given their results, I bet you'd see the standard error is wide enough to encompass not just zero, but also many possible values that are perfectly reasonable and perfectly in line with what scouts think height is worth.

The same thing for the other factors -- 40 yard dash speed, Wonderlic, and BMI. It was almost guaranteed that the regression wouldn't find small effects in that sample.

What about the overall results for the four factors? You might get none of the individual combine stats being significant, but the overall correlation might be. Was it?

We really need to see the estimates for the coefficients. How many of them are reasonable individually? If you add them all up, are they also reasonable? If they are, that's all the more reason to point out that the lack of significance doesn't prove anything.

Again, the authors don't show the results ... but they do give a little hint. They run a second regression, this time using rate statistics instead of playing time stats. In a footnote to that, the authors say,

"The adjusted R-squared from these regressions, though, is in the negative range and the F-statistic is statistically insignificant."


A negative adjusted R-squared ... at first glance, that seems to say no relationship.

Except ... I looked up "adjusted R-squared". And, it turns out, for a regression with 5 variables and 121 rows, you can have a negative adjusted R-squared even if the "real" R-squared is as high as .042. That's not as small as it looks. An r-squared of .042 is an r of around 0.2, which is nothing to sneeze at.

(That makes sense. According to this calculator, a single-variable regression on 121 rows needs to find a correlation of 0.178 to find statistically significance, and I think the "adjusted" is meant to make the 5-variable case comparable to the 1-variable case.)

But 0.2 is probably higher than the effect we're looking for. Or at least, on par with it.

Suppose you ranked all the QBs on their combine stats. And then you took a QB who was +1 in SD in combine stats, and compared him to one that was -1 SD in combine stats. What kind of difference would you expect in on-field performance between the two?

Well, to get a correlation of 0.2, you'd have to expect a difference of about 3 points of NFL Quarterback Rating, or 3 or 4 positions in the performance rankings. (To estimate that, I looked here, and added 3 points to the rating of a typical QB.)

Now, remember, QB performance is very noisy. 3-4 positions in the performance rankings probably means 5-6 positions in *talent* rankings.

That seems to me like it's too much. There's no way height, BMI, Wonderlic, and 40-yard-dash speed could be *that* important, could it, that it's 5 or 6 rankings?

If not, then we're looking for an effect that's too small to find with only 121 datapoints to look at.

So, I think Berri and Simmons' regression was doomed from the start. They were guaranteed not to find significance, even if the scouts were right.




Labels: , , , ,