Sunday, September 29, 2013

Do NFL underdogs consistently beat the spread?

I learned two things from investment blogger Eddy Elfenbein this week.  

First, I learned that if you were invested in the S&P 500 from 1932 to 2009, you'd have made a total return of 63,000% 14,000% (not including dividends).  But if you were invested only the middle 2/3 of each month, you'd have LOST money.  Wow.

Second, I learned that, in the NFL, heavy favorites consistently fail to beat the spread.

------

Since 1978, teams favored by 12 points or more were 220-275-9 against the spread.  Ignoring the nine pushes, that's a winning percentage of only .444.  The effect was easily large enough to turn a profit, even after the bookie's vigorish.

I wondered if, maybe, this was an anomaly that existed in earlier years, but that bookies eventually caught on to and erased.  But, since 2005, those heavy favorites are 64-94-2.

Does the effect disappear for "less heavy" favorites?  Again going back to 1978, and looking at teams favored by 6 to 11.5 points ... they were 1361-1475-57 (.480 excluding pushes).  Teams favored by 0.5 to 5.5 points were 2197-2330-156 (.485).  So, yes, it seems the effect is more pronounced for the heavy favorites.

I broke it down a bit further, into "one point" buckets from 8.5 up.  Teams favored by 8.5 or 9 points were 163-187 (.466).  Teams favored by 9.5 or 10 were 186-193-12 (.491).  And so on.

For every one of those groups, except one, the favorites had a losing record.  (The exception was teams favored by 15.5 or 16, who went 19-15 against the spread.)  No favorite above 20 points has ever covered (0-7). 

------

Is this a known anomaly?  Maybe it's just my ignorance, but I've never heard that this happens.  Well, actually, I should have known ... it was obvious in the numbers for the "home underdog" effect.  

But, actually, it seems to applies equally to home/road.  Home favorites (12+) were 194-236-8, while road favorites were 26-39-1.  

I'm very, very surprised.  

Add this to the list of arguments against NCAA basketball point shaving.  If favorites failing to cover is evidence of point shaving in the NCAA, then it must also be evidence of point shaving in the NFL too, right?

But hardly anyone argues that.  I still think it's just a case of bookies shading their lines towards the underdog favorite.  



(P.S.  Good discussion of bookies' lines in some of MGL's comments here.)

Labels: , , , ,

Sunday, September 22, 2013

Selective sampling could explain point-shaving "evidence"

Remember, a few years ago, when a couple of studies came out that claimed to have found evidence for point shaving in NCAA basketball?  There was one by Jonathan Gibbs (which I reviewed here), and another by Justin Wolfers (.pdf).  I also reviewed a study, from Dan Bernhardt and Steven Heston, that disagreed.  

Here's a picture stolen from Wolfers' study that illustrates his evidence.



It's the distribution of winning margins, relative to the point spread.  The top is teams that were favored by 12 points or less, and the bottom is teams that were favored by 12.5 points or more.  The top one is roughly as expected, but the bottom one is shifted to the left of zero.  That means heavy favorites do worse than expected, based on the betting line.  And, heavy favorites have the most incentive to shave points, because they can do so while still winning the game.

After quantifying the leftward shift, Wolfers argues,
"These data suggest that point shaving may be quite widespread, with an indicative, albeit rough, estimate suggesting that around 6 percent of strong favorites have been willing to manipulate their performance."

But ... I think it's all just an artifact of selective sampling.

Bookmakers aren't always trying to be perfectly accurate in their handicapping.  They may have to shade the line to get the betting equal on both sides, in order to minimize their risk.

It seems plausible to me that the shading  is more likely to be in the direction consistent with the results -- making the favorites less attractive.  Heavy favorites are better teams, and better teams have more fans and followers who would, presumably, be wanting to bet that side.  

I don't know whether that's actually true or not, but it's not actually necessary.  Even if the shading is just as likely to happen towards the underdog side as the favorite side, we'd still get a selective-sampling effect.

Suppose the bookies always shade the line by half a point, in a random direction.  And, suppose we do what Wolfers did, and look at games where a team is favored by 12 points or more.  

What happens?  Well, that sample includes every team with a "true talent" of 12 points or more --  with one exception.  It doesn't include 12-point teams where the bookies shaded down (for whom they set the line at 11.5).

However, the sample DOES include the set of teams the bookies shaded *up* -- 11.5-point teams the bookies rated at 12.

Therefore, in the entire sample of favorites, you're looking at more "shaded up" lines than "shaded down" lines.  That means the favorites, overall, are a little bit worse than the line suggests.  And that's why they cover less than half the time.  

You don't need to have point shaving for this to happen.   You just need for bookies to be sufficiently inaccurate.  That's true even if the inaccuracy is on purpose, and -- most importantly -- even if the inaccuracy is as likely to go one way as the other.

------

To get a feel for the size of the anomaly, I ran a simulation.  I created random games with true-talent spreads of 8 points to 20 points.  I ran 200,000 of the 8-point games, scaling down linearly to 20,000 of the 20-point games.  

For each game, I shaded the line with a random error, mean zero and standard deviation of 2 points.  I rounded the resulting line to the nearest half point.  Then, I threw out all games where the adjusted line was less than 12 points. 

(Oops! I realized afterwards that Wolfers used 12.5 points as the cutoff, where I used 12 ... but I didn't bother redoing my study.)  

I simulated the remaining games as 100 possessions each team, two-point field goals only.

The results were consistent with what Wolfers found.  Excluding pushes, my favorites went 355909-325578, which is a .478 winning percentage.  Wolfers' real-life sample was .483.  

--------

So, there you go.  It's not proof that selective sampling is the explanation, but it sounds a lot more plausible than widespread point shaving.  Especially in light of the other evidence:

-- If you look again at the graph of results, it looks like the entire curve moves left.  That's not what you'd expect if there were point shaving -- in that case, you'd expect to see an extraordinarily large number of "near misses".  

-- as the Bernhardt/Heston study showed, the effect was the same for games that weren't heavily bet; that is, cases where you'd expect point-shaving to be much less likely.

-- And, here's something interesting.  In their rebuttal, Bernhardt and Heston estimated point spreads for games that had no betting line, and found a similar left shift.  Wolfers criticized that, and I agreed, since you can't really know what the betting line would be. 

However: that part of the Bernhardt/Heston study perfectly illustrates this selective sampling point!  That's because, whatever method they used to estimate the betting line, it's probably not perfect, and probably has random errors!  So, even though that experiment isn't a legitimate comparison to the original Wolfers study, it IS a legitimate illustration of the selective sampling effect.

---------

So, after I did all this work, I found that what I did isn't actually original.  Someone else had come up with this explanation first, some five years ago.  

In 2009, Neal Johnson published a paper in the Journal of Sports Economics called "NCAA Point Shaving as an Artifact of the Regression Effect and the Lack of Tie Games."  

Johnson identified the selective sampling issue, which he refers to as the "regression effect."  (They're different ways to look at the same situation.)  Using actual NCAA data, he comes up with the result that, in order to get the same effect that Wolfers found, the bookmakers' errors would have to have had a standard deviation of 1.35 points.  

I'd quibble with that study on a couple of small points.  First, Johnson assumed that the absolute value of the spread was normally distributed around the observed mean of 7.92 points.  That's not the case -- you'd expect it to be the right side of a normal distribution, since you're taking absolute values.  The assumption of normality, I think, means that the 1.35 points is an overestimate the amount of inaccuracy needed to produce the effect.

Second, Johnson assumes the discrepancies are actual errors on the part of the bookmakers, rather than deliberate line shadings.  He may be right, but, I'm not so sure.  It looks like there's an easy winning strategy for NCAA basketball -- just bet on mismatched underdogs, and you'll win 51 to 53 percent of the time.  That seems like something the bookies would have noticed, and corrected, if they wanted to, just by regressing the betting lines to the mean.  

Those are minor points, though.  I wish I had seen Johnson's paper before I did all this, because it would have saved me a lot of trouble ... and, because, I think he nailed it.


Labels: , , , , , ,

Friday, August 10, 2007

NCAA point-shaving study convincingly debunked

There have been a couple of Justin Wolfers papers in the news recently. There was the “racial bias in the NBA” study a few months ago (co-authored with Joseph Price), which I reviewed here. And, last May, there was an article on NCAA point shaving, which I’ll talk about now.

Wolfers analyzed the outcomes and betting lines for over 70,000 NCAA Division I games. He found that, overall, the favorite covered the spread almost exactly half the time (50.01%), exactly as you would expect. But heavy favorites (defined as a spread of more than 12 points) covered only 48.37% of the time.

He then found that, for heavily favored teams, the results weren’t symmetrical. That is, for teams favored by, say, 4.5 points, they would cover by 0-4 almost exactly as often as by 5-8. But for large spreads, like 15.5 points, teams covered by 0-15 a lot more than by 16-31.



“Among teams favored to win by 12 points or more, 46.2% of teams won but did not cover, while 40.7% were in the comparison range of outcomes.”

The difference is about 6% of those games, which we can call the “Wolfers Discrepancy” for his dataset.

Wolfers concludes that 6% figure is evidence of cheating – that point shaving caused 3% of potential covers to become non-covers. Since only half of games can be shaved (games where the favorite is leading), that means that the proportion of corrupt games was double that. Conclusion: the fix was in in about 6% of 12-point-plus-spread games.

This study, and the 6% figure, got a lot of publicity.

Here’s a post from The Wages of Wins. Here’s an article from the New York Times. And you can find a lot more by Googling.

Not receiving anywhere near as much publicity is a rebuttal paper by Dan Bernhardt and Steven Heston. And that’s too bad, because it’s a great study, and it thoroughly debunks Wolfers’ conclusions. These guys nail it.

Bernhardt and Heston argue that there’s no reason to expect the symmetry that Wolfers assumes. For instance,



"Consider a 14 point favorite that is up by only seven points with five minutes to play. To maximize its chances of winning, the favorite will “manage the clock,” holding the ball to reduce the number of opportunities that the other team has to score ... In contrast, the same favorite up by 21 is sure to win, and has no need to manage the clock, raising the expected increment in winning margin, generating an asymmetric distribution in winning margins."

That’s very plausible, and the authors back it up with some clever analysis. They argue that if the Wolfers Discrepancy is caused by cheating, then it should be much higher than 6% in games that are more likely to be corrupt, and less then 6% in games that are less likely to be corrupt. And so they check – four different ways.

First, they note that a fixed game will attract heavy betting on the underdog, and all that money will move the betting line to reduce the spread. So they split games into two groups: games where margin moved the “wrong” way, towards the favorite, and all other games. It turns out that the Wolfers’ Discrepancy is almost exactly the same between the two groups, suggesting that it’s a natural part of the scoring distribution rather than an artifact of corruption.

Second, they note that if game fixing happens, it happens in games where the final score is very close to the spread. So, for a 14-game spread, instead of comparing finals of 1-14 points to finals of 15-28 points, they compare 8-14 to 15-21, to narrow the range. Now, the number of corrupt games in these smaller samples should be equal to the number in the bigger sample. But the denominator, the number of total games, is smaller. Therefore, the percentage of apparently corrupt games – the Wolfers Discrepancy -- should rise.

It didn’t. In fact, it *dropped*, by more than two-thirds.

As a third test, they check games where the line did move towards the underdog, games where corruption is plausible. You’d expect, in those games, that the more the line moved, the more likely the game is fixed, and so the larger the Wolfers Discrepancy for those games.

The result -- nothing statistically significant:

5.61% -- line moved towards the favorite
6.64% -- line didn’t move
7.17% -- line moved towards the underdog by 0 - 0.5 points
8.43% -- line moved towards the underdog by 0 - 1 point
6.83% -- line moved towards the underdog by 0 - 1.5 points.


Finally, they analyzed games where there was no betting line, which usually happens when there isn’t enough interest in the game for the sports books to bother. (For those games, Bernhardt and Heston had to estimate a spread using team strength ratings.)

If there is no betting, there’s no incentive to shave points. And so, if the Wolfers Discrepancy is really caused by corruption, it should be zero for those games.

But it wasn’t zero. In fact, it was 6% -- almost exactly the same as Wolfers found in his own study!

Truly outstanding work by Bernhardt and Heston. They took a statistical effect that Wolfers claimed showed corruption, and proved, four different ways, that the effect exists even when corruption is unlikely. In addition, they provided a plausible explanation of what the effect might be.


The NCAA should buy these guys a beer.

Hat tip:
Zubin Jelveh

Labels: , ,

Tuesday, July 31, 2007

Jonathan Gibbs study unconvincing on NBA point shaving

The Tim Donaghy NBA scandal has reignited interest in point shaving in general. In that light, a bunch of sources – a New York Times column, the Freakonomics blog, and TWOW, to name three – have mentioned a new study by Jonathan Gibbs, an undergraduate at Stanford. ESPN.com has an excellent overview article on its TrueHoop blog.

The study is called "
Point Shaving in the NBA: An Economic Analysis of the National Basketball Association's Point Spread Betting Market." Gibbs says he found reasonable grounds to suspect point shaving, and everyone (that I've read) seems to believe that the study does indeed show good evidence.

I disagree. Let me start by summarizing Gibbs' results (in between dotted lines), then I'll comment.

-----

Gibbs built a database of a large number of NBA games. He found that the favorite beat the spread almost exactly half the time, 49.95%. However, the larger the spread, the less likely the favorite would cover. If you bet underdogs of 10 or more points, you'll win 52.15% of your bets. At 13 points or more, 53.04% of underdogs covered. The betting market seems to be inefficient when it comes to games between unbalanced teams.

Taking home underdogs only, it's even worse. At 10 points or more, home underdogs had a .548 winning percentage against the spread. At 13 points or more, it was an unbelievably large .718 (although in only 42 games). You can apparently make very good money by betting on home underdogs. [There's a similar effect in the NFL, also – see
here.]

If you divide all games into three groups, based on the size of the spread (0-6 points, 6.5-12 points, 12.5 points and up), you find a continuum. Narrow favorites tend to cover more than 50% of the time. Medium favorites cover almost exactly 50% of the time. And, as mentioned, heavy favorites cover less than 50% of the time.

Moreover, there is an interesting "last minute" effect. After 47 minutes of play – or 46, 45, 44, and 43 -- favorites in games with small spreads have actually covered less than half the time. But in the last minute, they tend to outscore their opposition to the extent that they go from a couple of points below the spread, to a couple of points above it. That's because of the opposition's fouling strategy. If a team is up by 4 with less than a minute to play, the opposition will take deliberate fouls in order to get the ball back. On very rare occasions, the strategy works and they win. Usually, it just gives the team in the lead more points. And so, the favorite's 3 point lead often turns into 6 or 7 following the gift foul shots.

The same pattern appears for medium-spread games, but to a lesser extent – which makes sense, because fewer medium-spread games are within "fouling distance" towards the end.

But for games with large spreads, the effect goes the other way. Heavy favorites are a little bit ahead of the spread with a minute to go, but wind up behind the spread by the end of the game. The author believes this is evidence of point shaving by the favorite.

Finally, Gibbs runs a few regressions on the final score (against the spread) against a bunch of other variables. In several cases, the "point spread" variable is significant, which suggests to Gibbs that teams are aware of, and being affected by, the spread.

-----

Okay, now my comments.

First, the shift by the heavy favorites from covering the spread after 47 minutes, to not covering the spread after 48. Isn't there a more likely explanation than point shaving? Couldn't it just be normal "garbage time" strategy? Up by 15 with a minute to go, the team leading will perhaps want to seal the victory by wasting as much of the 24 seconds as it can, rather than rubbing it in by trying to score more points. The opposition has no such motivation, and they might gain a few points in the last minute on the couple of possessions they get.

Second, Gibbs is puzzled that the point spread variable should be significant in regressions. But, of course it should, because it's a proxy for team quality. Suppose team X is tied with over team Y with a minute to go. Who will win? From the information I gave, it's 50-50. But now suppose I tell you that X was a ten-point favorite over Y. Now who's more likely to win? Team X. Not because X cares about the spread, but because the spread means they're a better team than Y! Of course, it's only one minute, and so team X might only be a very slight favorite. But Gibbs' study includes an awful lot of games, which provides more than enough data for a slight advantage to become statistically significant.

Third, I think Gibbs misinterprets the results of his regressions. In one model, he tries to estimate the probability of the favorite covering by the betting line, and the five scores with 1 to 5 minutes to go. He finds that "as the ... point spread becomes ... larger, the likelihood of ... covering decreases." (And that this is evidence of shaving.)

But that assumes all dependent variables being equal – for instance, the amount of the lead (against the spread) after 47 minutes. Consider two situations where that margin is one point.

-- X is five points ahead of Y, against a 4-point spread, with one minute to go.
-- A is fifteen points ahead of B, against a 14-point spread, with one minute to go.

Which team is less likely to beat the spread? A, of course. They're 15 points up, so they’re going to kill time, let B score a few points, and win by 11. X is only 4 points up, so they can't afford to do that, so they'll try to score and increase their margin.

As Gibbs says, "the betting line is a significant determinant for whether the favored team covers." Yes, but that's only because the way he set up his regression makes it a proxy for actual margin of victory. It has nothing to do with the spread at all!

Fourth, there's something about these regressions I don't understand. Take the one just described, for instance. It turned out that all the variables were significant – margin with 5 minutes to go, margin with 4 minutes to go, margin with 3 minutes to go, margin with 2 minutes to go, and margin with 1 minute to go. Why is that? If A is ahead by 4 points with a minute to go, what difference does it make how much they led by three minutes ago? Shouldn't only the 1-minute-to-go number be significant? If the score is 93-89, why does it matter how it got that way?

I suppose you could come up with some hypothetical ... maybe if you were 10 points up before, but only 1 point up now, you might have benched the unfocused regulars who blew 9 points in three minutes, and that might make the last minute go differently. But that doesn't seem right, and I really wonder what's really going on in that regression.

Fifth, and very important: the fact that some teams beat the spread more or less often than 50% can't possibly be evidence of point shaving. Oddsmakers have studied thousands of games, and are the best in the world at what they do. If 6% of games were being fixed – or even 1% of games – the oddsmakers would have taken that into account. Even if they didn't know, or even suspect, that point shaving was happening, they would just notice that heavy favorites win by fewer points, and adjust accordingly.

Suppose that the oddsmakers figure that team X is good enough that they should beat team Y by 12 points. But, they know from studying past games that their model isn't good enough – that because of point shaving, the average is 11 points. So they drop the line to 11. Now, the corrupt player, seeing that, will try to lose by only 10. The oddsmaker figures that out, and drops the line to 10. The corrupt player now tries for 9. The oddsmaker matches, and the player drops to 8. On it goes, until the line has dropped so far that the player can just barely shave that many points without arousing suspicion. That's an equilibrium, and now the odds correct exactly for the probability of point-shaving.

So my argument is that looking at how often teams beat the spread can't possibly provide evidence of fixed levels of point-shaving. It can only show evidence of point shaving that's *unexpected* by the oddsmakers. Since Gibbs' sample of games is many seasons long, and oddsmakers are very, very good at what they do, the failure to beat the line can't automatically be attributed to cheating.

Finally, suppose you're a corrupt player trying to miss shots to come under the spread. Wouldn't you do as little as possible to ensure that result? Suppose the spread is 10, you're up by 9, and so you deliberately miss a shot with 30 seconds left. The opposition scores a three, and now you're only up by 6 with five seconds left. Wouldn’t you stop trying to lose? No matter what, the opposition is going to cover the spread. You won't risk getting caught just to lose by 6, when losing by 8 is perfectly acceptable.

And you're not going to cheat in the first or second quarter, when the game is close and you don't know how the score's going to wind up. For one thing, your team might lose outright because of your shaving. For another, the opposition might play so well that you don't have to take the risk of shaving at all.

To save your butt, you're not going to start deliberately missing until late in the fourth quarter. And even then, only when the game is close compared to the spread.

And so, if teams were regularly point shaving, you'd see an unusual shape when you plotted the results. You wouldn't expect to see the nice curve that Gibbs found. His graphs are smooth and symmetrical, just shifted left. Not only are –1 and -2 (against the spread) too high, but so is –5, and –10, and –15. But why would anyone cheat on a –15 game?

Instead of smoothness, wouldn't you expect a big hump only close to zero, at minus 1 and 2 and 3? And wouldn't all those extra games come from games that were plus 1 and plus 2 and plus 3? Gibbs claims that 6% of lopsided games may involve point shaving. That should create a strikingly huge growth immediately to the left of 0, and a similar valley immediately to the right of zero. But that's not what we see.

---

It's definitely possible that I'm missing something, especially considering that so many people have read Gibbs' paper and found it convincing. But I don't think it gives any real evidence of point shaving whatsoever.


Labels: , ,