Thursday, May 03, 2018

NHL referees balance penalty calls between teams


That finding, from Michael Lopez, shows that the next penalty in an NHL game is significantly less likely to go to the team that's had more penalties so far in the game.

That was a new finding to me. A few years ago, I found that the next penalty is more likely to go to the team that had the (one) most recent penalty -- but I hadn't realized that quantity matters, too.

(My previous research can be found here: part one, two, three.)

So, I dug out my old hockey database and see if I could extend Michael's results. All the findings here are based on the same data as my other study -- regular season NHL games from 1953-54 to 1984-85, as provided by the Hockey Summary Project as at the end of 2011.

-------

Quickly revisiting the old finding: referees do appear to call "make-up" penalties. The team that got the benefit of the most recent power play is almost 50 percent more likely to have the next call go against them. That team got the next penalty 59.7% of the time, versus only 40.3% for the previously penalized team.

39599/98167 .403 -- team last penalized
58568/98167 .597 -- other team

Now, let's look at total numbers of penalties instead. I've split the data into home and road teams, because road teams do get more penalties -- 52 percent vs. 48 percent overall.  (That difference is mitigated by the fact that referees balance out the calls. The first penalty of the game goes to the road team 54 percent of the time. The drop from 54 percent for the first call, down to 52 percent overall, is due to the referees balancing out the next call or calls.)

So far, nothing exciting. But here's something. It turns out that the *second* call of the game is much more likely than average to be a makeup call:

.703 -- visiting penalty after home penalty
.297 -- home penalty after home penalty

.653 -- home penalty after visiting penalty 

.347 -- visiting penalty after visiting penalty

Those numbers are huge. Overall, there are more than twice as many "make up" calls as "same team" calls.

In this case, quantity and recency are the same thing. Let's move on to the third penalty of the game, where they can be different.  From now on, I'll show the results in chart form:

.705 0-2 
.462 1-1
.243 2-0

Here's how to read the chart: when the home team has gone "0-2" in penalties -- that is, both previous penalties to the visiting team -- it gets 70.5% of the third penalties. When the previous two penalties were split, the home team gets 46.2%, similar to the overall average. When the home team got both previous penalties, though, it draws the third only 24.3% of the time (in other words, the visiting team drew 75.7%).

Here's the fourth penalty. I've added sample sizes, in parentheses.

.701 0-3 (755)
.559 1-2 (6951)
.373 2-1 (5845)
.261 3-0 (468)

It's a very smooth progression, from .701 down to .261, exactly what you would expect given that make-up calls are so common. 

Here's the fifth penalty:

.677 0-4 ( 195)
.619 1-3 (3244)
.465 2-2 (6950)
.351 3-1 (2306)
.316 4-0 ( 117)

That's the chart that corresponds to Michael Lopez's tweet, and if you scroll back up you'll see that these numbers are pretty close to his.

Sixth penalty:

.667 0-5 (  48)
.637 1-4 (1182)
.520 2-3 (4930)
.413 3-2 (4134)
.323 4-1 ( 773)
.226 5-0 (  31)

Again, the percentages drop every step ("monotonically," as they say in math).

Seventh penalty:

.692 0-6 (  13)

.585 1-5 ( 369)
.577 2-4 (2528)
.489 3-3 (4140)
.399 4-2 (1798)
.379 5-1 ( 219)
.200 6-0 (  13)

Eighth penalty:

.667 0-7 (   3)
.607 1-6 ( 122)
.588 2-5 ( 969)
.527 3-4 (2721)
.422 4-3 (2414)
.374 5-2 ( 652)
.412 6-1 (  68)
.000 7-0 (   1)

Still a perfect pattern.  It breaks up just a little bit here, for the ninth penalty, but that's probably just small sample size.

.000 0-8 (   1)
.553 1-7 (  38)
.586 2-6 ( 348)
.566 3-5 (1358)
.484 4-4 (2063)
.392 5-3 (1037)
.340 6-2 ( 191)
.333 7-1 (  21)

(This is getting boring, so here's a technical note to break the monotony. I included all penalties, including misconducts. I omitted all cases where both teams took a penalty at the same time, even if one team took more penalties than the other. In fact, I treated those as if they never happened, so they don't break the string. This may cause the results to be incorrect in some cases: for instance, maybe Boston takes a minor, then there's a fight and Montreal gets a major and a minor while Boston gets only a major. Then, Montreal takes a minor. In that case, the study will treat the Montreal minor as a make-up call, when it's really not. I think this happens infrequently enough that the results are still valid.)

I'll give two more cases. Here's the twelfth penalty:

.692 2-9 ( 13)
.623 3-8 ( 61)
.532 4-7 (250)
.506 5-6 (478)
.488 6-5 (459)
.449 7-4 (198)
.457 8-3 ( 35)
.200 9-2 (  5)

Almost perfect.  But ... the pattern does seems to break down later on, at the 14th to 16th penalty (I stopped at 16), probably due to sample size issues. Here's the fourteenth, which I think is the most random-looking of the bunch. You could almost argue that it goes the "wrong way":

.000  2-11 (  1)
.375  3-10 (  8)
.333  4- 9 ( 27)
.516  5- 8 ( 95)
.438  6- 7 (169)
.480  7- 6 (148)
.465  8- 5 ( 71)
.577  9- 4 ( 26)
.600 10- 3 (  5)

Still, I think the overall conclusion isn't threatened, that quantity is a factor in make-up calls.

------

OK, so now we know that quantity matters. But couldn't that mean that recency doesn't matter? We did find that the team with the most recent penalty was less likely to get the next one -- but that might just be because that team is also more likely to have a higher quantity at that point. After all, when a team takes three of the first four penalties, there's a 75 percent chance* it also took the most recent one. 

(* It's actually not 75 percent, because make-up calls make the sequence non-random. But the point remains.)

So, maybe the recency effect is just an illusion, by the quantity effect. Or vice versa.

So, here's what I did: I broke down every row in every table by who got the more recent call. It turns out: recency does matter.

Let's take that 3-for-4 example I just used:

.613 home team overall     (3244)
---------------------------------
.508 after VVVH            ( 486)
.639 after other sequences (2758)

From this, it looks like there's both aspects here. When the home team is "up 3-1" in penalty advantage, it gets only 51 percent of the penalties if its previous penalty was the last of the four. That's still more than the 46.1 percent it gets to start the game, or the 46.5 percent it would get if it had been 2-2 instead of 3-1.

This seems to be true for most of the breakdowns -- maybe even all the ones with large enough sample sizes. I'll just arbitrarily pick one to show you ... the ninth penalty, home team 3-5.

.392 home team overall     (1037)
---------------------------------
.362 when most recent was H (743)
.469 when most recent was V (294)

Even better: here's the entire chart for the eighth penalty: overall vs. last penalty went to home team ("last H") vs. last penalty went to visiting team "last V". 

overall   last H    last V
----------------------------------
 .607      .750      .596      1-6 
 .588      .477      .609      2-5 
 .527      .446      .584      3-4 
 .422      .372      .518      4-3 
 .374      .357      .466      5-2 
 .412      .406      .500      6-1 

Clearly, both recency and quantity matter. Holding one constant, the other still follows the "make-up penalty" pattern. 

Can we figure out *how much* is recency and *how much* is quantity?  It's probably pretty easy to get a rough estimate with a regression. I'm about to leave for the weekend, but I'll look at that next week. Or you can download the results (speadsheet here) and do it yourself.




Labels: , , ,

Monday, January 30, 2012

Do NHL teams get a boost after killing a two-man advantage?

In an OHL game I was watching the other day, one of the teams had a two-man advantage and didn't score. The announcer was disappointed that the shorthanded team to get a boost from having killed off the penalties, as conventional wisdom says they should.

Is conventional wisdom right? Now that I have access to a database of NHL games (thanks again to the Hockey Summary Project), I was able to check.

This study is basically the same format as the study I did on fights a few weeks back. I found all games from 1967-68 to 1984-85 where one team killed off a two-man advantage (of any length). Then, I found a random control game, which matched the score differential and the relative quality of the home and road teams. When I was done, I had two pools, each comprised of 1,703 games.

The teams that killed the penalties scored an average 0.26 more goals than their opponents from that point to the end of the game (actually, to the 17:00 mark of the third period). On the other hand, the control team scored only 0.12 more goals then their opponents.

That's statistically significant, at almost exactly 2 SDs.

I'll put that in chart form to make it easier to read, along with the SD. I use the term "killing teams" to mean the ones that actually killed off the two-man advantage.

Killing teams .... +0.26 goals (+/- 0.05)
Control teams .... +0.12 goals (+/- 0.05)
------------------------------------------
Difference ....... +0.14 goals (+/- 0.07)

At six goals per win, you'd have expected the extra goals to have resulted in around 40 extra wins. They actually resulted in 32 extra wins. Actually, 36 extra wins, minus 8 fewer ties:

Killing teams .... 836-604-263
Control teams .... 806-626-271
------------------------------------
Difference ....... +36 wins, -8 ties

So, should we conclude that killing off a two-man advantage causes a psychological boost? Well, not so fast. Because, after you take two consecutive penalties, the referee is very likely to try to even things up by giving future penalties to the other team.

The difference of +0.14 goals is almost exactly what you'd get from a single power play. So, if the result of surviving a two-man advantage is that you get one extra "free" power play in the remainder of the game, that would explain the results exactly.

As it turns out, it's not quite that high. It's only half that high. On average, the teams that survived being shorthanded two men got about half an extra power play in the remainder of the game:

Killing teams ... +.346 power plays rest of game
Control teams ... -.130 power plays rest of game
------------------------------------------------
Difference ...... +.476 power plays rest of game

That leaves about 0.07 goals per game as the unexplained difference. It's only 1 SD, which is no longer statistically significant. It's about the effect of half a power play. Or, with an average save percentage of .900, it works out to 7/10 of an additional shot on goal.

--------

We can also handle the penalty issue another way. We can insist that when we choose a control game for the real game, we make sure the control team was the lone who took the last penalty. That way, we'd expect some of the referee "evening up" difference to disappear. Perhaps not all of it, because a two-man advantage isn't the same as a one-man advantage -- but at least part of it.

The additional restriction reduced the sample size to 1,662 games; for the remaining 41 games, I couldn't find a suitable control.

As it turns out, the goal difference stays about the same, even though the penalty difference is significantly reduced:

Killing teams ... +0.25 goals (+/- 0.05)
Control teams ... +0.08 goals (+/- 0.05)
----------------------------------------
Difference ...... +0.17 goals (+/- 0.07)

Killing teams ... +.340 power plays rest of game
Control teams ... +.032 power plays rest of game
------------------------------------------------
Difference ...... +.308 power plays rest of game

The difference of .308 power plays accounts for around .04 goals of the observed .17 difference. That leaves .13, which is a little less than 2 SD from zero. Not statistically significant, but close. (Technically, it's even less than that, because the control games aren't completely independent. Also, when I ran the study a second time, I got +0.10 goals instead of +0.08, which lowers the difference. So think of the 1.9 SD as probably a bit too high.)

Strangely, though, there wasn't as much difference in game results; only the equivalent of 13.5 wins:

Killing teams ... 815-591-256
Control teams ... 807-610-245
------------------------------
Difference: +8 wins, +11 ties

Again at six goals per win, you'd expect 47 wins, not 13.5. What happened?

Well, it turns out that the "killing" teams spent a lot of their goals winning blowouts. For instance, in games won by six goals or more, they were 81-34. The control group was only 73-51.

In those games, the difference was 12.5 wins. That normally "costs" 75 goals, but, for these games, the difference was really around 150 goals. So, that accounts for 75 of the 282 goal difference right there.

The "killing" group also "wasted" goals in the 3- and 4-goal games. That was offset by the opposite effect in five-goal games, but not by much.

------

If you recall, we found the same effect when we looked at fighting: teams that started a fight appeared to score more goals, but not necessarily win more games.

What connects the two studies is ... penalties. It could be that teams that get penalized a lot win a lot of blowouts. Not necessarily because of cause-and-effect, but because it just so happened that, between 1967 and 1984, certain teams just happened to be high in both categories.

Or, it could be coincidence. Or, it could be something else.

------

For my bottom line, I'd say: after killing off a two-man advantage, teams did appear to benefit by about 1/7 of a goal. Half of that can be traced to referees calling fewer penalties against them in the remainder of the game.

The other half is unknown. It's not statistically significant, so you have to give serious consideration to the idea that it's just coincidence ... but the teams *did* appear to benefit, by around 0.07 goals.

Historically, the average size of the "boost" in a team's play after a two-man kill has been small: the equivalent of less than a single shot on goal over the remainder of the game.



Labels: , , ,

Sunday, January 15, 2012

Are more NHL penalties called in back-to-back games?

In a comment to one of the posts on "make-up" penalties, J.-P. Martel wrote,

"... blow-outs can easily lead to situations that get out of hand, so referees may call penalties on the leading team so that the trailing team still thinks it has a chance to come back, rather than resort to fighting to "prepare" the next game between the two teams.

Actually, you may want to check penalties in the second half of the third period when the teams' next game is (or may be, depending on outcome) against each other (particularly in the playoffs), as opposed to when it's not."


So I did. And, J.-P. is right, it looks like there's something there.

I found all cases from 1967-68 to 1984-85 where teams played back-to-back games (regular season only). Then, I formed three groups:

-- first game of back-to-back games
-- second game of back-to-back games
-- other games that year between those two teams

It turns out that, overall, there are more penalties than usual in the first game, and fewer penalties than usual in the second game:

First game .... 12.36
Second game ... 10.87
Other games ... 11.77

Broken down by periods:

-------------- Gm 1 --- Gm 2 --- Other
--------------------------------------
Period 1 ..... 4.78 ... 3.98 ... 4.37
Period 2 ..... 4.25 ... 3.75 ... 4.12
Period 3 ..... 3.32 ... 3.13 ... 3.26
--------------------------------------
Total ....... 12.36 .. 10.87 .. 11.77

So: there's 0.6 extra penalties in the first game, and 0.9 fewer penalties in the second game.

I thought the second game would be dirtier because the player are holding recent grudges from the previous game, but the numbers show the opposite. The players seem to be more aggressive early, rather than late. In fact, more than half the "first game" effect happens in the first period. By contrast, a large "second game" effect seems to last two periods rather than one.

Most of the differences are statistically significant, which suggests that they're all real. For those scoring at home, here are the standard errors:

-------------- Gm 1 --- Gm 2 --- Other
--------------------------------------
Period 1 ..... 0.16 ... 0.12 ... 0.06
Period 2 ..... 0.13 ... 0.11 ... 0.06
Period 3 ..... 0.15 ... 0.12 ... 0.07
--------------------------------------
Total ........ 0.30 ... 0.24 ... 0.13

Finally, coming back to J.-P.'s hypothesis about the second half of the third period of the first game, here are the numbers:

First game .... 1.76
Second game ... 1.63
Other games ... 1.67

So, yes, there's a small effect where, when the teams are going to meet again next game, the referee calls more penalties than normal in the last ten minutes of the third period. Whether that's because of the referee, or the players, we can't tell.

Taken alone, these differences aren't statistically significant. But, considering they match the pattern, and the broader picture is statistically significant, we can be fairly confident that this is a real effect we're seeing.

That's actually why I saved J.-P.'s scenario for last, so I could first show that the effect is probably real and not just random.

-----

UPDATE, 1/15/2012:

Technical note: the "other games" rows and columns are weighted by games, rather than matchups. Suppose teams A and B had back-to-back games, and so did C and D. But A and B met only 2 other times that year, while C and D met 4 other times. That means that C/D will be overrepresented in the "other games" column.

If I reweight that column so A/B and C/D get equal weight, the results change just a little bit. These are the revised "other" columns:

Overall ...... 11.48 (was 11.77)
1st period .... 4.24 (was 4.37)
2nd period .... 4.07 (was 4.12)
3rd period .... 3.15 (was 3.26)
Last 10 min ... 1.59 (was 1.67)



Labels: ,

Monday, January 09, 2012

Do NHL referees call "make up" penalties? Part IV

A couple of links to other similar studies on penalty-calling:

1. Commenter Jack linked to this article with some basketball foul-calling data. Turns out the more consecutive fouls against one team, the more likely the next will go to the other team.

2. Another reader pointed me to a 2009 hockey study (web version here, PDF here) by Jack Brimberg and William J. Hurley. They looked at the first three penalties of every game, and found results similar to what I found.


Labels: , , , , ,

Friday, January 06, 2012

Do NHL referees call "make up" penalties? Part III

The last two posts talked about how NHL referees are more likely to "even-up" their calls, issuing the next penalty to the opposing team 60% of the time.

This post, I'll show a regression I ran to quantify the effect a bit better. If you're not interested, just skip the technical parts (smaller font). If you're *really* not interested, you can probably just skip this entire post, since the results are pretty much the same as shown in the previous posts.

---

Technical notes 1:

Even though we're interested in whether the referee called the next penalty to the "other" team, I set up the regression to predict whether the referee called the next penalty to the *home* team. That just makes everything easier to interpret, but, as I'll describe, it still lets us estimate the "even-up" effect.

In the study, I ignored all misconduct penalties, all first penalties of the game, and all penalties where the other team had a player called at the same time. (I treated those penalties as if they didn't exist, so they didn't interrupt "consecutiveness" of the two surrounding penalties.)

Non-dummy variables I used: Time gone in game. Time since last penalty.

Dummy variables I used: PP goal on last penalty. SH goal on last penalty. Home team lead, from -3 to +3, except 0 (that is, six dummy variables), where anything more or less than 3 goals was coded as 3. All eight of the previous variables interacted with "whether the last penalty was to the home team." And, of course, the dummy for "whether the last penalty was to the home team" itself.

The regression shows the home team percentage diminishes during the game, by about 1 percentage point per period. In all the numbers in this post, I just used the beginning of the game. If you want the middle of the game, subtract about 1.5 percent from each "home team" percentage (or add 1.5 percent to each "visiting team" percentage) if you want to adjust to the middle of the second period.

Also, the regression says you have to subtract about 1 percentage point for every 20 minutes since the last penalty. I didn't bother for this post. If you assume penalties are usually around 5 minutes apart, feel free to subtract 0.25 percentage points from each of the "home team" percentages.

Those two time adjustments won't affect the "even-up" numbers, just the raw percentages of home team penalties.


-----

OK, first, let me show you the percentage of penalties taken by the home team, by game score. Clearly, teams are more likely to take penalties when they're ahead in the game.

After the visiting team took the last penalty, the home team took:

46.3% when down by 3+
49.2% when down by 2
53.1% when down by 1
61.2% when tied
64.3% when up by 1
67.9% when up by 2
66.8% when up by 3+

And after the home team took the last penalty, the home team took:

34.5% when down by 3+
32.8% when down by 2
32.3% when down by 1
35.0% when tied
42.6% when up by 1
47.6% when up by 2
52.9% when up by 3+

Obviously, you can do this for visiting teams just by subtracting all the percentages from 100.

-------

Now, we can calculate the "even-up" effect as the difference between the lines of the two tables. When the score was tied, the home team took:

61.2 percent after visiting team penalty
35.0 percent after home team penalty
-------------------------------------------------
26.2 percent difference


You can convert to visiting team just by subtracting the first two numbers from 100%. The difference has to come out the same. I'll do that anyway. When the score was tied, the visiting team took:

38.8 percent after visiting team penalty
65.0 percent after home team penalty
-----------------------------------------------------
26.2 percent difference

It turns out the "even-up" difference is highest for tie games. Here's the full breakdown:

11.8 percent difference down by 3+
16.4 percent difference down by 2
20.7 percent difference down by 1
26.2 percent difference tied
21.7 percent difference up by 1
20.3 percent difference up by 2
13.9 percent difference up by 3+

------

Tango suggested there might be an extra "compassion" effect when the team scores a PP or SH goal. He seems to have been right. The effect is small, relative to the overall effect, but still enough to affect the games:

-- If the home team scored a PPG on the previous penalty, add 1.2 percentage points to the above differences.

-- if the visiting team scored a PPG on the previous penalty, subtract 3.2 percentage points to the above differences.

-- if the home team scored a SHG on the previous penalty, add 2.8 percentage points from the above differences.

-- if the visiting team scored a SHG on the previous penalty, subtract 0.7 percentage points from the above differences.

The PPG numbers are statistically significant. The SHG numbers aren't, but they go in the right direction and are about the right magnitude, so I think it's reasonable to consider them as decent estimates.

------

So, as I promised in the second paragraph: the results of the regression seem to match what we found in the previous posts.

------

Technical notes 2:

For full disclosure, here are the coefficients for all the variables in the regression. I'll present them in "here's how to calculate the percentage of home-team penalties" format. (If you prefer a table, the full computer output is here (pdf). You'll be able to tell what the variables represent by matching the coefficients to what's below.)

Start with 0.6119 (constant).

Add -0.2619 if the home team took the last penalty.

Add -8.15E-06 for each second that's passed in the game. (About -.01 per period.)
Add -7.39E-06 for each second that's passed since the last penalty (not significant, p=.104, but magnitude is reasonable and has the right sign).

Add -0.1483 if the home team is down by 3 or more goals.
Add +0.1436 if the home team is down by 3+ and also took the last penalty.

Add -0.1196 if the home team is down by exactly 2 goals.
Add +0.0983 if the home team is down 2 and also took the last penalty.

Add -0.0807 if the home team is down by 1 goal.
Add +0.0545 if the home team is down 1 and also took the last penalty.

Add +0.0309 if the home team is up by 1 goal.
Add +0.0450 if the home team is up 1 and also took the last penalty.

Add +0.0668 if the home team is up by 2 goals.
Add +0.0591 if the home team is up 2 and also took the last penalty.

Add +0.0559 if the home team is up by 3 or more goals.
Add +0.1228 if the home team is up 3+ and also took the last penalty.

Add +0.0115 if a PP goal was scored on the last penalty
Add -0.0438 if a PP goal was scored on the last penalty and that penalty was to the home team.

Add -0.0072 if a SH goal was scored on the last penalty
Add +0.0350 if a SH goal was scored on the last penalty and that penalty was to the home team.


-----

Labels: , , ,

Tuesday, January 03, 2012

Do NHL referees call "make up" penalties? Part II

Last post, I found that referees are likely to "even up" their penalty calls: they're around 50% more likely to give the other team the next power play than to give the same team two power plays in a row.

I wasn't not convinced this is because of referee bias, or what Tango calls the "compassionate referee."

Tango suggested this experiment: check to see if a power play goal was scored on the first penalty. If the referee is indeed "compassionate" towards the other team, he should be more compassionate if the penalty actually cost them a goal, less so if there was no goal, and even less so if the penalized team *benefited* from the penalty by scoring shorthanded.

So I checked. I looked at all cases where there was a power play goal (PPG) on a first penalty, and then no more scoring until the next penalty was called. Indeed, that does appear to make the ref more compassionate.

After a penalty resulting a PPG, the next penalty was of the "even-up" variety 65.9% of the time. That's higher than the overall rate of 59.7%. Repeating that in a better font:

65.9% after a PPG
59.7% overall rate

And, the same effect appears for shorthanded goals (SHG):

52.5% after an SHG
59.7% overall rate

It's a large effect, and exactly in the direction Tango predicted.

-----

But wait! It might not be referee bias at all. Because, it turns out that teams with a lead take significantly more penalties than teams who are behind. For instance, when a penalty is called while you have a two goal lead, there's a 55.2% chance the penalty goes against you (and so a 44.8% chance the penalty goes against the other team). Full chart:

55.2% of penalties to team leading by 1
58.2% of penalties to team leading by 2
59.0% of penalties to team leading by 3
59.4% of penalties to team leading by 4
59.7% of penalties to team leading by 5

So, the score effect could explain what we're seeing. After a power play goal, the team has a bigger lead (or smaller deficit) than before. That would make it likely to take more penalties in future, even if the referee wasn't compassionate at all.

(Of course, the score effect might itself be due to referee "compassion," but that's a whole other argument.)

Specifically: a power play goal makes the team 6 percentage points more likely to take the next penalty. But scoring ANY tiebreaking goal in the first period makes a team 5 percentage points likely to take the next penalty. So how can we be sure there's a separate power-play effect, or how big it is?

-----

What might also complicate things is there's a "time of game" effect:

42,721 PPs came in the first period.
38.060 PPs in the second period.
26,705 PPs came in the third period.

There are fewer penalties in the third period than in the first. Is that a separate period effect? It might be.

Here's the score effect chart, again, but this time only for first-period penalties. The effect is more extreme than for the entire game:

55.6% of penalties to team leading by 1
60.1% of penalties to team leading by 2
61.0% of penalties to team leading by 3
66.5% of penalties to team leading by 4
58.7% of penalties to team leading by 5 (only 46 datapoints)

-----

It almost looks like we need a regression to sort all this out. But, wait! One more try before we turn to the dark side. Let's engineer a comparison where score and period won't screw things up.

I took every situation where:

1. It was the first period.

2. The game was tied at the time of the first penalty, and exactly one additional goal was scored before the second penalty.

3. The one extra goal was scored by the team that had the power play on the first penalty.

Then, I divided those situations into two groups.

The "Highest compassion" group is where the team scored the goal *on the power play*, presumably making the referee feel extra bad that he caused the goal. The "Typical compassion," is where the team scored the goal *after* the power play, and the referee's call wasn't the cause.

What percentage of the second penalties went to the other team?

Highest compassion: 71.6% (2163 datapoints)
Average compassion: 69.7% (1051 datapoints).

There's a small effect there, in the expected direction, of 1.9 percentage points. (That's less than 1 SD, so not statistically signficant.)

Here's the same result, but the other way, where it's the originally-penalized team that scored before the next penalty. When that goal was scored shorthanded, we can call that "Lowest compassion". When it wasn't, it's again "Average compassion."

Again, what percentage of the time did the second penalty even things out?

Lowest compassion: 62.5% (253 datapoints)
Average compassion: 58.2% (1006 datapoints).

This time the effect goes the "wrong" way, but there's too little data to draw any conclusions.

Doing the same thing for the second period instead of the first, we find a larger difference, but still not statistically significant (1.4 SD):

Highest compassion: 70.4% (568 datapoints)
Average compassion: 65.7% (271 datapoints).

And the shorthanded case, which really has too small a sample to take seriously:

Lowest compassion: 51.8 (83 datapoints)
Average compassion: 48.9% (268 datapoints).

-----

So, in summary: yes, there appears to be weak evidence for a small "compassion effect."

In the previous post, I considered three hypotheses:

1. Referee bias
2. Penalized teams play more carefully after the penalty
3. Power play teams play more aggressively after the penalty

Here's a fourth one, a variation of one suggested by commenter Wexler in the previous post:

4. Referees like to let the players play, and dislike calling penalties. But, sometimes they have to assert themselves to make sure the game doesn't get out of hand. Sometimes they're a bit too late, and they have to call a penalty on something that wasn't a penalty two minutes ago. This sends a message to the players, "OK, enough."

That might be necessary, but is obviously unfair to the penalized team. And, so, the referees know they have to call a "make up" penalty on those particular calls. Both teams understand what's happening, and won't object to either that call or the subsequent call.


I don't know if #4 is plausible or not ... but one of my co-workers is a soccer referee, and it's consistent with what he says about having to keep the game under control before it's too late.

As usual, I await comments from readers who know more about this stuff than I do.


------

UPDATE: Part 3 is here.




Labels: , , ,

Saturday, December 31, 2011

Do NHL referees call "make up" penalties?

Among NHL fans, there's a perception that referees like to call "make up" penalties. If a ref has just called a minor penalty on one team, it's very likely that the next penalty will go to the other team.

I was skeptical, until I downloaded a bunch of data from The Hockey Summary Project ... they're like Retrosheet for hockey. (Their website is here, and if you want data downloads, you can join their group by going here.)

I looked at all penalties from 1953-54 to 1984-85 (for which the HSP data is almost complete). I eliminated all cases where there both teams got penalties at the same time. Then, I checked what was left, to see if the team that got the current penalty was less likely to get the next one.

Absolutely, very much so. There's a 60% chance the next penalty will go to the other team -- 59.7%, to be more exact. (But, since I'm not sure that database is complete, and I forgot to remove misconducts, and I didn't consider situations where both teams got a penalty but one team got an extra one, I'm happier to drop the decimal and just go with 60%.)

The effect is reasonably consistent over time, although it was a little stronger back in the six-team era. Here's a too-long chart.

1953-54: 62.7 888/1416
1954-55: 61.3 857/1397
1955-56: 60.6 912/1506
1956-57: 61.2 833/1360
1957-58: 58.8 793/1348
1958-59: 61.4 801/1305
1959-60: 62.3 723/1160
1960-61: 61.7 740/1199
1961-62: 62.0 821/1324
1962-63: 62.0 797/1286
1963-64: 61.8 826/1337
1964-65: 61.7 841/1362
1965-66: 58.6 820/1399
1966-67: 58.6 710/1212
1967-68: 60.0 1515/2527
1968-69: 59.5 1666/2800
1969-70: 60.5 1793/2962
1970-71: 60.4 1944/3220
1971-72: 57.8 1918/3317
1972-73: 60.6 2200/3633
1973-74: 58.8 2135/3628
1974-75: 56.7 2873/5069
1975-76: 56.3 2890/5130
1976-77: 57.2 2371/4144
1977-78: 59.1 2316/3916
1978-79: 59.5 2337/3925
1979-80: 59.3 3021/5091
1980-81: 60.1 3780/6293
1981-82: 60.7 3613/5957
1982-83: 60.3 3479/5774
1983-84: 60.2 3788/6296
1984-85: 60.6 3542/5847
-------------------------
Overall: 59.7 39597/58543


Even though the effect is real, we can't say for sure that it's referee bias. It could just be that, after a penalty, the penalized team plays more cautiously, trying to avoid a second penalty. Or, it could be that the just had the power play decides to play more aggressively.

(As an aside: why did penalties drop so much between 1975-76 and 1976-77? At first I thought it might be bad data, but then I checked power-play opportunities on Hockey Reference, and it checked out.)

Here's what I think is some relevant evidence. I broke down the stats by referee (minimum 300 datapoints). The database only has the referee named for about a quarter of the total games (mostly older ones), but I figure it's probably good enough to at least look at.

The first column is the main number, the percentage of penalties called against the team who drew the last one.

Pctg Z-sc Size Ref
---- ---- ---- ----------------------------
59.3 00.0 0509 Andy Van Hellemond
60.4 +0.4 0846 Art Skov
59.3 -1.2 1128 Ashley
60.0 00.0 0460 Bill Friday
57.0 -1.1 0537 Bob Myers
58.9 -0.3 0878 Bruce Hood
57.1 -0.9 0580 Bryan Lewis
59.3 -1.6 1453 Buffey
64.6 +1.6 0933 Chadwick
59.4 -0.1 0567 Dave Newell
60.4 -0.3 0356 Farelli
60.1 -0.3 0511 Friday
59.9 -0.1 0696 John Ashley
61.3 +0.8 0359 Lloyd Gilmour
61.3 -0.2 0789 Macarthur
56.7 -1.4 0319 Mehlenbecher
63.0 +0.5 0327 Olinski
60.4 -0.8 1906 Powers
55.2 -2.2 0698 Ron Wicks
63.7 +1.7 1095 Skov
57.7 -3.1 2197 Storey
63.1 +2.7 4717 Udvari
59.9 +0.4 0709 Wally Harris

The least "biased" referee is 55%, and the most "biased" is 64%. If you think it's only referee bias that keeps the numbers from being 50%, you'd have to think that EVERY referee is biased almost exactly the same way. It's hard for me to accept that none of the referees noticed the bias and saw fit to try to eliminate it.

The second column of the table is the Z-score, the number of standard deviations the referee is from expected (which is normalized to the seasons he officiated). Normally, you concentrate on those with at least plus or minus 2 SD. That gives you Red Storey and Ron Wicks (less biased than most) and Frank Udvari (more biased than most).

The standard deviation of the Z-scores was 1.29. If every referee were the same, and differences were only random, it would be 1.00. This suggests that there are real differences between referees. Specifically, the SD of referee tendencies (or "talent", you might say) is 0.8 (since 1 squared plus 0.8 squared equals 1.29 squared).

In English, you can perhaps interpret that as saying that the differences in the table are about half real and half random, with a little more random than real (since 1.00 is a little higher than 0.8).

The observed range is 55 to 64. Regressing to the mean, the actual range of referee tendencies is probably 57 to 62, or something like that.

So if you think it's referee bias, you have to explain why all the referees seem to be biased within such a tight range, especially, when, presumably, they are all working hard to be as unbiased as possible.

--------

Here's another interesting breakdown, by time since the previous penalty:

0:01 to 1:00: 69.1 (7000)
1:01 to 2:00: 64.7 (9444)
2:01 to 3:00: 68.5 (11778)
3:01 to 4:00: 64.2 (10574)
4:01 to 5:00: 61.2 (8831)
5:01 to 6:00: 59.7 (7470)
6:01 to 7:00: 58.9 (6328)
7:01 to 8:00: 58.3 (5333)
8:01 to 9:00: 56.6 (4399)
9:01 to 10:00: 55.5 (3719)
10:01 to 11:00: 56.7 (3000)
11:01 to 12:00: 55.5 (2591)
12:01 to 13:00: 53.8 (2193)
13:01 to 14:00: 55.2 (1837)
14:01 to 15:00: 53.3 (1565)
15:01 to 16:00: 53.1 (1376)
16:01 to 17:00: 53.1 (1135)
17:01 to 18:00: 52.4 (1019)
18:01 to 19:00: 51.9 (807)
19:01 to 20:00: 53.6 (757)
20:01 to 99:99: 51.8 (3883)

The longer the interval since the previous penalty, the less likely the next penalty will go to the other team. That's consistent with many theories. The "referees are biased" theory would say that referees "forget" to even things up as the game goes on. The "other team wants revenge and plays aggressively" theory would say that if they don't get revenge early, they don't need it as much later. And the "penalized team takes fewer chances" theory would say that as time goes on, the players "forget" that they have to be more careful.

So, the data doesn't help us choose, but it's interesting nonetheless.

By the way, the 1:01 to 2:00 group is an exception to the pattern, but that's probably due to power plays, since the first penalty is probably still in effect. Actually, I'd have expected that part to go the other way, with the first two minutes being *more* than 50 percent, on the logic that the shorthanded team playing in the defensive zone is more likely to be forced to take a penalty. But, that doesn't happen.

And here's an interesting breakdown of the first half of the first group:

81.9% within 5 seconds
78.1% between 6 and 10 seconds
76.0% between 11 and 15 seconds
73.8% between 16 and 20 seconds
69.3% between 21 and 25 seconds
67.0% between 26 and 30 seconds.

-----

Finally, one more question: after one team gets, say, four straight penalties, what happens then? Is there an even stronger bias for the other team to take the next penalty?

Yup.

57.1 after exactly 1 in a row (64858 datapoints)
64.0 after exactly 2 in a row (23850)
66.6 after exactly 3 in a row (7042)
66.0 after exactly 4 in a row (1781)
63.8 after exactly 5 in a row (442)
60.6 after exactly 6 in a row (127)
67.5 after 7 or more in a row (40)

------

So: what's going on? Any ideas?



UPDATE: Part 2 is here. Part 3 is here.




Labels: , , ,

Monday, December 29, 2008

How much is an NHL power play worth?

Here's a piece by hockey analyst Alan Ryder, who tallies up some statistics on NHL power plays.

Ryder looks at teams' conversion rates with the man advantage, and finds that a full two-minute power play results in about .27 of a goal. That works out to 7.9 goals per 60 minutes, "the usual expression of scoring rates" when analyzing hockey. That 7.9 figure is "more than 300% the rate of scoring during even-handed play."

(I wish Ryder had also told us the even-handed rate directly.)

With a two-man advantage, there's less data, so Ryder had to smooth out the curve. His chart shows a 50% conversion rate after 80 seconds of 5-on-3 play (compared to about 17% for 4-on-3). So the second penalty, on a second-by-second basis, is worth twice as much as the first penalty. Of course, the two penalties are seldom simultaneous. If the second infraction comes a minute after the first infraction, it would be worth one-and-a-half times as much; that would be one minute of 4-on-3 followed by another minute of 5-on-3.

Ryder estimates that a two-man advantage results in about 25 goals per sixty minutes. Yesterday, in the World Juniors, Canada beat Kazakhstan 15-0. That's approximately the equivalent of playing with a two-man advantage for two periods.

If a (single man) power play is worth .27 goals, and that's three times the normal rate, then the one-man advantage is worth about .18 goals. However, add in the fact that the other team probably won't score, and you have to add back in the .09 defensive savings. That means the PP is indeed worth the full .27. At least approximately – this doesn't take into account that the average power play doesn't go the full two minutes, so maybe the .27 should really be only .25 or something. And there are shorthanded goals, so maybe it's now .24. But these are just estimates.

Back 20 years ago in the days of the Hockey Compendium, I think Klein and Reif adjusted players' points for power play goals against while they were in the box. It might once again be interesting to figure out how to do that once again. I'd use *expected* PP goals against rather than actual, and I'm not sure how I'd make the adjustment. But I'm curious as to whether players are costing their teams a lot of goals this way.

Good stuff as usual from Mr. Ryder.



Labels: , , , ,