Tuesday, June 04, 2013

The OBP/SLG regression puzzle -- Part V

(Links: Part I, Part II, Part III, Part IV)

When you run a regression to predict a team's runs per game based on OPS and SLG, you get that a point of OPS is about 1.7 times as important as a point of SLG.  When you predict *opposition* runs per game, you get 1.8.  But, when you try to predict run differential based on OPS differential and SLG differential, you get 2.1.

Why the difference?

It all hinges on the idea that the relationship between runs and OBP/SLG is non-linear.  

Suppose there are seasons where teams hit for a higher OBP and SLG are higher than normal -- steroid years, say.   And suppose those are the exact seasons where a single point of OBP or SLG is worth more in terms of runs.  That's not farfetched -- an offensive event seems like it would be worth more in years when there's more offense.  When there's lots of hitting, a walk has a better chance to score, and a double has more men on base to drive in.

So, it's a double whammy.  OBP/SLG are higher, and each point of OBP/SLG is worth more.  It's almost like a "squared" relationship, which is non-linear.  

-----

A good analogy would be something like, tickets sold vs. ticket revenue.  On one level, it's linear, because for each extra ticket you sell, your revenue goes up $25 or whatever.  But, then, the second whammy: attendance is much higher now than it was in the sixties.  So, if you sell a lot of tickets, it's more likely you're in 2011 than 1966, which means that, along with your higher attendance, you also have higher-priced tickets!  So, more tickets means more revenue because of more sales, but also more revenue because of more revenue *per ticket*.

Now, what happens when you switch to *differences*, so you're predicting "revenue over opposition" based on "tickets over opposition"?

Well, that depends.  The original source of the "double whammy" was that teams with high tickets sold were more likely to have higher ticket prices.  Is that still true for teams with high *differences* of tickets sold?

Maybe, or maybe not.  Suppose you have two teams, team A from 2002 that averaged 40,000 tickets, and team B from 1964 that averaged 10,000 tickets.  That's a ratio of 4:1, four times as many tickets sold when per-point OBP/SLG values were high.

After you take the difference, does the 4:1 ratio change?  If it stays at 4:1, nothing happens.  If it moves higher -- maybe team A outdrew its opposition by 10,000, but B outdrew its opposition by only 1,000, for a ratio of 10:1 -- the relationship becomes even more non-linear.  If it moves lower -- team A outdraws by 2,000, team B outdraws by 1,000, for a 2:1 ratio -- the relationship becomes *less* non-linear.

(Actually, I'm not sure if I should be dividing (to get the ratio, which I did) or subtracting (to get the raw difference) or something else.  But this is just an intuitive explanation anyway.)

Since there is no reason to expect that the differences will have the exact same ratio as the original attendance figures, we are almost *assured* that the nature of the relationship will change.  And, since, as a general rule, the coefficient increases with non-linearity, we expect the coefficients to change.

So it was a mistake, on my part, to originally assume that the "difference" regression should have the same ratio as the "single-team" regression.

-----

BTW, that was the "non-baseball" explanation that I promised you.  If tickets is still too basebally, just substitute any non-baseball relationship where a high X is correlated with a high value per X.  What works well is a time series featuring some commodity that sells more now, when per-unit prices are obviously higher because of inflation.  

Like, say, Starbucks coffee, or bicycle helmets.  If you want a "decreasing returns" example, I bet cigarettes is a perfect one -- lower cigarettes sold is strongly associated with higher prices per cigarette.

------

Now, to the actual baseball data.  

For the MLB teams and their oppositions in the dataset, we need to know: when we look at *differences* in OBP and SLG instead of the actual values, do high differences still correlate with times when individual points are more valuable?  That is: do the ratios get wider, or narrower?  It's hard to know, exactly, but we can get a rough idea, by looking at the spread of the data before and after.  

Let's start with OBP.

For the seasons in the study, the range of team OBP was .273 to .372, which is 99 points.  For opposition OBP, the range was 103 points.  In terms of standard deviations, it was .015 for the offenses, and .016 for the opposition offenses.

But, for the differences, the spread was wider.  The range was 123 points, and the SD was .019.

Teams:       spread  99 points, SD .015
Oppositions: spread 103 points, SD .016
Differences: spread 123 points, SD .019

That makes sense ... when you subtract opposition hitting, it's like adding team pitching.  The spread should be wider because if you take a team with awesome hitting, and they also have awesome pitching, they stick out from the average twice as much. 

In general, when you subtract one variable from another, if they're independent (or have a negative correlation), the SD and spread increase.  If hitting and opposition hitting were, in fact, independent, you'd expect the SD to increase by a factor of root 2 -- from .015 to .021.  It only increased to .019, because hitting and opposition hitting are, in fact, somewhat correlated.  They both are affected by whether you play in hitters' parks or pitchers' parks.  They also depend on what era you play in.  

Still, the difference in OBP seems to have increased the spread.  We don't know for sure, though, that the new difference numbers are still correlated with high values for each point OBP, but it seems reasonable to expect.  

-------

What about SLG?  

Well, the range for teams was .301 to .491 (190 points).  The range for opposition teams was .306 to .499 (193 points).  

But -- surprisingly -- the range for (team minus opposition) was almost the same.  It was 194 points.

And the SD of the differences was *smaller* than the SDs of the originals.  The teams were .034, the opposition was .033, but the SD of the differences was only .031.

Teams:       spread 190 points, SD .034
Opposition:  spread 193 points, SD .033
Differences: spread 194 points, SD .031

Why is the SD actually *lower* when you combine the two variables?  

Well, it's the same argument as for OBP: they're correlated with each other.  But, it seems, SLG correlates much more highly than OBP did.  Probably, park and era effects are slugging related more than on-base related -- after all, parks are better known for home runs than for walks, say.  

I wasn't expecting those effects to be so huge ... this seems to suggest that environment may be *more important* than team talent, at least for raw SLG!  

So, it would be reasonable to assume that, when we subtract opposition runs from team runs, the "non-linearness" of SLG to runs stays about the same, or decreases a bit.  

-------

What have we found?  We found that when we use "team minus opposition," we're increasing the non-linearity of OBP, but *decreasing* the non-linearity of SLG.  So, we'd expect the OBP coefficient to increase, and the SLG coefficient to decrease.  And that's what happens:

Teams:       16.62 OBP, 9.98 SLG   [ratio: 1.67]
Opposition:  17.38 OBP, 9.38 SLG   [ratio: 1.85]
Differences: 18.73 OBP, 8.97 SLG   [ratio: 2.09]

And that explains why the ratio for differences is higher than the ratio for individual teams.

By the way: in the above table, every higher coefficient in a column is associated with a higher SD of that variable, and every lower coefficient in a column is associated with a lower SD of that variable.  That doesn't have to be the case, but it's more likely that way than any other way, I would think.


Labels: , , ,

Sunday, June 02, 2013

The OBP/SLG regression puzzle -- Part IV

(Here's part 1, part 2, and part 3.)

---

A couple of posts ago, Alex and others suggested that I try a regression to predict runs per game (RPG) instead of winning percentage.  Maybe *that* regression would come out to the OBP/SLG ratio of 1.7 that we've been expecting.  

Jared Cross actually did that, in another comment, and it worked!  Here's my version of Jared's result:


RPG = (16.6*OBP) + (10.0*SLG) - 4.9  [ratio: 1.67]

And the same regression, but for opposition runs per game:

RPG = (17.4*OBP) + (9.4*SLG) - 4.9  [ratio: 1.85]

The two ratios, 1.67 and 1.85, are almost perfectly in line with the expected 1.7!

------

So, my first reaction was ... I wasted my time!  All those worries about walks and non-linearity and increasing returns ... well, they weren't necessary.  The issue was just that I used a different variable, winning percentage instead of runs!

But ... actually, on reflection, I don't think that's it.  I think it's just a bit of a coincidence that these regressions work out to 1.7.  Let me give you a couple of intuitive arguments that may or may not convince you.  

First argument: I redid the first regression above, the 1.67 one, but three times, with the dataset split based on walk tendencies (BB/(H+BB)).  "High" and "low" mean one percentage point above or below average; "medium" is everyone else.  The ratios:


High walks:   1.70
Medium walks: 1.25
Low walks:    1.89

It's strange: the teams with average numbers of walks had a much lower ratio than the high-walk and low-walk teams.  But the original 1.7 ratio, the one that Tango got, was based on a perfectly average team.  

So: if this is the answer, that this regression is the right one ... shouldn't Tango's result have come in at 1.25 instead of 1.7?  

Second argument: I repeated the regression, but this time I combined the team with its opponents.  That is, I predicted (RPG minus opposition RPG) from (OPS minus opposition OPS) and (SLG minus opposition SLG).

Since we're just subtracting the two equations, you'd expect that the ratio would be somewhere in between 1.67 and 1.85.  But, no:


RPG = (18.73*OBP) + (8.97*SLG)   [ratio: 2.09]

The ratio goes up to 2.09.

Why should that persuade you that the original 1.67 is coincidence?  This is a red herring: Tango's analysis didn't include opposition.

But ... it did, in a way.  Tango's logic and numbers work exactly the same way if you include opposition.

Instead of asking, "what happens if we add an event to an average batting line," you can ask, "what happens if we add an event to the zero batting line that's the difference of two average teams."  The calculation is exactly the same either way.  

Specifically, if adding a point of OBP to an average team gives 1.7 times as many additional runs as a point of SLG, then adding a point of OBP to an average team *with a given opponent* should add 1.7 times as many additional runs *over that opponent* as adding a point of SLG.  Right?

But ... here, the results are different.  We we added in the opponent, we got 2.09 instead of 1.67.  And the difference, I will argue next post, is real, not just a random artifact.  Actually, I don't even think the difference has anything to do with baseball.

-------

Talking baseball for now, though ... why are these results so much different from the winning percentage case?  Especially the walk breakdown.  With winning percentage, it seemed like walks increased the ratio, but, when it comes to runs, it seems like sometimes walks increase the ratio and sometimes they decrease it!

Well, it's going to sound like I'm just making this up, but here's my latest theory, for what it's worth:

The linear weights values of the various offensive events are based on an average team.  In real life, their actual values are different for good teams and bad teams, but we just assume the differences don't matter.  And they wouldn't, if they all changed roughly equally, because then the regression would adjust.

But, maybe they don't.

I ran a regression to predict the linear weights values based on runs.  Then, I looked at only the best offenses (based roughly on a 1.7 OPS stat), and the worst.  (I know it's not very accurate to use regression for this, but I was too lazy to do play-by-play data, and I only need roughly correct values anyway.)


event     1B   2B   HR    BB   out
----------------------------------
average  .53  .71  1.45  .34  -.10
----------------------------------
high     .58  .69  1.50  .34  -.13
low      .49  .72  1.47  .34  -.08

Almost all the difference is in singles and outs!  

In our sample, every team had roughly the same number of outs, so the out value doesn't matter that much.  But, not every team has the same number of singles, relative to the other events.  

For a given (1.7 * OPS + SLG), the team with the higher slugging will probably have more singles.  Therefore, they will score more runs than expected from their 1.7-weighted OPS.  Therefore, the regression will attribute that to the SLG, and weight it higher, reducing the ratio.

This is the reverse of what I thought happened in the "winning percentage" case, which is why you should assume I don't know what I'm talking about.  But, if you're still with me ... well, if I'm going to change my mind, I should probably come up with an explanation of why the cases are different. 

Here's an attempt.  I'll put it in block quotes to emphasize that it's a guess and I'm just throwing it out there ...


The singles thing was happening in the winning percentage case, too, obviously.  However, it was overshadowed by another factor.
 
When a team has a high SLG, much of that is caused by the park.  In that case, the opposition will also have a higher SLG.  So, much of the high SLG doesn't translate into a higher winning percentage (although it *does* translate into more runs).  

On the other hand, if a team has a high OBP, more of that is "real", since park factors don't affect walks and singles as much as, say, doubles and home runs.  

So that's why walks mattered more when we were looking at winning percentage.  When it comes to runs, walks are taken at face value.  But for winning percentage, we have to give them extra weight, because the relationship between slugging and winning percentage is too tainted by park. 

Is that right?  Who knows.  It sounds plausible.  But, so did everything else I argued earlier.  I'm probably full of crap.  Take this part with a grain of salt.

-------

For the record, though: when we combine a team and its opposition, we get a ratio of 2.33 for winning percentage, and 2.09 for run differential.  Those aren't too far off from each other, and still higher than 1.7.

-------

So, the question remains: if I've convinced you that the 1.7 is coincidence, and the 2.09 matters just as much ... then, why do we still have that difference?  I'm going to back off from specific theories, and stick to generalities.

The relationship between OBP/SLG and runs/wins is non-linear in many ways.  One way is that a point of OBP/SLG has different numbers of events depending on how good the offense is.  Another is that a point of OBP has different proportions of walks/hits depending on slugging.  A third is that the different basic events have different run values depending on offense.  And there are probably a lot more.  

So, the answer to the question is: because there is so much non-linearness going on, there's no reason to expect that the coefficient of the average team will equal the coefficient from a regression of all teams.  

-------

Next post, I'm going to make even a stronger argument, one that has nothing to do with baseball: I'm going to argue that the ratio we get from these regressions is almost useless anyway.  Tango's 1.7, for an average team, is meaningful -- but these other regression results that I've been doing, these 1.67s and 2.33s and 2.09s, don't tell us anything useful at all.

Here's that next post.

Labels: , , ,

Wednesday, May 29, 2013

The OBP/SLG regression puzzle -- Part III

I ran regressions in the previous posts to predict winning percentage from on-base percentage and slugging.  In those regressions, I had adjusted all teams to the league SLG and OBP.  I had to.  If you don't adjust, the results vary a lot.

Here's the regression completely unadjusted.  (It's all teams from 1961 to 2009, except strike seasons.)  Here's the equation.  (I'll put the OBP/SLG coefficient ratio in brackets too.)


wpct = (2.19*OBP) + (0.07*SLG) - .2405  [ratio: 30]

That's an OBP/SLG ratio of over 30!  We were expecting 1.7.  It seems like slugging barely matters at all!

Compare that to the "regular" regression, which adjusts for league-season: 


wpct = (2.70*OBP) + (0.89*SLG) - .7843   [ratio: 3]

OK, that's a bit better.  The ratio is down to 3.

Guy argued, in the comments to the first post, that I need to adjust for park, too.  He's right.

If I change winning percentage to what it would be if the team had posted those stats in a neutral park -- while still keeping the league adjustment -- I get this equation: 


wpct = (2.65*OBP) + (1.09*SLG) - .8504  [ratio: 2.43]

Even better: we're down to 2.43!

An easier way might be just to not adjust anything, but include the league and park in the regression:


wpct = (2.63*OBP) + (1.15*SLG) - (2.58*league OBP) - (1.16*league SLG) - (0.0029* BPF) + 0.091  [ratio: 2.1]

Now, the ratio is all the way down to 2.1.

What's going on?

This one's pretty simple.  When a team has a high OBP or SLG, it's a combination of two things:

-- batting talent, and
-- a high run environment for the league and park.

The first one actually has an impact on winning percentage.  The second one doesn't.  A high SLG doesn't help you if it's caused by the park, because the opposition benefits from it too.

The same is true for OBP.  But ... SLG should be affected *more*.  There are more high-HR and low-HR parks than there are, say, low-walk parks.  The "steroids era" was mostly home-run related.  

Comparing 1968 to 2001:


1968: OBP .299, SLG .340
2001: OBP .332, SLG .427

OBP increased 11 percent, but SLG increased 26 percent.

So, when you don't adjust, slugging doesn't matter as much, because it benefits the opposition too.  That makes OBP look more important, relatively speaking.

(All credit for this finding goes to Guy ... he actually explained all this to me in his comment.)

-----

As you'd expect, the problem goes away when you combine team offense with opposition offense in the same regression.  Even without adjusting, you don't have a big problem, because both teams are affected the same way.  

I used the *differences* between team OBP/SLG and opposition OBP/SLG, without any adjustument, and got


wpct = (2.09*OBP) + (0.897*SLG) + .5  [ratio: 2.33]

That's a ratio of 2.33.

-----

But why do we need to care about the opposition at all?  Commenter Alex suggested that if we try to predict "runs per game" instead of "winning percentage," we'll get even better results, because the opposition won't matter.  

I'm checking that out for a future post.

------


Update: that future post, part IV, is here.


Labels: , , ,

Thursday, May 23, 2013

The OBP/SLG regression puzzle

In the second "puzzle" last post, I noticed that, when you run a regression to predict winning percentage from on-base percentage and slugging percentage, you get that a point of OBP is worth between two and three times as much as a point of SLG.  That's different from the consensus value of 1.7 (which Tango derives here).  Why the difference?

When I wrote that post, I thought I knew the answer ... but I actually hadn't.  So I spent the last few days trying to figure it out.  I actually jumped around among a whole bunch of possibilities, and hit a lot of dead ends ... but I think I finally got somewhere.  Here's the current state of my thinking.  

As usual, I could be wrong.  I was wrong a few times in the course of working on all this ... 

------

I think the issue is one of non-linearity.  The regression assumes that runs/wins are linear in OBP and SLG, but they might not be.  In fact, in a bit, I'll show they're not.

Why does non-linearity matter?  Because, the "1.7" comes from adding events to an average team, and looking at the marginal impact.  If there's linearity, then we know that impact must be the same for all teams.  But, if there's not, that doesn't necessarily work.

To see why: consider a relationship that's actually cubic:


0, 1, 8, 27, 64, 125, 216

If you do a linear regression to predict x-cubed from x, for those seven values, you get

y = 34x - 39

The regression says that if you increase X by 1, Y increases by 34.

But ... that's not true for the *average* value of X.  The average value is 3.  The difference between 3.5 cubed and 2.5 cubed isn't 34 -- it's 27.25.  (To be more precise, we can take the first derivative of x-cubed at 3, and get 27.)

So the average coefficient is higher than the coefficient at the average.  

Sometimes the coefficient at the average will be "too high", and sometimes, like now, it will be "too low".  I'm not completely sure of the exact conditions for each.  The point is, though, that if you have non-linearity, the two coefficients will probably be different.

And that means if you have two variables, the ratio will be different.  Suppose you have Y = a^3 + b.  The regression will give you coefficients of 34 and 1.  But the values at the average will be 27 and 1.  So the ratio is 34 overall, but 27 at the average.

That might be what's happening here.  OPS and SLG are non-linear in separate ways, and that changes the ratio from 2.3 overall, to 1.7 at the average.


-------

OPS and SLG are indeed non-linear in a certain way.  In the way I'm going to show you, you don't need any baseball knowledge.  I'm going to show you that they're non-linear, not in terms of *runs*, but in terms of *raw events*.  

Suppose you have a batting line, and you want to add walks to raise the OBP by one point (.001).  How many walks do you have to add?  Well, it depends on your original OBP.

Suppose you're at .333 -- you have 333 "OBs" (walks or hits) in 1000 PA.  How do you get to .334?  You can't just add one walk, because that only brings you to .333666 (334/1001).  It turns out you have to add approximately 1.5015 walks.  That brings you to 334.5015 OBs in 1001.5015 PA, which brings you to .334.

But, now, suppose you're at .400, and you want to get to .401.  How many walks do you have to add now?  This time, it's 1.66945.  401.66945 divided by 1001.66945 equals .401.

(I did a little algebra to get the formula that, for every 1000 PA you start with, the number of additional walks you need is (1000 divided by (.999 minus OBP)).  That's where those two numbers came from.)

That is: points of OBPs give you increasing returns *in terms of number of walks*.  So the more OBPs you get, the more each additional one is worth.  Or, if you want to put it another way, the more OBPs you have, the harder it is to get another one, because you need more walks to get it.

Again, this is not a baseball observation.  The same thing applies, to, say, games of gin rummy.  If you're at .333 and you want to get to .334, you only need to win your next 1.5015 games.  If you're at .400 and you want to get to .401, though, you have to win your next 1.66945 games.  

-------

So: as we saw, a point OBP offers a higher return when OBP is already high.  That, by itself, is enough to make the coefficient of OBP different from the marginal value *for an average team*, which is where the 1.7 came from.  

But ... what about SLG?  If SLG also offers increasing returns, its coefficient will vary, too.  If it varies the same way, we should still get 1.7!  

Yes, indeed.  But, who knows if SLG *does* have increasing returns?  And who knows if it does, if it's exactly equal to the increasing returns of OBP?  That would be quite a coincidence, wouldn't it?

Since we have no reason to expect OBP and SLG to offer the exact same distortion caused by increasing returns ... we have no reason to expect the ratio OBP/SLG to be exactly 1.7.

This doesn't explain why it's at the level it's at -- "slightly higher than 1.7," we could call it.  From the logic we've seen so far, it could be anything: lower than 1.7, higher than 1.7, much different, a little different, whatever.

But: that's why, in theory, it won't be exactly 1.7.  If that's all you're looking for, an explanation of why it *could* be different, there it is.  I'm going to keep going, but it gets boring and technical and long for the next bit.

------

OK, so we talked about adding a point of OBP by adding walks.  Now, let's talk about adding a point of SLG by adding an extra base.

Adding extra bases doesn't change the denominator of SLG (which is at-bats).  So, if you want to add one point of SLG where there's 1,000 AB, you can just add one extra base.  Change a double to a triple, or something.

But: the denominator, the number of AB, is not the same for every team.  The more AB you have, the more valuable a point of SLG.  At 1,000 AB, you need only 1 extra base.  At 1,020 AB, you need 1.02 extra bases, which is 2% more valuable.

AB is hits plus outs.  In our regression, every team has roughly the same number of outs (since we did full seasons only), so the only difference is hits.  So, the more hits a team has, the more valuable a point of SLG from extra bases.  And hits correlates highly with OPS.

So: the more OPS a team has, the more valuable a point of SLG.  But ... well, it's a weak increase, compared to the OPS increase.  I'm almost willing to call this one linear.

------

What about adding a point of SLG by adding a single?  That's different, because a single affects both SLG and OBP.  So, we need to do this in two steps: we add enough singles to raise SLG by a point, and then subtract enough walks to lower OBP back to where it was before.

How many singles to we have to add to SLG?  That's the same formula as for how many walks we had to add to OBP.  For 1,000 AB, it's

1000 divided by (.999 minus SLG)

That increases OBP by that many "events", so we subtract that exact number of walks, and OBP is back to where it was before.  (Effectively, we've just converted walks to singles at the exact rate that OBP stays the same, but SLG goes up a point.)

The increase in runs is, therefore,

[1000 / (.999 - SLG)] * [value of single - value of walk]

We're assuming singles and walks have constant value -- .47 and .34, say -- so we get that adding one point of OBP adds

+.14 * [1000 / (.999 - SLG)] runs.

That's a higher number when SLG is higher, so we see that a point of SLG also has increasing returns.  (I'm not going to try to figure out by how much.)

------

The last case is adding a point of OBP by singles (and leaving SLG alone).

How many singles do we need to add?  Same formula: 


[1000 / (.999 - OBP)]

But, that will also increase SLG, so we have to subtract that enough "extra bases" from SLG to bring it back to where it was before.

Adding hits increased total bases by the same number as it increased AB.  But, to keep slugging the same when adding AB, we need to increase total bases by only SLG times the number of AB.  So, we need to subtract (1-SLG) total bases, for each single added.  

That is, we need to subtract


[1000 / (.999 - OBP)] * (1-SLG) bases for each hit.

Combining the the two steps, gives a batting line change of 

[1000 / (.999 - OBP)] cases of "add one single, and subtract (1-SLG) bases".

Assigning run values here -- say, .47 for a single, and .26 for a base -- gives a run increase of 

[1000 / (.999 - OBP)] * (.47 - .26 (1-SLG))

That gives increasing returns in OBP, and also increasing returns in SLG.  Again, I'm not going to try and quantify which is bigger.

------

Those are the only four cases I see of how to increase one of OBP and SLG at the margin.  (For extra-base hits, you just add the two cases -- add singles, and then add extra bases.  The math works out the same.)

That means, in terms of increasing returns, we have:


Increase SLG by bases -- roughly linear
Increase SLG by hits  -- increasing in SLG
Increase OBP by walks -- increasing in OBP
Increase OBP by hits  -- increasing in OBP and SLG

So, some ways are increasing in OBP, and some in SLG, and ... it looks like OBP and SLG are represented roughly equally.  It looks like we should expect a ratio that's not too far from 1.7.  It might not be *exactly* 1.7, but our gut says it should be not too different.  Which is about right -- it's in the 2s.

This is all theory.  Is there any evidence we can look at?

Well, it looks like teams with lots of walks should be different from teams with lots of hits.  The walking teams should see lots of increasing returns in OBP, and so a higher ratio.  And the hitting teams should see lots of increasing returns in SLG, and so a lower ratio.

So, I repeated the regression, but included only teams who were at least two percentage points higher than normal in their BB/H ratio.  This is the regression for those teams:


wpct = 2.69 OBP + 1.05 SLG - .845 (ratio: 2.5)

And for teams who walked two percentage points *less* than normal:

wpct = 1.62 OBP + 1.03 SLG - .491 (ratio: 1.6)

So, that seems to support the theory!  More walks = higher ratio, as hypothesized.

The results are similar if I use other point thresholds for higher/lower than average:


0 points: 2.6 low, 3.7 high
1 point : 2.0 low, 4.6 high
2 points: 1.6 low, 2.5 high
3 points: 0.8 low, 1.5 high
4 points: 7.1 low, 3.6 high

(The theory seems to fail in the extreme case ... but it's probably sample size.  If you up the SLG coefficient by 2 SDs, the ratio drops from 7.1 all the way to 1.6.)

Overall, I'd say, the test seems to support the theory.

-----

OK, now the bad news.  I don't think this is the real answer.  Yes, I think it's all correct, but I wonder if the effect is much too small to make such a big difference, from 1.7 to 2.3. 

Also, this occurred to me, another explanation that seems bigger: 

Walks get lumped in with singles in OBP.  Extra bases get lumped in with singles in SLG.  Which is worth more: a single, or the exact number of walks and extra bases that have the same impact on OBP and SLG?  Whichever is worth more, if the good teams get more of that one relative to the other, that will show up in a higher coefficient.  If the good teams get fewer of that one, the coefficient would be lower.  

This last explanation seems to me like the effect would be bigger.  Further research required, I guess.

-----

Part II is here.

Labels: , , ,