Wednesday, February 05, 2014

Ratings vs. measurements

Sabermetrics may have lot of fancy new methods, but there's no stat that rates who the best players are.  

That's not just because here's no sabermetric "holy grail" of the ultimate statistic.  The real reason is that such a thing isn't even possible.  

You can talk about which player is "better" than another, but the problem is that there's no objective definition of "better".  You can't use statistics to help measure something when you can't even define what it is you're measuring.  

What sabermetrics CAN do is provide statistics that bear on the question of "best".  It can provide objective data that can inform your arguments.  But those arguments, like your definition of "best," are always going to be subjective.

The very first chapter of the 1982 Bill James Baseball Abstract was about comparing hitters.  Who was better?  Johnny Pesky, who had little power but hit for a high average and lots of walks?  Or Dick Stuart, who didn't get on base a whole lot, but regularly hit 30 home runs?

Bill was able to discover that there is a mathematical connection between a batting line and the number of runs that score.  Using that "Runs Created" formula, Bill found that Pesky's three best years were almost perfectly identical to Stuart's three best years.  They created almost the exact same number of runs, in almost exactly the same opportunities (outs).

Doesn't that settle the question?  Doesn't that prove that Pesky and Stuart are equally as good as each other?

No, it doesn't.  "Equal performances" doesn't necessarily imply "equally good hitters."  

Part of the problem, as Bill mentioned in his essay, is that there are a lot of things that Runs Created doesn't consider: park effects, differing quality of opposing pitchers, etc.   We haven't corrected for those.  My experience has been that when people assume that there will eventually be that perfect statistic, this is what they're thinking of, that some day we'll have so much data that we can correct for everything, and be almost perfectly accurate in terms of run estimates.

But, that's not the real problem.  The real problem is that when you're rating players, you're not trying to figure out who created the most runs.  You're trying to figure out who is the *best*.

Who says the highest-rated batter should be the one who creates the most runs?  Sure, creating runs is a very big deal, and our ratings of "best" are tremendously better informed now that we have that information.  But, there are other factors to consider.

For instance: suppose it's obvious a player had a lucky "career year".  Wouldn't that influence your rating of who's better?  What if one of the players saw a platoon advantage much more than the other?  What if one player hit much better with runners on base, but the other hit much better when the score was close?

How do you deal with all those?  It's subjective, isn't it, how much weight you have to give each of those factors in a "best hitter" evaluation?  I don't see how there can possibly be a "right answer" when the question is so subjective and vague.

-----

There are two different ways someone could disagree with you about who's better.  They could disagree with your definition, or they could disagree with your measurement.

The measurement one is easy.  Suppose you think the better player is the one who creates the most runs.  

To measure that, you decide to use "Total Average," the Thomas Boswell stat that's basically bases divided by outs.  You'd calculate every player's stat, and point to the top ones, and say, "those are the best."  And a sabermetrician would come along and say, "well, that's not right.  You're counting a stolen base as equal to a single.  But the SB advances only one runner, while the single advances the batter and also every other baserunner."

And someone else would pull out some other evidence, like when Bill James showed that teams that steal lots of bases don't win as many games as teams who hit lots of singles.  And someone would come along with Pete Palmer's linear weights calculation, and show that a single is worth half a run, on average, while a steal is worth only a fifth of a run.

That's the easy part, critiquing and improving a measurement.  The hard part, and the subjective part, is the definition of "better".  

If you think better means "more runs per out," and I think better means "more runs above replacement," then how do we resolve that?  I guess, like any other debate on what's "better" -- like a political debate, arguing back and forth and appealing to principle.  

Does affirmative action make society better or worse?  Is legal abortion a good thing?  What's better, taking a strict interpretation of the first amendment, or protecting minorities from hate speech?  What's a better performance, a .300 hitter, or a .290 hitter who hits .325 in the clutch?

-----

Even in a simple, two-dimensional case, it can be all gut feel.

We all want the best players to be in the Hall of Fame, right?  Well, some players are among the best because they were very good for many, many years -- Phil Niekro, say.  And some players are among the best because they were superb, but for fewer years -- Sandy Koufax.

What's the tradeoff?  What's the definition of "best" that can tell you whether Koufax or Niekro is "better"?

If you think about it, and look up the stats, your gut will likely come to some firm conclusion about which player is better.  But your intuitive feeling will be different from my intuitive feeling.  And, neither of us can articulate just what the tradeoff is.  Maybe we can say, "well, if Niekro had a few more good years, I'd prefer him."  Or, "I'd be more comfortable saying Koufax is better if he hadn't had those mediocre years at the beginning."  But, we won't be able to say, "I'm weighting a Cy Young Sandy Koufax year as 2.32 times as important as an average Jim Kaat year."

We have a strong intuition, but we can't explain it.

One of my favorite things that Bill James ever did was his "Hall of Fame Monitor" method.  He figured out a point system that scores each player on his Hall of Fame qualifications.  The way the system works, if you have 100 points or more, you'll probably be voted in, and if you don't, you won't.

And it works pretty well -- it does a good job of separating the players who have been enshrined from the ones who haven't.   (It's a prediction, not a recommendation -- it separates on what the voters *did*, not what the voters *should have* done.)

What Bill James did, amazingly, was reverse-engineer the collective brain of the writers, to figure out what their internal mental "formula" was.  

If you look at the details of the system, how players score points for arbitrary-looking achievements, you'd say, "well, that's certainly not how *my* intuitive decision system is figuring it out!"  But ... you know, it's probably reasonably close.  For most of us, we generally agree with the voters.  There are exceptions -- many sabermetricians object to some of the weights the voters apparently apply to various stats -- but, we fans are generally in line with the writers.  Especially on the "peak value vs. career value" question.  

-----

Another way you can see how ratings are subjective is ... the use of the word "rating".  We "rate" the players by who's best, but we don't "rate" the players by who's tallest.  We "measure" the tallest players, or "order" the tallest players, or "rank" the tallest players.  All those words imply some objective criterion, which we can reliably measure.  Runs Created, too.  In his 1982 Pesky/Stuart article, Bill James wrote,


"... runs created is not a rating.  It is an estimated record.  A rating is something which tells you how good; a record is something which tells you how many. ... a record gives you specific information that you could use to move toward those evaluations ... If I *rated* one player at 90 and another at 88, then I would be saying that the player who was at 90 was a better hitter than the one at 88.  But in fact it is entirely possible -- indeed, commonplace -- that that a player who created 88 runs for his team could be considered a better hitter than a player who created 90 ... "

The same thing is true for "grades", like grades in school.  "Jimmy got a grade of 75 percent in math" makes sense.  "Jimmy got a grade of 95 pounds in weight" does not.  

Consumer Reports "assigned" the Tesla a "grade" of 99.  But they did not "assign" the Honda Civic a "grade" of 29 miles per gallon.  

-----

Let me quote Bill James one last time.  In the 1985 Abstract, Bill tried to figure out which great teams were the best of all-time.  He ran a bunch of objective criteria, then added them up to get an answer.  But then he said:


"I offer no proof of that; it is only a carefully worked-out opinion, which is very different from something that can be shown to be true."

Perfect.  If a ranking is not a measurement, then it's an opinion ... no matter how carefully you try to work it out.



Labels: ,

Sunday, April 21, 2013

Pythagorean good luck associated with Runs Created bad luck


I noticed recently that there's a negative correlation between certain measures that we think are random and independent.  For instance, outshooting Pythagoras tends to be associated with undershooting Runs Created.  I don't know why, and I'm looking for ideas.

----

Let me give you some background to what luck numbers I'm doing here.

Back in 2005, I did a study to estimate real teams' historical talent levels from their stats.  I figured that there were five mutually exclusive ways a team could perform differently from its talent:

1.  Its batters could have lucky or unlucky years, in terms of raw batting line.

2.  Its pitchers could have lucky or unlucky years, in terms of the opposition's raw batting line.

3.  It could create more or fewer runs than expected from its batting line (runs created).

4.  Its opponents could create more or fewer runs than expected from their batting line (runs created).

5.  It could over- or undershoot its Pythagorean projection.

The last three were easy -- I just compared them to their estimates.  The first two were harder.  How can you tell whether a player is having a career year?  What I did, for that, is I took the weighted average of the four surrounding seasons, and regressed that to the mean.  The results for players came out fairly reasonable.  

The results for teams came out reasonable too, IMO.  The luckiest team from 1960-2001 was the 2001 Mariners (who the study said "should have" won 89 games instead of 116), and the unluckiest was the 1962 Mets ("should have" won 61 instead of 40).  

[If you want more details, see my web page (search for "1994 Expos").  You can actually download the spreadsheet there that I'm using.  Also, I wrote up the findings for SABR's "Baseball Research Journal," and I found a repost here (.pdf).]

The "career year" estimates for teams seemed pretty good.  I had tweaked the formulas to make them close to unbiased.  For 1973 to 2001 -- the subset of seasons I'm using for this, less strike years -- the mean batting luck was +1.8 runs, and the mean pitching luck was -0.1 runs.  

So, I was pretty happy with the overall results.

-----

OK, so ... while I was working with the data yesterday, I noticed some correlations I didn't expect.  

First: it turns out there's a strong correlation between "Pythagoras luck" and "career year luck" (batting plus pitching).  That correlation is negative 0.1.  Why would that happen?

The only theory I can think of -- when a team plays well, it wins a lot of games.  That means it plays fewer ninth innings on offense, and more ninth innings on defense.  That artificially makes it look lucky in Pythagoras (which is based on run differential).  

But that should create a *positive* correlation with player performance luck, not negative!

Pythagoras luck had an SD of around 40 runs per season.  Career year luck was around 65.  So, every four extra Pythagoras wins is related to around negative 6 runs of "career year" effect.  Not a whole lot, but I still don't know what's going on.  

----

And, worse: there's a strong correlation between "Pythagoras luck" and "Runs Created luck".  This time, negative 0.15.  

So: for every win by which a team beats its Pythagoras, it's given up one-tenth of a win in Runs Created luck.  How would that happen?  The only thing I can think of is walkoff wins with runners on base: every one of those might lead RC to believe you were unlucky by ... what, half a run?  So that's not really enough.

-----

Finally ... there's a huge correlation (minus 0.2) between "career year luck" and "pythagoras + RC luck".  For every four wins a team gained due to Pythagoras/RC luck, they lost one back to player underperformance.  

For that, I have a hypothesis.  Runs Created is known for overestimating the best offenses.  So, when a team beats its RC estimate, it's less likely to be having a great year.  That means its batters are more likely to be underperforming.  

Here's something to support that idea: when I checked, I found almost all the correlation comes from comparing batting career years with batting RC luck, and from comparing pitching career years with pitching RC luck.  Comparing pitching to batting gives almost zero correlations.

I'm not sure if that explanation is enough to explain the -0.2, but it's something.

-----

So what's going on?  Shouldn't clutch hitting (which is what RC luck is) be uncorrelated with, say, scoring runs when you need them the most (which is what Pythagoras luck is)?  Shouldn't whether you get a few extra hits one season (which is career year luck) be uncorrelated with *when* those hits happen (which is RC luck)?

Why are these things associated?  It must be something about the way I'm measuring them, as opposed to, being lucky one way causes you to be lucky another way.  Right?


Any ideas?




Labels: , , ,

Wednesday, April 22, 2009

Has Runs Created stopped working?

Does Runs Created not work any more?

The reason I ask is that, if you take a look at the 2008 AL and NL pages on Baseball Reference, you'll see that RC overestimated actual runs for 28 of the 30 teams. The average discrepancy was a huge +58 runs in the NL, and +19 runs in the NL.

To emphasize: that's not the average after removing the signs, that's the average *including* the signs. If half the teams had been +58 and the other half had been –58, the average would have been zero. It wasn't.

So what I'm saying is, Runs Created now appears to be biased too high.

This has been happening since the mid 90s. Here is the average team discrepancy by season:

1985 -4
1986 –1
1987 +2
1988 –5
1989 –5
1990 +4
1991 –7
1992 +7
1993 +0
1994 +8
1995 +7
1996 +7
1997 +19
1998 +15
1999 +19
2000 +19
2001 +18
2002 +19
2003 +19
2004 +27
2005 +25
2006 +27
2007 +24
2008 +26

Now, we know that Runs Created is biased too high for higher run environments, so that might be part of it. But it's not all of it. In the three seasons 1994 to 1996, there were 4.92, 4.84, and 5.03 runs per game respectively. But in 2005 there were only 4.59 runs per game, and in 2008, only 4.65 runs per game.

Could it be that the pattern of offensive events is different? Maybe there are different patterns of offensive events than there used to be (maybe more walks per single, or something?), and Runs Created doesn't work well when that happens?

By the way, I tried Base Runs, using the first version found on page 18 here (.pdf) with X=.535; the results weren't as extreme, but they were similar.

Anyone know what's going on? Is this a well-known problem and I just missed it?

P.S. For the record, I think I'm using the "technical" version of Runs Created found on this Wikipedia page.

Labels: ,

Monday, April 20, 2009

A Diamond Mind simulation as baseball strategy research

A science column from Alan Schwarz a couple of weeks ago investigates the effects of various baseball strategies, using a simulation.

To check out batting orders, Schwarz got Luke Kraemer at Diamond Mind to simulate two sets of 100 seasons of the 2008 Yankees. In one set, A-Rod batted fourth; in the other set, he batted ninth. The difference was 42 runs; the regular Yanks scored 789 runs, while the A-Rod-at-the-bottom-of-the-order Yanks scored only 747.

Schwarz doesn't tell us how he checked intentional walks, but finds that they are a bad strategy, costing five runs per season. That's not a very useful result; there are times when the IBB makes more sense, and times when it makes less sense. Which did Diamond Mind simulate?

Stolen Bases: Diamond Mind took the 2008 Rays and the 2008 A's, and reversed their respective propensities to steal ("switched their mind-sets," is what the article says). The A's dropped by 20 runs, but the Rays *improved* by 47 runs, "suggesting that perhaps the Rays were running too often in real life."

As it turns out, the real Tampa Bay team stole 142 bases and were caught only 50 times, for a 74% success rate; that should put them well in the black, compared to the rule of thumb that you need to be successful 67% of the time to break even. So I'm at a loss to explain the 47 run difference.

The only thing I can think of is a sample size issue. I think the SD of a team's runs scored in a single game is about 3. So the SD of a season's worth of runs is 3 times the square root of 162, or about 38 runs. The SD of the average of 100 season's worth is one-tenth of that, or about 4 runs. The difference between two 100-season averages is the square root of 2 times that, or about 5.4 runs.

But 47 runs is almost 9 standard deviations. So I'm still not sure what's going on.

Finally, the sacrifice bunt. When the simulation forced the bunt-avoiding Red Sox (27 SH in 2008, compared to the league-average 34) to do it more often, they lost 19 runs. But when they got the bunt-loving Mets (73, league average 66) to do it less, the result was also a loss – 15 runs. Schwarz concludes that the Mets' real-life bunting was better than the Red Sox, that they chose to bunt in more favorable situations. But, weren't both these numbers based on the simulation? If so, the real-life situations should make no difference.

If the comparisons, however, *were* based on real life, then we have sample size issues based on the real-life sample, which is only 162 games, with an SD of about 38 runs. Maybe the 2008 Mets and Red Sox scored more or fewer than the simulation because of luck? We should be able to tell by looking at Runs Created – but, for some reason, almost all teams undershot their RC estimate in 2008 (and their Base Runs estimate too, at least for the versions I tried).

Anyway, while I like the simulation method, I wish the results had been presented more clearly. As it stands, I'll stick to "The Book"'s conclusions on these issues of baseball strategy.

P.S. Here's what Tony LaRussa thinks of these results:

“There’s way too much importance given to what you can produce from a machine,” he said. “These are human beings, and I don’t think any computer is going to model that close to what we deal with at this level.”

Hat Tip: Daniel Hamermesh at Freakonomics


Labels: , , ,