Monday, December 13, 2010

Replacement level talent vs. observations: a study

Recently, both JC Bradbury and King Kaufman expressed some skepticism about the concept of "replacement value". In theory, replacement value is the level of talent that can be obtained quickly at league minimum salary, like your best minor-leaguer, or the free agent who almost got an offer but didn't.

Generally, conventional sabermetric wisdom is that a replacement-level player is one who performs at a level 20 runs (or two wins) below average, if pro-rated to a full season. (Two wins below average is zero "wins above replacement," or "WAR".) By this standard, no team should play anyone expected to perform below this level.

In their arguments, Bradbury points to the fact that there were, in fact, many players who had performed at less than this level of performance. In a recent post on replacement value, King Kaufman checked, and found that

"In the major leagues in 2010, 24.5 percent of all innings were thrown by pitchers who ended the season with a negative WAR."

(UPDATE: In the original version of this post, I had originally incorrectly painted Kaufman as a replacement-value skeptic, which he is not. See his comment below.)


The explanation, of course, is that the fact that they *performed* below replacement doesn't mean teams *expected* them to perform below replacement. Teams might have overestimated their abilities, or, more likely, they just had bad years due to random chance. If you're bet on heads ten times, some coins are going to land heads only 4 times out of 10. It doesn't mean that there was anything wrong with your prior expectations of the fairness of the coin.

Anyway, I thought I'd run a little experiment.

I started with every batter in Jeff Sackmann's "Marcel" database from 2000-2009. ("Marcel" is a prediction method created by Tom Tango, which forecasts a player's performance this year based on his statistics the previous three years.) Assuming the Marcel predictions are reasonable, I counted how many player-seasons were expected to be below replacement level. If the theory is correct, it should be "zero."

It wasn't zero, but it was close. Over 10 years, only 152 players total were expected below replacement value that year. That's 15 players per year, spread among 30 teams. Half a player per team. That's not bad.

And, it's possible I had replacement value wrong. I didn't include fielding, just basic linear weights. And I used arbitrary position adjustments -- catchers and middle infielders had to be -40 runs per 600 PA to be replacement level, 1B and DH had to be 0, and everyone else had to be -20. It's very possible that, of the 152 players, many of them actually *weren't* below replacement level because of their defensive skills.

Also, there were playing-time limits. Jeff's database excludes players with low expected playing time (the smallest he forecasts is for 185 PA). And I left out all players who had fewer than 50 actual plate appearances that season. I figured that if a guy projected below replacement, but the team only gave him 10 AB, that's close enough to zero that we won't count it.

So, half a player per team per season seems reasonably consistent with the concept of replacement level. It's not like teams are signing these guys left and right.

If there were 15 players per season *projected* to be below replacement level, how many of them actually *performed* below replacement level?

The answer: 1,025 total, or 102 per season. There were about seven times as many players *observed* to be below replacement level than *predicted* to be below replacement level.

That makes sense -- if you flip 100 pennies, where replacement level is 0.5, there would be *infinitely* more coins observed below 0.5 than actually below 0.5. (Assuming all pennies are fair coins, that is.)

-----

Now, the experiment. For every player in the database, I decided to randomly simulate their season based on their Marcels. Basically, I treated their Marcel prediction like an APBA card, and ran off a bunch of plate appearances. I simulated the exact number of PA that they *actually* had that season, regardless of how Marcel predicted their playing time.

Then, to simulate the uncertainty about the player's talent, I chose an adjustment from a normal curve, with standard deviation of +5 runs, and added that to the performance. (UPDATE: the +5 was for 500 Marcel PA. I adjusted accordingly for fewer Marcel PA by the square root of the ratio, so for 125 PA, I used +10.)

If Marcels are good, unbiased predictors, and teams were indeed getting rid of players who fell below replacement, then we should see 1,025 below-replacement performances in the simulation, not just in real life.

Well, we don't.

I ran the simulation 10 times, and the average was 561 players, not 1,025. We got a little over half.


Why? After I ran this, I realized the reason is selective sampling. Suppose you have two players who have talent of -10. Six weeks into the season, and just by luck, one of them is awful, at a rate of -30, and the other one is doing OK, at a rate of +10.

What happens? The -30 guy is released, and winds up the season at -30 over 100 AB. The second guy is allowed to play the whole year, and winds up at -5 over 500 AB.

One out of two wound up having performed below replacement in real life. But, in the simulation, it'll be less than that.

In the simulation, there's less than a 50% chance that the first guy will wind up at less than -20 over 100 AB. And there's a much, much smaller than 50% chance that the -10 guy will wind up below -20 over a full 500 AB.


So the simulation will underestimate the number of below-replacement performances, because, in real life, once a marginal player is below replacement, he's not often given a chance to rise back out of it. But in the simulation, he gets his full number of PA regardless.

-----

in that light, I adjusted the simulation to add one new rule: if a player was expected to be +10 or less, and, a third of the way through his expected season, he's below replacement, he gets released. (If, after a third of the season, he's above replacement -- even a little bit -- he plays the entire rest of the season regardless of what happens afterwards.)

Now, the simulation goes from 561 below-replacement performances, to 800. Still less than 1,025, but better.

So, finally, I did one more thing: I changed the standard deviation of the uncertainty of the player's talent from 5 runs to 10.

Now, we get to 863. That's 84 percent of the way there.

-----

After all that, I'm not sure quite how much the simulation tells us. To do a proper comparison, we need a better model of how teams decide how much playing time to give a hitter based on expectations and performance.

What we *do* find out, though is:

-- If you trust Marcel, then it does seem that few teams are willing to keep a player who has performed below replacement.

-- Regardless, many players *do* perform below replacement.

-- Simple probability shows that, at a bare minimum, over half the players who perform below replacement do so because of luck.

-- With other not-too-unreasonable assumptions, we can get that percentage up into the 80s.

My view about all this that it's less than fully conclusive. Still, it should be fairly persuasive. If you didn't accept the "replacement player" hypothesis before, this little study should have enough in it to get you to reconsider.

What do you think?

-----

UPDATE, 12/14: King Kaufman posts in the comments that I misinterpreted his views on replacement value. My apologies to King, and I've revised the post accordingly.



Labels: ,

Wednesday, June 18, 2008

Replacement players, VORP, salaries, and MRP

In a post on his blog, J.C. Bradbury argues, again, that a player's free-agent player salary is equal to his "MRP," which means "marginal value of production." That is, a player should be paid exactly the amount by which his performance increases the team's revenue.

And he seems to think that revenue is exactly proportional to the player's performance, rather than the player's performance as measured against replacement level. Because of that, he dismisses the concept of replacement value (and
VORP). But I don't understand why he would do that.

The idea behind MRP is this: the more employees you hire, the less each one contributes to the bottom line. If you're running a Wal-Mart, you might want ten cashiers. If you hire an eleventh cashier, it might help a little bit: if there's a crowd of customers, fewer might leave the store if the lineups are shorter. But the eleventh is only useful in busy times, so he's worth less than the other 10.

The idea is this: suppose cashiers earn $30,000 a year. The first five or six might bring the company $70,000 in revenue each. The seventh might bring in only $60K. The eighth adds $50K, the ninth $40K, the tenth $30K, and the eleventh $20K. The eleventh cashier is actually losing the company money, so she never gets hired in the first place. And the last cashier hired brought in $30,000, which exactly matches his salary. Thus the equivalence: salary = MRP.

That works for Wal-Mart, but not for baseball. Why? Because in baseball, the number of employees is fixed, and so is the minimum salary. At Wal-Mart, if you have 11 cashiers, you figure that the 11th costs $30,000 but is bringing in only $20,000 in revenue. So you fire him. In baseball, you might figure that you're paying the 25th man $390,000, but he's contributing nothing to the bottom line (because he gets no playing time). So you want to release him. But you can't – there's a rule requiring you to have 25 men on your roster. And there's a minimum salary of $390,000. So you're stuck. In this case, the MRP of the 25th man is less than his salary.

It can work the other way around, too. Suppose you're a big-market team with lots of fans who love to win, and you figure that a 25th man, while costing only $390,000, is bringing in revenues of over a million. You'd like to hire a 26th player, who would bring in another $900,000 or so. But, unlike Wal-Mart, you can't go hiring that extra player. There are rules against that. In that case, the 25th man is earning less than his MRP.

So, in baseball, a player's salary could easily be more, or less, than his MRP.

The real-world equivalence between salary and MRP is

Salary = MRP

But that's a special case that just happens to apply to Wal-Mart. I would argue that the more general equivalence is

Salary over and above the alternative = MRP over and above the alternative

At Wal-Mart, the alternative is "nobody" – you just never hire the 11th cashier. That alternative has zero salary and zero MRP, so the second equation collapses into the first equation. But in baseball, the alternative is NEVER "nobody" – you have to fill the roster spot, whether you want to or not. The alternative is a player at minimum salary, creating a replacement-level MRP. It's one of the many freely available minor-leaguers. That means

Salary over and above the $390,000 minimum = MRP over replacement player

If you choose to define "VORP" in terms of dollars instead of runs, you get

Salary - $390,000 = VORP

Which, I think, is what's really happening in baseball.


J.C. doesn't agree with that formulation – he wants to stick with "salary = MRP". He wants to value marginal runs from zero, rather than from replacement value. But, I argue, that clearly leads to untenable conclusions.

For instance, suppose a marginal win is worth $5 million. Then a marginal run is worth about $500,000.

Suppose a replacement-level player creates 39 runs. At $50,000 per run, you'd expect him to cost $19.5 million. But you can pick up any one of these guys for $390,000! So "salary = MRP" just doesn't make sense.

------

If you don't buy that argument, here's another one. Suppose you really, truly believe that a player earns his MRP. And suppose the 25th guy on your roster earns $390,000, the MLB minimum.

Now, halfway through the season, the union and MLB agree to double the minimum salary. The 25th guy gets to keep his job – after all, the team has to have 25 guys, and this is still the best one available. But now he's making $780,000.

His salary doubled, but his MRP, obviously, is exactly the same! So even if his salary was equal to MRP before, it certainly isn't now. Which means that there's no reason to have expected them to be equal in the first place.

------

One of J.C.'s arguments is that not all replacement-level players are worth only $390,000. It could be that all the players eligible for the minimum are young draft choices, and you don't want to use up their "slave" years if your season is a lost cause – you'd rather save them for when your team is a contender. In that case, you might have to sign a replacement-level veteran for $1,000,000 or so.

To which I say: there is no shortage of mediocre veterans who can be had for $390K, that you would have to spend a million. If you DO spend a million, it's probably because you peg the veteran as better than replacement. The extra $610,000 is worth it if you expect about 12 1.2 runs better than replacement over a full season, which isn't a lot.

------

By the way, you could argue that a player is never paid less than his MRP. In a sense, even if the 25th guy on your roster never bats, his presence contributes more than $390,000. That's because if you released him, and didn't call anyone up, the commissioner would fine you a lot more than $390,000.

But that's a trick technicality, and it's not what J.C. is arguing here.

------

Finally, I am puzzled by J.C.'s dislike of VORP because, according to him, it's an insider term and hard to explain:

The big advantage of these is that I can have these conversations with people other than die-hard stat-heads ... I view VORP as an insider language, and by using it you can signal that you are insider. It’s like speaking Klingon at a Star Trek convention. I can signal to others who speak the language that I am one of you. But, the danger of VORP is that once you bring it up the discussion goes down the wrong path as the uninitiated have reason to feel they are being told they are not as smart as the person making the argument. It’s like constantly bringing up the fact that you only listen to NPR or watch the BBC news at dinner parties. The response is likely going to be the same, “well fuck you too, you pretentious asshole!”



But, as I think a commenter on J.C's site points out, "MRP" is also pretty jargon-y insider economist talk, isn't it? And it's a lot hard to explain than VORP. So I'm a bit confused by J.C.'s aversion to the term. How come sabermetric abbreviations are pompous, but economics abbreviations are not?

------


See also Tom Tango's comments,
here.

Labels: , , ,