Wednesday, October 29, 2014

Do baseball salaries have "precious little" to do with ability?

Could MLB player salaries be almost completely unrelated to performance? 

That's the claim of social-science researcher Mike Cassidy, in a recent post at the online magazine "US News and World Report."

It argues an "economics lesson from America's favorite pastime." Specifically: How can it be true that high salaries are earned by merit in America, when it's not even the case in baseball -- one of the few fields in which we have an objective record of employee performance?

The problem is, though, that baseball players *are* paid according to ability. The author's own data shows that, despite his claims to the contrary. 

------

Cassidy starts by charting the 20 highest-paid players in baseball last year, from Alex Rodriguez ($29 million) to Ryan Howard ($20 million). He notes that only two of the twenty ranked in the top 35 in Wins Above Replacement (WAR). The players in his list average only 2.2 WAR. That's not exceptional: it's "about what you need to be an everyday starter."  

It sounds indeed like those players were overpaid. But it's not quite so conclusive as it seems.

WAR is a measure of bulk contribution, not a rate stat. So it depends heavily on playing time. A player who misses most of the season will have a WAR near zero. 

In 2013, Mark Teixeira played only 15 games with a wrist injury before undergoing surgery and losing the rest of the season. He hit only .151 in those games, which would explain his negative (-0.2) WAR. However, it's only 53 AB -- even if Texeira had hit .351, his WAR would still have been close to zero. 

A-Rod missed most of the year with hip problems. Roy Halladay pitched only 62 innings as he struggled with shoulder and back problems, and retired at the end of the season.

If we take out those three, the remaining 17 players average out to around 2.6 WAR, at an average salary of $22 million. It works out to about $8.4 million per win. That's still expensive -- well above the presumed willingness-to-pay of $5 to $6 million per expected win.

If we *don't* take out those three, it's about $10 million per win. Even more expensive, but hardly suggestive of a wide disconnect between pay and performance. At best, it suggests that the one year's group of highest-paid players performed worse than anticipated, but still better than their lower-paid peers.

Furthermore: as the author acknowledges, many of these players have back-loaded contracts, where they are "underpaid" relative to their expected year's talent earlier in the contract, and "overpaid" relative to their expected year's talent later in the contract. 

Even a contract at a constant salary is back-loaded in terms of talent, since older players tend to decline in value as they age. I'm sure the Yankees didn't expect Alex Rodriguez to perform at age 37 nearly as well as he did at 33, even though his salary was comparable ($28MM to $33MM).

All things considered, the top-20 data is very good evidence of a strong link between pay and performance in baseball. Not as strong as I would have expected, but still pretty strong.

-------

As further evidence that pay is divorced from performance, the author notes that, even limiting the analysis to players who have free-agent status, "performance explains just 13 percent of salary."  It's not just a one-year fluke. For each of the past 30 years, the r-squared has consistently hovered in a narrow band between 10 and 20 percent.

That sounds damning, but, as is often the case, it's based on a misinterpretation of what the r-squared means. 

Taking the square root of .13 gives a correlation of .36. That's not too bad: it means that 36 percent of a player's salary (above or below average) is reflected in (above- or  below-average) performance.

Still, you do have to regress salary almost 64 percent to the mean to get performance. Doesn't that show that almost two-thirds of a player's salary is unrelated to merit?

No. It shows most of a player's salary is unrelated to *performance,* not that it's unrelated to *merit*. Performance is based on merit, but with lots of randomness piled on top that tends to dilute the relationship.

You might underestimate the amount of randomness relative to talent, especially if you're still thinking of those top-20 players. But most players in MLB are not far from the league minimum, both in salary and talent.

According to the article, the 358 lowest-paid players in baseball in 2013 made an average $534,000 each. 

With a league minimum of $500,000, those 358 players must be clustered very tightly together in pay. And the range of their talent is probably also fairly narrow. But the range of their performance will be wide, since they'll vary in how much playing time they get, and whether they have a lucky or unlucky year. 

For those 358 players alone, the correlation between pay and performance is going to be very close to zero, even if pay and talent correlate perfectly. (Actually, the author's numbers are based on only players with 6+ seasons in MLB, so it's a smaller sample size than 358 -- but the logic is the same.)

When you add in the rest of the players, and the correlation rises to 0.36 ... that's pretty good evidence that there's a strong link between pay and performance overall. And when you take into account that there's also significant randomness in the performances of the highly-paid players, it must be that the link between pay and *merit* is even higher.

------

The author has demonstrated the "low r-squared" fallacy -- the idea that if the number looks low enough, the relationship must be weak enough to dismiss. As I have argued many times, that's not necessarily the case. Without context or argument, the "13 percent" figure could mean anything at all.

In fact, here's a situation where you have an r-squared much lower than .13, but a strong relationship between pay and performance.

Suppose that player salary were somehow exactly proportional to performance. That is, at the end of the season, the r-squared turned out to be 100 percent, instead of 13 percent. (Or some number close enough to 100 percent to satisfy the author.)

In baseball, as in life, people don't perform exactly the same every day. Some days, Mike Trout will be the highest-paid player in baseball, but he'll still wind up going 0-for-4 with three strikeouts.

So even if the correlation between season pay and season performance is 100% perfect, the correlation between *single game* pay and *single game* performance will be lower.

How much lower?  I ran a test with real data. I compiled batter stats for every game of the 2008 season, and ran a regression between the player's on-base percentage (OBP) for that single game, versus his OBP for the season. 

The correlation was .016. That's an r-squared of .000265.

The r-squared of .13 the article found between pay and performance is almost *five hundred times* as large as the one I found between pay and performance. 

Even though my r-squared is tiny, we can agree that Mike Trout is still paid on merit, right? It would be hard to argue that there was a fundamental inequity in MLB pay practices for April 11, just because Mike Trout didn't produce that day.

Well, I suppose, on a technicality, you could argue that pay isn't based on merit for a game, but *is* based on merit for a season. But if you make that argument for game vs. season, you can make the same argument for season vs. expectation, or season vs. career. 

The r-squared might be only 13 percent for a single season, but higher for groups of seasons. Furthermore, if you could play the same season a million times over, luck would even out, performance would converge on merit, and the r-squared would move much closer to 100%.

And the article provides evidence of that! When the author repeated his regression by using the average of three seasons instead of one, the r-squared doubled -- now explaining "just a quarter of pay." An r-squared of 0.25 is a correlation of 0.5 -- half of performance now reflected in salary.

Half is a lot, considering the amount of luck in batting records, and taking into account that luck is much more important than talent for the bunch of players clustered at the bottom of the salary scale. 

Again, the article's own evidence is enough to refute its argument.

-------

I think we can quantify the amount of luck in a batter's season WAR. 

A couple of years ago, I calculated the theoretical SD of a team's season Linear Weights batting line that's due to luck. It came out to 31.9 runs. 

Assuming a regular player gets one-ninth of a team's plate appearances, his own SD would be 1/3 the team's (the square root of 1/9). So, that's about 10.6 runs. Let's call it 10 runs, or 1.0 WAR. 

That one-win figure, though, counts only the kind of luck that results from over- or undershooting talent. It doesn't consider injuries, suspensions, or sudden unexpected changes in talent itself. 

Going back to the top 20 players in the chart ... we saw that three of those had injuries. Another three, it appears, had sudden drops in ability after they were signed (Vernon Wells, Tim Lincecum, and Barry Zito). 

Removing those six players from the list (which might be unfair selective sampling, but never mind for now), the remainder averaged 3.4 WAR. That's about $6.4 million per win -- very close to the consensus number. It would be even lower if we adjusted for back-loaded contracts.

At an SD of 1 WAR per player, the SD of the average of 14 players is 0.27 WAR. Actually, that's the minimum; it would be higher if any of the 14 were less than full-time. Also, the list includes starting pitchers -- I don't know if the luck SD of 1 win is reasonable for starters as well as batters, but I suspect it's close enough.

So, let's go with 0.27. We'll add and subtract 2 SD -- 0.54 --from the observed average of 3.4. That gives us a confidence interval of 2.9 to 3.9 WAR.

At 3.9 WAR, we get $5.6 million per win: almost exactly the amount sabermetricians (and probably front offices) have calculated based on the assumption that teams want to pay exactly what the talent is worth.

That is: it appears the results are not statistically significantly different from a pure "pay for performance" situation.

------

When the US News article talks about luck, it's different from the kind of luck I'm calculating here. The author isn't actually complaining that the overpaid players got unlucky and underperformed their pay. Instead, he believes that the highly-paid players were overpaid for their true ability, because they were "lucky" enough to fool everyone by having a career year at exactly the right time:


"In America, we tend to think of income as a reward for skill and hard work. ...

"But baseball shows us this view of the world is demonstrably flawed. 
Pay has preciously little to do with performance. Instead, being a top earner means having a good season immediately preceding free agency in a year where desperate, rich teams are willing to award outsized long-term contracts. ... 

"In other words, while ability and effort matter, it’s also about good luck."

Paraphrased, I think he's saying something like: "I've shown that pay is barely related to performance. Why, then, are some players paid huge sums of money, while others make the minimum?  It can't be merit. It must be that some players have a lucky year at a lucky time, and GMs don't realize the player doesn't deserve the money."

In other words: baseball executives are not capable of evaluating players well enough to realize that they're throwing away millions of dollars unnecessarily.   

The article gives no evidence to support that; and, furthermore, it appears that the author himself doesn't try, himself, to evaluate players and factor out luck. Otherwise, he wouldn't say this:


"But among average players, salaries vary enormously. For every Francisco Cervelli (Yankees catcher, $523,000 salary, 0.8 WAR), there is a CC Sabathia (Yankees pitcher, $24.7 million salary, 0.3 WAR). Both contribute about the same to the Yankees’ success (or lack thereof), but Sabathia earns roughly 50 times more."

Does he really believe that Sabathia and Cervelli should have been paid as equal talents?  Isn't it obvious that their 2013 records are similar only because of luck and circumstance?

Francisco Cervelli earned his +0.8 WAR in 61 plate appearances. That's about one-and-a-half SDs above +0.3, his then-career average per 61 PA.

Sabathia's salary took a jump after the 2009 season, at a time where he was averaging around 4 WAR per season. From 2010 to 2012, he actually improved that trend, creating +15.6 WAR total in those three years. It wasn't until 2013 that he suddenly lost effectiveness, dropping to 0.3 as reported. 

So it's not that Sabathia was just lucky to be in the right place at the right time. It's that he was an excellent player before and after signing his contract, but he suffered some kind of unexpected setback as he aged. (Too, his contract was structured to defer some of his peak years' value to his declining years.)

And it's not that Cervelli was unlucky to be in the wrong place at the wrong time, unable to find a "desperate" team otherwise willing to pay him $20 million. He's just a player recognized as not that much better than replacement, who had a good season in 2013 -- a "season" of 61 plate appearances where he was somewhat lucky.

-------

In his bio, the author is described as "a policy associate at The Century Foundation working on issues of income inequality." That's really what the article is getting at: income inequality. The argument that MLB pay is divorced from performance is there to support the broader argument that inequality of income is caused by highly-paid employees who don't deserve it.

Here's his argument summarized in his own words:


"The first thing to appreciate is just how unequal baseball is. During the 2013 season, the eight players in baseball's 'top 1 percent' took home $197 million, or $6 million more than the 358 lowest-paid players combined. The typical, or 'median,' major league player would need to play 20 seasons to earn as much as a top player makes in one. ...

"But ... pay has preciously little to do with performance. ...

"In other words, while ability and effort matter, it’s also about good luck. And if that’s true of a domain where every aspect of performance is meticulously measured, scrutinized and endlessly debated, how much more true is it of our society in general?

"We end up with CEOs that make 300 times the average worker and 45 million poor people in a country with $17 trillion in GDP. And we accept it as fair."

Paraphrasing again, the argument seems to be: "Salary inequality in baseball is bad because it's caused by teams rewarding ability that isn't really there. If baseball players were paid according to performance instead of circumstance, those disturbing levels of inequality would drop substantially, and the top 1% would no longer dominate."

It sounds reasonable, but it's totally backwards. If the correlation between pay and performance were higher, players' pay would become MORE unequal.

Suppose salaries were based directly on WAR. At the end of the season, the teams pay every free agent $6 million dollars for every win above zero, plus the $500,000 minimum. (That's roughly what they're paying now, on expectation. Since expected wins equal actual wins, that would keep the overall MLB free-agent payroll roughly the same.)

Well, if they did that the top salary would take a huge, huge jump.

Among the top 20 in the chart, the top two WAR figures are 7.5 (Miguel Cabrera) and 7.3 (Cliff Lee. 

Under the new salary scale, both players would get sharp increases. Cabrera would jump to $45 million, and Lee to $44 million. The highest salary in MLB would go to Carlos Gomez, whose 2013 season was worth 8.9 WAR (4.6 of that from defense). Under the new system, Gomez would earn some $53 million. 

Under pay-for-performance, it would take only around 4.8 WAR to earn more than the current real-life highest salary, A-Rod's $29.4 million. In 2013, that would have been accomplished by 32 players. 

Carlos Gomez's salary would exceed the real-life A-Rod by 82 percent. Meanwhile, replacement players would still be making the minimum $500K. And Barry Zito, with his negative 2.6 wins, would *owe* the Giants $15 million. 

Clearly, inequality would increase, not decrease, if the connection between pay and performance became stronger. 

Mathematically, that *has* to happen. When luck is involved, and applies equally to everyone, the set of outcomes always have a wider range than the set of talents. As usual,

var(outcomes) = var(talent) + var(luck)

Since var(luck) is always positive, outcomes always have a wider range than expectations based on talent. 

In fairness to the author, he doesn't think teams are paid by talent. As we saw, he believes teams pay by misinterpreting random circumstances, a "right place right time" or "team likes me" kind of good luck. 

If that's really happening, and you eliminate it by basing pay directly on measurable performance, then, yes, it's indeed possible for inequality to go down. If Francisco Cervelli were being paid $100 million per season, because he was Brian Cashman's favorite, then instituting straight pay-by-performance would lower the top salary from $100 million to $53 million, and inequality would decrease.

But, as we saw, that's not the case: the real-life top salaries are much lower than the "pay-by-performance" top salaries. That means that teams aren't systematically overpaying. Or, at least, that they're not overpaying by anything near as much as 82 percent.

-------

Imagine an alternate universe in which players have always been paid under the "new" system, $6 million per WAR. In that universe, as we have seen, the ratio between the top and median salaries is much higher than it is now, maybe 50 times instead of 20.

Then, someone comes along and presents a case for more equality:


"MLB salaries aren't as fair as they could be. They're based on outcomes, where they should be based on talent. Francisco Cervelli gets credit for 0.8 wins in 61 PA, even though we know he's not that good, and he just happened to guess right on a couple of pitches. 

"Players should be paid based on their established and expected performance, by selling their services to the highest bidder, before the season starts. That eliminates luck from the picture, and salaries will be based more on merit. The salary ratio will drop from 50 to 20, the range will compress, and the top players will earn only what they merit, not what they produce by luck."

Isn't THAT the situation that you'd expect someone to advocate if they were concerned about (a) rewarding on merit, (b) not rewarding on luck, and (c) reducing inequality of salaries?

Why, then, is this author advocating a move in the exact opposite direction?




(Hat tip: T.M.)


Labels: , , ,

Saturday, August 30, 2014

Is MLB team payroll less important than it used to be?

As of August 26, about 130 games into the 2014 MLB season, the correlation between team payroll and wins is very low. So low, in fact, that *alphabetical order* predicts the standings better than salaries!

Credit that discovery to Brian MacPherson, writing for the Providence Journal. MacPherson calculated the payroll correlation to be +0.20, and alphabetical correlation to be +0.24. 

When I tried it, I got .2277 vs. .2346 -- closer, but alphabetical still wins. (I might be using slightly different payroll numbers, I used winning percentage instead of raw win totals, and I may have done mine a day or two later.)

The alphabetical regression is cute, but it's the payroll one that raises the important questions. Why is it so low, at .20 or .23? When Berri/Schmidt/Brook did it in "The Wages of Wins," they got around .40.

It turns out that the season correlation has trended over time, and MacPherson draws a nice graph of that, for 2000-2014. (I'll steal it for this post, but link it to the original article.)  Payroll became more important in the middle of last decade, but then dropped quickly, so that 2012, 2013, and 2014 are the lowest of all 15 years in the chart:






What's going on? Why has the correlation dropped so much?

MacPherson argues it's because it's getting harder and harder to buy wins. There is an "inability of rich teams to leverage their financial resources."  The end of the steroids era  means there are fewer productive free-agent players in their 30s for teams to buy. And the pool of available signings is reduced even further, because smaller-market teams can better afford to hang on to their young stars.


"Having money to spend remains better than not having money to spend. That might not ever change. Unfortunately for the Red Sox and their brethren, however, it matters far less than it once did."


------

My thoughts:

1.  The observed 2014 correlation is artificially low, because it's taken after only about 130 games (late-August), instead of a full season. 

Between now and October, you'd expect the SD due to luck to drop by about 12 percent. So, instead of 2 parts salary to 8 parts luck (for the current correlation of .20), you'll have 2 parts salary to 7.2 parts luck. That will raise the correlation to about .22.

Well, maybe not quite. The non-salary part isn't all binomial luck; there's some other things there too, like the distribution of over- and underpriced talent. But I think .22 is still a reasonable projection.

It's a small thing, but it does explain a tenth of the discrepancy.

------

2.  The lower correlation doesn't necessarily mean that it's harder to buy wins. As MacPherson notes, It could just mean that teams are choosing not to do so. More specifically, that teams are closer in spending than they used to be, so payroll doesn't explain wins as well as it used to.

Here's an analogy I used before: in Rotisserie League Baseball, there is a $260 salary cap. If everyone spends between $255 and $260, the correlation between salary and performance will be almost zero -- the $5 isn't enough of a signal amidst the noise. But: if you let half the teams spend $520 instead, you're going to get a much higher correlation, because the high-spending half will do much, much better than the lower-spending half.

That could explain what's happening here.

In 2006, the SD of payroll was around 42% of the mean ($32MM, $78MM). In 2014, it was only 38% ($43MM, $115MM). It doesn't look that much different, but ... teams this year are 10 percent closer to each other than they were, that has to be contributing to the difference.

(This is the first time I've done something where "coefficient of variation" (the SD divided by the mean) helped me, here as a way to correct SDs for inflation.

Also, this is a rare (for me) case where the correlation (or r-squared) is actually more relevant than the coefficient of the regression equation. That's because we're debating how much salary explains what we've actually observed -- instead of the usual question of much salary leads to how many more wins.)


------

3.  While doing these calculations, I noticed something unusual. The 2014 standings are much tighter than normal. 

So far in 2014, the SD of team winning percentage is .058 (9.4 games per 162). In 2006, the SD was larger, at .075 (12.2 games per 162). That might be a bit high ... I think .068 (11 games per 162) is the recent historical average.

But even 9.4 compared to 11 is a big difference.  It's even more significant when you remember that the 2014 figure is based on only 130 games. (I'd bet the historical average for late-August would be between 12 and 13 games, not 11.)

What's going on? 

Well, it could be random luck. But, it could be real. It could be that team talent "inequality" has narrowed -- either because of the narrowing of team spending (which we noted), or because all the extra spending isn't buying much talent these days.

I think the surrounding evidence shows that it's more likely to be random luck. 

Last year, the SD of team winning percentage was at normal levels -- .074 (12.04 games per 162). It's virtually impossible for the true payroll/wins relationship to have changed so drastically in the off-season, considering the vast majority of payrolls and players stay the same from year to year.

Also, it turns out that even though the correlation between 2014 payroll and 2014 wins is low, the correlation between 2014 payroll and 2013 wins is higher. That is: this year's payroll predicts last year's wins (0.37) better than it predicts this year's wins (0.23)! 

Are there other explanations than 2014 being randomly weird? 

Maybe the low-payroll teams have young players who improved since last year, and the high-payroll teams have old players who declined. You could test that: you could check if payroll correlates better to last year's wins than this year's for all seasons, not just 2013-2014.

If that happened to be true, though, it would partially contradict MacPherson's hypothesis, wouldn't it? It would say that the money teams spend on contracts *do* buy wins as strongly as before, but those wins are front-loaded relative to payroll.

We can see how weird 2014 really is if we back out the luck variance to get an estimate of the talent variance.

After the first 130 games of 2014, the observed SD of winning percentage is .058. After 130 games, the theoretical SD of winning percentage due to luck is .044.

Since luck is generally independent of talent, we know

SD(observed)^2 - SD(luck)^2 = SD(talent)^2 

Plugging in the numbers: .058 squared minus .044 squared equals .038 squared. That gives us an estimate of SD(talent) of .038, or 6.12 games per 162.

I did the same calculation for 2013, and got 10.2.

2013: Talent SD of 10.2 games
2014: Talent SD of  6.1 games

That kind of drop in one off-season pretty much impossible, isn't it? 

If that kind huge a compression were real, it would have to be due to huge changes in the off-season -- specifically, a lot of good players retiring, or moving from good teams to bad teams.

But, the team correlation between 2013 wins and 2014 wins is +0.37. That's a bit lower than average, but not out of line (again, especially taking the short season into account). 

It would be very, very coincidental if the good teams got that much worse while the bad teams got that much better, but the *order* of the standings didn't change any more than normal.

So, I think a reasonable conclusion is that it's just random noise that compressed the standings. This year, for no reason, the the good teams have tended to be unlucky while the bad teams have tended to be lucky. And that narrowed the distance between the high-payroll teams and the low-payroll teams, which is part of the reason the payroll/wins correlation is so low. 

------

4. We can just look at the randomness directly, since the regression software gives us confidence intervals. 

Actually, it only gives an interval for the coefficient, but that's good enough. I added 2 SDs to the observed value, and then worked backwards to figure out what the correlation would be in that case. It came out to 0.60. 

That's huge!  The confidence interval actually encompasses every season on the graph, even though 2014 is the lowest of all of them.

To confirm the 0.60 number, I used this online calculator. If the true correlation for the 30 teams is 0.4, the 95% confidence interval goes up to 0.66, and down to 0.05. That's close to my calculation for the high end, and easily captures the observed value of 0.23 in its low end. 

That's not to say that I think they really ARE all the same, that the differences are just random -- I've never been a big fan of throwing away differences just because they don't meet significance thresholds. I'm just trying to show how easy it is that it *could be* random noise.

I can try to rephrase the confidence interval argument visually. Here's the actual plot for the 2014 teams:




The correlation coefficient is a rough visual measure of how closely the dots adhere to the green regression line. In this case, not that great; it's more a cloud than a line. That's why the correlation is only 0.23.

Now, take a look at the teams between $77 million and $113 million, the ones in the second rectangle from the left.

There are eighteen teams in that group bunched into that small horizontal space, a payroll range of only $46 million in spending. Even at the historically high correlations we saw last decade, and even if the entire difference was due to discretionary free-agent spending, the true talent difference in that range would be only about 3 or 4 games in the standings. That would be much smaller than the effects of random chance, which would be around 12 games between luckiest and unluckiest. 

What that means is:  no matter what happens, that second vertical block is dominated by randomness, and so the dots in that rectangle are pretty much assured of looking like a random cloud, centered around .500. (In fact, for this particular case, the correlation for that second block is almost perfectly random, at -.002.)

So those 18 teams don't help much. How much the overall curve looks like a straight line is going to depend almost completely on the remaining 12 points, the high-spending and low-spending teams. In our case, the two low-spending teams are somewhat worse than the cloud, and the ten high-spending teams are somewhat better than the cloud, so we get our positive correlation of +0.23. 

But, you can see, those two bad teams aren't *that* bad. In fact, the Marlins, despite the second-lowest payroll in MLB, are playing .496 ball.

What if we move the Marlins down to .400? If you imagine taking that one dot, and moving it close to the bottom of the graph, you'll immediately see that the dots would get a bit more linear. (The line would get steeper, too, but steepness represents the regression coefficient, not the correlation, so never mind.)  I made that one change, and the correlation went all the way up to 0.3. 

Let's now take the second-highest-payroll Yankees, and move them from their disappointing  .523 to match the highest-payroll Dodgers, at .564. Again, you can see the graph will get more linear. That brings the correlation up to 0.34 -- almost exactly the average season, after mentally adjusting it a bit higher for 162 games.

Of course, the Marlins *aren't* at .400, and the Yankees *aren't* at .564, so the lower correlation of 0.23 actually stands. But my point is not to argue that it should actually be higher -- my point is that it only takes a bit of randomness to do the trick. 

All I did was move the Marlins down by less than 2 SDs worth of luck, and the Yankees by less than 1 SD worth of luck. And that was enough to bump the correlation from historically low, to historically average.

------

5. Finally: suppose the change isn't just random luck, that there's actually something real going on. What could it be?

-- Maybe money doesn't matter as much any more because low-spending teams are getting more of their value from arbs and slaves. They could be doing that so well that the high-spending teams are forced to spend more on free agents just to catch up. It wouldn't be too hard to check that empirically, just by looking at rosters.

-- It could be that, as MacPherson believes, there are fewer productive free agents to be bought. You couuld check that easily, too: just count how many free agents there are on team rosters now, as compared to, say, 2005. If MacPherson is correct, that careers are ending after fewer years of free agency, that should show up pretty easily.

-- Maybe teams just aren't as smart as they used to be about paying for free agents. Maybe their talent evaluation isn't that great, and they're getting less value for their money. Again, you could check that, by looking at free-agent WAR, or expected WAR, and comparing it to contract value.

-- Maybe teams don't vary as much as they used to, in terms of how many free-agent wins they buy. I shouldn't say "maybe" -- as we saw, the SD of payroll, adjusted for inflation, is indeed lower in 2014 than it was in 2006, by about 10 percent. So that would almost certainly be part of the answer. 

-- More specifically: maybe the (otherwise) bad teams *more* likely to buy free agents than before, and the (otherwise) good teams are *less* likely to buy free agents than before. That actually should be expected, if teams are rational. With more teams qualifying for the post-season, there's less point making yourself into a 98-win team when a 93-win team will probably be good enough. And, even an average team has a shot at a wild card, if they get lucky, so why not spend a few bucks to raise your talent from 79 games to (say) 83 games, like maybe the Blue Jays did last year?

-----

I'll give you my gut feeling, but, first a disclaimer: I haven't really thought a whole lot about this, and some of these argument occurred to me as I wrote. So, keep in mind that I'm really just thinking out loud.

On that basis, my best guess is ... that most of the correlation drop is just random noise. 

I'd bet that money buys free agents just as reliably as always, and at the usual price. The correlation is down not because spending buys fewer wins, but because more equal spending makes it harder for the regression to notice the differences.

But I'm thinking that part of the drop might really be the changing patterns of team spending, as MacPherson described. I wonder if that knot of 18 mid-range teams, clustered in such a small payroll range, might be a permanent phenomenon, resulting from more small-market teams moving up the payroll chart after deciding their sweet spot should be a little more extravagant than in the past. 

Because, these days, it doesn't take much to almost guarantee a team a reasonable shot at a wildcard spot -- which means, meaningful games later in the season than before, which means more revenue. 

In fact, that's one area where it's not zero-sum among teams. If most of the fan fulfillment comes from being in the race and having hope, any team can enter the fray without detracting much from the others. What's more exciting for fans -- being four games out of a wildcard spot alone, or being four games out of a wildcard spot along with three other teams? It's probably about the same, right? 

Which makes me now think, the price of a free agent win could indeed change. By how much? It depends on how increased demand from the small market teams compares to decreased demand from the bigger-spending teams.

------

Anyway, bottom line: if I had to guess the reasons for the lower correlation:

-- 80% randomness
-- 20% spending patterns

But you can get better estimates with some research, by checking all those things I mentioned, and any others you might think of.





Hat Tip: Craig Calcaterra


Labels: , , ,

Tuesday, October 16, 2012

Can money buy meat?

If you want to have meat in your diet, you have to spend money in the grocery store.  At least, that's the conventional wisdom.

But is that really true, or is it just a myth?  Let's look at the evidence.

I took a (made up) random sample of 30 shoppers in my local supermarket earlier this year.  I ran a regression to predict the total amount of meat they had, from the total amount of money they spent.  It did turn out that the York family, the one who spent the most money by far, did get the most meat.  And, that there was a positive slope, meaning that spending more money leads to more meat.

However, there was one very important issue: the link between meat and money was not statistically significant.  In other words, we can't argue that money spent and meat obtained are actually related to each other in 2012.

It's easy to understand why we got this result.  Some of the lowest-spending families wound up with a lot of meat -- one was stocking up for a BBQ, and one owned a cattle ranch.  And a few rich-spending families barely had any meat in their houses at all -- they paid a lot for only a few ounces of filet mignon.

But 2012 isn't typical.  When I (pretended that I) did the same experiment for other years, I got statistically significant results.  But even for those years, explanatory power is quite low.  Only 17 percent of the variation in meat over the last 25 years is explained by variation in spending.  So much of the variation in meat obtained is not explained by how much money was spent.

And if you look at each year individually -- as the following table (with made-up numbers) illustrates -- the power of money buying meat seems to vary quite a bit:

2012: not significant
2011: r-squared = .17, p = .01
2010: r-squared = .13, p = .04
2009: r-squared = .21, p = .02
2008: r-squared = .10, p = .06
2007: r-squared = .25, p = .00
2006: r-squared = .29, p = .00
2005: r-squared = .24, p = .00
2004: r-squared = .29, p = .00
2003: r-squared = .18, p = .02
2002: r-squared = .20, p = .01
2001: r-squared = .10, p = .04
2000: r-squared = .10, p = .04
1999: r-squared = .50, p = .00
1998: r-squared = .47, p = .00
1997: r-squared = .22, p = .01
1996: r-squared = .34, p = .00
1995: not significant
1994: r-squared = .16, p = .07
1993: r-squared = .09, p = .09
1992: not significant
1991: not significant
1990: not significant
1989: not significant
1988: r-squared = .18, p = .00


From 1996 to 2001, supermarket spending and meat were statistically linked each and every year.  However, explanatory power varied.  If we look at shoppers before 1993, we see four years where where the relationship was again not significant.

So here is the big question: Why is the relationship not stronger?  One would think that as shoppers spend more, they would wind up with more meat.  But, often, that's not what we see in the data.

One issue is that you can get meat at other places than the supermarket -- butchers, say, or gifts, or the slaughter of animals you own yourself.  Another issue is that it's hard to predict what shoppers will buy any given week. 

But, does the result from 2012 show that spending and meat will not be statistically related in future?  We don't know.  But what we *do* know is that spending does not guarantee a shopper more meat.

That's the nature of shopping.  Sometimes you don't get enough meat, and, it seems, no amount of spending can change that reality.

--------

So: do you believe me?  Do you believe that how much meat you have in 2012 doesn't depend on how much money you spend?  I hope not.

What, specifically, is wrong with the logic?  Lots of things, many of which I've written about before.

------

1.  Even if you don't get a statistically-significant relationship between spending and meat, that does NOT mean that you "can't argue that money spent and meat obtained are actually related to each other".  Of course you can!  Lack of significance just means that, in one specific, narrow, sense, you don't have enough grounds to assert a relationship *on this evidence alone*. 

But, of course, there's LOTS of other evidence that meat and spending are related.  For one thing, there's a big sign in front of the steaks, that says, "$7.99 per pound."  For another thing, millions of people will tell you that they have successfully exchanged money for meat.

You can only argue that there's no relationship if you choose to ignore all those things.  Which, I hope, you wouldn't.

2.  The implicit assumption in the argument is that every year is different.  That is: money bought meat in 1993 and 1994, but not in 1992 or 1995.  Why would you assume that, that the nature of shopping changes so often and so much that you can buy meat in 1994, but not 1995?  If we're going to assert that, we need some kind of explanation of how that could be plausible.

3.  Also, if you're interested in statistical significance, shouldn't you care about checking if 1994 and 1995 are actually significantly different?  What do you do if there's not statistically significant difference between them, as there probably isn't?  How can you say money bought meat one year, but not the other, when the p value of the difference is very high? 

You have a contradiction:

1994 is significantly different from zero
1995 is NOT significantly different from zero
1994 is NOT significantly different from 1995.


Isn't it just as reasonable to say there's no difference, than to say that money bought meat in 1994 but not 1995?  Even if you're depending on statistical significance, you still have to make an argument.

4.  Why use "different from zero" as your significance criterion anyway?  In this particular situation, there is no real reason to think that zero is more likely than any other value -- and, in fact, there's very, very good reason to believe it's different, unless you have good reason to believe that big spenders don't buy more meat than the guy in the express lane with one item. 

In some cases, like whether prayer cures cancer, a default of zero makes sense.  But not here.  Saying, "we'll assume money can't buy meat until we see strong evidence otherwise" ... well, that's just privileging your hypothesis.

5.  If you get a value that's significant in the real world sense, but isn't statistically significant, you need more data.  You can say, "I don't have enough evidence."  You can say, "there isn't enough evidence HERE."  But you can't just assume that there's no relationship.  Otherwise, it would be easy to argue that smoking is harmless.  You just do a double-blind study that's really small.  And then you say, "even though 40 percent of the five smokers got lung cancer, and only 20 percent of the five non-smokers got cancer, we got an r-squared of only .1, and that's not statistically significant.  So, there's no evidence that smoking causes cancer."

Yes, the evidence of THAT study is weak.  But that's because the study is too small.  Twice the risk of cancer is plausible, and important, and you can't just dismiss it because you deliberately designed your study the way you did.  And there are lots and lots of other studies showing a link, and a biological mechanism by which it happens.

If you did that study, and you deliberately ignore all the other evidence, than it's fair to say that YOU can't conclude that smoking causes cancer.  But WE can certainly conclude it. 

Similarly, if all you know is that within the dataset of your 30 individuals, the correlation between meat and spending is low ... YOU can conclude you don't have evidence that meat can be bought.  But WE cannot, because WE have other evidence: we've been to a supermarket.  We know something about how the market for meat works.

6.  Even noting that the r-squareds jump around a bit -- and that the jumping around is statistically significant -- that doesn't necessarily mean that the relationship between money and meat has changed.  The r-squared depends not just on the relationship, but on the scattering of the values in the actual dataset. 

So an increasing r-squared could simply indicate a larger variation in overall spending.  Think about it ... if some families spend $1, and some spend $1000, it should be easier to notice the relationship between spending and meat, which means a large r-squared.  But if everyone spends exactly $100, it's going to be harder -- a lower r-squared -- even if money buys the same amount of meat as always.

So when you see a changing r-squared, you can't really be sure what's going on.  It would be better to look at the coefficient estimate of the regression equation, rather than the r-squared.

In fact, for any arbitrarily low r-squared, I can construct a dataset where the coefficient is as statistically significant as you like, and meat costs any amount per pound you like.  (I thought I wrote about this fact before, but I can't find it.)

7.  Even though an r-squared less than .10 may look small intuitively, it probably isn't.  A low-looking r-squared can be very important in real life.  You can't just say ".10 is small".  You have to *argue* that, in context, it's small.

If you did a regression of suicide vs. life expectancy, the r-squared would be at least as small as the ones here.  But suicide and life expectancy are most definitely linked. 

You have to interpret the r-squared for what it is.  It's not really an indicator of how easily money buys meat.  It's a measure of how well you can predict meat from money, *relative to all the other things* that help you predict meat*. 

If cancer kills a million people, and suicide kills 10, the r-squared between suicide and life expectancy will be low, because suicide is being compared to cancer.  That's true even though a single suicide has a bigger effect on life expectancy than a single case of cancer.

8.  You'll notice how large the relationship really is if you look at the r, instead of the r-squared.  The square root of .17 is .41.  That means that for every standard deviation difference in money spent, you get 41% of a standard deviation in meat obtained.  That's a pretty strong association: if you move two inches to the right on the supermarket-spending bell curve, you move 2/5 of an inch to the right on the meat curve.

9.  The r-squared doesn't really tell you whether meat CAN be increased by increasing supermarket spending.  It tells you how much meat WAS increased with supermarket spending.  Obviously, you'd expect an imperfect correlation.  People use money on all kinds of things -- TVs, cars, tofu, vegetables.  They get meat from sources -- their own animals, gifts, butchers -- other than supermarkets.  And, they buy different kinds and forms of meat, at various prices: steaks, hamburger, spam, TV dinners, dog food, and so on.

Given all that variation, *of course* you're going to find a less-than-100% correlation between supermarket spending and meat purchased.  That doesn't mean that there's no cause-and-effect relationship of deliberately spending more money and getting more meat, at the margin.

This is easier to understand if we look at something other than meat -- say, hair. 

Hair CAN be bought for money.  If you're bald, and you want to have hair, you can write a check to Hair Club For Men, and they'll add hair to your head.  But if you look at whether hair HAS BEEN bought for money, very little of it has -- most of it we got free, from God.  The r-squared between "hairs on head" and "money spent" is low, because most hair is not bought for money, and most money is not spent on hair. 

But if you have money, and you choose to buy hair, you'll get it.  


Same for meat.

------

And so, the botttom line is: even if we get a legitimately small correlation, you CANNOT say that "no amount of spending can buy meat."  That's exactly like noting that the correlation between shooting yourself in the head and lifespan is small, and saying, "no amount of shooting yourself in the head can change your life expectancy." 

That's just not true, because it's just not what r-squared means.

-------







(Inspiration: this Freakonomics post.)



Labels: , , ,

Friday, December 09, 2011

A "Grantland" article on Moneyball effects

Here's a baseball salary article at Grantland, by economists Tyler Cowen and Kevin Grier. It’s a strange one ... the impression I get is that is that the authors are just going on the basics of the "Moneyball" story, but don’t really follow baseball discussions very much. And so some of their arguments are obviously behind the curve.

For instance, they talk about how closers used to be paid inefficiently, but aren't any more, except by free-spending teams like New York:

"This year, the Yankees' Mariano Rivera was ranked fifth in total saves with 44. At a salary of $14.9 million, that works out to be a hefty $338,600 per save. The four closers ranked ahead of him averaged 46.5 saves and a salary of $2.9 million, or $63,771 per save — quite the bargain."

The problem here is obvious to almost any serious baseball fan: closers aren’t normally evaluated by the number of saves, which is mostly a function of the opportunities the team provides. Rather, and like any other member of the roster, the closer is paid according to how many wins he can contribute to the team's record, as compared to a replacement player. For Rivera to be worth $15 million, he has to contribute about three extra wins (at a going rate of $4.5 million per win). Which means, basically, he has to blow three fewer saves, given his opportunities. Or, rather, he has to be *expected* to blow three fewer saves; there's still a lot of randomness there.

But Cowen and Grier don't mention randomness at all. And their only reference to blown saves is in one sentence that mentions the Twins' Joe Nathan and Matt Capps, who blew 12 saves out of 41 opportunities.

Another thing, too, is that the article doesn't mention one big difference between Rivera and the others: Rivera is a free agent, while young players like Neftali Feliz can be paid whatever the team wants. The Yankees might prefer Feliz to Rivera, but that’s not a choice they have open to them.

It's not a new "Moneyball" discovery that "slaves" make less money than established free-agent stars ... but the article seems to imply that teams don’t realize that the $400,000 stopper can be just as valuable, for the money, as the $15,000,000 stopper.

To me, it looks like the problem is that if you don’t know baseball that well, you tend to overrate the “Moneyball” possibilities, because that’s the story that you’ve heard the most.

-----

The authors then go on to say:

"The best-known Moneyball theory was that on-base percentage was an undervalued asset and sluggers were overvalued. At the time, protagonist Billy Beane was correct. Jahn Hakes and Skip Sauer showed this in a very good economics paper. From 1999 to 2003, on-base percentage was a significant predictor of wins, but not a very significant predictor of individual player salaries. That means players who draw a lot of walks were really cheap on the market, just as the movie narrates."

The authors imply that “walks were really cheap on the market,” means that the A’s had a huge hole to exploit.

But ... even if walks were indeed “really cheap,” it would still be a small hole. Walks are a significant part of a player’s value, but still in the sense of a small edge, not a huge one. Suppose teams valued walks at only half their actual value. If you can pick up a player with 60 walks, for the price of 30, you’ll gain about 10 runs, or one win. Not a big deal.

Of course, if you can do that nine times, that’s nine free wins. But the A’s didn’t. In 2002, they walked 609 times, third in the league. But that was only 157 more walks than Baltimore, second-worst in the league. If 157 was the number of walks they got at half-price, that’s still only two or three wins.

You could choose, instead, to compare the A’s to the 2002 Tigers, who walked only 363 times. That would be completely unrealistic, in my view, to assume the A’s would have been as bad as one of the worst recent teams ever. But if you do, you *still* only gain four wins.

----

The authors also put too much faith in the Hakes/Sauer paper. As I wrote a few years ago, it seems to me that the paper has a few problems, and I don’t think it shows what it purports to show.

The study found a huge increase in the correlation between salary and OBP between 2003 (when the "Moneyball" book was released) and 2004. The numbers for 2004 almost exactly matched the actual value of a walk, so the authors concluded that the market became efficient in the off-season, and teams wised up after reading the book..

But that conclusion doesn’t make sense. Since only a small percentage of players got new contracts between 2003 and 2004, for the overall average to move so much, the market would have had to overcompensate for walks by double, or triple their real value! That doesn’t sound like a reasonable possibility, and it’s certainly not consistent with GMs now learning to be efficient.

-----

Finally, on the subject of correlation:

"Here's something funny about the Moneyball strategy: It is bringing us a world where payroll matters more and more. Spotting undervalued players boosts their salaries and makes money more important for the general manager; little did Billy Beane know that in the long run he would be strengthening the hand of the large home-market teams, such as the Yankees. From 1986 to 1993, payroll explained 2.2 percent of the variation in team winning percentage, and that meant spending more money yielded little return in terms of quality on the field. In the 2004 to 2006 seasons, after the Moneyball revolution was under way, payroll explained 27.1 percent of the variation in team winning percentage, which means a stronger reason to spend more."

I've written about this before, and Tango’s written about it several times: a higher r-squared does NOT necessarily mean money is more important in buying wins. Rather, the r-squared is a combination of:

1. the extent to which money can actually buy wins;
2. the extent to which teams differ in spending, in real-life.

When the authors say, "spending more money yielded little return," they seem to be assuming it’s all the first thing, when it might be all the second thing.

As an example, take dueling, where two people go out at dawn, draw weapons, and one of them kills the other. Back when it was legal, dueling would explain a lot of the variation in death rates of people who didn’t like each other. Now that it’s illegal, it explains zero.

However, the fact that the r-squared dropped doesn’t mean that dueling is any less dangerous than it used to be (point 1) -- it just means that people no longer vary in how often they get killed in duels (point 2).

The same thing could be happening here. I did a Google search and found an article (.pdf) that gives some team payroll data for the period the article covers. From Table 1, the article shows that from 1985 to 1990, fourth quartile teams (the 25% of teams with the highest payrolls) outspend the first quartile teams by only about 2 to 1. From 1998 to 2002, the ratio jumped to 3 to 1. The paper only covers to 2002, but a glance at later numbers seems to show around 2.5 to 1 (but up to 3.1 to 1 for the 2011 season).

This is evidence that at least *some* of the difference is probably caused by teams being willing to spend more.

I may be unfair to the authors here ... that might be partly what they’re saying. If I read them right, they’re saying that, armed with "Moneyball" concepts, teams are realizing they can buy wins cheaper by evaluating players more accurately (1) -- and, that teams are therefore more likely to vary in how much they pay when they know it’s money well spent (2).

But ... well, I think these effects are pretty small. As I argued, walks are a small part of the overall equation, even if they were undervalued by half (which itself is probably an overestimate). It’s not like, in 1990, teams were paying Jose Oquendo as much as Wade Boggs. To be sure, teams weren’t perfect in evaluating players -- but they were still reasonably good. Any improvement since then has to be relatively small, at the margins.

So, the idea that teams would say, "hey, we can now evaluate players slightly more accurately, so let’s go on a spending spree" doesn’t seem all that plausible.

------

What actually *did* happen to tighten the relationship between payroll and wins? As usual, you guys probably know better than I do. I’ll give you my guess anyway, which is that it’s a combination of a bunch of things:

1. It became more "socially acceptable" for teams to pay big money to free agents. Remember, 1985 to 1990 includes the collusion year, and there was probably a significant amount of pressure to keep spending down. That pressure was probably more significant in discouraging headline-grabbing salaries, rather than routine signings, so maybe a player who was twice as valuable wouldn’t be able to sign for twice as much. That would help keep the correlation between salary and success low.

2. When baseball revenues exploded, they grew more in some cities than others. That meant that marginal wins would be extremely valuable to the Yankees, but not so much to the Pirates. That increased the variation in team spending, which pushed up the r-squared.

3. Teams got smarter, in line with Cowen and Grier’s theory. But I think that was a small part of what happened. Also, I’d guess that a lot of improvement in that regard would have happened well before Moneyball, as Bill James’ discoveries got around a bit. Conventional wisdom denies that baseball executives put any faith in what Bill James had to say, but ... I dunno, good ideas tend to get noticed, even if people say they don’t believe in them. Also, Bill James’ ideas showed up early in arbitration hearings, which affected the teams’ bottom lines pretty much immediately.

4. Randomness. In a team payroll to wins regression, Cowen and Grier give an r-squared of .022 for 1986 to 1993.

(By the way, I assume Cowen and Grier's regression adjusted for payroll inflation ... salaries more than doubled between 1986 and 1993. If they didn't adjust, that might explain the low correlation.)

I wonder if that .022 might just be an outlier. Here are equivalent numbers from Berri/Schmidt/Brook in "The Wages of Wins," page 40:

Wages of Wins:

1988 to 1994: r-squared = .062, r = .25
1995 to 1999: r-squared = .325, r = .57
2000 to 2005: r-squared = .176, r = .42

Cowen/Grier:

1986 to 1993: r-squared = .022, r = .15

The numbers sure do move around a lot! It probably doesn’t take much to knock the correlation down: you need a few teams to get lucky in exceeding their talent, and a few teams to get lucky and get some good slaves and arbs. Maybe I’ll try a simulation and see how common a .022 might actually be.



Labels: , , , ,

Friday, October 07, 2011

Would MLB salaries drop if all players were free agents? Part II

In my previous post, I said I had another argument for why free agent salaries would drop if all players were free agents. Here it is.

Right now, some players are free agents. Their value seems to be about $4.5 million per win.

Now, suppose that, instead of *more* players being free agents, *fewer* players become free agents. In the extreme case, suppose that only one player is a free agent, with a value of 1 WAR. What happens?

It seems like his price will be bid up. But why? It can't be not scarcity in and of itself. Because, even in the real world, when there are lots of free agents, there is still be a time when there's only one left, and HIS price still seems to be $4.5 million per win. There's something else going on. Here's what I think it is.

In in a world with few free agents, wins must be distributed without regard to where they can make the most money. They go to whichever teams made the best draft choices or trades. That's inefficient, financially, for the league as a whole.

The Yankees value wins highly, and would like to buy more, but they can't. The Pirates don't value them much, and would like to sell some to the Yankees. But MLB rules forbid that. So the Yankees are stuck with many fewer wins than they want, and the Pirates are stuck with many more.

So what happens when this one and only free agent goes on the market? Clearly, the Yankees and Red Sox are desperate for wins, with which they can make a lot of money. They'll easily outbid the Royals and Pirates.

But why will the price wind up higher than $4.5 million? Because $4.5 million is the price that results when teams have the ability to fill a lot more of their needs. The Yankees may have been blessed with only 80 wins from their farm system, but they've been able to sign free agents to bring themselves up to 95. And $4.5 million is the value of that 95th win. The wins before that, they valued much higher (or they wouldn't have bought them).

But, in this case, the Yankees are truly stuck at 80. That 81st win they're thinking of buying must be worth more to them than the 95th win (which is worth $4.5 million). And so, they'll be willing pay a lot more for it.

The same logic applies to the Red Sox, and the Phillies, and other teams, and so the price of the single free agent gets bid up well beyond $4.5 million.

--------

That argument shows that fewer free agents means higher costs. That means that more free agents means lower costs, which is what we were trying to prove.

--------

Another thing we can conclude is that the more free agents there are, the less competitive balance. Why? Because the fewer the number of free agents, the more wins are distributed haphazardly among teams. Since teams aren't allowed to sell those wins, small-market teams wind up with wins they otherwise wouldn't have bought. That means more competitive balance.

If all players were free agents, it's possible that some teams would not find it profitable to buy ANY wins, and would stay with replacement-value, minimum-salary players. Obviously, that means competitive balance suffers.

--------

At the risk of my usual overkill, here's another way to look at it:

Suppose that there is a fixed supply of BMWs. In a free market, only rich people would own them, because they're of little use to poor people (who can't drive them much because they can't afford much gas or maintenance). There might be 100,000 rich people in the city who might fancy a nice BMW, but only 10,000 actual cars. So the BMWs go to the people who bid the most for them, and maybe the auction price is $50,000 each.

Now, suppose MLB calls half of the cars "draft choices", airdrops them randomly on households, and prohibits selling them for what they're worth. The poor people are happy to have them, since they're free, but they don't get much benefit from them. On the other hand, the rich families who got them are thrilled: some of them were about to go out and spend $50,000 on one, and now here's one for nothing!

But now, that leaves only 5,000 cars left for auction. And there might still be 99,000 rich people who are interested. Obviously, with fewer cars available on the market, but not many fewer buyers, the price goes up, maybe to $100,000.

Also, "competitive balance" increases. It used to be that the rich owned 100% of the BMWs. Now the rich only own maybe 60% of the BMWs.

And, none of this would happen if the poor people were allowed to sell their cars to the rich people. In that case, the cars would still go for $50,000, same as before, and "competitive balance" wouldn't change. If Bud Selig changed the rules so that the poor could sell to the rich, both the poor and rich would benefit: the poor people would have more money, and the rich people would get their cars cheaper.

Who wouldn't benefit? The BMW fans rooting for poor people. These fans don't care how much money their poor friend has: all they live for is the day that their poor friend drives a BMW to the World Series! Before, when their poor friend had to keep his car, they had some hope. Now that their poor friend will almost always sell his airdropped car, they have little to no hope.

-------

Another thing we can conclude is that it must be true, right now, that there do exist poor teams who have BMWs they'd like to sell but can't. In previous posts, I assumed this didn't happen, to keep things nicely theoretical. I assumed that even the Royals can earn a little bit more by buying a win or two, at the going rate of $4.5 million.

But now we have evidence that's not true, that there are some teams who want to sell wins but can't.

Why do I say we have evidence? Because the previous post showed that, today, if all players were free agents, the price would come down. This proves that there must be at least some BMWs being held by poor households. Otherwise, it wouldn't matter if you airdropped all the BMWs, or none of them: either way, they'd go to the same rich people at the same price. The fact that free-agent restrictions are increasing prices proves that MLB could make more money redistributing wins to the rich teams.

But, that makes an additional assumption: that the fans only care about wins, and not about the fairness of a sport that organizes itself so the Yankees will always be great, and the Royals will always be bad. As I once wrote, I think that even though wins seem to drive revenues today, the fans might get sick of it in the future, and MLB might be better off sacrificing some short-term revenues in favor of keeping the fans interested in the long-term.

-------

In summary, I think these arguments lead to a few real conclusions about the current state of MLB:

1. More free agents would mean lower salaries for those free agents;
2. MLB could make more money by allowing small-market teams to sell players to big-market teams;
3. A marginal win is worth an equal $4.5 million not to all teams, but only to the middle-class and rich teams;
4. The "arbs" and "slaves" do, in fact, contribute to competitive balance.

Labels: , ,