Wednesday, August 14, 2019

Aggregate career year luck as evidence of PED use

Back in 2005, I came up with a method to try to estimate how lucky a player was in a given season (see my article in BRJ 34, here). I compared his performance to a weighted average of his two previous seasons and his two subsequent seasons, and attributed the difference to luck.

I'm working on improving that method, as I've been promising Chris Jaffe I would (for the last eight years or something). One thing I changed was that now, I use a player's entire career as the comparison set, instead of just four seasons. One reason I did that is that I realized that, the old way, a player's overall career luck was based almost completely on how well he did at the beginning and end of his career.

The method I used was to weight the four surrounding seasons in a ratio of 1/2/2/1. If the player didn't play all four of those years, the missing seasons just get left out.

So, suppose a batter played from 1981 to 1989. The sum of his luck wouldn't be zero:

(81 luck) = (81)                     - 2/3(82) - 1/3(83) 
(82 luck) = (82) - 2/5(81)           - 2/5(83) - 1/5(84) 
(83 luck) = (83) - 2/6(82) - 1/6(81) - 2/6(84) - 1/6(85) 
(84 luck) = (84) - 2/6(83) - 1/6(82) - 2/6(85) - 1/6(86) 
(85 luck) = (85) - 2/6(84) - 1/6(83) - 2/6(86) - 1/6(87) 
(86 luck) = (86) - 2/6(85) - 1/6(84) - 2/6(87) - 1/6(88) 
(87 luck) = (87) - 2/6(86) - 1/6(85) - 2/6(88) - 1/6(89)
(88 luck) = (88) - 2/5(87) - 1/5(86) - 2/5(89) 
(89 luck) = (89) - 2/3(88) - 1/3(87) 
---------------------------------------------------------
total luck = 13/30(81) +1/6(82) - 7/30(83) - 1/30(84) - 1/30(86) - 7/30(87) - 1/6(88) + 13/30 (89)

(*Year numbers not followed by the word "luck" refer to player performance level that year).

(Sorry about the small font.)

If a player has a good first two years and last two years, he'll score lucky. If he has a good third and fourth year, or third last and fourth last year, he'll score unlucky. The years in the middle (in this case, 1985, but, for longer careers, any seasons other than the first four and last four) cancel out and don't affect the total.

Now, by comparing each year to the player's entire career, that problem is gone. Now, every player's luck will sum close to zero (before regressing to the mean).

It's not that big a deal, but it was still worth fixing.

--------

This meant I had to adjust for age. The old way, when a player was (say) 36, his estimate was based on his performance from age 34-38 ... reasonably close to 36. Although players decline from 34 to 38, I could probably assume that the decline from 34 to 36 was roughly equal to the decline from 36 to 38, so the age biases would cancel out.

But now, I'm comparing a 36-year-old player to his entire career ... say, from age 25 to 38. Now, we can't assume the 25-35 years, when the player was in his prime, cancel out the 37-38 years, when he's nowhere near the player he was.

---------

So ... I have to adjust for age. What adjustment should I use? I don't think there's an accepted aging scale. 

But ... I think I figured out how to calculate one.

Good luck should be exactly as prevalent as bad luck, by definition. That means that when I look at all players of any given age, the total luck should add up to zero.

So, I experimented with age adjustments until all ages had overall luck close to zero. It wasn't possible to get them to exactly zero, of course, but I got them close.

From age 20 to 36, for both batting and pitching, no single age was lucky or unlucky more than half a run per 500 PA. Outside of that range, there were sample size issues, but that's OK, because if the sample is small enough, you wouldn't expect them close to zero anyway.

---------

Anyway, it occurred to me: maybe this is an empirical way to figure out how players age! Even if my "luck" method isn't perfect, as long as it's imperfect roughly the same way for various ages, the differences should cancel out. 

As I said, I'm still fine-tuning the adjustments, but, for what it's worth, here's what I have for age adjustments for batting, from 1950 to 2016, denominated in Runs Created per 500 PA:

      age(1-17) = 0.7
        age(18) = 0.74
        age(19) = 0.75
        age(20) = 0.775
        age(21) = 0.81
        age(22) = 0.84
        age(23) = 0.86
        age(24) = 0.89
        age(25) = 0.9
        age(26) = 0.925
        age(27) = 0.925
        age(28) = 0.925
        age(29) = 0.925
        age(30) = 0.91
        age(31) = 0.8975
        age(32) = 0.8775
        age(33) = 0.8625
        age(34) = 0.8425
        age(35) = 0.8325
        age(36) = 0.8225
        age(37) = 0.8025
        age(38) = 0.7925
     age(39-42) = 0.7
       age(43+) = 0.65

These numbers only make sense relative to each other. For instance, players created 11 percent more runs per PA at age 24 than they did at age 37 (.89 divided by .8025 equals 1.11).

(*Except ... there might be an issue with that. It's kind of subtle, but here goes.

The "24" number is based on players at age 24 compared to the rest of their careers. The "37" number is based on players at age 37 compared to the rest of their careers. It doesn't necessarily follow that the ratio is the same for those players who were active both at 24 and 37. 

If you don't see why: imagine that every active player had to retire at age 27, and was replaced by a 28-year-old who never played MLB before. Then, the 17-27 groups and the 28-43 groups would have no players in common, and the two sets of aging numbers would be mutually exclusive. (You could, for instance, triple all the numbers in one group, and everything would still work.)

In real life, there's definitely an overlap, but only a minority of players straddle both groups. So, you could have somewhat of the same situation here, I think.

I checked batters who were active at both 24 and 37, and had at least 1000 PA combined for those two seasons. On average, they showed lucky by +0.2 runs per 500 PA. 

That's fine ... but from 750 to 999 PA, there were 73 players, and they showed unlucky by -3.7 runs per 500 PA. 

You'd expect those players with fewer PA to have been unlucky, since if they were lucky, they'd have been given more playing time. (And players with more PA to have been lucky.)  But is 3.7 runs too big to be a natural effect? (And is the +0.2 runs too small?)

My gut says: maybe, by a run or two. Still, if this aging chart works for this selective sample within a couple of runs in 500 PA, that's still pretty good.

Anyway, I'm still thinking about this, and other issues.)

---------

In the process of experimenting with age adjustments, I found that aging patterns weren't constant over that 67-year period. 

For instance: for batters from 1960 to 1970, the peak ages from 27 to 31 all came out unlucky (by the standard of 1950-2015), while 22-26 and 32-34 were all lucky. That means the peak was lower that decade, which means more gentle aging. 

Still: the bias was around +1/-1 run of luck per 500 PA -- still pretty good, and maybe not enough to worry about.

---------

If the data lets us see different aging patterns for different eras, we should be able to use it to see the effects of PEDs, if any.

Here's luck per 500 PA by age group for hitters, 1995 to 2004 inclusive:

-1.75   age 17-22
-0.74   age 23-27
+0.61   age 28-32
+0.99   age 33-37
+0.45   age 38-42

That seems like it's in the range we'd expect given what we know, or think we know, about the prevalence of PEDs during that period. It's maybe 2/3 of a run better than normal for ages 28 to 42. If, say 20 percent of hitters in that group were using PEDs, that would be around 3 runs each. Is that plausible? 

Here's pitchers:

-1.22   age 17-22
-0.51   age 23-27
+1.36   age 28-32 
+1.42   age 33-37 
+1.07   age 38-42 

Now, that's pretty big (and statistically significant), all the way from 28 to 42: for a starter who faces 800 batters, it's about 2 runs. if 20 percent of pitchers are on PEDs, that's 10 runs each.

By checking the post-steroid era, we can check the opposing argument that it's not PEDs, it's just better conditioning, or some such. Here's pitchers again, but this time 2007-2013:

-0.06   age 17-22
+1.01   age 23-27
+0.30   age 28-32
-1.67   age 33-37
+0.59   age 38-42

Now, from 28 to 42, pitchers were *unlucky* on average, overall.

I'd say this is pretty good support for the idea that pitchers were aging better due to PEDs ... especially given actual knowledge and evidence that PED use was happening.







Labels: , ,

Wednesday, December 07, 2016

Charlie Pavitt: Steroids and the Hall of Fame

This guest post is from occasional contributor Charlie Pavitt. Here's a link to some of Charlie's previous posts.

------

I am writing today about a much-discussed topic, performance enhancing drugs and Baseball Hall of Fame enshrinement. My goal is not to defend a particular opinion about it, but rather to attempt to lay out five possible positions and some strengths and weaknesses each has. In fact, one reason why I will not defend a particular opinion is that, given these strengths and weaknesses, I am torn among several of the options.

But before I start, a few preliminaries. First, research of which I am aware provides strong evidence that steroid use significantly increases offensive performance, whereas there is little if any evidence that human growth hormone has any impact.

Second, none of this is new. Ancient Greek athletes took then-known stimulants before competitions, and nobody back then batted an eye.
                
Third, one must be careful throwing rocks when one’s own house could potentially, in a different context, be made of glass. When I was in graduate school, if someone had come to me and whispered, "Hey man, I have this pill you can take every day that will make you read, write, and think more quickly and efficiently," I would have been sorely tempted to partake.  In fact, one of my grad school cohort-mates imagined a situation in which you took a pill that provided you with the information you are supposed to learn from assigned reading, with lighter doses for undergraduate students and heavier doses for us grad students. Mighty tempting fantasy.
                
Fourth, and this is critical: Before throwing rocks, one needs to defend the claim that there is something wrong with taking performance enhancing drugs.  The fact that it may be illegal is, in my view, irrelevant, as many illegal items are not only harmless but helpful. For example, without getting into the marijuana debate, it is the case that any use of hemp has been illegal in some places, despite its many many positive applications. And taking something into the body to improve athletic performance is often a good thing. After all, a person can improve athletic performance by eating better, and perhaps taking supplements of necessary vitamins and minerals in moderation. So what’s the difference?  Here’s an argument for that difference; eating better and taking supplements in moderation promotes overall health whereas taking steroids (and overdosing on vitamins/minerals) does quite the opposite. The early deaths of many professional rasslers (I reserve the word "wrestlers" for the real sport), perhaps some football players (Lyle Alzado?), and two well-known baseball players (more on this later) has been linked with steroid use. 

One could then make the claim that it is the use of a substance that causes bodily harm that warrants rejection from the HOF. After all, the criteria for entry include "Integrity, sportsmanship, and character" along with "record, playing ability," and "contributions to the team(s) on which the player played."  

So, the argument continues, PED use is contrary to the former three criteria.  I think the best angle for this argument is that it sets the wrong example for others, particularly young people, whereas eating well and getting one’s vitamins/minerals sets the right example. Fair enough. But: Lots of HOF players were smokers or used chewing tobacco, and Babe Ruth certainly did not set a good dietary example by reportedly eating multiple hot dogs just before games.  And speaking of setting bad examples, if there is anybody enshrined who does not deserve it for absence of integrity etc., it is Adrian "Cap" Anson, who was proactive in the successful attempt to get Moses Fleetwood Walker, the first African-American major league baseball player, banned for the color of his skin.

So this argument leads to a slippery slope. But let us assume that we accept it.  Here are five possible responses, ordered from most lenient to most strict.

Position Number One: Let everyone in. The argument here is that great performers deserve entry no matter why they performed greatly. Buttressing this position is the seeming fact that until the public response to Jose Canseco’s confession among other events forced action, the powers-that-be in MLB’s establishment knew what was going on and intentionally turned a blind eye to it. After all, fans like offense, particularly home runs, and attendance was swelling, so all seemed right with the world. So if that was what baseball was in those days, goes the argument for this position, one must accept it and its great performers no matter what.  ne strength of this argument is that one must not make always-problematic non-performance-related judgments about players, as we will must for the other positions on this issue to be discussed in turn below.  The problem with this argument is that it is contrary to the "integrity, sportsmanship, and character" criteria, condones unhealthy behavior, and as such if anything encourages youthful copy cats.
                
Position Number Two: Let everyone in who deserves enshrinement independently of PED use. This implies that one likely accepts Alex Rodriguez, Roger Clemens, and Barry Bonds, because if one mentally subtracts the PED-fed "value-added" part of their performance, they are still HOF material. But one rejects those who would not have reached performance criteria otherwise: Mark McGwire and Rafael Palmeiro among others come to mind.  Perhaps this makes some sense, but one is still condoning bad behavior by allowing in known users while making questionable judgments about whose performance would have been "good enough" without PEDs.
                
Position Number Three: Ban known users. So Bonds, Clemens, McGwire, Palmeiro, Sammy Sosa, Manny Ramirez, and some others who reached supposed HOF performance levels are out. Also some who approached HOF levels and might otherwise deserve consideration (Miguel Tejada, Jason Giambi) get none. In so doing, we clear the deck of those guilty of poor integrity etc.  Also, it might allow us to consider those whose performance would have reached criteria in another era; think Fred McGriff, who hit as many homers as Lou Gehrig.  But what about those suspected of use? Take Jeff Bagwell for an example. Although there is no clear evidence of his use and he has steadfastly denied it, he did get a lot bigger fairly quickly, hit way more homers than anyone originally expected, and associated closely with the first known user alluded to earlier, Ken Caminiti, whose early death has been partly linked with use. If we lower our performance criteria to allow for McGriff, it also allows for Bagwell.  So now we’ve admitted someone who may have been as guilty as Bonds et al. but whose usage (if any) has not been publicly verified. Setting the in versus out boundary is a pretty questionable judgment call.
                
Position Number Four: Ban everyone either known or rumored to be users.  So Bagwell and Gary Sheffield and perhaps Juan Gonzalez if you think he reached HOF performance levels and maybe even David Ortiz are out, and Mike Piazza should not have been recently admitted. Now we know we’ve kept the HOF free of those with poor integrity etc. But at least in the U.S. court of law one is considered innocent until proven guilty. Take Jeff Bagwell.  Although he got a lot bigger fairly quickly, hit way more homers than anyone originally expected, and associated with a known user, there is no clear evidence of his use and he has steadfastly denied it. Again, setting the in versus out boundary is a pretty questionable judgment call.
                
Position Number Five: Not only ban everyone suspected, but kick out anyone currently in who is suspected. Now we are sure everyone in baseball had the proper integrity etc.  Out goes Mike Piazza. Further, and this is the second player I alluded to at the beginning of this essay, out goes Kirby Puckett.  Jose Canseco fingered him, plus the physical problems that ended his career along with those that ended his life along with his violent post-baseball behavior sure seem to be signals of steroid use. In addition, in a 2002 article, statistician Scott Berry calculated that Puckett’s jump from no home runs his rookie year (1985) to four his sophomore year to 31 his junior year was the most unlikely performance increase in the history of MLB, with an odds of one in 100 million, much greater than similar jumps made by other known or suspected users. But this is all indirect evidence, there are other explanations for all of it (maybe Kirby started taking his vitamins or radically changed his swing between the 1986 and 1987 seasons). And if we kick out Kirby, should we kick out Adrian Anson? (Actually, I think we should, but that’s a side issue here.)  How about Ty Cobb for his racism (to be expected, natural attitudes for a Georgian in his time)?  Babe Ruth for eating all those hot dogs? I do not believe I have heard anyone support this position, but I suppose someone could.
                
So – as I noted at the top, I am frankly torn among several of these options. If I had a vote, my heart would point me toward Position Three, but my head would tell me that it would be hard to rationally defend relative to some of the others (particularly Four). Anyway, I hope that I’ve laid out at least some of the arguments on either side well enough that readers can have an intelligent discussion about it and maybe even add some arguments to my list, and that those who are SURE that their position, whichever it is, is obviously correct think twice about its weaknesses along with its strengths.

-- Charlie Pavitt





Labels: , , ,