Tuesday, August 26, 2014

Sabermetrics vs. second-hand knowledge

Does the earth revolve around the sun, or does the sun revolve around the earth?

The earth revolves around the sun, of course. I know that, and you know that.

But do we really? 

If you know the earth revolves around the sun, you should be able to prove it, or at least show evidence for it. Confronted by a skeptic, what would you argue?  I'd be at a loss. Honestly, I can't think of a single observable fact that I could use to make a case.

I say that I "know" the earth orbits the sun, but what I really mean by that is, certain people told me that's how it is, and I believe them. 

Not all knowledge is like that. I truly *do* know that the sun rises in the east, because I've seen it every day. If a skeptic claimed otherwise, it would be easy to show evidence: I'd make sure he shared my definition of "east," and then I'd wake him up at 6 am and take him outside.

But that sun/earth thing?  I can only I only say I "know" it because I believe that astronomers *truly* know it, from direct evidence.

------

It occurred to me that almost all of our "knowledge" of scientific theories comes from that kind of hearsay. I couldn't give you evidence that atoms consist, roughly, of electrons orbiting a nucleus. I couldn't prove that every action has an equal and opposite reaction. There's no way I could come close to figuring out why and how e=mc^2, or that something called "insulin" exists and is produced by the pancreas. And I couldn't give you one bit of scientific evidence for why evolution is correct and not creationism. 

That doesn't stop us from believing, really, really strongly, that we DO know these things. We go and take a couple of undergraduate courses in, say, geology, and we write down what the professors tell us, and we repeat them on exams, and we solve mathematical problems based on formulas and principles we are told are true. And we get our credits, and we say we're "knowledgeable" in geology. 

But it's a different kind of knowledge. It's not knowledge that we have by our own experience or understanding. It's knowledge that we have by our own experience of how to evaluate what we're told -- how and when to believe other people. We extrapolate from our social knowledge. We believe that there are indeed people, "geologists," who have firsthand evidence. We believe that evidence gets disseminated among those geologists, who interact to reliably determine which hypotheses are supported and which ones are not. We believe that, in general, the experts are keeping enough of a watchful eye on what gets put in textbooks and taught at universities, that if Geology 101 was teaching us falsehoods, they'd get exposed in a hurry.

In other words, we believe that the system of scientists and professors and Ph.D.s and provosts and deans and journals and textbook publishers is a reliable separator of truth from falsehood. We believe that, if the earth really were only 6,000 years old, that's what scientists would be telling us.

------

Most of the time, it doesn't matter that our knowledge is secondhand. We don't need to be able to prove that swallowing arsenic is fatal; we just need to know not to do it. And, we can marvel at Einstein's discovery that matter and energy are the same thing, even if we can't explain why.

But it's still kind of unsatisfying. 

That's one of the reasons I like math. With math, you don't have to take anyone's word for anything. You start with a few axioms, and then it's all straight logic. You don't need geology labs and test tubes and chemicals. You don't need drills and excavators. You don't actually have to believe anyone on indirect evidence. You can prove everything for yourself.

The supply of primes is infinite. No matter how large a prime you find, there will always be one larger. That's a fact. If you like, you can look it up on the internet, or ask your math teacher, or find it in a textbook. It's a fact, like the earth revolving around the sun.

If you do it that way, you know it, but you don't really KNOW it. You can't defend it. In a sense, you're believing it on faith. 

On the other hand, you can look at a proof. Euclid's proof that there is no largest prime number is considered one of the most elegant in mathematics. The versions I found on the internet use a lot of math notation, so I'll paraphrase.

-----

Suppose you have a really big prime number, X. The question is: is there always a prime bigger than X?  

Try this: take all the numbers from 1 to X, and multiply them together: 1 times 2 times 3 .... times X. Now, add 1. Call that really huge number N. That huge N is either prime, or is the product of some number of primes. 

But N can't be divisible by X, or anything less than X, because that division has to always leave a remainder of 1. Therefore: either N is prime, or, when you factor N into other primes, they're all bigger than X. 

Either way, there is a prime bigger than X.

------

I may not have explained that very well. But, if you get it ... now you know that there is no highest prime. If you read it in a book, you "know" it, but if you understand the proof, you KNOW it, in the sense that you can explain it and prove it to others.

In fact ... if you read it in a textbook, and someone tells you the textbook is wrong, you may have some doubt. But once you see the proof, you will *never* have doubt (except in your own logic). Even if the greatest mathematician in the world tells you there's a largest prime, you still know he's wrong. 

-----

In theory, everything in math is like that, provable from axioms. In practice ... not so much. The proofs get complicated pretty quickly. (When Andrew Wiles solved Fermat's Last Theorem in 1993, his proof was 200 pages long.)  Still, there are significant mathematical results where we can all say we know from our own efforts. For years, I wondered why it was that multiplication goes both ways -- why 8 x 7 has to equal 7 x 8. Then it hit me -- if you draw eight rows of seven dots, and turn it sideways, you get seven rows of eight dots.

There are other fields like math that way ... you and I can know things on our own, fairly easily, in economics, and finance, and computer science. Other sciences, like physics and chemistry, take more time and equipment. I can probably prove to myself, with a stopwatch and ruler, that gravitational acceleration on earth is 9.8 m/s/s, but there's no way I could find evidence of what it is on the moon. 

But: sabermetrics. What started me on all this is realizing that the stuff we know about sabermetrics is more like infinite primes than like the earth revolving around the sun. Active researchers don't just know sabermetrics because Bill James and Pete Palmer told us. We know because we actually see how to replicate their work, and we see, all the way back to first principles, where everything came from. 

I can't defend "e equals mc squared," but I can defend Linear Weights. It's not that hard, and all I need is play-by-play data and a simple argument. Same with Runs Created: I can pull out publicly-available data and show that it's roughly unbiased and reasonably accurate. (I can even go further ... I can take partial derivatives of Runs Created and show that the values of the individual events are roughly in line with Linear Weights.)

DIPS?  No problem, I know what the evidence is, there, and I can generate it myself. On-base percentage more important than batting average?  Geez, you don't even need data for that, but you can still do it formally if you need to without too much difficulty. 

For my own part -- and, again, many of you active analysts reading this would be able to say the same thing --  I don't think I could come up with a single major result in sabermetrics that I couldn't prove, from scratch, if I had to. Even the ones from advanced data, or proprietary data, I'm confident I could reproduce if you gave me the database.

For all the established principles that are based on, say, Retrosheet-level data ... honestly, I can't think of a single thing in sabermetrics that I "know" where I would need to rely on other people to tell me it's true. That might change: if something significant comes out of some new technique -- neural nets, "soft" sabermetrics, biomechanics -- I might have to start "knowing" things secondhand. But for now, I can't think of anything.

If you come to me and say, "I have geological proof that the earth is only 6,000 years old," I'm just going to shrug and say, "whatever."  But if you come to me and say, "I have proof that a single is worth only 1/3 of a triple" ... well, in that case, I can meet you head on and prove that you're wrong. 

I don't really know that creationism isn't right -- I only know what others have told me. But I *do* know firsthand what a triple is worth, just as I *do* know firsthand that there is no highest prime. 

------

And that, I think, is why I love sabermetrics so much -- it's the only chance I've ever had to actually be a scientist, to truly know things directly, from evidence rather than authority.

I have a degree in statistics, but if nuclear war wiped out all the statistics books, how much of that science could I restore from my own mind?  Maybe, a first-year probability course, at best. I could describe the Central Limit Theorem in general terms, but I have no idea how to prove it ... one of the most fundamental results in statistics, one they teach you in your first statistics class, and I still only know it from hearsay.

But if nuclear war wipes out all the sabermetrics books ... as long as someone finds me a copy of the Retrosheet database, I can probably reestablish everything. Nowhere near as eloquently as Bill James and Palmer/Thorn, and I'd probably wouldn't think of certain methods that Tango/MGL/Dolphin did, but ... yeah, I'm pretty sure I could restore almost all of it. 

To me, that's a big deal. It's the difference between knowing something, and only knowing that other people know it. Not to put down the benefits of getting knowledge from others -- after all, that's where most of our useful education comes from. It's just that, for me, knowing stuff on my own ... it's much more fulfilling, a completely different state of mind. As good as it may be to get the Ten Commandments from Moses, it's even better to get them directly from God.



Labels: , ,

Wednesday, October 19, 2011

How much does "Moneyball" help a team?

How much is sabermetrics worth to a team?

That's probably a hard question to answer. Every team uses statistics to some extent. Even before sabermetrics, teams were looking at player statistics to decide who to play and who not to play. They may not have had any fancy formulas, but they had a pretty good idea of how to weight the relative contributions of players. Nobody ever released a 30-HR guy because he was only hitting .240, and nobody ever released a .330 hitter because he had no power. Intuitive evaluations weren't perfect, of course, but they were pretty reasonable most of the time.

Where sabermetrics helps, I think, is not in evaluating actual performance, but in helping figure out *future* performance. How to extrapolate minor-league performance in to major league performance ... how to take luck out of a player's batting or pitching line ... figuring how different kinds of players age ... that sort of thing.

Suppose you took a team management right out of the early 1970s, and gave them a team today without letting them learn anything discovered after 1977. How much would that team underperform compared to the rest of MLB? I don't have an answer to the question, but I'd be interested in hearing yours.

Anyway, here's a narrower question. How much can a more sabermetric approach *today* benefit a team, compared to, say, the typical team's sabermetric approach? For instance, how much did Billy Beane really mean to the A's?

A couple of weeks ago, Tango did a study to figure out which teams did better or worse than expected, given their payroll. The A's were the team that outperformed the most over the last decade -- about 7 games per season, it looks like. That's a lot, but there's probably a whole bunch of luck there, since we're cherry-picking them as the best of the lot. Also, it's possible that much of their outperformance came in the early years, when, as many critics of "Moneyball" hype have pointed out, they had three underpriced ace starters.

So, we'd have to regress that 7 games to the mean a fair bit. If you made me make an arbitrary guess, I'd be willing to bet that less than half of that seven game advantage came from sabermetrics. (But, I have no real basis for that guess without studying it.)


Anyway, with the Cubs signing Theo Epstein, we now have a market estimate for what sabermetrics might be worth today. Epstein's new agreement is for about $4 million per season. He still had one year to go on his contract with the Red Sox, for which they will receive some sort of compensation from the Cubs. Let's say that compensation will be worth $1 million. So Epstein's value is around $5 million. I don't know how much an average replacement level GM makes, by comparison. To be conservative, let's say it's $500,000, although it's probably more than that. That means that Epstein's excess value is $4.5 million, exactly what it costs in free agent players to gain one extra win.

It looks like that's what Epstein is worth: one win per season.

Is that a lot? Frankly, I don't know. It's a competitive market for players these days, with lots of money on the line, and there's lots of random luck in who makes it and who doesn't. In that light, it could be that one win per season is an exceptional, genius-level performance.

If that's the case, doesn't it mean that the "Moneyball" approach is overrated? I mean, one win a year. At that rate, it would take decades, even centuries, to have good statistical evidence that the sabermetric approach works.

Of course, you have to remember that that's compared to other teams ... and, nowadays, those other teams are doing a fair amount of statistical work themselves. Maybe it's three or four games over a team that won't look at anything new at all, that never heard of Voros McCracken and winds up overpaying pitchers with lucky BABIPs. And, maybe Epstein took less pay than he was worth in order to become a Cub. Maybe it's a win and a quarter, or a win and a half.

Still ... to me, one game doesn't seem that unreasonable. The point might not be that an you can win pennants just by embracing sabermetrics. The point might be that, with every team in a sabermetric arms race against every other team, you certainly can *lose* pennants if you persist in living in the 70s.

But, again ... one game. Doesn't that mean that if a team does well, and someone credits "Moneyball," they're probably just blowing smoke?


UPDATES:

1. In the comments, Bill Waite suggests that sabermetrically-savvy managers might have a significant impact, too. He says that just rejigging the lineup is worth almost half a game a season, and says that the difference between best and worst could be as much as eight games.

Food for thought. It would be interesting to consider how to try to look for this in the historical record (if indeed that is possible), since we know that some managers are indeed more numbers-oriented than others.

2. Matt Swartz e-mailed me about a study where he found a positive correlation between sabermetric management and team performance. It's here.


Labels: , ,

Tuesday, March 29, 2011

"Sabermetrics" and "Analytics"

What is sabermetrics?

We sabermetricians think we know what it means ... one definition is that it's the scientific search for understanding about baseball through its statistics. But, like a lot of things, it's something that's more understood in practice than by strict definition. I think a few years ago Bill James quoted Potter Stewart: "I know it when I see it."

But how we see it seems to be different from how the rest of the world sees it. The recent book "Scorecasting" is full of sabermetrics, isn't it? There are studies on how umpires call more strikes in different situations, on how hitters bat when nearing .300 at the end of the season, on how hitters aren't really streaky even though conventional wisdom says they are, and on how lucky the Cubs have been throughout their history.

So why isn't "Scorecasting" considered a book on sabermetrics? It should be, shouldn't it? None of the reviews I've seen have called it that. The authors don't describe themselves as sabermetricians either. In fact, on page 120, they say,

"Baseball researchers and Sabermetricians have been busily gathering and applying the [Pitch f/x] data to answer all sorts of intriguing questions."


That suggests that they think sabermetricians are somehow different from "baseball researchers".

Consider, also, the "MIT Sloan Sports Analytics Conference," which is about applying sabermetrics to sports management. But, no mention of "sabermetrics" there either -- just "analytics".

What's "analytics"? It's a business term, about using data to inform management decisions. The implication seems to be that the sabermetrician nerds work to provide the data, and then the executives analyze that data to decide whom to draft.

But, really, that's not what's going on at all. The executives make the decisions, sure, but it's the sabermetricians who do the actual analysis. Sabermetrics isn't the field of creating the data, it's the field of scientifically *analyzing* the data in order to produce valid scientific knowledge, both general and specific.

For instance, here's a question a GM needs to consider. How much is free agent batter X worth?

Well, towards that question, sabermetricians have:

-- come up with methods to turn raw batter statistics into runs
-- come up with methods to turn runs into wins
-- come up with methods to estimate future production from past production
-- come up with methods to quantify player defense, based on observation and statistical data
-- come up with methods to compare players at different positions
-- come up with methods to estimate the financial value teams place on wins.

But isn't that also what "analytics" is supposed to do? I don't understand how the two are different. I suppose you could say, the sabermetricians figure out that the best estimate for batter X's value next year will be, say, $10 million a season. And then the analytics guy says, "well, after applying my MBA skills to that, and analyzing the $10 million estimate the sabermetricians have provided, I conclude that the data suggest we offer the guy no more than $10 million."

I don't think that's what the MIT Sloan School of Management has in mind.

Really, it looks like everyone who does sabermetrics knows that they're doing sabermetrics, but they just don't want to call it sabermetrics.

Why not? It's a question of signalling and status. Sabermetrics is a funny, made-up, geeky word, with the flavor of nerds working out of their mother's basements. Serious people, like those who run sports teams, or publish papers in learned journals, are far too accomplished to want to be associated with sabermetrics.

And so, an economist might publish a paper with ten pages of analysis of sports statistics, and three paragraphs evaluating the findings in the light of economic theory. Still, even though that paper is sabermetrics, it's not called sabermetrics. It's called economics.

A psychologist might analyze relay teams' swim times, discover that the first swimmer is slower than the rest, and conclude it's because of group dynamics. Even though the analysis is pure sabermetrics, the paper isn't called sabermetrics. It's called psychology.

A new MBA might get hired by a major-league team to find ways of better evaluate draft prospects. Even though that's pure sabermetrics, it's not called sabermetrics. It's called "analytics," or "quantitative analysis."

I think that word, sabermetrics, is costing us a lot of credibility. My perception is that "sabermetrics" has (incorrectly) come to be considered the lower-level, undisciplined, number crunching stuff, while "analytics" and "sports economics" have (incorrectly) come to symbolize the serious, learned, credible side. If you looked at real life, you might come to the conclusion that the opposite is true.

My perception is that there isn't a lot of enthusiasm for the word "sabermetrics." Most of the most hardcore sabermetric websites -- Baseball Analysts, The Hardball Times, Inside The Book, Baseball Prospectus -- don't use the word a whole lot. Even Bill James, who coined the word, has said he doesn't like it. From 1982 to 1989, Bill James produced and edited a sabermetrics journal. He didn't call it "The Sabermetrician." He called it "The Baseball Analyst." It was a great name. About ten years ago, I suggested resurrecting that name for SABR's publication, to replace "By the Numbers" (.pdf, see page 1). I was voted down (.pdf, page 3). I probably should have tried harder.

In light of all that, I wonder if we should consider slowly moving to accept MIT's word and start calling our field "analytics."

It's a good word. We've got a historical precedent for using it. It will help correct misunderstandings of what it is we do. And it'll put us on equal footing with the MIT presenters and the JQAS academics and the authors of books of statistical analysis -- all of whom already do pretty much exactly what we do, just under a different name.


Labels: ,