Channel closed, case closed: obituary for a manager-value estimator

I’ve made it through draft 426 (or at least 6) of “Against Pythagorean Benchmarking”. . . . I think I feel good about it now; so am soliciting comments from those who will tell me I shouldn’t—pile on!—before I make the revisions necessary to submit it to the ordeal of formal submission.

One thing the paper contains is an obituary.

The deceased is the claim that managers can be meaningfully evaluated based on how their records stack up against their teams’ Pythagorean Expectation (PE) winning percentage.

By far the most popular member of the PE benchmarking family, it is actually the one that can most readily be shown incapable of empirical confirmation.

The problem is that managers can’t break through the noise barrier of PE.

PE is a random variable: it’s an estimate—a really good one but subject to the kind of error that any model-based projection is.

As a result, PE projections come with an irreducible band of noise surrounding them. That noise defines how big effects have to be before they can be distinguished from chance deviations. We want to know, then, whether manager PE deviations meet that threshold or not.

In the paper, I measure how much the career record of every manager since 1910 has deviated (positively or negatively) from the record PE predicts his teams should have compiled.

Plenty of them appear to have done so by a “statistically significant” margin. The problem is that the normal “null hypothesis test” apparatus used to measure “significance” is innocent of what a genuine null looks like and has no way to separate out “false non-nulls”—ones that occur by chance as more and more tests are performed—from “true” ones.

Empirical Bayes supplies tools for fixing these defects.

One thing we can do is calculate a “False Discovery Rate.” Using a valid empirical null baseline (in this case, one constructed by a simulation that recreated an empirically realistic replay of every manager’s games devoid of the impact he could have had on the “run scoring variance channels” that control PE deviation), we can test each manager against a 20,000-member ensemble of zombies to see whether he deviated any more often than they did.

The FDR was 0.77, indicating that the odds were > 3:1 that every apparently “significant” result was a false positive.  If we do the calculations, that implies that at most 1.5% of the managers over AL/NL history have “beaten” or been “beaten by” PE to a degree that differs from chance.

But the real point of an analysis like this isn’t to figure out whether any manager has ever “really” deviated from PE—a sort of goofy exercise. It is to determine whether there is enough signal of real manager effects in the data to justify the inference that any particular manager record that appears to deviate from PE isn’t really just a statistical mirage.

An FDR of 0.77 basically tells you the answer is “no”: if you find a manager who looks like he is a PE-beater (or -beatee), you will always be more likely to be wrong than right. You’d be a fool, in other words, to base any judgments on such data.

But using empirical Bayes, we can sharpen the point of this inferential bayonet even further.

Bayesian analysis, unlike normal null-hypothesis testing, uses techniques that tease apart the contributions that real differences and measurement error make, respectively, to variance in the quantity being estimated—here manager skill as reflected in PE deviations.  The “real difference” part is referred to as τ, or “tau”—like the restaurant but much more satisfying than the overpriced servings of meh served there.

When that analysis was applied to the manager PE data, poor τ collapsed to a value of 0—meaning that no genuine signal of manager heterogeneity could be detected at all.  The effect of managers when measured this way was completely entombed in PE noise.

So unless something really fundamental changes in the role that managers play in baseball, no matter how big a manager’s apparent PE deviation (positive or negative) appears to be, it will never give us any basis for inferring that he has made a difference.

But that doesn’t mean that managers are irrelevantIn my recently published JSA paper, I used the same Bayesian methods to assess managers against a team WAR benchmark, and got a very different result.

The FDR—effectively, false-positive rate—for those who evince “significant” deviations from their teams’ WAR-predicted records is substantially below 0.50.

And the Bayesian analysis shows τ = 2.4 wins per 162. That’s effectively the standard deviation for manager effects.

On this basis, we can form probability estimates that enable us to be confident that dozens of managers over the course of AL/NL history have compiled career records between -2.1 and +6.4 wins per season over their careers.

In sum, the manager PE benchmark result doesn’t show that managers are irrelevant. It just shows that the PE benchmark test is a defective instrument for measuring how much managers matter.

The obituary, then: “Devoured by hungry zombies, the manager PE test finally succumbed to a case of inferential bareness. Fortunately, though, it is survived by a new generation of estimators that can be used to measure baseball value with a degree of practical assurance previously unattainable in baseball history. . . .”

Leave a Reply

Your email address will not be published. Required fields are marked *