As all 4,326,379 regular readers of this blog/journal know, I’m perplexed—haunted—by an anomaly in Gould’s conjecture.
Gould’s conjecture is essentially that as a sport matures, differences in the performances of the top competitors will become more compressed—variance will shrink—as a result of a growing talent pool and selection pressures.
This phenomenon has been observed across a wide range of sports, from track and field records to U.S. and European football performances to NHL goal-scoring.
And of course it has been observed in baseball. That was Gould’s focus. He used his account to try to explain the extinction of the .400 hitter (I’m actually not sure why it would predict extinction as opposed to just a denser concentration of them but leave that aside).
What has bothered me is that Gould’s conjecture doesn’t seem to apply to baseball pitching. There, FIP, the best measure of pitcher skill, shows an increase in variance over time, a pattern that can be observed, too, in less discerning but still okay measures like ERA and WHIP.
One thought I had is that maybe it is wrong to see pitching as just being “one thing” over time. Clearly pitching usage has changed dramatically—with teams reducing their reliance on starters and with the advent of role-specialization in relievers.
If it is a mistake to think of pitcher as a unitary thing, maybe it is a mistake to look for declining variance in a conglomeration of the roles it comprises.
Maybe the position should be disassembled over time and investigated that way for shrinking skill differentials.
I decided I should give that a try.
For this purpose, I changed my testing strategy. Usually Gould’s conjecture is assessed by charting trends in variance over time.
But I thought I’d try empirical-Bayes measurement of latent player skill.
As pitching roles get specialized, lower workloads for starters and a greater profusion of work by short-stint relievers will magnify measurement error and make isolating compression of skill differences harder to detect.
But separating those two things is at the core of Bayesian methods, which yield and make use of a “true” latent skill heterogeneity measure, known as τ, to be separated from sampling noise—exactly what is growing as pitcher specialization reduces the innings that pitchers accumulate over a season.
So I constructed empirical-Bayesian models (ones largely Efron-Morris in nature) for both starters and relievers. Among the latter, I distinguished between ones summoned to appear in game–consequential situations—ones in which their teams either were protecting leads or maintaining their teams’ chances when they weren’t but victory was still within reach. (The exact criteria are spelled out in my paper “Against Pythagorean Benchmarking” if you want details.)
In the models, I was especially interested to measure τ separately over multiple pitcher usage periods. I wanted to see if skill differences among both starters and consequential relievers were contracting, consistent with the Gould conjecture, over periods in which their functions changed. I was thinking, as I said, that maybe declining variance in FIP was being obscured by the conflation of the evolution of pitching roles. The periods were also identified by Bayesian means—and are described in more detail in “Against Pythagorean Benchmarking.”
The currency in which I was measuring performance was pitching-runs saved derived from eFIP, an empirical variant of FIP that I’ve discussed in multiple previous posts (this is the first one, though, in which I’ve used an empirical-Bayesian formulation of it).
For comparison, I did the same for hitters. For them I used a measure of run production that was based on OPS, or really on the two components of OPS—slugging and on-base percentage—treated as separate indicators of batting proficiency. Previous posts supply details on how OPS and like measures of hitter skill can be turned into latent “run productivity” estimators. For hitters, I formed separate τ estimates for each “run-scoring period” in AL/NL history (again identified with Bayesian breakpoint analysis).
I also rescaled both the pitching-runs-saved and batting-runs-produced estimators to remove the confound associated with variance attributable to fluctuations in run scoring over different baseball eras.
I didn’t use season-specific z-scores–that by design would have extinguished the historical changes in between-player variance that I was trying to measure.
Instead, I used a variance normalization technique: dividing the performance scores by the season-specific standard deviations (more specifically the ETEL-implied standard deviations) corresponding to the mean number of runs per game for the season in which those performances were being measured. Doing that removes only the model-generated multiplicative effect of the prevailing run environment on performance-metric dispersion levels while leaving differences in between-player variance free to vary over time.
So what did I find?

Well, as expected, the analysis basically confirmed that batting latent skill has been becoming more concentrated, more compressed, exactly as the Gould conjecture predicts. BTW, I removed pitchers from the sample before estimating the period-specific τs, since obviously the advent of the DH has removed the anemic hitting of pitchers as a source of variance in hitting talent.
The trend is strongly downward over time, with a τ–basically the standard deviation in latent skill, once it has been separated from measurement error–dropping from about 15.7 runs produced per 500 plate appearances to 11.5 in the “modern era.”
There is an interesting short-term reversal in the steroid era, where heterogeneity increased to τ = 13.3. Innovations in sports—advances in equipment, evolution in techniques, better training methods, etc.—are like an external jolt that temporarily spreads performers out until they adapt and start to converge again. It is plausible to think of the advent of steroids operating that way in baseball—until they were effectively purged from the sport.
But what about pitching?
The anomaly was not dispelled!
As is clear, there has been no pattern of decline in the heterogeneity of skill levels, as measured by τ, either for starters or relievers over the course of baseball history.

Indeed, such heterogeneity is much more pronounced today than in previous eras: in the case of starters, τ = 0.73 runs saved per 9 IP now vs. 0.54 in the first decade of the 20th century; for consequential relievers, τ = 0.60 today vs. 0.44 from 1910-1953! (Non-consequential relievers—basically pitchers in role of “mopping up”—unsurprisingly displayed no meaningful heterogeneity in any era.)
Remember, too, that these are values that have been rescaled to remove the impact that era differences in run-scoring make. The τs are era specific, but have been calibrated to remove the effect that those differences make in the spread of player run-generating or -suppressing skill.
So what is there to say?
Nothing, except that I remain as mystified as ever!
If anyone has any thoughts, please speak up—or we can all expect to be haunted by Gould’s ghost for the rest of our lives.