Every fan carries a quiet theory of which numbers to trust. Velocity is real; a hot April is noise. Strikeout stuff is a skill; a gaudy RBI total is a mirage. We’ve leaned on that theory here more than once — in the Jump Tax, the Adjustable Swing, the Bullpen Ledger — each time finding the same shape: the underlying trait carries over, the result it produces mostly doesn’t. It is a tidy story. It is also, we now think, the wrong explanation for the right observation.
Because the trait-versus-outcome line breaks the moment you look across enough stats at once. So we did the boring, load-bearing thing: we took 44 different measurements — pitching and hitting, “skills” and “results” — lined up each one’s 2024 against its 2025, and asked which ones actually repeat. The answer is not the one the theory predicts.
1. An outcome that is more real than a skill
Start with the pair that broke our own frame. A batter’s strikeout rate is an outcome — the end of a plate appearance, the thing that shows up in the box score. A pitcher’s called-and-swinging-strike rate (CSW%) is about as close to raw “stuff” as a rate gets — it lives on the pitch, before anything is decided. The trait theory says the stuff metric should be the sticky one.
It isn’t. Batter strikeout rate repeats year to year at 0.73; pitcher CSW% at 0.49 — and the strikeout rate does it on about 367 plate appearances against CSW’s 978 pitches. The outcome out-repeats the input, on fewer chances. Now widen the lens to all 44.
Every stat, sorted by how well it repeats, each bar colored green for a trait/input or red for an outcome/output. If “is it a skill or a result?” governed reliability, the colors would separate — green on top, red on the bottom. They don’t. Batter strikeout rate, ground-ball rate and contact quality (all outcomes) sit above pitcher CSW%, zone rate and chase rate (all inputs). The category doesn’t sort the ladder. Something else does.
2. What actually sorts it: how much each event tells you
The missing variable is not glamorous. It is the amount of real, between-player signal packed into a single event of the stat — one pitch, one swing, one batted ball, one plate appearance. Call it the per-event signal. A radar-gun reading of velocity is almost pure signal: two pitchers who differ in velocity differ on essentially every fastball. A single called strike is almost pure noise about CSW% skill: whether one borderline pitch is taken for a strike barely separates two pitchers at all, so you need to stack up hundreds before a real difference emerges.
That single quantity — per-event signal — times how many events you get is, by a century-old result in measurement theory, what fixes reliability. It is not a new law; it is the reliability formula, applied across a baseball ladder instead of a psychology test. What it buys you is a concrete question for any leaderboard: how many events does this stat need before it’s even a coin flip’s worth reliable?
The range is enormous — note the log scale. A pitcher’s velocity crosses the 0.5 line on a single fastball; his CSW% needs about 445 pitches to get there; a batter’s BABIP needs some 590 batted balls, more than a full season. Same “pitching” and “hitting” buckets, four orders of magnitude of difference — set entirely by per-event signal, not by category. (Observed year-over-year numbers run a little under the pure-sampling line, because players also genuinely change between seasons; the ranking is what per-event signal nails.)
A fair objection: velocity is near the top because a radar gun is a near-perfect instrument — that reliability is half measurement precision. True, and worth saying plainly. But the effect is not confined to the trivial top of the ladder. It reorders the messy middle too, where CSW% — a real, hard-won stuff metric — lands below a batter’s strikeout rate purely because each of its events carries so little.
3. It predicts leaderboards it has never seen
A story about “per-event signal” is only worth telling if the signal can be measured on some players and then correctly forecast how a stat behaves for different players. That is the difference between a re-description and a finding, so we built it as a hard out-of-sample test. We split every player pool in half at random. On one half we measured each stat’s per-event signal and nothing else. On the other, entirely separate half, we measured how well each stat actually repeated. Then we asked whether the first predicts the second.
Each dot is one of the 44 stats: predicted reliability from per-event signal (measured on players held completely out) along the bottom, observed reliability up the side. The points hug the diagonal at 0.96 — per-event signal, learned from one set of players, forecasts how a leaderboard repeats for a disjoint set. And the two colors are scattered all along the line, not stacked at opposite ends: once you know a stat’s per-event signal, being told it’s a “trait” or an “outcome” adds nothing.
That last clause is the whole result, and we pre-registered it before looking. In a locked, out-of-sample race between models, per-event signal cut prediction error roughly in half over knowing only the event count (−0.072 in mean error, interval clear of zero). Adding the trait/outcome label on top of per-event signal did nothing — the improvement was +0.003, its interval straddling zero from both directions, and a flexible machine-learning model asked the same question found the same nothing. Per-event signal absorbs the category; the category does not absorb it.
4. The exceptions were the rule all along
The nice thing about the right variable is that the apparent counterexamples stop being embarrassing and start being predictions. We have published two of them ourselves. In Reading a Team at the Break, a team’s run differential — a pure box-score outcome — forecasts its second half at 0.56, better than any fancy expected-stat we tried. In The Fielder’s Fingerprint, a fielder’s total Outs Above Average — the literal result — repeats at 0.53, twice as well as the directional tendency buried inside it.
Under the trait theory these are contradictions. Under per-event signal they are exactly what you’d expect: run differential aggregates some seven hundred runs across half a season, total OAA a few hundred fielding chances. Pile up enough events and even a noisy-per-event outcome becomes one of the most reliable numbers on the board. Meanwhile the stats our earlier pieces flagged as “noise” — a reliever’s win-probability ledger, a runner’s stolen-base value — are the ones measured over mere dozens of high-leverage events. We kept calling that pattern “traits repeat, results don’t.” It was a lossy shorthand for “a lot of events, or a lot of signal per event.”
5. Why you can trust this: two methods, made to fight
This one nearly fooled us, so the safeguards matter. We ran the analysis twice, with two models that share almost no assumptions and never spoke during the work — one built per-event signal from interpretable variance components and a linear reliability model; the other estimated it by brute-force resampling and let gradient-boosted trees do the prediction. Each locked its predictions in a hashed, timestamped file before computing a single reliability number.
The honest part: in a first pass, the two methods reached opposite verdicts — one read the ladder as trait-versus-outcome, the other as pure event-count. Only when each was forced to hold its per-event-signal estimate out of sample, widen the panel, and referee the other’s code line by line did they converge — both landing on the same answer, each re-running the other’s scripts to check rather than take them on faith. A finding that survives its two authors starting on opposite sides and trying to break each other is a finding worth printing. One difference remains, and it is instructive: in a wide panel that includes fielding and baserunning, the trait/outcome label does beat raw event-count — because in that panel the label is standing in for per-event signal it hasn’t been told. Add the signal and the label goes inert again. The category is a shadow of the real thing; useful, lossy, and not the cause.
6. What this is, and what it is not
The takeaway is a single question you can put to any leaderboard on any broadcast: how many events is this measured on, and how much does each one tell you? A league-leader in a stat that clears its coin-flip line by opening day is telling you something durable. A league-leader in a stat that never reaches that line all season is telling you mostly about this season. That is a more useful instrument than sorting stats into “skills” and “luck.”
And it is worth being clear about what we are not claiming. Per-event signal is a property of a measurement, not a verdict on a player — a stat that doesn’t repeat at a half-season isn’t “random,” it’s undersampled, and more events measurably close the gap. High per-event signal is not the same as something a coach can teach; we measured what persists, not what responds to instruction. And our hard test held players out, not seasons — every number here lives inside the same 2024–2025 tracking era. Whether per-event signal forecasts a leaderboard a decade forward, or on a different measurement system, is the next test, not this one. What this piece establishes is narrower and, we think, sturdier: across a wide ladder of baseball stats, reliability is set by how much each event tells you — and the skill-versus-result frame we all reach for, ourselves included, is a shadow of that, not the substance.
Methodology
How we built and stress-tested this
Data. Full-season Statcast for 2024 and 2025, aggregated to per-player-season values for 44 stats spanning pitching and hitting, inputs and outputs, at four event scales (pitch, swing, batted ball, plate appearance). A second, independent analysis widened the panel to 50 by adding fielding, baserunning and team stats with proxy per-event definitions. Every stat’s “event” is the denominator its season value averages over, so its sampling variance is the within-player variance divided by the event count.
Per-event signal. For each stat we estimated λ, the ratio of true between-player variance to per-event noise variance, from a one-way random-effects decomposition. Reliability implied at n events is the Spearman-Brown form nλ/(1+nλ); the “events to a coin flip” count is 1/λ. This is classical test theory, reported as itself — not a new hypothesis.
The out-of-sample test. Players were split into two disjoint halves by a fixed hash of player id. λ and event counts were estimated on one half only; observed reliability (year-over-year rank correlation, 2024→2025) was computed on the other. Four nested models — event count only; count + category; count + λ; count + λ + category — were compared by leave-one-stat-out prediction error, with the whole race pre-registered (hashed, timestamped) before any reliability number was computed. The decisive contrast is whether adding category to a model that already knows λ lowers out-of-sample error: it did not (+0.003 [−0.008, +0.014] in one analysis, −0.003 [−0.019, +0.015] in the other — both straddle zero, under both a linear model and gradient-boosted trees). λ beat count alone by −0.072 [−0.097, −0.048].
Two divergent methods. Agent A (interpretability): analytic variance components, a linear reliability model on a logit link, bootstrap intervals over stats. Agent B (machine learning): empirical resampling curves for λ, gradient-boosted models with held-out-stat cross-validation. They reached opposite first-round verdicts and converged only after each held λ out of sample, widened the panel, and reviewed the other’s code. Intervals bootstrap the player, never the row; different estimators are labeled and never pooled.
Limitations. The hold-out is across players within one tracking era, not across future seasons or measurement systems — a future-season transport test is follow-up, not something this piece claims. The two panels disagree on whether category beats raw event-count (a panel-composition effect: the wider panel’s category label proxies for per-event signal it isn’t given). A handful of pitcher contact-quality and win-value stats repeat across seasons noticeably worse than their within-season sampling would predict; we report that gap rather than smoothing it, and we do not attribute it to a specific cause. Velocity and release metrics sit near the top partly because tracking measures them almost perfectly; the reordering of the middle of the ladder does not depend on them.
Cite this analysis
CalledThird. "The Coin-Flip Line." CalledThird.com, July 16, 2026. https://calledthird.com/analysis/the-coin-flip-line
All CalledThird analysis is original research. If you reference our findings, data, or charts in your work, please link back to the original article. For data inquiries: hello@calledthird.com