There’s a ritual in baseball analysis that mostly doesn’t happen: going back. A finding ships in April on six weeks of data, the season moves on, and nobody checks whether it survived. We think the going-back is the job — it’s why every early piece we publish carries a due date on a public queue. This week we paid all of them at once: five re-tests plus a checkpoint, each one re-running the original, already-reviewed methodology on roughly twice the data. No new modeling choices, no moved goalposts — the scripts from the original rounds, pointed at the break.
The honest headline: the scorecard cuts both ways, and that’s the point. Here it is, worst news first.
1. The walk-spike names didn’t survive — and our caveat was the finding
In May, The Walk Spike Is Fading named three relievers as the cleanest winners and losers of the new ABS zone. At 2.2× the data, none of the three clears the stability filter we ourselves pre-registered — their walk-rate swings regressed hard toward the mean (the most extreme, Riley O’Brien’s −8.3 points, is now −1.7). The zone-attribution share also softened, from +26% to +18% with an interval that now spans zero. We’ve posted the correction on the article itself, retired the named list, and declined to print the new names that emerged at 2× — naming a fresh trio after watching the first one dissolve would be learning nothing.
Two parts of that piece got stronger: the fading itself (the year-over-year walk gap is down to +0.50 points, and just +0.34 in the entirely new weeks) and the zone’s mechanical signature (top edge +0.33, bottom edge −0.51). And the one name we held out in May for failing the stability bar — Mason Miller — failed it again. The filter worked; we should have trusted it on the others too.
2. The bat-speed effect graduated
The Bat-Speed Arms Race shipped with a deliberate hedge: within a hitter, adding bat speed looked worth about +0.27 runs per 100 PA, but the interval touched zero, so we said “no treadmill” and refused to say “swinging harder wins.” At 1.6× the swings, the effect is +0.43 [+0.12, +0.74] — clear of zero at every sample floor we pre-registered, surviving the attrition correction, still with no whiff cost. The hedge is retired; the update is on the article. Delightfully, the best counter-anecdote also got stronger — Cam Smith added the most bat speed of any in-news hitter (+3.0 mph) and declined further — and it still doesn’t bend the group trend. That’s what a real but modest effect looks like: it doesn’t rescue every individual.
3. The April Sell List went 11-for-11
The most testable thing we published all spring: on April 25, a list of six hot starts to sell, six cold starts to buy, with frozen rest-of-season projections for each. Every one of the eleven names with a scoreable sample landed where the projection said — all six sells regressed (mean April-to-rest drop: 69 points of wOBA), and the scored buys gained (Caglianone +.109 and Basallo +.103 over their priors, both confidence-separated).
Each dot is a named hitter: our frozen April 25 projection along the bottom, what he actually hit afterward up the side. The cloud hugs the diagonal. Quantified: the projections beat “just extrapolate his April” as a forecast by a factor of three (mean error .015 vs .048 of wOBA, Wilcoxon p=.016) — and also beat “assume he’s exactly his old self” (.015 vs .039). April told you almost nothing; the projection knew that, and knew which Aprils were the exceptions.
Honest fine print: two names (Judge, House) have truncated windows — no MLB plate appearances since late May — and the one buy we couldn’t score (Pereira, 16 PA) looks bad in the tiny sample that exists. Murakami’s .391 since the list leans toward the “signal” read, but the NPB-translation caveat from the original piece still stands.
4. Umpires really are calling better games under ABS
Our early-season checkpoint claimed every umpire with 3+ games was beating his 2025 accuracy baseline, by about 2.6 points — on ten umpires. At the break it’s 79 of 80 umpires improved, +2.35 points [+2.17, +2.52] — and the early cohort, re-measured on its later games, didn’t fade at all (+0.21 late-vs-early, indistinguishable from zero). The league-wide accuracy level shifted from 92.5% to 94.9%, which regression to the mean cannot produce. The one umpire who declined? He was 2025’s most accurate — exactly the exception regression predicts, sitting right where it should.
One promised item converts to a note: the CB Bucknor re-check had a pre-registered gate of 20 games. He’s had two home-plate assignments all season — against 28 by this point last year. His accuracy question is unanswerable at that sample; the disappearance from plate duty is the more interesting question, and it isn’t answerable from our data. We’re holding it until we can source it properly.
5. The arm-angle null held — and its tax cleared the bar
The Arm-Angle Gambit found that dropping your arm slot buys nothing on average — a convergent null — with one suggestive cost: dropped slots lose four-seam carry, significant on one test but not the other. At the fuller window the null holds (slot credit +0.02 [−0.05, +0.07] runs per 100, dropper-vs-stable), and the carry tax now clears both tests (14 of 16 droppers lost carry; both p < .005). The quiet surprise: the dropper pool barely grew — 16 to 18, when we expected it to double. The gambit itself has stalled league-wide, which is its own kind of verdict on the trade.
And the 7-hole tax — which we failed to find six different ways — survived its two open threads: the stricter per-umpire model the review demanded flags 0 of 90 umpires with any spot-7 calling quirk, and the low-chase interaction — the one thread that pointed anywhere — still doesn’t clear zero (−0.17 [−0.42, +0.08]). Still not there, now four more ways.
What this buys you
Not every re-test flattered us, and publishing the one that didn’t is what makes the other five worth believing. The pattern across all six is the one our reliability work predicts: mechanisms and traits replicated (the zone’s shape signature, the bat-speed effect, umpire accuracy, the carry tax); small-sample leaderboards of names did not (the walk-spike trio). When you see us name individuals on a half-season of anything, hold us to the stability filter — it was right both times it spoke.
The queue continues: the umpire question gets its definitive end-of-season answer, the free-vs-bought bat-speed split gets a season-end look, and the reliability law itself faces its across-seasons test. Due dates are on the queue, as always.
Methodology
How the re-tests were run
Ground rule: every re-test re-ran the original round’s reviewed scripts and conventions on the extended window (2026 opening day through July 8), with no new modeling choices. Sample multipliers: walk spike 2.2× (405,511 pitches), bat speed 1.6× (186,496 tracked swings, 311-hitter primary panel), umpires 8× the umpire count (n=80 with 2025 baselines), sell list scored on post-list play April 26–July 8 under the project’s pre-existing holdout rule (50+ PA floor, 10,000-rep bootstrap). Verdict labels (held / strengthened / weakened) were assigned against each project’s pre-registered gates, not eyeballed. Full artifacts — scripts, logs, per-name and per-umpire tables — are banked per-project in the research tree (retest_2026break/), and the corrections described here are posted on the affected articles themselves.
Limitations: half-to-break windows throughout; the umpire improvement cannot be fully decomposed into challenge-system effect vs. secular trend without a no-ABS control season; two sell-list names have truncated windows; and the arm-angle dropper pool (18) remains small — that null is stable, not yet precise.
Cite this analysis
CalledThird. "The Re-Test Report: We Promised Six Follow-Ups. Here's What Held." CalledThird.com, July 25, 2026. https://calledthird.com/analysis/the-retest-report
All CalledThird analysis is original research. If you reference our findings, data, or charts in your work, please link back to the original article. For data inquiries: hello@calledthird.com