Every July slump gets the same autopsy on television: the clubhouse is tight, the lineup is pressing, they’ve forgotten how to win close games. We wanted the boring version instead. Take every bad three-week stretch a team has had since 2021 — there are 491 of them — measure everything about each one (was it the bats or the arms? real decline or bad sequencing? close losses or blowouts?), and then check what actually happened over the next thirty games. If slumps have anatomy, the anatomy should predict something.

It mostly doesn’t. Across 491 slumps, no detail of the slump itself — its depth, its luck, its close-loss count, its bullpen meltdowns — added any out-of-sample forecasting power beyond two numbers the team already owned: its season record and its season-long expected-contact differential. We tried to break that null two independent ways, including turning a gradient-boosted model with 33 engineered features loose on it. The honest measured lift from everything slump-shaped: +0.000, interval [−.019, +.021]. The slump is the noise. The season is the signal.

Everything this piece measures now runs as a living tool: The Form Table re-computes the 15-game form-vs-luck decomposition for all 30 teams every night, with forecasts for whichever teams are in the slump window that morning.

Four slumps, three diseases

Start with July 2026. The standings said four teams were collapsing at once. The process data says they were doing four different things.

Fifteen games in July: who actually got worse-0.08-0.040.00+0.04+0.08-0.06-0.030.00+0.03+0.06playing better + luckyplaying better, unluckydeclined + luckydeclined + unluckyBAL — 9-6 in window; form +0.047, luck -0.019SD — 7-8 in window; form +0.052, luck -0.033LAD — 8-7 in window; form -0.013, luck -0.043DET — 11-4 in window; form +0.002, luck +0.009SF — 6-9 in window; form +0.031, luck -0.010AZ — 11-4 in window; form +0.047, luck +0.025CIN — 8-7 in window; form +0.042, luck +0.009WSH — 6-9 in window; form +0.032, luck -0.000MIL — 9-6 in window; form -0.016, luck +0.006NYY — 9-6 in window; form -0.015, luck -0.002NYM — 7-8 in window; form -0.007, luck -0.020PIT — 9-6 in window; form -0.013, luck +0.029CWS — 9-6 in window; form -0.018, luck -0.007HOU — 8-7 in window; form -0.006, luck +0.021TB — 8-7 in window; form +0.000, luck +0.050TEX — 7-8 in window; form -0.017, luck -0.004LAA — 5-10 in window; form -0.007, luck +0.015MIN — 8-7 in window; form -0.007, luck -0.001ATL — 9-6 in window; form -0.037, luck +0.042CHC — 9-6 in window; form -0.024, luck +0.037KC — 7-8 in window; form -0.032, luck +0.046COL — 6-9 in window; form -0.019, luck -0.003BOS — 14-1 in window; form +0.075, luck +0.010BOSCLE — 7-8 in window; form +0.046, luck -0.031CLEMIA — 5-10 in window; form +0.018, luck +0.002MIASEA — 6-9 in window; form -0.010, luck -0.003SEAPHI — 6-9 in window; form -0.019, luck -0.044PHITOR — 5-10 in window; form -0.011, luck -0.025TORSTL — 5-10 in window; form -0.013, luck -0.048STLATH — 3-12 in window; form -0.065, luck -0.020ATHProcess change vs season baseline (xwOBA diff) →Results vs process (luck) →Right of center, the underlying process actually improved over the last 15 games; below center, results ran behind that process.

Each dot is a team’s last 15 games through July 24, measured two ways: did the underlying play change (horizontal), and did the results run ahead of or behind that play (vertical). The interesting teams are the ones far from where the standings put them.

Miami’s streak is the strangest object in baseball right now: eleven straight losses through July 25, the first ten all by three runs or fewer, and an 0-4 record in streak games they led or were tied after six innings. The bats genuinely froze — expected wOBA fell .313 → .284 during the streak, bottom-five in baseball over that window — but the pitching never broke: the staff’s expected contact allowed barely moved (.310 → .316) while its results ran 22 points worse than that. Half real cold spell, half coin-flips landing tails eleven times.

Seattle is the ordinary version: hitters running 22 points of expected wOBA below their own established level, run prevention fine. (At our data freeze the Mariners had, absurdly, outscored opponents while going 6-9; Saturday’s 1-7 loss in Texas finally pushed the window run differential negative. That fact was never stable enough to build on — which is rather the point of this piece.) Toronto is the counterintuitive one: the Jays’ blowout losses look like a pitching collapse, but the staff’s underlying contact quality improved during the slide — results simply ran 24 points hot against it while the offense supplied nothing. And the Athletics are what an actual collapse looks like: the largest process decline in baseball (−.065), driven by a bullpen whose expected wOBA allowed went .315 → .395.

The control case matters too: Boston’s 14-1 was the league’s best process differential over the window (+.067), both rotation and bullpen improving, with almost no luck in it (+.010) — a whole-roster leap, not a heater. (Baseball being baseball, they were promptly shut out by Toronto the night we froze the data.)

What predicts the next 30 games

So which slumps do teams escape? We anchored on every 15-game window since 2021 in which a team won six or fewer (491 events, all 30 teams represented), measured each one the same way we measured July’s, and raced models on the next 30 games:

ModelOut-of-sample R²
Season record at the time of the slump.156
+ season expected-wOBA differential.170
+ how cold the bats were vs. their norm.175
+ every other slump detail (13 features).158worse
33-feature machine-learning attack.170 — no lift

That last row deserves a sentence, because a null is only as good as the effort to kill it. Two frontier models built this analysis independently — separate code, separate methods, forbidden from reading each other until done — and one of them existed purely to attack the null with engineered features (window volatility, opponent strength, momentum, home/away mix) under nested cross-validation. It found nothing transportable. When both the structure-first model and the kitchen-sink model land on the same shrug, we believe the shrug.

Forecast honesty, up front: even the winning model explains about 17% of the variance in a team’s next 30 games, with a typical miss of ±.086 of win percentage. That’s the ceiling. A month of baseball is mostly noise, and anyone selling you a sharper slump forecast is selling.

The things that predict nothing

The useful product of a strong null is the list of stories it retires.

What a slump does NOT tell youWas the slump ‘unlucky’? Doesn’t matter.0.25.50.75Most unlucky: bounce rate 53.7% (n=123), 95% CI 44.9%–62.2%most unluckyQ2: bounce rate 51.2% (n=123), 95% CI 42.5%–59.9%q2Q3: bounce rate 40.2% (n=122), 95% CI 31.9%–49.0%q3Most earned: bounce rate 49.6% (n=123), 95% CI 40.9%–58.3%most earnedAll slumps: bounce rate 48.7%all slumps.487Q1−Q4 gap: 4pp, CI spans zero;Q3 is the worst — noiseshare ≥.500 over next 30 →slump luck: most unlucky → most earnedThe ‘contender gap’ is identity, not information010203040Pre-registered gate: a gap must clear 10pp to count as informationpre-registered gateNaive splitNaive split: gap 23.3pp, 95% CI 11.8 to 34.9pp+23.3One event per team-seasonOne event per team-season: gap 21.4pp, 95% CI 4.8 to 37.3pp+21.4Hierarchical (identity pooled)Hierarchical (identity pooled): gap 10.9pp, 95% CI -0.6 to 22.9pp; 46% of bootstrap draws fall below the 10pp gate+10.9contender − non-contender bounce gap (pp)How a slump happened predicts nothing: teams losing close games (n=93) bounced back 48.4% of the timevs 48.7% for teams getting blown out (n=398) — identical.

Left: slumps sorted by how “unlucky” they were (results running behind underlying play). Bounce-back rates are flat — and non-monotonic, with the third quartile worst, which is what noise looks like. Right: the famous “good teams bounce back” gap, shrinking as you control for the fact that good teams are simply good.

“They’re in every game” is worth nothing. Teams whose slump losses were mostly close bounced back at 48.4%; teams getting blown out, 48.7%. Identical. “They’re due” is worth nothing. The most process-unlucky quartile of slumps outbounced the most “earned” quartile by four points with an interval spanning zero. And the one we half-expected to survive — contenders bounce back — turns out to be an accounting trick: the raw gap is 24 points, but a hierarchical model that stops counting the 2024 White Sox six separate times cuts it to 10.9 [−0.6, 22.9], and a within/between decomposition shows the information lives entirely at the level of which team you are, not that you’re slumping. Good teams bounce because they’re good. The slump added nothing. One sobering corollary survives everything: slumping contenders return to winning, but not to their prior pace — next-30 records run .038 below their pre-slump season mark. Regression, not restoration.

The lone survivor: cold bats revert. Slumps driven by an offense running below its own established level — Seattle’s kind, Miami’s kind — predicted better next-30s in every specification we and the attacking model ran (direction negative in 14 of 14, roughly 1 point of win percentage per standard deviation of bat-chill). It’s small, and its magnitude wobbles across eras, so we’ll say it the careful way: if you must buy a slumping team, buy the one whose hitters stopped hitting. Broken pitching carries no such promise.

The streak cliff, read honestly

Here is where we tell on ourselves. Our first pass at the losing-streak question produced a terrifying table: teams whose streaks reached nine-plus games went on to play .366 ball with a 13% bounce rate. We nearly led with it. It’s also survivorship arithmetic — that table sorts streaks by how long they eventually got, which no one standing inside a streak gets to know. Both of our independent analyses killed it within a day of each other.

The streak cliff is mostly composition.45.50.55.60.65.700123456789+Upper bound: the largest 9+ streak effect the data can't rule out is +7.2pp over the no-streak baseline of .497Identity-controlled P(lose next): .497 at streak 0 to .520 at 9+ with same team, opponent & venue controls (joint p=.43)Streak 0: P(lose next game) = .483 (95% CI .473–.492), 10,643 gamesStreak 1: P(lose next game) = .506 (95% CI .492–.520), 5,129 gamesStreak 2: P(lose next game) = .516 (95% CI .497–.535), 2,604 gamesStreak 3: P(lose next game) = .516 (95% CI .490–.543), 1,340 gamesStreak 4: P(lose next game) = .547 (95% CI .510–.584), 694 gamesStreak 5: P(lose next game) = .522 (95% CI .472–.572), 379 gamesStreak 6: P(lose next game) = .528 (95% CI .458–.596), 199 gamesStreak 7: P(lose next game) = .654 (95% CI .560–.738), 107 gamesStreak 8: P(lose next game) = .549 (95% CI .434–.659), 71 gamesStreak 9+: P(lose next game) = .664 (95% CI .574–.743), 116 gamesraw P(lose next game)largest effect thedata can't rule outsame team, opponent& venue controlsraw: .48 → .66controlled:.50 → .52 (p=.43)Current losing streak (consecutive losses) →P(lose next game) →The raw climb is teams that were already bad — conditioned on who's playing, streak length adds nothing the data can resolve, though a small −0.46pp-per-loss drag on the next 30 games survives.

Red: the raw chance of losing your next game, by how many you’ve already lost — it climbs from .48 to .66 and looks like doom. Dashed: the same question after controlling for who’s playing whom, and where. The cliff is mostly who has long streaks, not what streaks do.

Conditioned on team quality, opponent, and venue, streak length adds no statistically resolvable change to the odds of losing the next game (p .29–.43 across control sets; the data can’t rule out effects bigger than ~7 points at nine losses, so we say “unresolved,” not “disproven”). What does survive, in every clean specification, is a small drag on the month: about −0.46 points of next-30 win percentage per consecutive loss. Eleven losses in, that’s a real but modest tax — on the order of one extra loss over the next thirty games, not a death sentence. Beyond twelve losses the historical record simply runs out: 39 team-games from seven team-seasons, most of them the 2024 White Sox. Nobody knows what L13 does, including us.

The forecast, on the record

Which brings us to the payoff. Here’s the identity model’s next-30 forecast for every team in July’s slump conversation, with uncertainty bands sized so that 80% of historical forecasts landed inside them — and a public promise to re-grade this chart in thirty days.

The next 30 games, forecast honestly80% band = ±.145 · ordering mostly unresolved:P(PHI beats SEA next 30) = .57.30.40.50.60.70.500Phillies 6-9Phillies (PHI) 55-48 — 6-9 over the window; predicted next-30 win% .545, 80% band .400–.690Phillies (PHI) 55-48 — 6-9 over the window; predicted next-30 win% .545, 80% band .400–.690.545Mariners 6-9Mariners (SEA) 51-52 — 6-9 over the window; predicted next-30 win% .522, 80% band .377–.668Mariners (SEA) 51-52 — 6-9 over the window; predicted next-30 win% .522, 80% band .377–.668.522Cardinals 5-10Cardinals (STL) 51-51 — 5-10 over the window; predicted next-30 win% .502, 80% band .357–.647Cardinals (STL) 51-51 — 5-10 over the window; predicted next-30 win% .502, 80% band .357–.647.502Marlins 5-10Marlins (MIA) 52-52 — 5-10 over the window; predicted next-30 win% .489, 80% band .344–.634; cross-model consensus .440–.520Marlins (MIA) 52-52 — 5-10 over the window; predicted next-30 win% .489, 80% band .344–.634; cross-model consensus .440–.520consensus .44–.52Marlins cross-model consensus: point .480, band .440–.520, bounce range .350–.500.489Guardians 7-8 *Guardians (CLE) 53-51 — 7-8 over the window; predicted next-30 win% .485, 80% band .340–.630; outside slump support (shown for context)Guardians (CLE) 53-51 — 7-8 over the window; predicted next-30 win% .485, 80% band .340–.630; outside slump support (shown for context).485Red Sox 14-1 *Red Sox (BOS) 52-49 — 14-1 over the window; predicted next-30 win% .485, 80% band .340–.630; outside slump support (shown for context)Red Sox (BOS) 52-49 — 14-1 over the window; predicted next-30 win% .485, 80% band .340–.630; outside slump support (shown for context).485Blue Jays 5-10Blue Jays (TOR) 47-57 — 5-10 over the window; predicted next-30 win% .478, 80% band .333–.623Blue Jays (TOR) 47-57 — 5-10 over the window; predicted next-30 win% .478, 80% band .333–.623.478Athletics 3-12Athletics (ATH) 44-59 — 3-12 over the window; predicted next-30 win% .474, 80% band .329–.619Athletics (ATH) 44-59 — 3-12 over the window; predicted next-30 win% .474, 80% band .329–.619.474Predicted next-30 win% →Dots are point forecasts from the identity model (record + xwOBA diff + cold-bats term); bands show where80% of past forecasts landed; * = outside the model's slump support, shown for context.

Forecasts frozen on July 24 data. Hollow dots sit outside the model’s slump definition (Cleveland was 7-8; Boston is a surge, scored for fun) — context, not forecasts. The ordering is soft: the model says PHI over SEA at only 57/43.

Philadelphia gets the best forecast (.545) for reasons that are almost embarrassingly simple: it’s a good team, its season-long process is intact, its staff’s ugly three weeks were the most results-behind-process stretch in baseball, and its bats are the cold kind that revert. Seattle (.522) is the same story at a lower altitude. The Athletics (.474) get the worst read because theirs is the one slump the process data confirms.

And Miami? The model that can’t see the streak says .489 — a coin-flip team. The streak-aware analyses, once we fixed both of our first attempts (one simulator flunked its own history exam; one tree model was extrapolating past its training data — details in the methodology), converge on .48, honest band .44–.52, with a 35–50% chance of playing .500-or-better ball the rest of the month. The closest historical comparables — 24 team-seasons that hit seven-plus straight losses while holding a solid record — played .482 afterward and bounced 11 times out of 24. The eleven-game streak is real, agonizing, franchise-record-tying — and worth about one win of forecast. What it’s actually done is what results always do: moved the deadline calculus. That’s a decision about this month’s evidence, and this month’s evidence is, on the record above, the least informative kind there is.

(The nightly-updated version of this scoreboard lives on The Form Table.)

One clear takeaway for the fan watching tonight: when your team drops its fourth straight, check two numbers — the season record and the season expected-wOBA differential. If those are intact, you are watching variance wearing a story. If the bats are the problem, history is mildly on your side. Everything else in the postgame autopsy is, as far as five seasons of data can tell, narration.

Methodology

Data, definitions, models, and the two numbers we withdrew

Data. Pitch-level Statcast, 2021 through July 24, 2026, aggregated to a per-team-game panel (27,368 team-games): wOBA and expected wOBA (xwOBA) for and against, split by starter/reliever, plus scores and results. Expected wOBA uses Statcast’s contact-quality estimate with actual values for non-batted-ball events. 2026 is scoring-only; no 2026 games enter any training fold. Game results verified against the MLB Stats API (zero mismatches on audited spans).

Slump events. A slump window is the first game at which a team’s trailing 15-game record hits 6 wins or fewer, with at least 20 games of season before the window and 20 after, and 20-game spacing between events (491 events, 2021–2025). Streak events anchor at a sixth consecutive loss (138). Outcome: win percentage over the next 30 games. Three fully independent implementations of this extraction (a pilot and both research agents) matched to the event and the decimal.

Process. This piece ran as an adversarial two-model study: one interpretability-first analysis (hierarchical pooling, discrete-time hazard models, splines, cluster bootstraps) and one ML-engineering analysis (gradient boosting under nested grouped cross-validation, SHAP with permutation agreement, calibration audits), built blind to each other from a pre-registered brief with kill criteria, then cross-reviewed. The model race table reports grouped out-of-fold R² (folds grouped by team-season, so one bad team never trains on its own future).

Two withdrawn numbers, for the record. Our first Miami forecast (.518, 65% bounce) came from a forward simulator that failed a historical backtest — at eleven-loss states it predicted .446 where history delivered .345 — and was withdrawn. An alternative ML estimate (.438–.453) was extrapolating a single tree split beyond its training support (no team has ever been at eleven straight losses with Miami’s season record in the sample) and was withdrawn symmetrically. The published .48 [.44–.52] triangulates the window model, the identity-fed variants of both corrected models, and matched historical comparables. New house rule either way: no simulator publishes here again without passing its own history.

Limitations. A 30-game outcome is mostly noise even for the best model (R² ≈ .17; 80% of forecasts land within ±.145 of win percentage — the bands on the chart). The cold-bats effect’s direction is robust (14/14 specifications, both agents) but its magnitude is era-unstable and should not be priced precisely. The deep-streak tail (12+ losses) is unestimable from 2021–2025. The contender-gap and archetype tables are descriptive; their information collapses into team identity under pooling. Forecasts froze on July 24 data; games since are noted in the text but not re-fit. We will re-grade the scoreboard, publicly, 30 days from the freeze.

Share Twitter/X Bluesky

Cite this analysis

CalledThird. "The Anatomy of a Slump: What Eleven Straight Losses Actually Tell You." CalledThird.com, July 26, 2026. https://calledthird.com/analysis/the-anatomy-of-a-slump

All CalledThird analysis is original research. If you reference our findings, data, or charts in your work, please link back to the original article. For data inquiries: hello@calledthird.com