← all backtests  ·  original BandGrid study (2026-07-03)

Band family rolling walk-forward — final report

Four studies of the band strategy family (three variants + a dev-slope-gate follow-up) · Binance BTCUSDT 5-min venue-native prints · perp fees 2/4.5 bps · protocol: optimize on trailing 1 month → trade the winner 1 week → roll weekly (10 folds; 15 for the dev-gate study) · ext swept 2→10 · OOS 2026-04-20 → 06-29 (dev gate: 2026-03-16 → 06-29) · all KPIs from LEAN result JSON.

VERDICT: ALL FOUR STUDIES FAIL — family closed under the rolling re-optimization protocol
Mean-reversion grid: fee-eaten (gross PF 1.03) with a −1.70% tail week. Momentum scale-out: gross-negative, nothing to harvest. Momentum pyramid: −6.78% (pnl rule); its sharpe-rule "+5.25%" is one +15.35% trend week masking losses in 8 of 9 other weeks — fails the pre-registered single-week gate. Weekly re-optimization added no selection skill anywhere; filters were immaterial in every variant where testable — including the user-specified dev-slope gate (15-week follow-up): its dev-rising arm won 12/15 folds in-sample and still lost \u22122.27% forward.

Out-of-sample forward weeks — the actual test

Each bar: OOS return of trading that fold's in-sample winner for the following week (max-PnL rule). Hover for details; sharpe-rule differences are in the KPI table.

IS net PnL by config × fold — where the optimizer looked

· gray = 0 · values beyond range clamped · ring = fold winner

Fold-averaged IS net PnL by length × ext

Mean net PnL % per cell across the 10 fold windows. High ext bleeds less on average only because it trades less — it was never the in-window optimum for the mean-reversion grid, and for momentum there was nothing to find.

Table view — per-fold winners with forward-week results & ext summary

Written report

Background

The band family was previously killed twice: the original BandGrid mean-reversion grid's apparent edge (+13.7% OOS on Alpaca/Coinbase prints) failed cross-venue replication — its fills lived in wide wicks that never print on Binance — and the kill proved fee-insensitive (zero-fee full-span still −16.2%). The two inverse breakout constructions were killed on venue-native data (5-min BTC breakouts revert to the SMA ~98–100% of the time). One question remained open: could rolling re-optimization — refitting parameters every week on the trailing month, with a much broader expansion factor (ext 2→10) — track regimes fast enough to find pockets the fixed-parameter tests missed? This study was commissioned to answer exactly that.

Protocol

For each weekly fold: sweep the full config grid over the trailing calendar month of Binance BTCUSDT 5-minute venue-native prints at real USDT-M perp fees (2 bps maker / 4.5 bps taker), pick the winner under three pre-registered selection rules (max net PnL; max Sharpe with ≥5 trades; sharpe-or-flat which may sit the week out), then trade that winner for the following week. Stitch the forward weeks per rule. Pass requires: net-positive, PF>1, ≥30 trades, and not driven by a single week. In total 2,760 in-sample and 71 out-of-sample LEAN backtests across the four studies; every figure from LEAN result JSON.

Findings

1. Weekly re-optimization has no selection skill. In-sample winners were positive in nearly every fold window; their forward weeks netted negative under every rule for every variant. The optimizer cannot see regime changes coming — late-fold pyramid picks showing +9.8..+11.8%/month in-sample lost −1.2..−3.0% forward (29 of its 30 positive IS cells sit in folds whose trailing month contains the June trend).

2. Broad ext does not rescue the grid — it only dilutes it. Mean in-sample PnL improves monotonically from −0.54%/month at ext 2 to −0.02% at ext 10, but purely because trade count falls 410 → 55 per month. A high-ext cell was never the in-window optimum; max-IS selection therefore structurally loads the highest-churn ext-2 configs — the most fee-exposed corner of the grid.

3. The mean-reversion harvest is real but sub-cost; momentum has nothing to harvest. The grid's gross PF was 1.03 on closed trades with ~₮420 of fees against it (fee-eaten) — and the independent review found the equity-basis gross (including truncated open cycles) is ≈zero-to-negative, strengthening the kill. Both momentum variants were gross-negative (PF<1 before the fee story); scale-out's sharpe-or-flat rule chose to sit out 9 of 10 weeks.

4. One trending week is the family's fingerprint. Week 2026-06-01: mean-reversion −1.70%, scale-out +0.11%, pyramid +15.35%. The pyramid's headline +5.25% (sharpe rule) is that single week masking losses in 8 of the other 9 — it fails the pre-registered single-week gate. The only money made anywhere in this study was accidental trend exposure, consistent with the one edge already on the books (BTC daily trend), not a 5-minute band edge.

Addendum — dev-slope gate study (user-directed follow-up, 15 weeks)

A pre-registered follow-up tested the last untested filter: gating the grid's entries on the direction of the stdev's slope, measured over the same window as the SMA length, in both directions plus a no-gate control, ER/MTF off, over 15 weekly folds (2026-03-16 → 06-29, June trend week inside the span). Result: FAIL on all three rules (−2.27% / −2.07% / −2.45%). The optimizer preferred the dev-rising arm in 12 of 15 folds — a genuine in-sample improvement that did not transfer: forward weeks still netted negative. The June trend week was not dodged (the in-window winner that fold was a rising arm, −1.53% OOS; the falling arm that would have blocked it was never in-window optimal when it mattered). The control arm's bleed rate matched the original 10-week study, confirming its stability. This closes the family's last open filter question: in-sample-selected volatility gates are as untransferable as everything else here.

Independent review (2026-07-10)

A quant reviewer with no involvement in the studies recomputed every stitched total, Sharpe, PF and trade count from the raw ledgers: all 12 registry records tie out exactly; no fold-boundary lookahead; picks match the in-window argmax in 50/50 folds; the dev-gate implementation is clean. Errors found were cosmetic or kill-strengthening (registry fees understated ~3–4% by a currency-symbol parse bug; “gross PF 1.03” is closed-trades-only — equity-basis gross is negative). Reviewer’s verdict: FAILs confirmed at high confidence, further testing in this family is negative-EV — the binding constraint is a ≈0 gross harvest, which no gate, sizing or venue change multiplies. Full review: backtests/BandGridRollingWF/QUANT_REVIEW.md.

Conclusion

All four studies FAIL all three selection rules. Parameter instability and volatility gating were the family's last open questions and both are now answered: the band family (both directions, with and without filters, including dev-slope gates) is closed on 5-minute BTC free/venue data. A genuinely new edge requires different raw material — paid tick/order-book or funding/basis data, a different timeframe, or a different asset — not further variants of this family.

Reproducibility

Registry records {BandGridSmaExit,BandBreakoutScaleOut,BandBreakoutPyramid}__binance_wf1mo_w10_ext2-10_{pnl,sharpe,sharpeflat}__OOS and BandGridSmaExit__binance_wf1mo_w15_devgate_{pnl,sharpe,sharpeflat}__OOS in lean-strategy/backtests/registry · study artifacts & pre-registered PLAN in backtests/BandGridRollingWF · ledgers bwf/bbs/bbp/bdg_wf_weekly_{is,oos}.tsv · harness bandgrid_driver.py (wf mode).