MES · full history May 2019 – Aug 31 2026 · 1,887 trading days · five P1 books · walk-forward tested · Codex verdict: refuted as an edge, confirmed as arithmetic

P1 Regime Overlay

A simulator was built that lets a daily rule choose, at each close, whether to pause or which P1 method to trade next day. About 2,600 candidate rules were searched. This is what survived honest out-of-sample testing, what did not, and why the independent reviewer refused to call the best rule a proven edge.

The best rule the search found, in words
Trade P1 Runner M3. At the close, look at the 20-day realized volatility. If it is above roughly 17%, take no trades tomorrow. Otherwise trade normally.
  1. After the 16:55 ET flat, compute the annualized standard deviation of the last 20 daily log returns (prop-firm trading days, full session).
  2. Above the cut (walk-forward fits landed between 16.3% and 19.8%; full-history fit 17.2%): tomorrow is a no-trade day. No hysteresis, no minimum hold; those were tested and lost.
  3. Below the cut: take every Runner M3 signal tomorrow as usual.
Full history, total R
vs Runner M3 alone every day. Same signals, 35% fewer trades.
Out-of-sample, 2021–2026
Anchored walk-forward, rule re-fitted each year on prior data only. Hurdle = best single method with hindsight.
Worst drawdown
Peak to trough in R, full history. Year by year the cut held in 2022, 2023 and 2025, tied in 2021 and 2024, and was slightly worse in 2026.
Chance this is luck
Within its own family. Repeating the choice among 14 families on shuffled histories gives 0.02 to 0.05; a bootstrap with year-sized blocks gives 0.17. Not established.

In dollars at the books' own sizing ($800 risk per trade, 40-contract cap): over seven years against for Runner M3 unfiltered, with the largest end-of-day drawdown falling from to . Sim fills, not live.

Independent verdict, in plain words
The numbers are real arithmetic. The edge is not proven. Nine tenths of the extra R comes from one year, 2022, and the search itself was shaped by looking at the test years.
  1. What holds: no complex regime rule (maps, trend filters, momentum, a learned model) survived walk-forward. A volatility pause on a fixed method is the only thing that did not collapse out-of-sample.
  2. What holds: the pause removes 2022-style bleeding. In that year it turned a −31.5 R year for Runner M3 into −3.6 R, with a 4 R drawdown instead of 32 R.
  3. What does not hold: extra R in normal years. Outside 2022 the pause adds about +3 R over seven years, and in 2026 so far its drawdown is slightly worse than unfiltered. Treat it as insurance against a bad regime, not as a way to make more.

What was built and how it was tested

The simulator

A feature layer computes ~25 market measures per trading day using only bars up to the prior day's close (proven by a test that alters day D's bars and checks day D's features do not move). A policy picks PAUSE or one of five P1 books each day. Trading a single book every day reproduces its stored DB result to the cent.

What was searched

  • Pause thresholds on 8 features × 15 levels × 5 methods
  • Regime maps: 1–2 features cut into bins, each bin assigned pause or a method
  • Volatility × trend maps, with and without minimum hold
  • "Trade whatever has been working" momentum rules
  • A small ridge-regression model
  • Plus the fixed rules from the earlier study

The honesty layer

  • Anchored walk-forward: fit on years up to Y−1, test on Y, 2021–2026
  • Rolling 24-month fit / 6-month test
  • 200-permutation p-value against month-shuffled histories
  • Fragility: result when the threshold moves one grid step
  • Result without the best 5 trades
  • Codex (GPT 5.6 Sol) adversarial review at two milestones

What survived and what did not

Every rule family, in-sample fit against out-of-sample walk-forward. The gap between the two bars is overfitting. Complex rules fit history beautifully and then fail. The one thing that holds is a volatility pause on a fixed method.

In-sample fit (whole history)Out-of-sample (anchored walk-forward)Hurdle: best single method with hindsight (+54.7 R)Fair hurdle: best single method chosen without hindsight (+22.4 R)
FamilyCandidatesIn-sample ROOS R (anchored)OOS R (rolling)OOS max DDOOS ex-top-5Perm pFragility −1 / +1
Read the ridge model row. It fits +88.7 R in-sample and delivers +22.7 out-of-sample, below even the fair hurdle. The regime maps (+143.7 in-sample) do the same. Whenever a rule has many knobs, history is generous and the future is not. The champion family examined 15 thresholds, of which 8 passed the pause cap.

Equity curves

Cumulative R by month, full history. Shaded columns are months where the champion paused on more than half the days. Hover for values.

Champion: Runner M3 + vol pauseRunner M3 alwaysMaster M2 alwaysBest single method chosen each year without hindsightChampion mostly paused

Walk-forward, year by year

Each test year, the rule was fitted on data through the prior year only, then run blind. The threshold it chose, how many days it paused, and what the best single method that year would have made (a hindsight number the rule did not have).

Test yearThreshold fittedDays pausedChampion RRunner M3 RMaster M2 RBest single that year

The R edge is concentrated in 2022, the rate-hike bear, where the pause avoided most of a −31 R year. In other years the rule is within a few R of Runner M3, sometimes below. Codex checked the drawdown claim year by year: the pause cut it in 2022 (4.0 vs 31.9 R), 2023 (9.8 vs 10.8) and 2025 (6.3 vs 11.7), tied it in 2021 and 2024, and was slightly worse in 2026 (9.0 vs 6.7). Leaving 2022 out, the worst drawdown of both is the same 16.4 R. That is the honest shape of this rule: it does nothing in good years and stops the bleeding in one kind of bad year.

Yearly results, all methods

YearChampionRunner M3Master M2Fast reBreakFast no-rearmFast RSI-off

How sensitive is the threshold?

Total R (full history) for Runner M3 and Master M2 with a fixed pause threshold from 10% to 24%, no fitting. A rule that only works at one exact number is a coincidence. This one has a broad plateau from about 15% to 21%.

Outlined = full-history fitted cut (17%)Colour: green positive, deeper = more R

Single methods, for reference

MethodTradesFull RFull max DDNet $ (at $800 risk)R pre Aug-2023R post Aug-2023OOS-window ROOS max DD

Runner M3 and Master M2 are the two strongest books and the only two where the pause rule clears the hurdle. P1 Fast reBreak with the pause improves a lot on itself (+4 to +49 R) but still does not beat Runner M3 unfiltered, so it is not the champion for a live account that can pick its method.

Independent review

Milestone 1 · foundations (no lookahead, clock, engine reproduces books)
VERDICT: REFUTED - vol_tier_fixed embeds full-sample 2011–2026 ES tertiles, violating the literal prior-day-only claim.

The one refuted point: a helper feature carried volatility cut-offs computed from the whole 2011–2026 ES sample. The champion never uses that feature, but it was removed from the searchable set and kept only as a labelled reference. Every other check passed: the clock mapping across both DST regimes, the momentum columns, the five book reproductions, and determinism.

Milestone 2 · validation and champion
VERDICT: REFUTED - the arithmetic is reproducible, but the selected champion has no untouched OOS validation and its p-values do not account for adaptive family selection or year-scale dependence.

What Codex confirmed: no fit-to-test leakage, the champion and every baseline cover the same 1,460 days, the database reconciliation reproduces the out-of-sample R, trade count and drawdown exactly, the five largest winners have entry and exit prices inside their one-minute bars, and the database verify passed 16 of 16. What it refuted: the claim that the advantage is a validated edge. The ten findings and the responses to each are recorded in the study's CODEX-REVIEWS.md.

What to do with this

  1. This is a study result, not a live change, and not a proven edge. Per the account protocol, the pause would become a pre-registered engine knob and go through the grid and kill bars before touching any deployed method. Given the reviewer's verdict, the case for it is drawdown insurance, not extra profit.
  2. The rule is one number on one feature. If you adopt it, use a fixed cut around 17% and do not tune it monthly. The plateau is wide; the exact value is not the edge.
  3. Expect three weeks of lag. A 20-day measure notices a shift about three weeks in. This rule never saves you from the first bad week. It saves you from weeks two to twelve.
  4. Runner M3 first, Master M2 as the lower-drawdown alternative. Master M2 with the same pause makes less but has the smallest drawdown of anything tested (about 12 R).
  5. Do not add features. Trend filters, method switching, momentum and learned models were all tested and all lost out-of-sample. More knobs made it worse every time.

Caveats