MES · P1 Runner M3, Fast reBreak, Fast no-rearm · May 2019 – Aug 2026 · $1,000 and $5,000 starts · 4,032 cells · Codex-critiqued plan, three implementation reviews

Flexible-R Sizing Study

What happens when the risk per trade is a slice of the account instead of a fixed dollar amount: one micro contract at a time, starting from $1,000, every trade re-sized to the current balance. The question was "what is the perfect R?". The honest answer is a survival test, and it is mostly about account size.

The answer in one paragraph
At $1,000 the account cannot express a sensible risk fraction. At $5,000 it can, on paper: about 1/40th to 1/50th of equity per trade, only with the volatility pause on. But once fills are realistic, the edge that would be compounded does not survive. The account size question is secondary to that.
$1,000 · Runner M3 + pause
Chosen N under the pre-registered test. Median 7-year outcome across 2,000 resampled histories, and the historical path.
$5,000 · Runner M3 + pause
Same policy, five times the capital: the rounding tax disappears and the same rule compounds.
The fast lane, and its price
N = 10 at $1,000: historical end balance, and the bootstrap chance of a 50% drawdown on the way.
Trades skipped at $1,000, N = 20
Signals a $1,000 account could not afford at 1/20th risk. The account size, not the method, decides which trades you take.

The arithmetic of one micro

MES is $5 a point. A round trip costs $1.48 in commission. The median P1 stop, rounded outward to the tick as a live stop must be, is 5.5 points, so one micro risks about $29 all-in, which is 2.9% of a $1,000 account, roughly 1/34th. You can only take a trade when one contract fits inside your risk budget of equity ÷ N. That single fact drives everything below.

Your stopOne micro risksAffordable at N = 10 if equity ≥N = 20N = 40
3 points$16.50$165$330$660
5.5 points (median, tick-rounded)$29.00$290$580$1,160
6 points$31.50$315$630$1,259
10 points$51.50$515$1,030$2,059
25 points (widest)$127.00$1,270$2,540$5,079
Rounding is a filter, not a rounding error. A small account systematically drops the wide-stop trades and keeps the narrow-stop ones. At N = 20 on $1,000, 16% of the 7 to 10-point-stop trades were skipped and none of the sub-3-point ones. Some of what looks like "better sizing" at large N is this accidental filter.

What was tested

Sizing rule

  • Risk budget = equity immediately before the entry ÷ N
  • Contracts = floor(budget ÷ all-in cost per micro), where all-in = stop × $5 + $1.48 commission (+ slippage in stress arms)
  • Also capped by margin ($50 per micro, 20% buffer) and the 40-micro engine cap
  • Zero contracts = trade skipped
  • Daily stop: once the day's loss reaches 3 × (day-start equity ÷ N), no more trades that day

Grid

  • N ∈ {5, 8, 10, 12, 15, 20, 25, 30, 40, 50, 75, 100}
  • Start $1,000 / $2,500 / $5,000 / $10,000
  • Runner M3, Fast reBreak, Fast no-rearm, each with and without the RV20 pause (skip tomorrow if 20-day vol > 17.18%, fixed from the overlay study)
  • Master M2 excluded: its half-off exit cannot be executed with one or two micros and its stored trades cannot be rescaled

Survival test, written before the run

An N is allowed only if, across 2,000 resampled histories (intact month blocks, whole equity path recomputed every draw):

  • P(drawdown ≥ 50%) ≤ 5%
  • P(account destroyed) ≤ 1%
  • P(end below start) ≤ 35%

Among allowed N, take the highest median growth. If none is allowed: cash.

What the survival test picked

Chosen N per policy, with the bootstrap odds at that N. Rows without the pause pass only because at N = 75 to 100 they barely trade; do not read them as good sizing.

PolicyChosen NRisk / tradeMedian end5th pct95th pctMedian CAGRP(DD ≥ 50%)P(end < start)P(destroyed)Year-block choice
The block size moves the answer. Bootstraps with quarter blocks, which respect a regime that lasted all of 2022, push Runner M3 + pause at $1,000 over its own 5% line (P(DD ≥ 50%) 3.1% with month blocks, 5.2% with quarter blocks) and the choice moves to N = 75, the barely-trading region; year blocks bring it back to N = 40. At $5,000 the same policy sits at 2.0% to 4.9% depending on block size, right against the line. Read the region, N ≈ 40 to 75, never the exact number.

Growth versus survival, every N

The full picture for the three paused policies: what each N did on the historical path (growth multiple and worst drawdown) and what the bootstrap says about a 50% drawdown. The outlined cell is the survival test's pick. Hover for the numbers.

Top row of each policy: historical growth multiple (green above 1×, red below). Bottom row: bootstrap P(drawdown ≥ 50%) (red = high).

What the paths look like

Equity by month, historical path, Runner M3 + pause. Three sizes: the fast lane (N = 10), a middle setting (N = 20), and the survival test's pick. Log scale, because the fast lane goes to about $130 before it goes past $30,000.

Re-scored on independent draws

The survival test was applied to 288 policy × N × balance cells, so the winners were selected on the same 2,000 draws that scored them. Each chosen cell was re-scored on a second, independent set of 2,000 draws. This is a Monte Carlo stability check on the same history, not a selection-bias correction and not a holdout. A Bonferroni-style level for 288 cells would be about 0.0002; the test uses 5% and no correction is applied, so treat every pass as selection-conditional.

Chosen cellP(DD ≥ 50%) on selection drawsOn fresh drawsMedian end, independentStill passesRank of chosen N by growth (of 12)

The chosen N ranks 8th to 10th of 12 by growth for the paused books, stable across block sizes. That is the point: the constraint deliberately picks a low-growth N, not the best one. One cell, Fast reBreak + pause at $5,000, fails its own test on the independent draws.

Kelly, as a sanity anchor

PolicyMean RKelly fraction (exact)Kelly NHalf-Kelly N

Full Kelly lands at N ≈ 11 to 14 for the paused books, exactly the region that produced the 26 to 31× replays and the 80 to 90% drawdowns. Half-Kelly (N ≈ 22 to 28) sits near the N = 40 to 50 the survival test picked, which is stricter because it also pays the rounding tax. Kelly ignores rounding and uses R at the book's large sizing, so treat it as an anchor, not advice.

Stress arms

StressEffect on the chosen N
TPT commission ($0.75 vs $0.74 per side)None. $0.02 a round trip proves nothing either way.
$150 margin per microNo change at these balances.
Prior-day-close equity as the sizing baseNo change to the choice.
Winners capped at +5 RNo change for the paused books at $1,000.
2 ticks of slippage on every tradeCash or N = 75 to 100 almost everywhere. $2.50 is 10% of a $25 risk budget.
4 ticks of slippageN = 100 or cash, the degenerate "don't trade" answer.
Gap arm: every stop-order exit fills 4 ticks through the stop (no tail)Cash for all six policies at both balances. Applied to 1,450 of 1,798 exits: every stop-loss fill and Runner M3's 221 trailing-stop fills; limit and close exits untouched.
Gap arm plus a tail: 1 in 20 stop exits fills 20 ticks throughCash in all 12 cells.
Raw computed stop level instead of tick-rounded (the old primary arm)Slightly more optimistic; moved 4 of 12 headline cells one step more aggressive. Tick-rounded is now primary because a live stop must sit on a tick.
PnL halvedPicks a smaller, more aggressive N in some cells (Runner M3 + pause at $5,000: 50 to 30). A weakness of the objective: percentage-drawdown constraints stop binding when the path flattens, so the growth tiebreak reaches for the aggressive end.
Oversizing allowed (1.5× target)Cash at $1,000. Shown only to prove that rounding up is not free.

Follow-up: your champion and the other management modes, current regime

Master M2 could not be rescaled from its stored book, so it was re-run through the engine at fixed lots of 1, 2, 3 and 4 micros (in-bar audit clean, database verify 16 of 16) and each trade in the compounding replay uses the book that matches the lot it could afford. The Master stack with M-EMA management (the engine's fourth mode; there is no "M4") was run the same way for the first time. M1 is the "Fast no-rearm" book: the Master signal stack with fixed-target management and breakeven.

$1,000 · N = 20 · Jan 2024 to Aug 2026 · 2 ticks on stop exits · commissions · volatility pause on

ManagementEndWorst drawdownLow pointP(DD ≥ 50%)P(end < start)Median end (1,000 draws)
M2 · half off at +1R, swing trail (champion)$5,09621%$9067%6%$5,939
M3 · ATR chandelier trail$5,32438%$83432%10%$4,771
M1 · fixed target, breakeven$4,05327%$91713%8%$3,056
M-EMA · trail on the 9 EMA$61043%$57427%95%$628

Master M2 by N, same window and costs

NRisk / tradeEndWorst drawdownP(DD ≥ 50%)P(end < start)Median endMedian lots
1010%$15,69841%64%3%$16,29513
156.7%$7,23429%21%2%$9,3695
205%$5,09621%7%6%$5,9393
254%$4,33118%1%27%$2,6932
402.5%$1,03818%0%17%$1,1781
Two things sit under the M2 numbers. First, M2 needs three or more micros to be M2: at one contract there is no half to take and it becomes a breakeven-then-trail runner, at two it takes one off and trails one. The fixed-lot books make $2,400 and $2,700 over seven years at one and two micros against $5,100 and $5,400 at three and four. At $1,000 and N = 20 the median position is three micros, and only on the tighter stops. Second, the full history reverts it: the same N = 20 cell over 2019 to 2026 ends at $4,567 through a 55% drawdown with a 68% chance of halving and a 24% chance of ending below the start; without the pause it ends at $472. The table above is the good years, chosen after the fact.

M-EMA on this signal stack loses money at every lot size over the full history (profit factor 0.85) and never enters a sizing decision.

Companion table
Every tested cell at $1,000 over the last 32 months, with 2 ticks of slippage, in one sortable table →

How to use this, if at all

  1. Do not compound a $1,000 account with these methods. One micro at a median stop is already 2.9% of the account, more than the survival test allows, and the fast settings are unholdable.
  2. At $5,000 or more, 1/40th to 1/50th per trade is the region on paper, only with the volatility pause. With realistic fills (two ticks of slippage, or the gap arm) the survival test returns cash almost everywhere, so treat this as a sizing exercise, not a plan.
  3. Fix N now and judge it on future data. Everything here was measured on the same seven years the methods were selected on. There is no untouched holdout.
  4. This says nothing about Master M2, the live champion. Covering it needs an engine replay at one and two contracts, which is separate work.

Independent review

Plan critique · Codex Sol, medium reasoning, before the build
PLAN VERDICT: REVISE - The source data cannot support the proposed rescaling, M2 outcomes change with integer size, and the validation cannot justify a "perfect R" for a $1,000 live account.

Nine required changes, all implemented: reconstruct stop and per-contract PnL from stored prices and prove it (max error 0.000 on the three single-leg books); exclude Master M2; size on realized equity immediately before entry; no oversizing tolerance in candidate selection; all-in stop budget plus margin and a liquidation buffer; replace "perfect R" with a pre-registered survival-first objective including cash; separate failure states with insolvency absorbing; recompute the whole path inside every bootstrap draw on intact blocks, with quarter and year sensitivity and slippage, winner-cap and PnL stresses; count the search surface and label everything historical.

Implementation review · Codex Sol, after the first build
VERDICT: REFUTED - The implementation omits required intratrade drawdown, gap stress, and selection adjustment, while the report contains several materially false statements.

Ten findings. Fixed and re-run: the survival test now uses the full equity path including the intratrade trough (a trough at or below zero is a forced liquidation and absorbs the path); a gap stress arm was added; every chosen cell is re-scored on independent draws with the selection surface counted; tick-rounded stops are the primary arm; the margin-utilisation denominator was corrected; and the report's false or contradictory statements ("byte-identical", "no daily-loss rule", "more conservative") were rewritten. The two corrections together moved 4 of 12 headline cells, every one toward more conservative. Pre-registration cannot be proven from self-reported hashes; only a git commit before the run would do that, and none was made.

Second review · Codex Sol, after the fixes
VERDICT: REFUTED - The gap arm omits 221 Runner-M3 trailing-stop exits, "selection-adjusted" is overstated, and several report headlines contradict the results or basic sizing arithmetic.

Seven of the ten original findings confirmed resolved, one honestly disclosed (pre-registration cannot be proven without a commit). Fixed in a third round and re-run: the gap arm now charges every stop-order exit including Runner M3's trailing stops, a plain four-tick arm was added, the "selection-adjusted" wording was replaced, the below-zero trough is clamped for drawdown metrics, and the report's N = 5 failure description, one-micro arithmetic and "every correction more conservative" sentence were corrected. Including the trailing stops removed the one cell that had survived the gap arm. The third run was not sent for a further review; its changes are the ones listed here.

Caveats