Topstep Open-Fleet Grid — which configuration extracts the most

Every grid cell of the topstep_openfleet simulation: Topstep 50K vs 150K, Standard vs Consistency payout paths, with and without the DLL (RTA) ruleset, across the eval and funded method variants plus M2 management — run on the latest-regime window only, under an open-fleet protocol.

Written report

Jump to the interactive layers: headline cardsrecommended pickshonest rampthe full gridstress armsupporting readscaveats

This is the prose layer over the tables below. Every figure in it is interpolated from topstep_openfleet_grid.json, topstep_openfleet_summary.json and topstep_openfleet_ramp1.json at build time, so it cannot drift away from the run it describes. Nothing here is TradingView-verified — see §9.

1. What was tested

The study prices 8 Topstep account variants against each other: the 50K and the 150K sizes (assumed $95/mo and $199/mo respectively), each crossed with the Standard and the Consistency payout path, and each of those with and without the DLL / RTA ruleset. Across those variants it runs 2 eval method books (eval_fast_taper, eval_fast_rebreak), 4 funded books on the 50K and 8 on the 150K, plus the M2 management arm — 450 cells in total once the stress arm and the same-basis trough_control clones are counted.

The protocol is an open fleet: 4 evaluations are held from day one (target 4), at most 4 funded accounts run at once, and a funded account is never retired — the only way one leaves the fleet is a rule breach. Every payout is requested at first eligibility and approval is assumed instant. The 50K keeps $51,000 in the account after a withdrawal, the 150K keeps $153,000. The metric everywhere is bankNet: trader cash after the profit split, minus every fee.

The window is 2025-01-02 → 2026-07-08 (390 market days) and it is read through three arms. The h12 arm is 133 rolling twelve-month cohorts starting 2025-01-02 through 2025-07-09 — that is the comparative arm, and every ranking, pick and verdict on this page is made on it. The h6 arm is the same construction at a six-month horizon: 263 rolling cohorts starting 2025-01-02 through 2026-01-09, each carrying the identical metric set, added so the page can answer "what did a six-month commitment actually look like". The full arm is 21 January-2025 starts run to the end of the data; because those starts overlap almost completely it is closer to one observation than to 21, and this report treats it as a direction check only, never as a second sample.

Do not read h6 as "half of h12". A six-month campaign is dominated by the ramp: the fleet has to buy an eval, pass it, get funded and only then earn its way to payout eligibility before a single dollar of cash exists. On the h6 arm the median cohort's first payout lands on day 76–87 of a roughly 182-day window, so most of an h6 cohort is subscription bill against an account that has not paid yet. h6 medians are therefore structurally much smaller than h12 medians and far more sensitive to the start date; the ratio between the two arms is not 0.5 and carries no interpretation. Use h12 to compare configurations and h6 to size a six-month commitment.

Two things about trust. First, this study is engine-trusted by explicit instruction: the TradingView referee gate that normally has to clear before a number is actionable was deliberately skipped, so nothing below is TV-verified unless a row says so. Second, it is deterministic: the runner builds the entire grid 2 times independently and the two builds hash identically (ed627a22e1af07fe…, match confirmed), and the ramp run does the same (d227d704fceafcb8…, match confirmed) while pinning the grid file it re-simulated against (ab4a4c1b3cefc2ab…).

2. Headline results

Because the DLL cells sit on a different (and more forgiving) breach basis than the no-DLL cells, there is no single "best cell" — there is a best cell per basis. The table gives both, per size and per arm. The strict, path-basis winner is the number to quote to anyone; the trough-basis DLL winner is the number the study's own recommendation engine ranks first, and it carries an optimism the other row does not.

size · armbest cellbasisTVmedian netpayoutsfeespass days
50K · h6topstep50k_cons|headline|eval_fast_taper|fund_rebreak_cap700path · strict · no DLLTV ok$12,11212$2,28030d
50K · h6topstep50k_rta_cons|headline|eval_fast_rebreak|fund_rebreak_cap700trough · optimistic · DLL (RTA)no TV$13,25212$2,28028d
50K · h12topstep50k_cons|headline|eval_fast_taper|fund_rebreak_cap700path · strict · no DLLTV ok$34,84324$3,23050d
50K · h12topstep50k_rta_cons|headline|eval_fast_rebreak|fund_rebreak_cap700trough · optimistic · DLL (RTA)no TV$35,98324$2,85042d
50K · fulltopstep50k_cons|headline|eval_fast_taper|fund_rebreak_cap700path · strict · no DLLTV ok$58,35636$3,23095d
50K · fulltopstep50k_rta_std|headline|eval_fast_rebreak|fund_rebreak_cap700trough · optimistic · DLL (RTA)no TV$58,79940$2,47028d
150K · h6topstep150k_std|headline|eval_fast_rebreak|fund_rebreak_r1200c100_cap1500path · strict · no DLLno TV$27,87312$4,77623d
150K · h6topstep150k_rta_std|headline|eval_fast_rebreak|fund_rebreak_r1200c100_cap1500trough · optimistic · DLL (RTA)no TV$29,58012$4,77623d
150K · h12topstep150k_cons|headline|eval_fast_rebreak|fund_rebreak_r1200c100_cap1500path · strict · no DLLno TV$95,31424$9,15436d
150K · h12topstep150k_rta_std|headline|eval_fast_rebreak|fund_rebreak_r1200c100_cap1500trough · optimistic · DLL (RTA)no TV$108,70624$7,56222d
150K · fulltopstep150k_cons|headline|eval_fast_rebreak|fund_rebreak_r1200c100_cap1500path · strict · no DLLno TV$148,88436$11,54252d
150K · fulltopstep150k_rta_cons|headline|eval_fast_rebreak|fund_rebreak_r1200c100_cap1500trough · optimistic · DLL (RTA)no TV$166,34236$11,94055d
Source: summary.best_by_basis. "pass days" is the cell's eval_pass_days_med — see the caveat box on why that column carries the reserve-queue wait. Rows on different bases are NOT comparable to each other; the difference between a trough row and a path row is not a DLL cost.

On the 50K, the h12 arm puts the best trough/DLL cell at $35,983 against $34,843 for the best strict-basis cell — a gap of +$1,140 that is part ruleset and part basis flattery, in unknown proportion. On the 150K the same two rows are $108,706 and $95,314. Note that the winning cell on both sizes runs the same funded family as the recommended picks in §6, and that fees are a rounding error next to the spread: $7,562 of fees against $108,706 of net.

One column in the grid below needs its label read carefully: “med payout / paid acct (pooled)”. It is the median lifetime payout total among accounts that reached at least one payout, pooled across all overlapping cohorts, and it counts only accounts that reached ≥1 payout — zero-payout funded accounts are excluded. On the strict-basis 150K flagship (topstep150k_cons|headline|eval_fast_rebreak|fund_rebreak_r1200c100_cap1500) that median is $8,670 over 1,004 accounts drawn from 133 overlapping cohorts, against a cohort-level median gross of $116,960. It is NOT a within-cohort figure and not comparable to cohort-level gross: a reviewer proved a within-cohort median this high would be arithmetically impossible, which is exactly why the pooling must be labelled.

3. The 150K gap is leverage, not account size

The 150K books look overwhelmingly better, and the first instinct — "the bigger account is worth 3.02×" — is wrong. The matched-analog rollup separates the two effects. Pairing every 50K cell with its canonical 150K sibling (same eval book, risk-scaled funded book, same basis, same start dates), the analogs whose funded risk step is show a paired median of only +$5,159, while the analogs whose funded risk step is show +$62,102. Same account size on both rows. The only thing that changed is how much risk the funded book takes.

arm · rollupfunded books pairedpairscohorts/pairpaired median (150K − 50K)
h6 · 2x funded-risk analogsmaster_fund400 -> master_fund_r800c1008263−$3,500
h6 · 3x funded-risk analogsfund_rebreak -> fund_rebreak_r1200c100
fund_rebreak_cap700 -> fund_rebreak_r1200c100_cap1500
16263+$16,778
h6 · all matched analogsevery canonical pair24263+$9,203
h12 · 2x funded-risk analogsmaster_fund400 -> master_fund_r800c1008133+$5,159
h12 · 3x funded-risk analogsfund_rebreak -> fund_rebreak_r1200c100
fund_rebreak_cap700 -> fund_rebreak_r1200c100_cap1500
16133+$62,102
h12 · all matched analogsevery canonical pair24133+$51,845
full · 2x funded-risk analogsmaster_fund400 -> master_fund_r800c100821+$19,852
full · 3x funded-risk analogsfund_rebreak -> fund_rebreak_r1200c100
fund_rebreak_cap700 -> fund_rebreak_r1200c100_cap1500
1621+$84,739
full · all matched analogsevery canonical pair2421+$73,050
Source: summary.size_compare_matched_analogs.headline. Paired per cohort on identical start dates, same breach basis on both sides.

The best-vs-best comparison people reach for first is worse than this one, and the page flags it everywhere: +$18,631 (3.02×) between topstep50k_std|headline|eval_fast_rebreak|fund_rebreak and topstep150k_std|headline|eval_fast_rebreak|fund_rebreak_r1200c100_cap1500 — but those two cells were selected independently, so they run different books, and the 150K side carries 3× the per-trade eval risk ($2,400 vs $800) against a maximum loss limit only 2.25× larger ($4,500 vs $2,000), with doubled payout caps and a $2,496 larger fee bill. Both numbers measure the 150K package. Neither isolates account size, and neither should ever be quoted as "the 150K is worth X more".

4. DLL verdict

The DLL (RTA) ruleset has to be priced against a control that sits on the same basis, which is exactly what the trough_control clones exist for. On that honest comparison the DLL is worth +$627 on the 50K and +$796 on the 150K at the twelve-month median — both inside the noise floor for a window this short and this overlapping. The interquartile spread makes the point better than the medians do: on the 50K it runs −$3,278 to $4,902 across only 16 pairs, with the DLL ahead in 9 of them and behind in 7.

size · armpairsmedian Δ (DLL − control)meanp25 / p75DLL better/worse/tiedpairs where basis is free
50k · h616$0−$304−$217 / $3417 / 7 / 215 / 16
150k · h632$0$141$0 / $04 / 3 / 2532 / 32
50k · h1216+$627$1,086−$3,278 / $4,9029 / 7 / 012 / 16
150k · h1232+$796$4,522$0 / $7,44523 / 0 / 932 / 32
50k · full16+$620$1,232$164 / $2,35912 / 4 / 012 / 16
150k · full32+$7,127$9,305$398 / $17,75026 / 0 / 628 / 32
Source: summary.dll_headline. Every row is (RTA firm on trough) minus (the same no-DLL firm cloned onto trough) — same eval book, same funded book, same price.

Two limits on this verdict, both material. Price is not in it. The study charges every variant the same assumed monthly fee, but data.json's list prices say the RTA products are cheaper — $85 vs $95 on the 50K and $199 vs $229 on the 150K, i.e. $10 and $30 a month. A real-world DLL account is therefore better than these numbers by roughly that discount per month per account, which on a four-account fleet is the same order of magnitude as the entire measured effect. Basis is in it, on one side only. The measurement is same-basis by construction, but the basis the pair sits on is the optimistic trough, so both sides are flattered on funded deaths; on the 50K the strict-basis cost is $0 at the median but reaches −$2,033 in the worst pair, and it is free in only 12 of 16 pairs — which is luck in this window, not a licence.

5. Standard vs Consistency

Consistency (3 trading days, largest day ≤ 40% of net profit, higher request cap) is not free, and its cost is size-dependent. Across 48 matched cell pairs on the h12 arm — each pair holding size, DLL flag, eval book and funded book fixed and swapping only the payout path — Consistency comes in at −$1,588 overall. Split by size, that is −$270 on the 50K and −$4,456 on the 150K. The overall figure is a blend weighted toward the 150K (32 of 48 pairs, because the 150K runs twice as many funded books), so quote the split, never the blend.

scopepairsmedian Δ (cons − std)meanp25 / p75cons wins / losses
h6 · all48−$468−$718−$2,828 / $2,23519 / 29
h6 · 50K16−$726−$541−$1,821 / $5296 / 10
h6 · 150K32−$80−$806−$2,828 / $2,23513 / 19
h12 · all48−$1,588−$1,932−$5,776 / $2,29821 / 27
h12 · 50K16−$270−$304−$1,891 / $2,0718 / 8
h12 · 150K32−$4,456−$2,746−$6,611 / $3,07413 / 19
full · all48−$4,704−$11,217−$19,523 / $5,58214 / 34
full · 50K16−$4,704−$8,272−$16,518 / −$3,4782 / 14
full · 150K32−$7,505−$12,690−$32,211 / $9,61412 / 20
Source: summary.std_vs_cons_headline. Cell-level pairing (cons cell median minus std cell median), not per cohort.

The full arm agrees on direction and is harsher — −$4,704 overall and −$7,505 on the 150K — but with 21 overlapping starts it is one observation, so read it as a sign check rather than a magnitude. The takeaway is: take Standard on the 150K, where the 3× per-trade risk and the doubled request cap make the 40%-of-profit rule bind hard; treat the two paths as a wash on the 50K, where the gap (−$270 across 16 pairs, split 8/8) is smaller than the noise floor. That last point matters because the best single 50K cell in §2 and §6 happens to be a Consistency cell — that is a single-cell artefact of a coin-flip distribution, not evidence that Consistency wins.

6. Recommended picks

Inside the winning variant for each size, these are the two best eval × funded combos by h12 median net. The interactive recommended picks section carries the full ranking, the variant chips and the per-column definitions; this is the summary of it.

pickeval bookfunded bookTVtrue pass→fundedresetspayoutsfeesnet (h6)net (h12)net (full)
50K #1eval_fast_rebreakfund_rebreak_cap700no TV19d42d024$2,850$13,252$35,983$58,356
50K #2eval_fast_rebreakfund_rebreakno TV19d33d024$2,850$8,546$32,817$44,533
150K #1eval_fast_rebreakfund_rebreak_r1200c100_cap1500no TV22d22d024$7,562$29,580$108,706$157,491
150K #2eval_fast_rebreakfund_rebreak_r1200c100no TV22d22d024$7,562$28,784$101,988$147,196
Source: summary.recommended_picks. "true pass" is buy_eval → eval_pass; "→funded" additionally carries the passed_reserve queue wait and is NOT time-to-pass. Ranked by h12 median bankNet (133 rolling 12-month cohorts), deterministic cell_id tiebreak — the h6 column is a REPORT of the same cell over 263 six-month cohorts and never moves the selection.

The TV column is the warning, and the runner states it plainly: “these combos are SIM-ONLY unless the row's tv_status says otherwise. Only eval_fast_taper + fund_rebreak_cap700 has ever reproduced on TradingView; master_fund_r800c100 and fund_rebreak_r1200c100_cap1400 FAILED per-trade dollar parity. A pick here is a candidate to port and verify, not a cleared method.” The consequence is that the one combo on this page that has ever reproduced on TradingView (eval_fast_taper × fund_rebreak_cap700) is not the top-ranked combo on either size. Note also that the 150K picks sit on the trough basis while that TV-verified 50K combo sits on the strict path basis, so the ranking is partly a basis ranking.

Whichever pick you read, do not size a plan off the grid's “med payout / paid acct (pooled)” column. As §2 spells out, it is pooled across overlapping cohorts and counts only accounts that reached ≥1 payout — it is not a within-cohort figure and not comparable to cohort-level gross. The per-cohort number to plan against is the bankNet median on the h12 arm, after the haircut discussion in §8.

7. Honest ramp — start with 1 eval

The base study buys 4 evaluations on day one, which nobody actually does. The ramp run asks what starting honestly costs. Its policy: day one buys exactly one eval and one subscription; the purchase target then grows by one — capped at 4 — on every terminal account outcome: an eval pass, an eval breach, or a funded death. Payouts never grow the fleet, which is the deliberate asymmetry against the engine's native pass-and-payout ratchet. Buying pauses at 4 evals plus 4 funded, and the firm's real ten-account pool cap stays in force. Mechanically this is engine v2.11.0 with program.growOn="terminal" in place of the default pass_payout. Everything else is identical to the base study, and the base column in the ramp tables is that day-one-4 policy re-simulated in the same process and asserted equal to the frozen grid cell — so every Δ below is a same-start-date paired number.

cellh6 paired Δh6 winsh12 paired Δramp winsmeanp25worst cohortfull paired Δfull winswinners_half Δ
a · 50K RTA/DLL + Consistency
50K recommended pick #1 · fund_rebreak_cap700
−$2,30243.7%+$47558.7%−$1,409−$3,447−$8,162−$4,6260.0%+$760
b · 50K RTA/DLL + Consistency
50K recommended pick #2 · fund_rebreak
−$86843.7%+$47558.7%−$1,039−$2,954−$9,680−$2,9540.0%
c · 50K Consistency
50K best PATH-basis cell (strict basis) · fund_rebreak_cap700
−$2,49039.9%+$47555.6%−$2,417−$4,621−$13,524−$4,6210.0%
d · 150K RTA/DLL + Standard
150K recommended pick #1 · fund_rebreak_r1200c100_cap1500
−$7,72027.0%+$99552.6%−$3,741−$7,720−$26,649−$6,1810.0%+$1,592
e · 150K RTA/DLL + Standard
150K recommended pick #2 · fund_rebreak_r1200c100
−$7,11827.8%+$99554.1%−$3,476−$8,587−$29,905−$2,8970.0%
f · 150K Consistency
150K best PATH-basis cell (strict basis) · fund_rebreak_r1200c100_cap1500
−$7,95435.7%+$99557.9%−$4,379−$8,772−$23,144−$8,1750.0%
Source: topstep_openfleet_ramp1.json, paired_delta_ramp1_minus_base. Every column is a per-cohort (ramp1 − base) figure on identical start dates — never a difference of two medians. The mean / p25 / worst-cohort columns are the h12 arm — and they are NOT a worst case: they are extrema over heavily overlapping cohorts inside one selected regime.

The h6 arm is where the ramp actually bites, and it is the one arm that disagrees with the verdict below. Over 263 six-month cohorts the paired median runs −$868 to −$7,954 — the ramp is behind on every cell, and behind hardest on the 150K, where the forgone slots are leveraged. That is exactly what the arm is measuring: six months is mostly the ramp phase itself, so a fleet that starts at one account has not finished catching up by the time the window closes, and the subscription saving is too small to cover the difference. By twelve months the same cells have converged and the saving dominates. If the commitment really is six months, the ramp is not free.

At the twelve-month median the ramp is free — and the reason is almost funny. On all three 50K cells the paired median is exactly +$475, which is five months of the $95 subscription; on all three 150K cells it is exactly +$995, five months of $199. The median cohort's trading outcome does not change at all: median payouts are identical between the two policies on 5 of 6 cells and the median days-to-first-payout is identical on 6 of 6. What the ramp changes at the median is the subscription bill, nothing else. The ramp is also ahead in a clear majority of matched starts — 52.6% to 58.7% of cohorts depending on the cell.

The asymmetries are real and they are all in the upper tail. The paired mean is negative on 6 of 6 cells, p25 runs from −$2,954 to −$8,772, and the worst single cohort runs from −$8,162 to −$29,905. That shape says the ramp usually saves a little and occasionally forfeits a lot: the cohorts it loses are the ones where an early four-account fleet compounds into extra funded slots that the one-account start never reaches. The full arm makes the same point at full volume — 6 of 6 cells have the ramp behind on every single one of the 21 starts, from −$2,897 to −$8,175, with the 150K cells costing disproportionately more because what is forgone there is leveraged. One observation, but an unambiguous one.

Under stress the ramp flips to a hedge. On the cells the runner stressed with winners_half, the ramp wins — +$760 to +$2,786 paired — because when the edge is not there, the only thing four day-one evals buy is four subscription bills.

One trap, spelled out, because it is the easiest way to misread this table. Take cell d (150K RTA/DLL + Standard, fund_rebreak_r1200c100_cap1500). Its ramp median net is $100,120 and its base median net is $108,706. Eyeballing those two medians gives −$8,587 and the conclusion "the ramp costs eight and a half thousand dollars". The paired median — the median of the per-cohort differences on identical start dates — is +$995. Both numbers are correct; they answer different questions. The medians are of different cohorts' outcomes (the two policies' distributions have different shapes), while the paired figure is what the same trader, starting on the same day, would have experienced. Only the paired number is a decision input.

Verdict: start with one eval — if the horizon is twelve months. It is free at the twelve-month median, it is cheaper if the edge is dead, and the only thing it costs is compounding in the lucky tail — a tail that this window's overlapping cohorts are the least qualified part of the study to estimate. The six-month arm is the exception and it is not a small one (−$868 to −$7,954 paired): over six months the ramp has not finished catching up, so a short commitment pays for the honesty in forgone funded slots rather than saving on subscriptions.

8. Fragility — and what this study does NOT claim

Every headline number above depends on the edge surviving. It does not survive the standard shock. Halving every winning trade (winners_half) takes every stressed cell negative, and not by a little.

cell (h12)base netwinners_half netΔpayoutseval pass ratecohorts paid
topstep50k_rta_cons|headline|eval_fast_rebreak|fund_rebreak_cap700$35,983−$6,460−$42,44324 → 058.7% → 9.0%100.0% → 33.8%
topstep50k_rta_std|headline|eval_fast_rebreak|fund_rebreak_cap700$35,156−$6,460−$41,61624 → 054.4% → 9.0%100.0% → 33.8%
topstep50k_cons|headline|eval_fast_taper|fund_rebreak_cap700$34,843−$5,700−$40,54324 → 050.6% → 13.7%100.0% → 47.4%
topstep150k_rta_std|headline|eval_fast_rebreak|fund_rebreak_r1200c100_cap1500$108,706−$13,532−$122,23824 → 046.2% → 10.7%100.0% → 0.0%
topstep150k_rta_cons|headline|eval_fast_rebreak|fund_rebreak_r1200c100_cap1500$103,162−$13,532−$116,69424 → 050.9% → 10.7%100.0% → 0.0%
topstep150k_rta_std|headline|eval_fast_rebreak|fund_rebreak_r1200c100$101,988−$13,532−$115,52024 → 046.8% → 10.7%100.0% → 0.0%
Source: summary.stress_delta, h12 arm. winners_half halves every winning trade and re-runs the same cohorts.

The mechanism is worth naming, because "negative" here does not mean "blown accounts": FUNDED BUT NEVER PAYS — accounts do get funded and mostly survive; they never earn to payout eligibility, so bankNet collapses to roughly minus the fee bill. On the flagship 50K cell that lands at −$6,460 against $6,460 of fees — the loss is the fee bill.

The fleet is not as full as "4 evals" suggests. On the median cohort, the fleet holds four evals on only 34.1% of days and holds zero evals on 51.0% of them, ending with a median of 6 accounts stuck in passed_reserve. That is the firm's ten-account pool binding, not a trader choice: passed evals cannot be promoted while four funded accounts are alive and never retired, and they crowd out further purchases.

The sample is one to two effective observations. 263 h6, 133 h12 and 21 full-arm cohorts over eighteen months overlap almost entirely — the h6 arm has the most starts and the shortest horizon, so it has the most cohorts and still nothing like 263 independent ones; on the full arm 95 of 96 cells have p25 equal to the median and the typical cell shows only 4 distinct bankNet values. Median gaps under about $1,000 are noise.

A realistic planning number — reviewer judgement, not measurement. An external plausibility review (Kimi K3, 2026-08-01) checked the study's six plausibility claims and confirmed all six as internally consistent engine ceilings. Its conclusion is that the coherent arithmetic is not the issue: regime selection is the largest unconditional inflatorwinners_half flips every top cell negative, so edge decay is a sign-flip risk, not a haircut — and fill realism the largest measured one (−44% at one tick per side in this project's earlier work, worse at 100-micro clips). The reviewer's defensible planning haircut is ×0.15–0.30 on the PATH-basis cells only: it turns the TV-verified 50K cell's $34,843 into roughly $6–$9k/yr and the 150K path cell's $95,314 into roughly $15–$25k/yr (the multiplier applied literally gives $5,000–$10,500 and $14,500–$28,500; the quoted bands are the reviewer's own rounding) — with explicit probability mass on “edge dead → net = −fees”. Those bands are one reviewer's opinion, not an engine output, and no haircut has been measured on this window; the trough-basis and DLL cells are deliberately excluded from it.

What this page therefore does not claim. Not a live expectation: fills are touch fills with zero net slippage and no within-minute entry-bar stop check. Not a payout schedule: approval is instant here and takes 24–72 hours in reality, and can be denied. Not a clean DLL price: the RTA rows sit on the optimistic trough basis and the study charges list-price-free assumed fees. Not list pricing: $95/mo and $199/mo are the user's assumptions. Not a holdout: 2025-01-02 to 2026-07-08 is the exact regime the live methods were built on. And not a diversified fleet: it is N clones of one book on one MES path, so accounts pass, pay and die together and the tails are understated. A smaller true number beats a bigger fragile one — the honest reading of this grid is the strict-basis, TV-verified, one-eval-at-a-time corner of it — not the $166,342 that the biggest, most leveraged, most optimistically-based cell in the table prints on an arm worth one observation.

9. Verification trail

The engine suite passes 246 checks on the current tree, including the program.growOn suite added by the ramp run. Both runs are deterministic over 2 independent full builds: the grid hashes ed627a22e1af07fe71ab59756abe0bed36f6ff09e82e193a4ca3143217e85b1a twice, the ramp payload hashes d227d704fceafcb8f66205e774ebbe514aa088b55923487b94d15a4439d7743d twice, and the ramp pins the grid file it re-simulated against at ab4a4c1b3cefc2ab1122f12f23b616c6a1e95f19638d5892e69ab776c601ce8e. Input data is pinned at bbe8a5ce1e8a2e74f3e7efc1f7766d17b65393dcaf6d3e2ff7c8053091408f61 and is asserted unchanged before and after each run.

Beyond the runner's own asserts, the base study was re-verified by a fresh-context verifier working only from the source and the output JSONs, which independently re-simulated seven cells across 133 cohorts each and matched every published median to the cent — the full verdict is quoted verbatim in the supporting reads section below, with bracketed annotations where the current tree has moved on. Both the base study and the ramp extension were additionally reviewed by an independent cross-model reviewer (Codex, gpt-5.6-sol, medium reasoning); both rounds came back REFUTED first, the findings were fixed, and both re-reviews then returned CONFIRMED. A third, independent external plausibility review (Kimi K3, 2026-08-01) then re-checked the page's six plausibility claims and returned all six CONFIRMED as internally consistent engine ceilings. Its qualitative findings are recorded in §8 and are reviewer judgement, not measurement: the coherent arithmetic is not the issue, regime selection is the largest unconditional inflator (winners_half flips every top cell negative — a sign-flip risk, not a haircut) and fill realism the largest measured one; the reviewer's defensible planning haircut is ×0.15–0.30 on the path-basis cells only, with explicit probability mass on “edge dead → net = −fees”. What none of that buys is external reproduction of the trading itself: the TradingView referee gate was skipped by instruction, and it remains the gating next step for any cell here.

Headline reads

Recommended picks

Honest ramp — start with 1 eval

The grid — every headline cell

Click any column header to sort (numeric, click again to flip). Default sort = h12 median net PnL, descending. Rows are headline cells only — the same-basis trough_control clones are in their own collapsed section below, because they exist to price the DLL and must never be ranked against headline cells.

Stress arm — winners_half

Supporting reads

trough_control clones (same-basis no-DLL controls — cells)

DLL same-basis paired detail (per size × arm)
Standard vs Consistency paired detail
50K vs 150K — size comparison and the leverage confound
Fleet occupancy, arm dispersion, and the identical-day-cap arms
Full caveat list as recorded by the runner
    Verifier verdict — archived base-study round (pre-ramp1), quoted verbatim with bracketed notes

    Read this before you act on any number above

    1. This is a simulation, not a live expectation. Every figure on this page is an engine ceiling produced by replaying one historical MES path. Treat it as a ranking device between configurations, not as a forecast of cash.
    2. No TradingView referee verification. By explicit instruction this study skips the TV parity gate. Only eval_fast_taper + fund_rebreak_cap700 has ever reproduced on TradingView; master_fund_r800c100 and fund_rebreak_r1200c100_cap1400 FAILED per-trade dollar parity; everything else is not TV-ported at all. The TV column on each row carries this — the winning cells are mostly not TV-verified.
    3. The fill model is optimistic — and the biggest single piece of it is “entry-bar SL/TP not checked intra-minute”. Naming it separately because it is exactly the bug class that once fabricated five figures of phantom profit in this project: on the entry bar the engine does not walk the minute to see whether the stop or the target was touched first, so ambiguous entry bars resolve in the trade's favour. On top of that: touch fills and zero net slippage. Real fills will be worse, and the whole grid shifts down together — the project's own measurement is −44% at one tick per side.
    4. No capacity or market-impact model exists in the engine. The 150K cells trade clips of up to 100 micros, and up to 4 correlated accounts fire the same signal — that is roughly 400 MES micros at one touch price inside one minute, including overnight sessions where top-of-book is thin. The engine fills all of it at the same price with zero impact. Nothing on this page prices the liquidity required to actually do that.
    5. Payouts are assumed instantly approved. The engine books cash at first eligibility. Real approval takes 24-72 hours, can be denied, and no processing pause is modelled — so every "days to first payout" is a floor.
    6. DLL cells run on the optimistic trough breach basis (an engine limitation: it refuses a path breach check with an intraday-flatten daily-loss rule). That flatters every DLL row on this page relative to the no-DLL rows next to it. The honest DLL comparison is the same-basis paired delta in the summary cards and in the DLL detail section — never derive a DLL cost by subtracting one headline row from another.
    7. "days→funded" is not time-to-pass. The column measures buy_eval → funded_start. Under this protocol (max 4 funded, never retire) an eval that passes while all four funded slots are occupied is parked in passed_reserve and cannot be promoted until a funded account breaches the MLL — which can take months. The published figure therefore carries that reserve-queue wait on top of the real pass time: measured on the h12 arm it runs from 1.0× up to 2.2× the true buy→pass time (worst sampled cell: 42 days to funded vs 19 days to pass; longest single queue wait in the sample 243 days). Never read it as eval difficulty or compare it to a firm's advertised pass time.
    8. "150K vs 50K" is a package comparison, not an account-size one. The best-vs-best card pairs each size's independently chosen best cell, so the two sides usually run different books. The matched canonical-analog card fixes that, but the 150K still runs 3× the per-trade eval risk against an MLL only 2.25× larger, with doubled payout caps and a $199-vs-$95 fee. Both numbers measure the 150K package; neither isolates account size.
    9. Consistency's cost is size-dependent — quote the split. Across matched pairs on the h12 arm, Consistency minus Standard is only −$270 on the 50K (inside the ~$1,000 noise floor for this window) but −$4,456 on the 150K, where the larger risk and doubled request cap make the "largest day ≤ 40% of net profit" rule bind much harder. The overall −$1,588 figure is a blend of the two and is weighted toward the 150K (32 of 48 pairs).
    10. One 18-month window, heavily overlapping cohorts. The 263 h6, 133 h12 and 21 full-arm starts collapse to roughly one to two independent event chains — the full arm is essentially a single observation with a thin upper tail. Median gaps under about $1,000 are noise. Quote the h12 arm for anything comparative.
    11. The dispersion columns are not a worst case. Every min / p25 / “worst cohort” figure on this page is an extremum over heavily overlapping cohorts, not a downside scenario. The 133 h12 starts (2025-01-02 → 2025-07-09) span only 6.2 months of an 18.1-month window, so at the very most they contain 6 non-identical twelve-month samples — by the study's own dispersion measures, closer to one or two — and every one of them sits inside the single regime the methods were selected on. A p25 or a minimum drawn from that is the unluckiest start date in a good regime, not the unluckiest regime. The real downside scenario on this page is the winners_half stress arm, and it is negative everywhere.
    12. The h6 arm is ramp-dominated — it is not "half of h12". Six months is mostly the phase before any cash exists: buy the eval, pass it, get funded, then earn to payout eligibility. On the h6 arm the median cohort's first payout lands on day 76–87 of ~182, so a large share of every h6 cohort is subscription bill against an account that has not paid yet. h6 medians are structurally much smaller than h12 medians and much more start-date-sensitive; the ratio between the arms is not 0.5 and carries no meaning. Every ranking, pick and verdict on this page is still made on h12 — the h6 columns are a report of the same cells, never a selection input. The one place h6 changes a conclusion is the honest-ramp section, and it is called out there.
    13. Prices are the user's assumptions, not list prices: $95/mo on every 50K, $199/mo on every 150K, $0 activation, applied identically to DLL and no-DLL variants. That deliberately erases the real DLL discount ($85 vs $95, $199 vs $229), so the DLL verdict here is a pure ruleset + payout-cap read, not a price read. The pricing itself is an idealisation worth money: $0 activation plus the assumed $95/$199 schedule understates the real bill by an amount an external reviewer puts at roughly $1–2k per fleet-year (reviewer judgement, not measured here) — on top of the unmodelled copier/data fees.
    14. The "honest ramp" section is a separate run with its own caveats. It comes from topstep_openfleet_ramp1.json (engine v2.11.0, mode program.growOn="terminal" instead of the default pass_payout), re-simulated against grid sha ab4a4c1b3cef…. Its base column is the day-one-4 policy re-run in the same process and asserted equal to the frozen grid cell, and every Δ there is a same-start-date paired per-cohort number — never a difference of two medians. Every caveat on this list applies to those rows unchanged (optimistic fills, trough basis on the DLL cells, one MES path, overlapping cohorts, instant payouts, assumed prices, selection window), and the ramp adds nothing that fixes any of them.
    15. Selection, not holdout. Jan-2025 to Jul-2026 is the exact regime the live methods were built on. The fleet is N clones of one book, so accounts pass, pay and die together — real diversification is absent and the tails are understated. Topstep's live call-up, copier fees, news windows and inactivity rules are unmodelled.