SPECIMEN
A public laboratory
for adaptive intelligence
Simulation

Lineage

Every specimen runs a numbered strategy version with its parameters written down. A new generation is a curated change with a stated reason, evaluated on its own period. Nothing here is produced by an optimiser, and a generation that performs worse than its parent stays on the record.

Candidates on recordNone promoted
CandidateChangesDatasets aheadTypicalBy chanceStanding
MOTH-002Confirm Bars95 / 158+0.179.4%Not separable from chance
MOTH-003Max Extension84 / 139+0.0910%Not separable from chance
MOTH-004Exit Threshold81 / 154+0.0257%Not separable from chance
MOTH-005Trend Slow85 / 149+0.2728%Not separable from chance
MOSS-002Min Mean Slope42 / 102-0.6028%Not separable from chance
MOSS-003Min Mean Slope Vol30 / 80-0.2416%Not separable from chance
MOSS-004Exit ZUntested
MOSS-005Stop LossUntested
ECHO-002Epoch Bars40 / 108-0.677.2%Not separable from chance
ECHO-003Carry Across Epochs17 / 49-0.3318%Not separable from chance
ECHO-004Position FractionUntested
ECHO-005Breakout Exit LookbackUntested

Every figure here is the corpus result corrected for the 12 candidates tested against it. A candidate that fails keeps its slot, its parameters, its reason and its numbers; one that is supported is still not promoted, because evidence is not a decision.

MOTH-001No parent
Momentum · Confirmed momentum

First generation. Enter only when 12-bar momentum clears +1.2% and the 8-bar mean is above the 24-bar mean, so a single spike cannot trigger an entry on its own. The confirmation gate is open: one qualifying close is enough and there is no limit on how far price may have already travelled.

Momentum Period12Entry Threshold1.20%Exit Threshold-0.40%Trend Fast8.00Trend Slow24.00Stop Loss6.00%Reduce Gain5.00%Position Fraction50.00%Confirm Bars1Max Extensionno limit
Fills in this run22Return+1.39%Max drawdown-3.52%

Not yet measured against random trading. This runs in the background after the candidate work.

Candidate MOTH-002 is recorded against this version. It has not been promoted; see below for what changed and how it did on bars this version was not read against.

MOSS-001No parent
Mean reversion · Sigma reversion

First generation. Buy closes stretched 1.4 sigma below a 36-bar mean while RSI confirms weakness, release as price returns through the mean. The regime guard is open: it will buy a dip regardless of what the mean itself is doing.

Z Period36Entry Z-1.40σExit Z0.20σReduce Z-0.20σRsi Period14Max Rsi45.0Stop Loss8.00%Position Fraction50.00%Mean Slope Period36Min Mean Slope-99.00Min Mean Slope Vol-99.00
Fills in this run24Return+0.02%Max drawdown-1.34%

Not yet measured against random trading. This runs in the background after the candidate work.

Candidate MOSS-002 is recorded against this version. It has not been promoted; see below for what changed and how it did on bars this version was not read against.

ECHO-001No parent
Explorer · Seeded rotation

First generation. Rotate between three published rules every 72 bars using the run seed, so each rule accumulates its own trades under identical conditions. The rotation is fixed by the seed, not chosen by performance. Every open position is closed at the epoch boundary to keep that attribution clean.

Epoch Bars72Drift Period6Drift Entry0.60%Drift Exit-0.20%Rebound Period24Rebound Entry Z-1.00σRebound Exit Z0.00σBreakout Lookback20Breakout Exit Lookback10Stop Loss7.00%Position Fraction40.00%Carry Across Epochs0
Fills in this run38Return+0.56%Max drawdown-3.50%

Not yet measured against random trading. This runs in the background after the candidate work.

Candidate ECHO-002 is recorded against this version. It has not been promoted; see below for what changed and how it did on bars this version was not read against.

Candidate MOTH-002 · not promotedParent MOTH-001

Candidate, not promoted. MOTH-001 commits on the first close that crosses its threshold. This requires momentum to hold above +1.2% for two consecutive closes and changes nothing else. It was originally written together with an extension limit, which meant no result could be attributed to either parameter; the two were split so each can be tested on its own. The value here is the one it was written with, not one read off a grid.

What changed · 1 of 10 parameters
Confirm Bars1 2
Corpus · 480 windows · 160 datasetsNot separable from chance
Windows ahead240 / 437
Datasets ahead95 / 158
Untouched42 / 480
Effect (95% CI)0.03 … 0.31points per dataset
p over datasets0.013the independent unit
p corrected0.09412 candidates tested
Could have found±0.00points, at 0.05
Fills compared7020 / 6134
By market family · holding across these is worth more than holding across more of one
crypto / USDahead on 62 of 108 datasetscrypto / other fiatahead on 20 of 24 datasetscrypto / cryptoahead on 10 of 20 datasetsoutside cryptoahead on 3 of 6 datasets
By timescale · fires on · a change that only works on one is not a change that works
15m barsfires on 79% · ahead on 24 of 39 datasets1h barsfires on 96% · ahead on 21 of 39 datasets4h barsfires on 95% · ahead on 25 of 40 datasets1d barsfires on 95% · ahead on 25 of 40 datasets

MOTH-002 was run against its parent on 480 held-out windows across 160 datasets, 7020 parent fills against 6134. It came out ahead on 240 of 437 decisive windows (median difference +0.07 points). 42 windows were untouched by the change and are excluded. Windows from one dataset share an instrument and their indicator history, so the test is taken over datasets: 95 of 158 went the candidate's way. A split that lopsided would happen 1.3% of the time by chance. That is evidence worth having, and still not a decision: it is one family of instruments over one stretch of history, and promotion remains a deliberate act.

Typical difference per dataset +0.17 points, 95% of resamples between 0.03 and 0.31. That range excludes zero, which narrows the size of the effect without settling whether it survives another corpus.

With 158 datasets this corpus would have detected a consistent edge of about 0.00 points per dataset. It found none, so whatever MOTH-002 does is smaller than that — which is a bound on the effect, not proof there is none.

On its own this reads as 0.013, under the 0.05 line. It is not on its own: 12 candidates are tested against the same corpus, and asking that many times makes one low number ordinary. Corrected across the family it is 0.094, and MOTH-002 stays a candidate.

Readings over time · 34 corpora
WhenCorpusDatasets aheadEffectBy chanceRead as
13 Sept 202677,760 bars62 / 108+0.1459%Not separable from chance
13 Sept 202677,760 bars62 / 108+0.1449%Not separable from chance
13 Sept 202677,760 bars62 / 108+0.1445%Not separable from chance
13 Sept 202677,760 bars62 / 108+0.1458%Not separable from chance
13 Sept 202677,760 bars62 / 108+0.1445%Not separable from chance
13 Sept 202677,760 bars62 / 108+0.1448%Not separable from chance
13 Sept 202677,760 bars62 / 108+0.1445%Not separable from chance
13 Sept 202677,760 bars62 / 108+0.1459%Not separable from chance
13 Sept 202677,760 bars62 / 108+0.1447%Not separable from chance
13 Sept 202677,760 bars62 / 108+0.1445%Not separable from chance
13 Sept 202677,760 bars62 / 108+0.1449%Not separable from chance
13 Sept 202677,760 bars62 / 108+0.1459%Not separable from chance
13 Sept 202677,760 bars62 / 108+0.1452%Not separable from chance
13 Sept 202677,760 bars62 / 108+0.1445%Not separable from chance
13 Sept 202677,760 bars62 / 108+0.1454%Not separable from chance
13 Sept 202677,760 bars62 / 108+0.1445%Not separable from chance
13 Sept 202677,760 bars62 / 108+0.1446%Not separable from chance
13 Sept 202677,760 bars62 / 108+0.1445%Not separable from chance
13 Sept 202677,760 bars62 / 108+0.1445%Not separable from chance
13 Sept 202677,760 bars62 / 108+0.1457%Not separable from chance
13 Sept 202677,760 bars62 / 108+0.1447%Not separable from chance
13 Sept 202677,760 bars62 / 108+0.1445%Not separable from chance
13 Sept 202677,760 bars62 / 108+0.1449%Not separable from chance
13 Sept 202677,760 bars62 / 108+0.1445%Not separable from chance
13 Sept 202677,760 bars62 / 108+0.1447%Not separable from chance
13 Sept 202677,760 bars62 / 108+0.1459%Not separable from chance
13 Sept 2026115,200 bars95 / 158+0.179.4%Not separable from chance
13 Sept 2026115,200 bars95 / 158+0.178.0%Not separable from chance
13 Sept 2026115,200 bars95 / 158+0.178.4%Not separable from chance
13 Sept 2026115,200 bars95 / 158+0.178.0%Not separable from chance
13 Sept 2026115,200 bars95 / 158+0.178.8%Not separable from chance
13 Sept 2026115,200 bars95 / 158+0.179.4%Not separable from chance
13 Sept 2026115,200 bars95 / 158+0.179.3%Not separable from chance
13 Sept 2026115,200 bars95 / 158+0.179.4%Not separable from chance

Judged 34 times as the corpus grew by 37,440 bars. The corrected probability has ranged across 0.51 while the standing held, so the reading is steadier than the number behind it.

Sensitivity grid pending: 5 settings to run, each a full held-out simulation. They are computed a couple at a time so the engine keeps serving.

Waiting for a held-out test. It runs once the current dataset has been replayed in full.

Candidate MOTH-003 · not promotedParent MOTH-001

Candidate, not promoted. The other half of the change that was originally bundled into MOTH-002: refuse an entry when the close already sits more than 1.5% above the 8-bar mean, on the reasoning that the move has run before the signal arrived. Confirmation is left at the parent's setting so this parameter is the only thing that moves. A sibling of MOTH-002, not a successor to it.

What changed · 1 of 10 parameters
Max Extensionno limit 1.50%
Corpus · 480 windows · 160 datasetsNot separable from chance
Windows ahead204 / 346
Datasets ahead84 / 139
Untouched132 / 480
Effect (95% CI)0.03 … 0.19points per dataset
p over datasets0.017the independent unit
p corrected0.10312 candidates tested
Could have found±0.00points, at 0.05
Fills compared7020 / 6175
By market family · holding across these is worth more than holding across more of one
crypto / USDahead on 51 of 96 datasetscrypto / other fiatahead on 22 of 24 datasetscrypto / cryptoahead on 10 of 15 datasetsoutside cryptoahead on 1 of 4 datasets
By timescale · fires on · a change that only works on one is not a change that works
15m barsfires on 31% · ahead on 16 of 25 datasets1h barsfires on 70% · ahead on 20 of 36 datasets4h barsfires on 92% · ahead on 23 of 38 datasets1d barsfires on 98% · ahead on 25 of 40 datasets

It fires 67 points more often at one timescale than another, so it is not being tested the same way at both ends of the corpus.

MOTH-003 was run against its parent on 480 held-out windows across 160 datasets, 7020 parent fills against 6175. It came out ahead on 204 of 346 decisive windows (median difference +0.14 points). 132 windows were untouched by the change and are excluded. Windows from one dataset share an instrument and their indicator history, so the test is taken over datasets: 84 of 139 went the candidate's way. A split that lopsided would happen 1.7% of the time by chance. That is evidence worth having, and still not a decision: it is one family of instruments over one stretch of history, and promotion remains a deliberate act.

Typical difference per dataset +0.09 points, 95% of resamples between 0.03 and 0.19. That range excludes zero, which narrows the size of the effect without settling whether it survives another corpus.

With 139 datasets this corpus would have detected a consistent edge of about 0.00 points per dataset. It found none, so whatever MOTH-003 does is smaller than that — which is a bound on the effect, not proof there is none.

On its own this reads as 0.017, under the 0.05 line. It is not on its own: 12 candidates are tested against the same corpus, and asking that many times makes one low number ordinary. Corrected across the family it is 0.103, and MOTH-003 stays a candidate.

Readings over time · 32 corpora
WhenCorpusDatasets aheadEffectBy chanceRead as
13 Sept 202677,760 bars51 / 96+0.05100%Not separable from chance
13 Sept 202677,760 bars51 / 96+0.0575%Not separable from chance
13 Sept 202677,760 bars51 / 96+0.05100%Not separable from chance
13 Sept 202677,760 bars51 / 96+0.0578%Not separable from chance
13 Sept 202677,760 bars51 / 96+0.05100%Not separable from chance
13 Sept 202677,760 bars51 / 96+0.0585%Not separable from chance
13 Sept 202677,760 bars51 / 96+0.0578%Not separable from chance
13 Sept 202677,760 bars51 / 96+0.0563%Not separable from chance
13 Sept 202677,760 bars51 / 96+0.0578%Not separable from chance
13 Sept 202677,760 bars51 / 96+0.0596%Not separable from chance
13 Sept 202677,760 bars51 / 96+0.0578%Not separable from chance
13 Sept 202677,760 bars51 / 96+0.05100%Not separable from chance
13 Sept 202677,760 bars51 / 96+0.0585%Not separable from chance
13 Sept 202677,760 bars51 / 96+0.0578%Not separable from chance
13 Sept 202677,760 bars51 / 96+0.0588%Not separable from chance
13 Sept 202677,760 bars51 / 96+0.0578%Not separable from chance
13 Sept 202677,760 bars51 / 96+0.0592%Not separable from chance
13 Sept 202677,760 bars51 / 96+0.0578%Not separable from chance
13 Sept 202677,760 bars51 / 96+0.0565%Not separable from chance
13 Sept 202677,760 bars51 / 96+0.0561%Not separable from chance
13 Sept 202677,760 bars51 / 96+0.05100%Not separable from chance
13 Sept 202677,760 bars51 / 96+0.0591%Not separable from chance
13 Sept 202677,760 bars51 / 96+0.0578%Not separable from chance
13 Sept 202677,760 bars51 / 96+0.0561%Not separable from chance
13 Sept 202677,760 bars51 / 96+0.0578%Not separable from chance
13 Sept 202677,760 bars51 / 96+0.05100%Not separable from chance
13 Sept 202677,760 bars51 / 96+0.0578%Not separable from chance
13 Sept 202677,760 bars51 / 96+0.0581%Not separable from chance
13 Sept 202677,760 bars51 / 96+0.0578%Not separable from chance
13 Sept 202677,760 bars51 / 96+0.0585%Not separable from chance
13 Sept 202677,760 bars51 / 96+0.0578%Not separable from chance
13 Sept 2026115,200 bars84 / 139+0.0910%Not separable from chance

Judged 32 times as the corpus grew by 37,440 bars. The corrected probability has ranged across 0.90 while the standing held, so the reading is steadier than the number behind it.

Sensitivity grid pending: 6 settings to run, each a full held-out simulation. They are computed a couple at a time so the engine keeps serving.

Waiting for a held-out test. It runs once the current dataset has been replayed in full.

Candidate MOTH-004 · not promotedParent MOTH-001

Candidate, not promoted. MOTH-001 waits for momentum to turn negative by 0.4% before leaving, so every exit is taken after the move has already reversed and the give-back is built into the rule. This releases while momentum is still positive but decaying, at +0.2%. The expected cost is leaving early in a move that pauses and resumes, which is the same move the parent is paid for.

What changed · 1 of 10 parameters
Exit Threshold-0.40% 0.20%
Corpus · 480 windows · 160 datasetsNot separable from chance
Windows ahead196 / 391
Datasets ahead81 / 154
Untouched87 / 480
Effect (95% CI)-0.07 … 0.13points per dataset
p over datasets0.573the independent unit
p corrected0.57312 candidates tested
Could have found±0.07points, at 0.05
Fills compared7020 / 7602
By market family · holding across these is worth more than holding across more of one
crypto / USDahead on 61 of 105 datasetscrypto / other fiatahead on 7 of 24 datasetscrypto / cryptoahead on 10 of 20 datasetsoutside cryptoahead on 3 of 5 datasets
By timescale · fires on · a change that only works on one is not a change that works
15m barsfires on 74% · ahead on 21 of 38 datasets1h barsfires on 97% · ahead on 17 of 39 datasets4h barsfires on 91% · ahead on 24 of 40 datasets1d barsfires on 66% · ahead on 19 of 37 datasets

MOTH-004 was run against its parent on 480 held-out windows across 160 datasets, 7020 parent fills against 7602. It came out ahead on 196 of 391 decisive windows (median difference +0.00 points). 87 windows were untouched by the change and are excluded. Windows from one dataset share an instrument and their indicator history, so the test is taken over datasets: 81 of 154 went the candidate's way. A split at least that lopsided happens 57% of the time when a change does nothing at all, so this corpus does not separate it from chance. MOTH-004 stays a candidate.

Typical difference per dataset +0.02 points, 95% of resamples between -0.07 and 0.13. That range contains zero, so on this corpus MOTH-004 is consistent with making no difference at all.

With 154 datasets this corpus would have detected a consistent edge of about 0.07 points per dataset. It found none, so whatever MOTH-004 does is smaller than that — which is a bound on the effect, not proof there is none.

Readings over time · 40 corpora
WhenCorpusDatasets aheadEffectBy chanceRead as
13 Sept 202677,760 bars1 / 1100%Not separable from chance
13 Sept 202677,760 bars2 / 2100%Not separable from chance
13 Sept 202677,760 bars3 / 3+0.1475%Not separable from chance
13 Sept 202677,760 bars3 / 4+0.10100%Not separable from chance
13 Sept 202677,760 bars4 / 5+0.0678%Not separable from chance
13 Sept 202677,760 bars4 / 6+0.04100%Not separable from chance
13 Sept 202677,760 bars4 / 7+0.03100%Not separable from chance
13 Sept 202677,760 bars5 / 8+0.04100%Not separable from chance
13 Sept 202677,760 bars5 / 9+0.03100%Not separable from chance
13 Sept 202677,760 bars6 / 10+0.04100%Not separable from chance
13 Sept 202677,760 bars7 / 11+0.06100%Not separable from chance
13 Sept 202677,760 bars7 / 12+0.04100%Not separable from chance
13 Sept 202677,760 bars8 / 13+0.06100%Not separable from chance
13 Sept 202677,760 bars9 / 14+0.1085%Not separable from chance
13 Sept 202677,760 bars10 / 15+0.1478%Not separable from chance
13 Sept 202677,760 bars11 / 16+0.1563%Not separable from chance
13 Sept 202677,760 bars11 / 17+0.1478%Not separable from chance
13 Sept 202677,760 bars11 / 18+0.1096%Not separable from chance
13 Sept 202677,760 bars12 / 19+0.1078%Not separable from chance
13 Sept 202677,760 bars13 / 20+0.0878%Not separable from chance
13 Sept 202677,760 bars13 / 21+0.0678%Not separable from chance
13 Sept 202677,760 bars13 / 22+0.05100%Not separable from chance
13 Sept 202677,760 bars13 / 23+0.03100%Not separable from chance
13 Sept 202677,760 bars14 / 24+0.05100%Not separable from chance
13 Sept 202677,760 bars15 / 25+0.0685%Not separable from chance
13 Sept 202677,760 bars16 / 26+0.0578%Not separable from chance
13 Sept 202677,760 bars16 / 27+0.0388%Not separable from chance
13 Sept 202677,760 bars17 / 28+0.0378%Not separable from chance
13 Sept 202677,760 bars17 / 29+0.0392%Not separable from chance
13 Sept 202677,760 bars18 / 30+0.0378%Not separable from chance
13 Sept 202677,760 bars19 / 31+0.0478%Not separable from chance
13 Sept 202677,760 bars20 / 32+0.0565%Not separable from chance
13 Sept 202677,760 bars21 / 33+0.0659%Not separable from chance
13 Sept 202677,760 bars22 / 34+0.0849%Not separable from chance
13 Sept 202677,760 bars23 / 35+0.1036%Not separable from chance
13 Sept 202677,760 bars24 / 36+0.1026%Not separable from chance
13 Sept 202677,760 bars25 / 37+0.1022%Not separable from chance
13 Sept 202677,760 bars26 / 38+0.1117%Not separable from chance
13 Sept 202677,760 bars27 / 39+0.1214%Not separable from chance
13 Sept 202677,760 bars27 / 40+0.1119%Not separable from chance

Judged 40 times as the corpus grew. The corrected probability has ranged across 0.86 while the standing held, so the reading is steadier than the number behind it.

Sensitivity grid pending: 6 settings to run, each a full held-out simulation. They are computed a couple at a time so the engine keeps serving.

This run’s own window · 289 bars · 01 Sept 2026 13 Sept 2026 · both start at 10,000 · one window, far weaker than the corpus above
MOTH-001MOTH-004
Return+0.45%-0.66%
Max drawdown-1.53%-2.53%
Fills68
Bars held4230
Fees paid30.2840.17

Both versions ran the same engine over the same bars, from the same capital, with the same fee and slippage. A difference over one window of 289 bars is not evidence that the change works, so MOTH-004 stays a candidate and MOTH-001 keeps running.

Candidate MOTH-005 · not promotedParent MOTH-001

Candidate, not promoted. MOTH-001 measures momentum over 12 bars and then filters it with a 24-bar mean, which is built from largely the same bars: the filter mostly agrees with the signal it is supposed to check. A 60-bar mean is far enough away to disagree. The expected cost is missing early turns, since a slower mean confirms later.

What changed · 1 of 10 parameters
Trend Slow24.00 60.00
Corpus · 480 windows · 160 datasetsNot separable from chance
Windows ahead242 / 422
Datasets ahead85 / 149
Untouched57 / 480
Effect (95% CI)-0.04 … 0.50points per dataset
p over datasets0.101the independent unit
p corrected0.27612 candidates tested
Could have found±0.05points, at 0.05
Fills compared7020 / 5733
By market family · holding across these is worth more than holding across more of one
crypto / USDahead on 60 of 105 datasetscrypto / other fiatahead on 11 of 22 datasetscrypto / cryptoahead on 11 of 18 datasetsoutside cryptoahead on 3 of 4 datasets
By timescale · fires on · a change that only works on one is not a change that works
15m barsfires on 63% · ahead on 22 of 32 datasets1h barsfires on 92% · ahead on 14 of 38 datasets4h barsfires on 98% · ahead on 21 of 39 datasets1d barsfires on 100% · ahead on 28 of 40 datasets

MOTH-005 was run against its parent on 480 held-out windows across 160 datasets, 7020 parent fills against 5733. It came out ahead on 242 of 422 decisive windows (median difference +0.32 points). 57 windows were untouched by the change and are excluded. Windows from one dataset share an instrument and their indicator history, so the test is taken over datasets: 85 of 149 went the candidate's way. A split at least that lopsided happens 10% of the time when a change does nothing at all, so this corpus does not separate it from chance. MOTH-005 stays a candidate.

Typical difference per dataset +0.27 points, 95% of resamples between -0.04 and 0.50. That range contains zero, so on this corpus MOTH-005 is consistent with making no difference at all.

With 149 datasets this corpus would have detected a consistent edge of about 0.05 points per dataset. It found none, so whatever MOTH-005 does is smaller than that — which is a bound on the effect, not proof there is none.

Sensitivity grid pending: 6 settings to run, each a full held-out simulation. They are computed a couple at a time so the engine keeps serving.

This run’s own window · 289 bars · 01 Sept 2026 13 Sept 2026 · both start at 10,000 · one window, far weaker than the corpus above
MOTH-001MOTH-005
Return+0.45%+1.27%
Max drawdown-1.53%-0.63%
Fills62
Bars held4215
Fees paid30.2810.13

Both versions ran the same engine over the same bars, from the same capital, with the same fee and slippage. A difference over one window of 289 bars is not evidence that the change works, so MOTH-005 stays a candidate and MOTH-001 keeps running.

Candidate MOSS-002 · not promotedParent MOSS-001

Candidate, not promoted. MOSS-001 measures how far price has stretched from its mean but never asks what the mean is doing, so in a sustained decline it buys a level that keeps moving away from it. This changes one parameter: entries are refused while the 36-bar mean has itself fallen more than 3% over the last 36 bars. It narrows when an entry is allowed and adds no new signal. The expected cost is missing the genuine bottom, where the mean is always falling hardest.

What changed · 1 of 11 parameters
Min Mean Slope-99.00 -0.03
Corpus · 480 windows · 160 datasetsNot separable from chance
Windows ahead83 / 204
Datasets ahead42 / 102
Untouched276 / 480
Effect (95% CI)-0.71 … 0.26points per dataset
p over datasets0.092the independent unit
p corrected0.27612 candidates tested
Could have found±0.71points, at 0.05
Fills compared6138 / 4538
By market family · holding across these is worth more than holding across more of one
crypto / USDahead on 34 of 79 datasetscrypto / other fiatahead on 6 of 12 datasetscrypto / cryptoahead on 2 of 9 datasetsoutside cryptoahead on 0 of 2 datasets
By timescale · fires on · a change that only works on one is not a change that works
15m barsfires on 7% · ahead on 4 of 8 datasets1h barsfires on 23% · ahead on 5 of 19 datasets4h barsfires on 51% · ahead on 20 of 36 datasets1d barsfires on 89% · ahead on 13 of 39 datasets

It fires 83 points more often at one timescale than another, so it is not being tested the same way at both ends of the corpus.

MOSS-002 was run against its parent on 480 held-out windows across 160 datasets, 6138 parent fills against 4538. It came out ahead on 83 of 204 decisive windows (median difference -0.57 points). 276 windows were untouched by the change and are excluded. Windows from one dataset share an instrument and their indicator history, so the test is taken over datasets: 42 of 102 went the candidate's way. A split at least that lopsided happens 9.2% of the time when a change does nothing at all, so this corpus does not separate it from chance. MOSS-002 stays a candidate.

Typical difference per dataset -0.60 points, 95% of resamples between -0.71 and 0.26. That range contains zero, so on this corpus MOSS-002 is consistent with making no difference at all.

With 102 datasets this corpus would have detected a consistent edge of about 0.71 points per dataset. It found none, so whatever MOSS-002 does is smaller than that — which is a bound on the effect, not proof there is none.

Readings over time · 21 corpora
WhenCorpusDatasets aheadEffectBy chanceRead as
13 Sept 202677,760 bars34 / 79-0.5878%Not separable from chance
13 Sept 202677,760 bars34 / 79-0.5875%Not separable from chance
13 Sept 202677,760 bars34 / 79-0.5878%Not separable from chance
13 Sept 202677,760 bars34 / 79-0.5863%Not separable from chance
13 Sept 202677,760 bars34 / 79-0.5878%Not separable from chance
13 Sept 202677,760 bars34 / 79-0.5865%Not separable from chance
13 Sept 202677,760 bars34 / 79-0.5859%Not separable from chance
13 Sept 202677,760 bars34 / 79-0.5852%Not separable from chance
13 Sept 202677,760 bars34 / 79-0.5858%Not separable from chance
13 Sept 202677,760 bars34 / 79-0.5852%Not separable from chance
13 Sept 202677,760 bars34 / 79-0.5859%Not separable from chance
13 Sept 202677,760 bars34 / 79-0.5852%Not separable from chance
13 Sept 202677,760 bars34 / 79-0.5859%Not separable from chance
13 Sept 202677,760 bars34 / 79-0.5852%Not separable from chance
13 Sept 202677,760 bars34 / 79-0.5854%Not separable from chance
13 Sept 202677,760 bars34 / 79-0.5852%Not separable from chance
13 Sept 202677,760 bars34 / 79-0.5857%Not separable from chance
13 Sept 202677,760 bars34 / 79-0.5852%Not separable from chance
13 Sept 202677,760 bars34 / 79-0.5878%Not separable from chance
13 Sept 202677,760 bars34 / 79-0.5859%Not separable from chance
13 Sept 202677,760 bars34 / 79-0.5878%Not separable from chance

Judged 21 times as the corpus grew. The corrected probability has ranged across 0.26 while the standing held, so the reading is steadier than the number behind it.

Sensitivity grid pending: 7 settings to run, each a full held-out simulation. They are computed a couple at a time so the engine keeps serving.

This run’s own window · 288 bars · 01 Sept 2026 13 Sept 2026 · both start at 10,000 · one window, far weaker than the corpus above
MOSS-001MOSS-002
Return-1.41%-1.41%
Max drawdown-1.95%-1.95%
Fills1313
Bars held130130
Fees paid44.7044.70

The change never triggered on these 288 bars: MOSS-002 produced exactly MOSS-001's trades, fills and fees. That makes this window inconclusive rather than level — it does not test the change either way, and MOSS-002 stays a candidate until a window arrives that does.

Candidate MOSS-003 · not promotedParent MOSS-001

Candidate, not promoted. MOSS-002 asks the same question as this one but in fixed percent, and the corpus showed it untouched on half its windows: a 3% fall over 36 bars is routine on daily bars and almost unheard of on 15-minute bars, so the guard is dead at one timescale and dominant at another. This measures the same fall against the instrument's own volatility instead, which travels across both. The threshold is one standard deviation of the move, chosen for being the obvious unit rather than for anything it scored — the point of the change is that the guard engages at all, not that this number is better than another.

What changed · 1 of 11 parameters
Min Mean Slope Vol-99.00 -1.00
Corpus · 471 windows · 157 datasetsNot separable from chance
Windows ahead45 / 115
Datasets ahead30 / 80
Untouched356 / 471
Effect (95% CI)-0.96 … -0.05points per dataset
p over datasets0.033the independent unit
p corrected0.16512 candidates tested
Could have found±1.12points, at 0.05
Fills compared6041 / 5654
By market family · holding across these is worth more than holding across more of one
crypto / USDahead on 21 of 60 datasetscrypto / other fiatahead on 5 of 9 datasetscrypto / cryptoahead on 2 of 8 datasetsoutside cryptoahead on 2 of 3 datasets
By timescale · fires on · a change that only works on one is not a change that works
15m barsfires on 20% · ahead on 12 of 20 datasets1h barsfires on 9% · ahead on 2 of 8 datasets4h barsfires on 28% · ahead on 9 of 25 datasets1d barsfires on 41% · ahead on 7 of 27 datasets

MOSS-003 was run against its parent on 471 held-out windows across 157 datasets, 6041 parent fills against 5654. It came out ahead on 45 of 115 decisive windows (median difference -0.27 points). 356 windows were untouched by the change and are excluded. Windows from one dataset share an instrument and their indicator history, so the test is taken over datasets: 30 of 80 went the candidate's way. A split that lopsided happens 3.3% of the time by chance, and it runs against the candidate: on this corpus MOSS-003 is worse than the version it was meant to improve. That is a real finding, and the useful kind — the change is rejected on evidence rather than left open.

Typical difference per dataset -0.24 points, 95% of resamples between -0.96 and -0.05. That range excludes zero, which narrows the size of the effect without settling whether it survives another corpus.

With 80 datasets this corpus would have detected a consistent edge of about 1.12 points per dataset. It found none, so whatever MOSS-003 does is smaller than that — which is a bound on the effect, not proof there is none.

On its own this reads as 0.033, under the 0.05 line. It is not on its own: 12 candidates are tested against the same corpus, and asking that many times makes one low number ordinary. Corrected across the family it is 0.165, and MOSS-003 stays a candidate.

Readings over time · 6 corpora
WhenCorpusDatasets aheadEffectBy chanceRead as
13 Sept 202677,760 bars21 / 60-0.2416%Not separable from chance
13 Sept 202677,760 bars21 / 60-0.2414%Not separable from chance
13 Sept 202677,760 bars21 / 60-0.2416%Not separable from chance
13 Sept 202677,760 bars21 / 60-0.2416%Not separable from chance
13 Sept 202677,760 bars21 / 60-0.2416%Not separable from chance
13 Sept 202677,760 bars21 / 60-0.2419%Not separable from chance

Judged 6 times as the corpus grew. The standing has held throughout.

Sensitivity grid pending: 7 settings to run, each a full held-out simulation. They are computed a couple at a time so the engine keeps serving.

This run’s own window · 288 bars · 01 Sept 2026 13 Sept 2026 · both start at 10,000 · one window, far weaker than the corpus above
MOSS-001MOSS-003
Return-1.41%-1.41%
Max drawdown-1.95%-1.95%
Fills1313
Bars held130130
Fees paid44.7044.70

The change never triggered on these 288 bars: MOSS-003 produced exactly MOSS-001's trades, fills and fees. That makes this window inconclusive rather than level — it does not test the change either way, and MOSS-003 stays a candidate until a window arrives that does.

Candidate MOSS-004 · not promotedParent MOSS-001

Candidate, not promoted. MOSS-001 holds until price is 0.2 sigma past its mean, which asks the position for an overshoot after the reversion it was bought for has already happened. This releases at the mean. The expected cost is giving up the overshoots that do arrive, which on a strategy this patient may be where the return lives.

What changed · 1 of 11 parameters
Exit Z0.20σ 0.00σ

Corpus evidence pending. 160 of 160 datasets fetched (115,200 bars); every candidate is scored on each of them once they land.

Sensitivity grid pending: 6 settings to run, each a full held-out simulation. They are computed a couple at a time so the engine keeps serving.

This run’s own window · 289 bars · 01 Sept 2026 13 Sept 2026 · both start at 10,000 · one window, far weaker than the corpus above
MOSS-001MOSS-004
Return-1.12%-1.12%
Max drawdown-1.95%-1.95%
Fills1313
Bars held132127
Fees paid44.7044.70

Both versions ran the same engine over the same bars, from the same capital, with the same fee and slippage. A difference over one window of 289 bars is not evidence that the change works, so MOSS-004 stays a candidate and MOSS-001 keeps running.

Candidate MOSS-005 · not promotedParent MOSS-001

Candidate, not promoted. MOSS-001 buys into falling prices by design and then allows an 8% loss before admitting the level did not hold, which is a wide stop for a rule whose entries are all into weakness. This tightens it to 5%. The expected cost is being stopped out of reversions that were going to work, since a stop that fires sooner fires more often on noise.

What changed · 1 of 11 parameters
Stop Loss8.00% 5.00%

Corpus evidence pending. 160 of 160 datasets fetched (115,200 bars); every candidate is scored on each of them once they land.

Sensitivity grid pending: 6 settings to run, each a full held-out simulation. They are computed a couple at a time so the engine keeps serving.

This run’s own window · 289 bars · 01 Sept 2026 13 Sept 2026 · both start at 10,000 · one window, far weaker than the corpus above
MOSS-001MOSS-005
Return-1.12%-1.12%
Max drawdown-1.95%-1.95%
Fills1313
Bars held132132
Fees paid44.7044.70

The change never triggered on these 289 bars: MOSS-005 produced exactly MOSS-001's trades, fills and fees. That makes this window inconclusive rather than level — it does not test the change either way, and MOSS-005 stays a candidate until a window arrives that does.

Candidate ECHO-002 · not promotedParent ECHO-001

Candidate, not promoted. Each rule gets few trades before the rotation moves on, so experiments run for 120 bars instead of 72 and nothing else changes. This was originally written bundled with a carry flag; the corpus rejected the bundle and the sensitivity grid showed the carry flag producing identical results at four of five epoch lengths, so the two were split and this is the half that moves anything.

What changed · 1 of 12 parameters
Epoch Bars72 120
Corpus · 324 windows · 108 datasetsNot separable from chance
Windows ahead137 / 324
Datasets ahead40 / 108
Untouched0 / 324
Effect (95% CI)-1.43 … -0.12points per dataset
p over datasets0.009the independent unit
p corrected0.07212 candidates tested
Could have found±1.49points, at 0.05
Fills compared7137 / 7468
By timescale · fires on · a change that only works on one is not a change that works
15m barsfires on 100% · ahead on 12 of 27 datasets1h barsfires on 100% · ahead on 6 of 27 datasets4h barsfires on 100% · ahead on 13 of 27 datasets1d barsfires on 100% · ahead on 9 of 27 datasets

ECHO-002 was run against its parent on 324 held-out windows across 108 datasets, 7137 parent fills against 7468. It came out ahead on 137 of 324 decisive windows (median difference -0.49 points). Windows from one dataset share an instrument and their indicator history, so the test is taken over datasets: 40 of 108 went the candidate's way. A split that lopsided happens 0.91% of the time by chance, and it runs against the candidate: on this corpus ECHO-002 is worse than the version it was meant to improve. That is a real finding, and the useful kind — the change is rejected on evidence rather than left open.

Typical difference per dataset -0.67 points, 95% of resamples between -1.43 and -0.12. That range excludes zero, which narrows the size of the effect without settling whether it survives another corpus.

With 108 datasets this corpus would have detected a consistent edge of about 1.49 points per dataset. It found none, so whatever ECHO-002 does is smaller than that — which is a bound on the effect, not proof there is none.

On its own this reads as 0.009, under the 0.05 line. It is not on its own: 12 candidates are tested against the same corpus, and asking that many times makes one low number ordinary. Corrected across the family it is 0.072, and ECHO-002 stays a candidate.

Readings over time · 2 corpora
WhenCorpusDatasets aheadEffectBy chanceRead as
13 Sept 202677,760 bars40 / 108-0.676.3%Not separable from chance
13 Sept 202677,760 bars40 / 108-0.677.2%Not separable from chance

Judged 2 times as the corpus grew. The standing has held throughout.

Sensitivity grid pending: 6 settings to run, each a full held-out simulation. They are computed a couple at a time so the engine keeps serving.

This run’s own window · 288 bars · 01 Sept 2026 13 Sept 2026 · both start at 10,000 · one window, far weaker than the corpus above
ECHO-001ECHO-002
Return-0.11%-1.54%
Max drawdown-1.45%-1.58%
Fills2121
Bars held109132
Fees paid84.2583.40

Both versions ran the same engine over the same bars, from the same capital, with the same fee and slippage. A difference over one window of 288 bars is not evidence that the change works, so ECHO-002 stays a candidate and ECHO-001 keeps running.

Candidate ECHO-003 · not promotedParent ECHO-001

Candidate, not promoted. The other half of the bundle: an open position is carried into the next experiment rather than closed because the calendar said so. It costs clean attribution, since a trade can no longer be assigned to the rule that opened it, which was the boundary exit's whole purpose. The grid suggested it changes almost nothing at the epoch lengths tested, which is a reason to test it alone rather than a reason to assume it.

What changed · 1 of 12 parameters
Carry Across Epochs0 1
Corpus · 324 windows · 108 datasetsNot separable from chance
Windows ahead22 / 70
Datasets ahead17 / 49
Untouched254 / 324
Effect (95% CI)-0.46 … -0.01points per dataset
p over datasets0.044the independent unit
p corrected0.17812 candidates tested
Could have found±0.54points, at 0.05
Fills compared7137 / 7055
By timescale · fires on · a change that only works on one is not a change that works
15m barsfires on 12% · ahead on 7 of 9 datasets1h barsfires on 15% · ahead on 3 of 11 datasets4h barsfires on 4% · ahead on 2 of 3 datasets1d barsfires on 56% · ahead on 5 of 26 datasets

It fires 52 points more often at one timescale than another, so it is not being tested the same way at both ends of the corpus.

ECHO-003 was run against its parent on 324 held-out windows across 108 datasets, 7137 parent fills against 7055. It came out ahead on 22 of 70 decisive windows (median difference -0.42 points). 254 windows were untouched by the change and are excluded. Windows from one dataset share an instrument and their indicator history, so the test is taken over datasets: 17 of 49 went the candidate's way. A split that lopsided happens 4.4% of the time by chance, and it runs against the candidate: on this corpus ECHO-003 is worse than the version it was meant to improve. That is a real finding, and the useful kind — the change is rejected on evidence rather than left open.

Typical difference per dataset -0.33 points, 95% of resamples between -0.46 and -0.01. That range excludes zero, which narrows the size of the effect without settling whether it survives another corpus.

With 49 datasets this corpus would have detected a consistent edge of about 0.54 points per dataset. It found none, so whatever ECHO-003 does is smaller than that — which is a bound on the effect, not proof there is none.

On its own this reads as 0.044, under the 0.05 line. It is not on its own: 12 candidates are tested against the same corpus, and asking that many times makes one low number ordinary. Corrected across the family it is 0.178, and ECHO-003 stays a candidate.

Readings over time · 18 corpora
WhenCorpusDatasets aheadEffectBy chanceRead as
13 Sept 202677,760 bars17 / 49-0.3322%Not separable from chance
13 Sept 202677,760 bars17 / 49-0.3318%Not separable from chance
13 Sept 202677,760 bars17 / 49-0.3319%Not separable from chance
13 Sept 202677,760 bars17 / 49-0.3322%Not separable from chance
13 Sept 202677,760 bars17 / 49-0.3322%Not separable from chance
13 Sept 202677,760 bars17 / 49-0.3318%Not separable from chance
13 Sept 202677,760 bars17 / 49-0.3318%Not separable from chance
13 Sept 202677,760 bars17 / 49-0.3318%Not separable from chance
13 Sept 202677,760 bars17 / 49-0.3320%Not separable from chance
13 Sept 202677,760 bars17 / 49-0.3318%Not separable from chance
13 Sept 202677,760 bars17 / 49-0.3321%Not separable from chance
13 Sept 202677,760 bars17 / 49-0.3322%Not separable from chance
13 Sept 202677,760 bars17 / 49-0.3318%Not separable from chance
13 Sept 202677,760 bars17 / 49-0.3322%Not separable from chance
13 Sept 202677,760 bars17 / 49-0.3322%Not separable from chance
13 Sept 202677,760 bars17 / 49-0.3318%Not separable from chance
13 Sept 202677,760 bars17 / 49-0.3322%Not separable from chance
13 Sept 202677,760 bars17 / 49-0.3327%Not separable from chance

Judged 18 times as the corpus grew. The standing has held throughout.

Sensitivity grid pending: 2 settings to run, each a full held-out simulation. They are computed a couple at a time so the engine keeps serving.

This run’s own window · 288 bars · 01 Sept 2026 13 Sept 2026 · both start at 10,000 · one window, far weaker than the corpus above
ECHO-001ECHO-003
Return-0.11%-0.38%
Max drawdown-1.45%-1.71%
Fills2121
Bars held109111
Fees paid84.2584.13

Both versions ran the same engine over the same bars, from the same capital, with the same fee and slippage. A difference over one window of 288 bars is not evidence that the change works, so ECHO-003 stays a candidate and ECHO-001 keeps running.

Candidate ECHO-004 · not promotedParent ECHO-001

Candidate, not promoted. ECHO trades more than either sibling and pays the most in fees, while committing 40% of equity to whichever rule the rotation happened to draw — a rule it has no reason to believe in, since the draw is seeded rather than chosen. Sizing at 25% matches the commitment to the confidence. The expected cost is a smaller share of whatever the rotation gets right.

What changed · 1 of 12 parameters
Position Fraction40.00% 25.00%

Corpus evidence pending. 160 of 160 datasets fetched (115,200 bars); every candidate is scored on each of them once they land.

Sensitivity grid pending: 6 settings to run, each a full held-out simulation. They are computed a couple at a time so the engine keeps serving.

This run’s own window · 289 bars · 01 Sept 2026 13 Sept 2026 · both start at 10,000 · one window, far weaker than the corpus above
ECHO-001ECHO-004
Return-0.12%-0.07%
Max drawdown-1.45%-0.91%
Fills2222
Bars held110110
Fees paid88.2455.09

Both versions ran the same engine over the same bars, from the same capital, with the same fee and slippage. A difference over one window of 289 bars is not evidence that the change works, so ECHO-004 stays a candidate and ECHO-001 keeps running.

Candidate ECHO-005 · not promotedParent ECHO-001

Candidate, not promoted. ECHO's breakout rule enters on a 20-bar high and leaves on a 10-bar low, so it demands a level of strength to get in and accepts half as much weakness to get out — the exit is twice as easy to trigger as the entry. This matches them. The expected cost is holding failed breakouts longer, which is exactly what the tighter exit was protecting against.

What changed · 1 of 12 parameters
Breakout Exit Lookback10 20

Corpus evidence pending. 160 of 160 datasets fetched (115,200 bars); every candidate is scored on each of them once they land.

Sensitivity grid pending: 6 settings to run, each a full held-out simulation. They are computed a couple at a time so the engine keeps serving.

This run’s own window · 289 bars · 01 Sept 2026 13 Sept 2026 · both start at 10,000 · one window, far weaker than the corpus above
ECHO-001ECHO-005
Return-0.12%-0.46%
Max drawdown-1.45%-1.78%
Fills2222
Bars held110117
Fees paid88.2488.09

Both versions ran the same engine over the same bars, from the same capital, with the same fee and slippage. A difference over one window of 289 bars is not evidence that the change works, so ECHO-005 stays a candidate and ECHO-001 keeps running.

Evidence base160 / 160 datasets · 115,200 bars

Every candidate is scored against its parent on all of this. Kraken serves roughly 720 bars at whatever interval is asked for and will not page further back, so depth comes from coarser bars and breadth from more instruments. They are all crypto and they move together, so the instrument count is worth rather fewer independent experiments than it looks; the timescale axis separates more than the symbol axis does.

BarsInstrumentsCandlesCoveringFetched
15m4028,80006 Sept 2026 → 13 Sept 202613 Sept 2026
1h4028,80014 Aug 2026 → 13 Sept 202613 Sept 2026
4h4028,80016 May 2026 → 13 Sept 202613 Sept 2026
1d4028,80023 Sept 2024 → 12 Sept 202613 Sept 2026
Run conditionsRPL-XBTUSD-60m-MU00AIT6
Data modeHISTORICAL REPLAYSourceKraken public OHLC (REST, no key)InstrumentBTC / USD · 60m barsWindow14 Aug 2026 → 13 Sept 2026 UTCBars processed611 / 720Starting capital10,000 quote units per specimenFee0.10% per fillSlippage0.05% against the takerMax position50% of equityWarm-up60 bars before any decisionSeedSPECIMEN-001ExecutionDecide on a closed bar, fill at the next bar's open

Fetched 720 committed 60m bars from Kraken. Dropped 0 malformed, 0 duplicate, 0 gap(s) in the series.