Retire a setup when your own logged numbers show the edge has gone, not when it feels cold. That needs three things: a recent sample of at least thirty to fifty occurrences, an expectancy measured across it, and a structural reason the setup should have stopped working. A drawdown by itself is never evidence.
Traders get this decision wrong in both directions, and they get it wrong at the worst possible moment — mid-drawdown, when the loss is loud and the analysis is quiet. Most abandon perfectly healthy setups during entirely normal losing streaks. A smaller group clings to a genuinely dead one for a year because it used to work. The only defence against both is deciding what evidence would change your mind before you need it.
A drawdown is not a dead setup
Start with the statistic that resolves most of these arguments: long losing streaks are a routine feature of profitable strategies, not a warning sign.
A setup that wins 40% of the time loses six in a row with probability 0.66 — about 4.7%, or roughly once in every twenty-one sequences of six trades. Take fifty trades in a quarter and a six-loss streak somewhere inside it is closer to expected than exceptional. Ten in a row at the same win rate happens about 0.6% of the time, which across a career of thousands of trades means it will happen to you, more than once, while nothing whatsoever is wrong.
This is why the streak is the wrong instrument. It measures variance, which is loud, rather than expectancy, which is what you actually care about. The arithmetic behind this — and why a deeper drawdown demands a disproportionately larger recovery — is set out in drawdown explained.
The three numbers that actually reveal decay
Pull the last thirty to fifty occurrences of the setup out of your log and compare them against the same setup's earlier record. Three numbers do the work.
| What to compare | What a decline signals |
|---|---|
| Expectancy in R (avg win × win rate − avg loss × loss rate) | The headline. A fall from +0.3R to −0.1R across fifty trades is a real change, not a bad month. |
| Average win size in R | Falling first, before win rate, is the classic decay pattern: the setup still triggers correctly but the move no longer travels. Usually a volatility or participation change. |
| Frequency (occurrences per month) | A setup that stopped appearing has not decayed — the condition it depends on has left the market. That is a regime question, not a validity one. |
Measuring in R rather than dollars is what makes the comparison legitimate, because it strips out every position-size change you made in between. If your trading journal does not record outcomes in R, this entire diagnosis is unavailable to you, and that is the actual problem to fix first.
Why edges decay in the first place
There are only a handful of real reasons, and naming which one applies is most of the decision.
- It was never an edge. The most common case by a distance. The setup was fitted to a sample, looked good, and reverted to nothing when it met fresh data.
- Crowding. Enough participants trade the same pattern that the move gets front-run and the follow-through disappears.
- Regime change. Volatility compresses, a trending market turns rotational, a correlation breaks. The setup is intact; the conditions it needs are absent.
- Structural change. Contract specifications change, a session's participation shifts, an instrument's liquidity profile alters permanently. This is the one that genuinely kills a setup.
- You changed. The setup is fine and your execution slipped. Painful, common, and the log usually shows it as a rising count of trades marked “rule followed: no”.
On how often the first case applies, the academic record is bracing. In the Review of Financial Studies, Kewei Hou, Chen Xue and Lu Zhang re-tested a library of 452 published stock-market anomalies and found that, once microcap distortions were handled with NYSE breakpoints and value-weighted returns, 65% failed to clear even the ordinary significance hurdle of a t-value of 1.96 — a figure that rises to 82% under a stricter multiple-testing threshold (Hou, Xue & Zhang, “Replicating Anomalies”, RFS vol. 33 no. 5, 2020). These were peer-reviewed findings by professionals with far more data than you have. The base rate for “this pattern was real and then stopped” is much lower than the base rate for “it was never there”.
Write the retirement rule before you need it
The decision has to be pre-committed, for the same reason a stop does: you will not make it well while it is happening. A workable rule has four parts, and it belongs in the written document described in how to build a trading plan.
- A review trigger. Not a retirement trigger — a trigger that starts the analysis. Something like “whenever this setup is net negative over its last thirty occurrences”.
- A size reduction on review. Cut the setup to half size the moment the review triggers, and keep it there until the review concludes. This is the step that makes the whole process survivable, because it caps the cost of being slow.
- A defined evidence bar. Write down now what would make you retire it: for example, negative expectancy across fifty occurrences and a named structural reason.
- A decision deadline. A date by which you will either retire it, restore full size, or formally extend the review once. Setups left indefinitely “under observation” at half size quietly become a permanent drag.
Note the deliberate asymmetry: the review triggers easily and cheaply, while retirement requires real evidence. That ordering is what stops you from killing a good setup on a normal streak while still limiting the damage if the streak turns out to be something worse.
Retire, or repair?
The tempting middle path is adjustment: widen the stop, add a filter, skip the first hour. Sometimes that is right. Usually it is fitting to the noise that just hurt you.
| Repair is defensible when… | Retire when… |
|---|---|
| You can name the structural change first, and the adjustment follows from it | You are searching for a parameter that would have avoided the recent losses |
| The change is one parameter, decided once, and then re-tested | You are on your third adjustment this quarter |
| Frequency held up and only the follow-through faded | The setup no longer appears at all |
| The original logic still describes what the market is doing | You can no longer explain in one sentence why the setup should work |
The non-negotiable part: any repaired setup is a new setup. It goes back through a fresh forward test at reduced size before it earns full allocation again. Skipping that step is how a working strategy gets slowly tuned into a fitted one over the course of a year.
What to do with a retired setup
Do not delete it. Archive the rules, the full log and a short note recording the date, the numbers at retirement and the reason you named. Then set a calendar date — six months is reasonable — to look at it again.
Setups that died of regime rather than structure genuinely do come back. A volatility-expansion setup that stops working in a quiet range is not broken; it is waiting. When you revisit it, the archived log gives you something most traders never have: a clean before-and-after comparison, with the rules unchanged, across two different market environments. That is worth more than the setup itself.
And when you retire one, resist the urge to immediately replace it. The reflex to fill the gap is how traders end up running six half-tested setups at once — a problem covered in how many setups one trader should run. A smaller playbook you can actually measure beats a larger one you cannot.
Frequently Asked Questions
How many losing trades in a row means a setup has stopped working?
No number of consecutive losses is evidence on its own. A setup with a 40 percent win rate produces six losses in a row roughly once every twenty-one sequences of six, which means a long streak is an expected feature of a working strategy rather than a symptom. Judge on expectancy across at least thirty to fifty recent occurrences, not on the streak.
Do trading setups really stop working?
Some do, and the wider evidence is sobering. Hou, Xue and Zhang re-tested 452 published stock-market anomalies and found that 65 percent could not clear an ordinary significance hurdle once microcaps were handled properly, rising to 82 percent under a stricter multiple-testing threshold. Many apparent edges were never edges. Others are real and decay as conditions change.
Should you retire a setup or adjust it?
Adjust only when you can name the structural change that broke it and the adjustment follows from that change rather than from the losses. If the honest answer is that you are searching for a parameter that would have avoided the recent drawdown, you are fitting to noise. Retire it instead, and treat any revision as a new setup that has to be tested from scratch.
Can a retired trading setup start working again?
Yes, particularly when it failed because of a market regime rather than a permanent structural change. A volatility-expansion setup that dies in a quiet range often returns when volatility does. Archive the rules and the log rather than deleting them, set a date to re-examine it, and require it to pass a fresh forward test before it trades size again.
Bottom line
The question is never “does this feel like it stopped working”, because during a normal losing streak it always does. It is whether expectancy has actually fallen across a sample big enough to mean something, and whether you can name a reason it should have. Write the review trigger, the size cut and the evidence bar into your plan while the setup is performing, so the decision is already made when it is not. Then retire on evidence, archive rather than delete, and require anything you revise to earn its size back the same way it earned it the first time.
