Skip to content
STATISTICSK·M·F
StatisticsSeptember 1, 2026·11 min read
Statistics

How Many Trades Before You Know Your Strategy Works?

Share:

You are 23 trades into a new strategy and you are down. Two things could be true. The strategy has no edge and you are watching it prove that. Or the strategy is one of the best systems you will ever trade and you are standing in an ordinary rough patch. At 23 trades those two possibilities look identical — not similar, identical — and no amount of chart study will separate them. This is not a motivational point about patience. It is arithmetic, and the arithmetic produces an actual number. That number is almost never 30, it is different for every strategy, and for most retail systems it is considerably larger than the entire time most traders give a method before replacing it.

216
trades before P/L alone confirms
a genuine +0.20R edge
18%
of the time a profitable strategy
is still losing after 30 trades
9%
of the time it is still losing
after 100 trades

Where "30 Trades" Came From, and Why It Does Not Answer This

The 30-trade rule is real, but it answers a different question. It comes from a textbook convention about the central limit theorem: with roughly 30 or more observations, the distribution of the sample mean starts behaving like a normal distribution, which lets you use a particular set of statistical tools. It is a threshold for which math you are allowed to use. It was never a threshold for when a small edge becomes visible through noise.

Those are wildly different problems. Detecting a large effect in a quiet dataset takes very few observations. Detecting a small effect in a violently noisy one can take thousands. Trading is the second case. Your average trade might be worth a fifth of your risk unit, while individual outcomes swing between −1R and +3R. The signal is small and the noise around it is roughly seven times larger. Thirty observations of that is not a sample. It is an anecdote with a spreadsheet.

What the Real Answer Depends On

Two numbers determine how long the test must run, and neither of them is your win rate.

  • Your edge: expectancy per trade, measured in R. How much the average trade is worth as a multiple of the amount you risked. This is the signal.
  • Your variability: the standard deviation of your R outcomes. How widely individual trades scatter around that average. This is the noise.

The rough sample size you need is the square of the ratio between them, doubled:

The formula

n ≈ (2 × SD ÷ E)²  — where E is your expectancy per trade in R and SD is the standard deviation of your trade results in R. Two standard errors is the usual rough stand-in for "unlikely to be luck." Halve your edge and the required sample quadruples.

That last sentence is the part worth sitting with. The relationship is quadratic, not linear. A trader with a thin edge does not need somewhat more evidence than a trader with a strong one — they need dramatically more. And thin edges are exactly what most retail strategies have after spread, commission and slippage are subtracted. If you have never calculated your expectancy, you have no way of knowing which category you are in, which means you have no way of knowing how long to wait.

One caveat, stated plainly: this formula assumes your trades are independent and that market conditions stay comparable across the sample. Neither is perfectly true. Both assumptions are generous to you, which means the real requirement is usually higher, not lower.

Three Strategies, Three Very Different Answers

Here are three systems, all genuinely profitable, all plausible. The only differences are the shape of the edge and how much the results scatter.

Strategy profileEdge per tradeVariability (SD)Trades neededVerdict
33% win rate, 1:3 R:R+0.32R1.88R139Provable in months
40% win rate, 1:2 R:R+0.20R1.47R216Provable in about a year
55% win rate, 1:1 R:R+0.10R0.99R396Most traders quit first

All three make money. All three are worth trading. But the third one — the high win rate system, the one that feels most reassuring day to day because you win more often than you lose — needs nearly three times the evidence of the first before its P/L proves anything. Traders systematically prefer the strategy that is hardest to validate, because a 55% win rate feels like confirmation on a weekly basis while a 33% win rate feels like failure on a weekly basis. See profit factor vs win rate for why the comfortable metric is the misleading one.

A Winning Strategy That Looks Broken

The formula tells you when the P/L becomes conclusive. It does not tell you what the P/L looks like on the way there. That is the part that actually destroys accounts, so it is worth simulating rather than assuming.

The table below comes from Monte Carlo simulation: 300,000 runs per cell, each run drawing outcomes from the strategy's own distribution. It answers one question — after n trades, how often is a genuinely profitable system still showing a loss?

Trades completed+0.32R edge+0.20R edge+0.10R edge
20 trades16% still losing25% still losing25% still losing
50 trades11% still losing16% still losing20% still losing
100 trades3% still losing9% still losing14% still losing
200 trades0.6% still losing2.5% still losing7% still losing
300 trades0.1% still losing0.7% still losing4% still losing

Read the top row. One in four traders running a genuinely profitable system is underwater at trade 20. Those traders are not unlucky in some exotic sense — they are the ordinary tail of a normal outcome. They will conclude the strategy does not work. They will be wrong, and there was no information available to them at the time that could have told them so.

Now run the same logic backwards, because it is symmetrical and nobody likes this half. A system that loses at the same rate — expectancy of −0.19R per trade — finishes its first 20 trades in profit about 28% of the time. That is how traders end up funding a method with no edge whatsoever: they ran a short test, it printed green, and the green was noise. Small samples do not merely fail to confirm good strategies. They actively promote bad ones.

The window where traders quit

Between trade 15 and trade 40 sits a zone where a profitable strategy has a very real chance of looking like a failure and a losing strategy has a very real chance of looking like a winner. Almost every strategy switch in retail trading happens inside this window — which means the decision is being made precisely where the data is least capable of supporting it.

What 20 Trades Can Actually Tell You

None of this means the first 20 trades are wasted. It means they answer a different question. Your P/L needs 200 trades. Your execution needs about 20, and execution is where most strategies actually die.

  • Plan adherence. What percentage of the 20 followed your written rules exactly? If it is under 80%, you did not test your strategy — you tested an improvised variant of it, and the result is meaningless regardless of sample size.
  • Entry-criteria drift. Do your logged entries still match the criteria you wrote down, or have you started taking setups that are "close enough"? This shows up in 20 trades with total clarity.
  • Loss discipline. Are your losses capped near −1R, or are several of them −1.6R because you widened a stop? A single unplanned loss can erase the edge of fifteen well-executed trades.
  • Cost load. Add up spread, commission and slippage across the 20 trades and express it in R. If your expected edge is +0.15R and your costs run 0.10R per trade, the version you are trading has almost no edge even though the backtested version does.
  • Life compatibility. Did you actually take the signals, or did you miss a third of them because they fired while you were at work? No sample size fixes a strategy you cannot be present for.

Every item on that list is measurable from a journal after two or three weeks. None of them requires you to know whether the strategy is profitable yet. And this is why the gap between backtest and live results — covered in full in our piece on the execution gap — is usually visible long before the edge is.

Three Legitimate Reasons to Kill a Strategy Early

Waiting for the full sample is the default, not a religion. There are exactly three conditions that justify stopping before the sample completes, and notice that none of them is "it lost money."

1. You cannot execute it in your real life

The signals fire during your working hours. The system needs four screens and you have a phone. Holding overnight keeps you awake and you close positions early to sleep. This is not a strategy failure and it is not fixable with discipline — it is a mismatch between the method and the life you actually have. Kill it at trade 10 and lose nothing.

2. Costs consume the edge

If your measured trading costs are a large fraction of your expected edge per trade, the live version of the system is a different system from the tested one. This is arithmetic and it does not improve with more trades. Either move to a cheaper instrument, take fewer and larger setups, or drop the method. Our breakdown of market vs limit orders covers the largest recoverable piece of that cost.

3. Your adherence is too low to be measuring anything

Below roughly 80% adherence, the sample is not a test of your strategy. It is a test of a hybrid between your strategy and your impulses, and you cannot draw conclusions about either one from it. The fix is not a new strategy — a new strategy will be executed exactly as badly. The fix is covered in why you break your own rules.

You Don't Have 500 Trades. You Have 25 Samples of 20 Strategies.

Here is the reason most traders never resolve this question, and it has nothing to do with statistics.

A trader with three years and 800 logged trades sounds like someone with abundant data. Then you look at the journal: a breakout system for six weeks, an ICT-flavoured variant for two months, a moving-average pullback method until a bad fortnight, then a supply and demand approach, then back to breakouts with a new filter. Eight hundred trades, twenty-five different systems, no sample longer than 40. There is nothing in that journal to draw a conclusion from — not one system in it was ever given enough trades to prove or disprove anything.

Meanwhile a trader who ran one unchanged system for 400 trades knows, with reasonable confidence, whether it works. They may have discovered it does not — which is itself a definitive, valuable answer that the first trader will never obtain about anything.

Worse, the switching is not random in its timing. Traders abandon a method during its drawdown, which is by definition its worst stretch. That means you systematically exit every system at its low point and adopt the next one after seeing its good stretch. Executed repeatedly, this converts a series of viable edges into a guaranteed loss. If that pattern sounds familiar, the sequence for interrupting it is in recovering from a losing streak.

Version your strategy, not just your trades

A journal only accumulates evidence if the strategy label stays constant. Tag every trade with a strategy name and a version number, and treat any rule change as a new version with a fresh counter. K.M.F. Trading Journal tracks R-multiples, plan adherence and per-strategy statistics, so the sample builds itself while you trade instead of being reconstructed from memory afterwards.

How to Run the Test Properly

The whole thing collapses into five decisions, four of which are made before the first trade.

  • Freeze the rules in writing and give them a version number. Any change to any rule creates version 2 and resets the counter to zero. This is the rule that does the most work.
  • Compute your target sample in advance from the backtest expectancy and standard deviation. No backtest? Use 150 to 250 trades for a roughly 1:2 system and accept that the number is a guess.
  • Halve your position size for the duration of the test. A sample you cannot afford to finish is not a sample, and reduced size also removes most of the emotional pressure that causes adherence to collapse mid-test.
  • Set a process checkpoint at trade 20 to 30 — adherence, criteria match, average loss, cost load. This is where you kill the three early-exit failures. It is not where you judge profitability.
  • Evaluate the edge only when the sample is complete, then choose one of three outcomes: keep it unchanged, change exactly one variable and restart the counter, or kill it and move on with an actual answer.

A useful habit alongside this: review on a fixed sample rather than a fixed calendar. Reviewing every Sunday guarantees you will eventually respond to a red week that means nothing. Our weekly review template is built around process metrics for precisely this reason — the weekly rhythm is for maintaining execution quality, not for judging edges.

The Uncomfortable Conclusion

The honest summary is that most retail traders will never accumulate enough trades on any single method to know whether it worked. Not because the number is unreachable — 200 trades is a few months for a day trader and about a year for a swing trader — but because staying with one unchanged system through the stretch where it looks broken requires tolerating exactly the discomfort that strategy-switching exists to relieve.

There is a compensation, though, and it is a large one. Once you accept that the P/L cannot answer the question for another 200 trades, the pressure comes off the P/L entirely. There is nothing to read in it, so you stop reading it. What is left to work on is the part you can measure now and control now: adherence, execution quality, cost, consistency. Those improve on a timescale of weeks. And a trader who spends the waiting period improving those has a materially better strategy at trade 200 than the one they started with at trade 1 — regardless of which way the sample eventually lands.

Key Takeaways

  • The 30-trade rule answers a different question. It is a threshold for which statistical tools apply, not for when a small edge becomes visible through large noise.
  • The real requirement is n ≈ (2 × SD ÷ E)², both in R. A +0.32R edge needs about 139 trades; +0.20R needs about 216; +0.10R needs close to 400. Halve the edge and the required sample quadruples.
  • A genuinely profitable strategy is still losing money after 30 trades roughly 18% of the time, and after 100 trades roughly 9%. Being down at trade 30 is evidence of almost nothing.
  • The symmetry cuts both ways: a losing system finishes its first 20 trades in profit about 28% of the time. Small samples do not just fail to confirm good strategies — they actively promote bad ones.
  • Twenty trades cannot evaluate your edge but can fully evaluate your execution: adherence, entry drift, loss discipline, cost load and whether the method fits your actual life.
  • Only three things justify killing a strategy early — you cannot execute it in your life, costs consume the edge, or adherence is too low to be testing anything. "It lost money" is not on the list.
  • Every rule change restarts the counter. Eight hundred trades across twenty-five systems contains no conclusions; four hundred trades on one system contains a definitive answer, even if the answer is no.

Found this useful? Share it with a fellow trader.

Share

K.M.F. Trading Research

Trading Education

The K.M.F. team builds tools and writes content for serious traders — focused on evidence-based psychology, risk management, and performance analysis.

Track These Metrics Automatically

K.M.F. Trading Journal calculates profit factor, R-multiple, expectancy and more — so you can focus on trading, not spreadsheets.
Download free with a 7-day Premium trial.

Download on Google Play