Back to Blog

Trading Journal

How to Categorize Trades by Setup

By The TradeReveal TeamOctober 8, 2025

You run three or four setups. You have a hunch about which one carries you and which one bleeds. But it is only a hunch, because your trades sit in one undifferentiated pile. Your win rate is a blend. Your average win is a blend. Every number you look at is an average of things that behave nothing alike.

Categorizing trades by setup fixes that. You attach a setup label to each trade, then read your stats one setup at a time. Suddenly the pullback that felt profitable shows a small negative expectancy, and the breakout you almost stopped trading turns out to carry the account. You cannot see any of that in the blended number.

The hard part is not tagging. It is defining a setup precisely enough that the grouped stats mean something. A label that means five different things is worse than no label, because it looks like data. This post is about drawing those lines tight.

TL;DR

  • A setup is a repeatable, pre-defined pattern of conditions you enter on. If you cannot write down the entry rule, you have a feeling, not a setup.
  • Blended stats hide the truth. Two setups with opposite expectancies average out to a mediocre middle number that describes neither.
  • Define each setup by a small set of required conditions, name it consistently, and tag every trade with exactly one primary setup.
  • Keep the taxonomy small. Five to eight setups is plenty. Splitting too finely leaves each bucket with too few trades to trust.
  • You need a real sample per setup before you rank them. A handful of trades tells you almost nothing, because small samples are noisy by nature.

Why blended stats lie to you

Say you trade two setups. Breakouts: 45% win rate, average winner +2.5R, average loser -1R. Pullbacks: 60% win rate, average winner +0.8R, average loser -1R. Run the expectancy on each.

Breakout expectancy per trade, in R:

(0.45 × 2.5) + (0.55 × -1) = 1.125 - 0.55 = +0.575R

Pullback expectancy per trade, in R:

(0.60 × 0.8) + (0.40 × -1) = 0.48 - 0.40 = +0.08R

The breakout returns roughly +0.58R per trade. The pullback barely clears zero at +0.08R. If you trade both in equal numbers and look only at the blended average, you see something around +0.33R and conclude you have a decent, uniform edge. You do not. You have one strong setup and one that is close to break-even. The blend describes a trader who does not exist.

Expectancy here means the average result per trade over a series, expressed in units of your initial risk (R). That framing comes from trading psychologist Van K. Tharp, who popularized measuring each trade as a multiple of the risk taken rather than in raw dollars (Van Tharp Institute). The practitioner version of the formula, win rate times average win minus loss rate times average loss, is the same idea in dollars (TradeZella).

The arithmetic is beside the point. What matters is that averaging across setups destroys the one comparison you actually need: which pattern pays and which one does not. You can only recover that by splitting the pile first.

Here is the shape of the problem.

What actually counts as a setup

A setup is a repeatable pattern of conditions that triggers your entry. The working test is simple. Could you hand your entry rule to another trader and have them flag roughly the same trades you would? If yes, it is a setup. If the honest answer is "you have to feel it," you are still working from a discretionary judgment you have not written down.

A usable setup definition has a few required conditions and nothing vague:

  • A trigger. The specific event that puts the trade live. A break above the prior day's high. A pullback to a rising 20-period moving average that holds. A failed breakdown that reclaims a level.
  • A context filter. The condition the trigger only counts inside. Trend direction, session, volatility regime, whether it is an earnings day. Most setups only work in one context, and mixing contexts is the quiet way a clean setup turns into mush.
  • An invalidation. Where the idea is wrong. This is usually your stop, and it is what defines your R in the first place.

Notice what is not on the list: the outcome. A setup is defined entirely by what was true at entry. If your definition sneaks in "and it worked," you have stopped classifying and started cherry-picking, and the stats you build on top will flatter you.

Write each definition as one or two sentences and keep them somewhere you will actually reread. The definitions are the contract. When you tag a trade later, you are asking whether it met the contract, not whether it felt similar.

Draw the lines before you tag, not after

The failure mode is tagging by vibe. You close a trade, glance at the chart, and think "that was kind of a breakout," so you tag it breakout. Do that a hundred times and your breakout bucket is a grab-bag of clean breakouts, marginal breakouts, and things that only rhyme with a breakout. The bucket's stats are now an average of unrelated trades, which is exactly the disease you were trying to cure.

The fix is to decide the taxonomy first, as rules, and tag against the rules. Two guardrails make this hold up:

One primary setup per trade. A trade can rhyme with two patterns. Force yourself to pick the one that best describes why you entered. If you routinely cannot decide, your definitions overlap and need sharper edges. You can add a secondary tag for texture, but the primary setup is what you group and rank by, so it has to be single-valued.

A real "other" bucket. Every honest trader takes trades that fit no defined setup: the revenge trade, the bored click, the one-off tip. Do not force those into a real setup to keep it tidy. Tag them "unplanned" or "off-book" and keep them out of your setup comparison. Contaminating a setup with impulse trades is how a good setup gets falsely convicted. Deciding what is worth logging at all, and what belongs in an off-book bucket, is the same discipline covered in what to log in a trading journal.

The flow, start to finish, looks like this.

How granular should each setup category be

There is a real tension here. Split too coarsely and one setup swallows several distinct behaviors, so its stats blur again. Split too finely and each setup ends up with six trades, which is not enough to conclude anything.

Sane defaults:

  • Start with five to eight setups. That is usually enough to capture how you actually trade without shattering your sample into dust. Most retail traders who think they have fifteen setups really have five, plus ten labels for market conditions.
  • Split a setup only when you have a reason and the sample to support it. If your "breakout" bucket has a hundred trades and you suspect morning breakouts behave differently from afternoon ones, that is a legitimate split, because each half will still have enough trades to read. Splitting a twenty-trade bucket into two ten-trade buckets just doubles your noise.
  • Use tags for the dimensions you might slice by later. Session, day of week, and market regime are better handled as separate tags than as brand-new setups. That way you keep one clean "breakout" setup and can still filter it by morning versus afternoon when the sample is large enough. A disciplined approach to what you log on each trade is what lets one setup answer many questions without multiplying the buckets.

The guiding idea: the setup is the strategy, and tags are the conditions you run that strategy in. Keep those two layers separate and your taxonomy stays legible even as it grows.

You need a real sample before you rank

This is where most setup analysis quietly goes wrong. You tag twelve trades of a new setup, see nine winners, and decide it is your best pattern. It might be. It might also be a coin that happened to land heads nine times. At small sample sizes, you genuinely cannot tell the two apart.

This is not a trading opinion. It is a well-documented feature of how people misread randomness. Amos Tversky and Daniel Kahneman named it "belief in the law of small numbers": we expect a small sample to be as representative of the true pattern as a large one, when in fact small samples swing wildly around it (Tversky & Kahneman, 1971, Psychological Bulletin, doi.org/10.1037/h0031322). A twelve-trade streak is exactly the kind of small sample that fools the intuition.

Practitioners who study trade evaluation land in the same place from the data side. A common rule of thumb is that a setup's numbers only start to firm up somewhere past 30 to 50 trades and get reasonably solid past 100, because below that, ordinary variance can produce almost any win rate from a strategy with no real edge at all (Pipup). These are heuristics, not hard thresholds, and the exact number depends on how large and how skewed your winners are. But the direction is not in doubt: more trades, more trust.

Two practical habits follow:

  • Report the trade count next to every setup's stats. A +0.9R expectancy on 8 trades and a +0.3R expectancy on 140 trades are not comparable claims. Always look at the sample beside the number so you are not seduced by a big figure built on nothing.
  • Treat thin setups as provisional. A new or rarely-traded setup goes on a watch list, not into your size decisions. Keep trading it at a normal, cautious size while the sample builds, then rank it once it has the trades to earn a verdict.

Reading the ranked table

Once every trade carries a primary setup and each real setup has a decent sample, the payoff is a single sorted view: setups ranked by expectancy, with win rate, average R, and trade count beside each. That table answers the questions that were invisible in the blend. It only holds up if the per-trade numbers feeding it are honest, which is why scaled and partial-fill trades need their average price and R recorded correctly before you group them by setup.

Look for three things:

  • The clear winners. Setups with positive expectancy and enough trades to trust. These deserve more of your attention, your preparation, and, once you are confident, your size. This is also where a per-trade confidence score pays off, because you can check whether your best-ranked setups are the ones you actually felt sure about.
  • The quiet losers. Setups with negative or near-zero expectancy that you keep trading out of habit. This is often the highest-value finding, because cutting a losing setup improves your account without you having to get better at anything.
  • The mismatches. A high win rate paired with negative expectancy means your losers are far bigger than your winners, so you are winning often and still bleeding. A low win rate paired with strong positive expectancy is the opposite and completely fine. Judge by expectancy, not by how often a setup feels right.

The move after that is rarely to overhaul everything. Make one change: trade the best setup more deliberately, or stop trading the worst one, then let the next stretch of trades tell you whether the ranking held. Categorization works best as a standing lens you keep looking through, not a one-time audit.

If your trades already live in a journal or a portfolio tool, this is mechanical rather than manual. In TradeReveal, a setup is a strategy or tag you attach at entry, and the analytics and Explorer views let you slice performance by that field, so the ranked table falls out of data you already logged instead of a spreadsheet you rebuild by hand.

Frequently Asked Questions

How many setups should I track?

Start with five to eight. That is usually enough to describe how you actually trade without splitting your history into buckets too small to read. If you think you have more, you are probably confusing setups (the strategy) with conditions like session or regime (better handled as tags). Add setups only when you have both a real behavioral difference and the trade count to support the split.

What is the difference between a setup and a strategy?

For tagging purposes you can treat them as the same layer. If you want the distinction: a strategy is the broad approach (momentum, mean reversion), and a setup is a specific, rule-defined pattern within it (opening-range breakout, failed breakdown reclaim). What matters is that the label maps to a written entry rule, not a vibe.

How many trades do I need before a setup's stats mean anything?

There is no single magic number, but the direction is settled. Below roughly 30 trades the figure is barely more than a guess, it firms up past 50, and gets reasonably solid past 100, because small samples swing far from the true pattern by nature (Tversky & Kahneman, 1971). Big-winner, low-win-rate setups need more trades than steady, high-win-rate ones because their results are more skewed. Always read the trade count next to the expectancy.

Should I put a trade in two setups if it fits both?

Pick one primary setup, the one that best describes why you entered, and group by that. Frequent inability to choose is a signal your definitions overlap and need sharper edges. A secondary tag is fine for extra texture, but keep the primary setup single-valued so your grouped stats stay clean.

Where do impulse trades and one-offs go?

Into a dedicated off-book bucket, tagged "unplanned" or similar, and excluded from your setup comparison. Forcing a revenge trade or a random tip into a real setup contaminates its stats and can get a genuinely good pattern falsely convicted.

Final Thoughts

Categorizing by setup is less about extra journaling work and more about a single distinction: the gap between one blurry number that describes an imaginary average trader and a ranked list that tells you which of your patterns to feed and which to starve. The whole method rests on two disciplines: define each setup tightly enough that the bucket holds one behavior, and wait for a real sample before you crown or cut anything. Do both, and your setup table stops being a story you tell yourself and starts being a decision you can act on.

Sources

Start your free TradeReveal account today

Happy Trading,

The TradeReveal Team