How to Evaluate Copy Trading Strategy Managers: A 430-Manager Due Diligence Study

July 27, 2026

A transparent, evidence-first framework for assessing cTrader Copy and XM Copy Trading strategy managers, built on an observed sample of 430 managers — what the visible statistics do and do not tell you, and what happened when we tested one manager’s observable edge against real broker execution.


This is research and due diligence, not investment advice. Nothing here is a recommendation to copy any strategy manager, to use any platform, or to trade at all. Copy trading carries substantial risk of loss. Past performance does not indicate future results. No strategy manager is identified. See the full disclaimer at the end.

About the sample. The 430 managers in this study are the managers observed during the study period — not a claim to cover every strategy manager available on either platform. Copy-trading directories are dynamic: managers can join, leave, pause, restrict visibility, become unavailable, or change behaviour after observation. Every figure below describes this observed sample, under this study’s own definitions, at the time it was observed.


Key takeaways

  • We analysed the visible, platform-displayed evidence for an observed sample of 430 strategy managers across two copy-trading venues — roughly 112,500 closed trades. Of the 419 with enough closed-trade evidence to score, 276 carried at least one critical risk pattern under this study’s definitions — patterns the headline statistics do not reveal.
  • A high win rate is often a warning sign, not a virtue. The most common hazard in the sample was deferred-loss accumulation: closing winners while holding losers open. It produces a near-perfect closed-trade record and a drawdown figure that only appears in a different field on the page.
  • Evidence quality differs enormously between platforms. One venue publishes a complete, reconcilable history; the other exposes a moving window of the most recent trades. The same analysis applied to both without adjustment would systematically favour the one you can see least of.
  • We ran two independent analyses of the same evidence, deliberately blind to each other. They agreed on very little at the top of the ranking — a single manager appeared in both top fives — and the disagreements were more informative than the agreements.
  • The final illustrative shortlist was three managers, referred to here as Manager Alpha, Manager Beta and Manager Gamma. Manager Alpha was the strongest final research case and the only manager independently accepted by both pipelines. This is an illustrative due-diligence outcome, not a recommendation to copy anyone.
  • One manager was wrongly written off because a single availability check failed. Separating "we cannot currently see this record" from "this strategy has deteriorated" restored it to the shortlist. That distinction is a general lesson, not a footnote.
  • We then asked a harder question: does one observable behavioural edge survive real execution costs? On smoothed vendor price data it looked profitable. On real bid/ask data it was negative at every stop width tested. The difference was not the strategy; it was the data.
  • That unresolved question is now running as an instrumented demo experiment, not as live capital.

Table of contents

  1. The question we set out to answer
  2. Why picking a copy trading manager is genuinely hard
  3. Why headline statistics mislead — three worked examples
  4. Scope: what evidence we actually had
  5. The evidence-quality asymmetry between platforms
  6. The metrics, defined properly
  7. Pipeline one: forensic manager analysis
  8. Pipeline two: an independent blind audit
  9. Where the two pipelines disagreed
  10. The availability correction
  11. The final shortlist and how allocation was reasoned
  12. The risk features that actually mattered
  13. From observed behaviour to a testable hypothesis
  14. Two independent verdicts on the strategy question
  15. The forward-demo probe
  16. Limitations and assumptions
  17. A practical framework you can apply
  18. Frequently asked questions

Advertisement

1. The question we set out to answer

Copy trading platforms present a ranked list of strategy managers with a handful of attractive numbers next to each: return, win rate, drawdown, number of investors, a risk score. The implicit promise is that these numbers are sufficient to choose.

We wanted to test that promise. Specifically:

Given only the evidence a platform makes visible to a prospective investor, can you reliably distinguish a strategy manager with a durable edge from one whose record merely looks good?

And a second, harder question that follows from it:

If you can identify a behavioural pattern in a manager’s observable trade history, does that pattern still make money once real spreads, commissions and slippage are applied?

This paper reports what we found. It is a due diligence study: an independent audit of available evidence, conducted on visible dashboard data and platform-disclosed statistics. It is not an attempt to obtain, replicate or acquire any manager’s proprietary method, and we make no claim to know what any manager’s actual strategy is.

2. Why picking a copy trading manager is genuinely hard

Three structural problems make this different from picking a fund.

The record you see is not the record that exists. A platform decides what to display. It may show every trade ever closed, or only the most recent few hundred. It may show realised profit while open losing positions sit outside the calculation. It may report a drawdown computed on account balance, which by construction excludes unrealised losses on positions still open. None of these are deceptions; they are design choices. But they mean the number labelled "drawdown" on two different platforms can measure two different things.

You inherit the shape of the return stream, not its size. When you copy a manager, your positions are scaled to your capital. What transfers is the pattern — the sequence of wins and losses, the holding times, the overlapping exposure, the behaviour after a loss. The manager’s absolute profit in their account currency is almost irrelevant to you. This is why every metric in a serious evaluation should be scale-free.

The metrics that look most reassuring are the easiest to manufacture. A 95% win rate is trivial to produce: never close a loser. A 2% drawdown is trivial to produce: measure drawdown on realised balance and keep the losses unrealised. High profit factor follows automatically from both. A manager doing this is not necessarily dishonest — plenty of grid and averaging strategies work this way by design — but the resulting statistics describe the strategy’s bookkeeping, not its risk.

The consequence is that naive ranking by displayed metrics does not merely fail to find good managers. It actively selects for the hazardous ones, because hazardous mechanics produce better-looking numbers than honest ones.

3. Why headline statistics mislead — three worked examples

These are real records from our dataset, anonymised. Each was eliminated by our analysis. Each looked excellent on the surface.

Example 1 — the near-perfect win rate. This manager’s visible history showed 199 wins in its last 200 closed trades. The same page reported a 27.5% equity drawdown. Both figures were accurate. The winners had been realised and appeared in the trade history; the losses were sitting in open positions and appeared only in the equity calculation. A prospective investor reading the win rate would see a 99.5% success rate. A prospective investor reading the equity drawdown would see a strategy that had lost more than a quarter of its value at some point. The second number is the true one.

Example 2 — realised profit against floating loss. This record showed approximately +125 in realised closed profit while carrying roughly −126 in floating losses across 23 open positions. Net, the strategy was slightly underwater. The closed-trade record — the part that populates most performance statistics — showed only the +125.

Example 3 — the spectacular young account. This one reported over +35,000% return on an account less than two months old. That is not evidence of an edge. It is evidence of leverage, and a copier inherits leverage in full. Short, extreme records are the easiest to produce and the least informative.

None of these problems is visible from the headline figures alone. All three become obvious the moment you cross-reference the closed-trade record against the equity-based figures and the open position list — which is precisely what a systematic evaluation should do.

4. Scope: what evidence we actually had

Strategy managers observed during the study period 430
Split by venue 76 / 354
Managers retaining usable metadata or trade evidence 428
Managers with enough closed-trade evidence to score 419
Closed trades in the consolidated record 112,498
Split by venue 50,855 / 61,643
Open positions observed at final reading 1,913
Managers surviving every disqualifying gate (primary pipeline) 41
Managers eliminated from the study’s research shortlist 378

The trade record was built from platform-visible strategy statistics and trade history, deduplicated across repeated readings so that a manager observed several times contributes each trade once. Of the observations that appeared more than once, 43,125 repeated trade keys were checked for consistency and none had conflicting values — a useful internal check that the record is stable over time.

What this sample is, and is not. These 430 are the managers visible to the study during the observation period. They are not the complete live directory of either platform, and no figure in this paper should be read as a statement about every manager available to you today. Copy-trading directories change continuously: managers join, leave, pause, restrict their visibility, become temporarily unavailable, or change how they trade. A manager absent from this sample is not thereby better or worse than one in it.

Two further caveats, both surfaced by the independent audit rather than by the pipeline that produced the numbers:

  • 419 is not even the whole observed sample. It is the subset with sufficient realised-profit evidence to score. Nine managers — seven on one venue with too few profit observations, two on the other with metadata only — fall outside it. "419 scored" is accurate; "419 analysed" would overstate it, and "all managers" would be plainly wrong.
  • The consolidated record is not a perfect union. One trade observed repeatedly in the underlying evidence is absent from the consolidated ledger: 112,498 against 112,499 distinct keys. It changes nothing in any ranking, and we report it because a study that claims perfect data integrity without checking is not a study you should trust.

5. The evidence-quality asymmetry between platforms

This is the most important methodological finding in the study, and the one most likely to be useful to a reader evaluating managers on their own.

Venue 1 publishes complete history. For managers here, the full closed-trade record is available and reconciles against the platform’s own reported lifetime figures. In one verified case the net profit we computed from the trade record matched the platform’s reported figure to the cent. When your independently recomputed total agrees with the platform’s total, you have strong evidence that you are seeing the whole record.

Venue 2 publishes a moving window. Here the visible trade history is capped at the most recent 200 closed positions. Anything older has left the window. You can accumulate more by observing repeatedly over time, and we did, but you cannot recover what closed before observation began.

Now combine that with the deferred-loss pattern from §3. A manager that closes winners promptly and holds losers open will, in a moving window of recent closes, appear systematically better than it is — because the window fills up with the trades the strategy chose to close, and the trades it chose not to close are invisible to it by construction.

That is not a hypothesis. In our sample, 71 of 335 managers on the windowed venue showed a visible window materially more flattering than their own lifetime figures: an observed win rate more than five points above the lifetime figure, or a window containing less than half of the worst loss the platform itself reported.

The methodological response was to classify every manager by evidence quality before scoring anything:

Class Meaning
Authoritative Full history; independently recomputed totals reconcile with platform-reported figures
Representative Visible window broadly consistent with lifetime figures
Window-biased Visible window materially more flattering than the lifetime record
Insufficient Fewer than 30 observed trades

Evidence quality then scales the final score. A flattering window cannot buy a top ranking. Every surviving manager on the complete-history venue was authoritative; no manager on the windowed venue can be, by construction — a structural ceiling on confidence, not a criticism of any individual manager.

The practical lesson for readers: before comparing two managers, ask whether you are seeing the same kind of evidence for both. If one record is complete and one is a recent window, the window needs to clear a higher bar, not the same bar.

6. The metrics, defined properly

Copy trading pages use these terms loosely. Here is what each actually measures and where each fails. Plain-English definitions first, formula second.

Win rate — the share of closed trades that made money.

win rate = winning trades / total closed trades

Fails when: losers are held open rather than closed. Win rate says nothing about the size of wins versus losses, so it is uninformative on its own and actively misleading when paired with asymmetric holding behaviour.

Payoff ratio — how big the average win is relative to the average loss.

payoff ratio = average win / average loss

Win rate and payoff ratio are only meaningful together. A strategy needs

win rate > 1 / (1 + payoff ratio)

to break even before costs. A 75% win rate sounds excellent, but if the average loss is four times the average win, break-even sits at 80% — and the strategy loses money.

Profit factor — gross profit divided by gross loss, over the same period.

profit factor = sum of winning trades / |sum of losing trades|

Above 1.0 means the record made money. Fails when: the losing side is incomplete — unclosed losers are simply not in the denominator. A profit factor computed from a window that excludes open losing positions can be arbitrarily high and mean nothing.

Maximum drawdown — the largest peak-to-trough decline.

drawdown(t) = (peak equity up to t − equity at t) / peak equity up to t
max drawdown = maximum of drawdown(t) over the period

The critical subtlety: if "equity" means realised account balance, unrealised losses on open positions never enter the calculation. A grid or averaging strategy can report a 1–2% balance drawdown while carrying far larger open exposure. If it means mark-to-market equity, it captures the real experience. Platforms differ. Always check which one you are reading.

Recovery factor — net profit divided by maximum drawdown; how much return was earned per unit of worst-case pain.

recovery factor = net profit / max drawdown

Fails when: the drawdown denominator is understated, which inflates it in exactly the cases where you most want it to be honest.

Evidence coverage — how much of a manager’s stated lifetime activity you can actually see.

evidence coverage = observed closed trades / platform-reported lifetime trades

A manager with 200 visible trades against 252 reported has ~79% coverage. One with 61 visible trades over 69 days has a short record regardless of coverage. Both matter, and they are different questions.

Active-day correlation — whether two managers lose money at the same time.

Take each manager’s daily profit or loss, standardise it by that manager’s own typical trade size so the two are comparable, and correlate the two series across days when both were active:

r = corr( standardised daily P/L of A , standardised daily P/L of B )

Fails when: the overlap is thin. A correlation computed over 14 shared active days is not an estimate you should allocate on. We required at least 20 shared days before trusting a figure, and even then treated it as weak evidence — because the correlation that matters is the one during a crisis, and crises are rare in any short sample.

Risk per trade and open-risk cap — the two numbers that govern position sizing.

position size = (account equity × risk fraction) / (stop distance in pips × pip value)
total open risk = sum over open positions of (entry − stop) × size

An open-risk cap limits the second quantity, so that several simultaneous positions cannot combine into an exposure none of them would have created alone.

Advertisement

7. Pipeline one: forensic manager analysis

The primary analysis was a deterministic pipeline: same inputs, same outputs, byte-identical across runs, with every result regenerable from the stored evidence.

Roughly 50 features per manager, all scale-free. Every profit was normalised by that manager’s own median absolute trade result, so a manager trading small and one trading large are directly comparable. The features spanned:

  • performance — profit factor, expectancy, payoff ratio;
  • risk and tails — drawdown, ulcer index, skewness, worst-loss ratios, losing streaks;
  • statistical strength — bootstrap confidence floors with a fixed seed, t-statistic, daily-aggregated Sharpe;
  • stability over time — first half versus second half, rolling-block profitability, trend slope;
  • execution behaviour — holding times, concurrency, size escalation after losses, grid clustering.

Behavioural classification. Each manager was assigned a data-derived archetype describing what the trade record is consistent with — not what the manager says they do, and not a claim to know their actual method. The taxonomy came from where managers genuinely cluster in the data: deferred-loss accumulator, martingale/recovery, grid basket, high-frequency scalper, intraday directional, swing positional, trend runner, mean-reversion fader, news-event burst, low-activity selective, diversified systematic, single-instrument specialist.

Crucially, measured mechanics and inferred conclusions were stored in separate fields, each inference carrying its own confidence. A directly measured fact — median holding time, share of same-side add-ons, maximum concurrency — is never presented as equivalent to a reasoned inference about strategy type.

Scoring used five weighted components against absolute anchors rather than peer percentiles: statistical edge (24%), risk-adjusted return (20%), drawdown control (20%), consistency (18%), tail safety (18%). Using absolute anchors matters: a manager’s score depends only on its own evidence, so it does not change when the peer set changes, and the method never manufactures a "top 10" out of a weak field.

Elimination gates were disqualifying, not deductions. A manager was removed outright for any of:

  • fewer than 40 observed trades, or under 30 days of history;
  • profit factor at or below 1.05, or non-positive expectancy;
  • a bootstrap 5th-percentile profit factor below 1 — the record does not survive resampling;
  • an edge that decays to negative in its second half;
  • any critical risk flag — hidden drawdown behind a near-perfect win rate, martingale sizing, profit offset by floating losses, no realised losses at all, unsustainable implied leverage, or a reported equity drawdown of 40% or more.

What "eliminated" means here. It means removed from this study’s research shortlist, under this study’s gate definitions, on the evidence visible at the time. It is not a judgement that a manager is objectively bad, unsuitable for every investor, or doing anything improper. Many eliminated records belong to strategies that are working as designed — grid and averaging methods, for instance, are legitimate approaches whose risk simply does not fit the profile this study was screening for. A different set of gates, or a different investor’s risk tolerance, would produce a different shortlist.

The elimination breakdown, in that light:

Primary elimination reason Managers
Short observed history 113
Not robust to resampling 59
Hidden drawdown behind near-perfect win rate 41
Martingale / recovery sizing 39
Too few observed trades 38
No realised edge (profit factor ≤ 1.05) 37
Severe reported equity drawdown (≥ 40%) 26
Unsustainable implied leverage 19
Profit offset by floating losses 3
Decaying edge 2
No realised losses ever 1

378 of the 419 scored managers were eliminated by this study’s disqualifying gates, and 276 of those 419 carried at least one critical risk pattern under this study’s definitions. Grid mechanics were the primary classification for 109 managers and recovery/martingale sizing for 99. Seventy showed the deferred-loss signature directly.

Note what dominates that table: the two largest categories — 113 and 38 — are about insufficient evidence, not bad behaviour. A manager with a short record has not failed a performance test; it has simply not yet produced enough evidence to assess.

Did the score actually predict anything? This was tested by chronological holdout: compute the score using only trades closed before a cutoff, then measure what happened afterwards. Platform statistics and open positions were withheld from the training view to prevent look-ahead. Four cutoffs were used, spaced across roughly the final two months of the study window:

Cutoff Managers Rank vs forward (Spearman) Top-quintile forward Bottom-quintile forward
1 (earliest) 108 0.31 1.07 0.27
2 131 0.39 1.15 0.30
3 160 0.40 1.33 0.34
4 (latest) 156 0.39 1.60 0.77

The continuous score carried a positive rank correlation with forward performance at every cutoff, and the top quintile earned two to five times the forward return of the bottom. Under equal, risk-dominant, return-dominant and consistency-dominant weightings, the top ten retained 9–10 of 10 members — the ranking is not an artefact of the chosen weights.

*But — and the independent audit was right to press on this — the gates were not validated the same way.* Forward realised return cannot detect losses deferred into open positions, which is the exact blind spot the gates exist to cover. We return to this in §9.

8. Pipeline two: an independent blind audit

The second analysis was run deliberately blind. It consumed only the raw evidence, and did not open the first pipeline’s outputs, scores, rankings, reports or code until its own results had been computed, fingerprinted and committed. That commitment is what makes the comparison meaningful: neither pipeline could quietly converge on the other.

Its method differed substantively. Rather than weighted components, it used an equal-component geometric aggregate of five factors — realised edge and bootstrap evidence, chronological persistence, combined realised/platform/open-position risk, evidence completeness, and copyability/execution characteristics. A geometric aggregate has a useful property here: a manager cannot compensate for one very weak dimension with strength elsewhere, because a near-zero factor drags the whole product down.

It also kept eligibility separate from ranking, and penalised deferred losses, stale open positions, observable recovery and grid behaviour, concentration, fragile profits, weak recent results, incomplete status and minimum-investment mismatch.

It produced 40 eligible managers — 17 and 23 across the two venues — against the first pipeline’s 41. Similar count, different membership.

9. Where the two pipelines disagreed

This is where the study earns its keep.

Agreement was thin at the top. Only one manager appeared in the top five of both — Manager Alpha, the long-record mean-reversion specialist. Of the blind pipeline’s top three, all three were subsequently rejected or downgraded once the forensic hazard features were applied:

  • Manager Zeta ranked first blind on an extreme annualised return with a reported drawdown under 3%. It was downgraded to demo-observation-only: on the available fields, that result cannot be separated from leverage, and a copier inherits leverage in full.
  • Manager Eta ranked second blind. The forensic pipeline measured position sizing roughly 3× larger following losses — a recovery-sizing signature — alongside an implied annualised return in the four figures. Rejected.
  • Manager Theta ranked third blind. It combined extreme annualisation with 52.6% same-side add-ons and concurrency of eight. Rejected.
  • Manager Iota ranked seventh blind and was rejected on 1.63× post-loss sizing.

The lesson generalises: a well-constructed score built on returns and risk-adjusted ratios will still rank hazardous managers highly, because their mechanics produce good ratios. Only explicit behavioural detection — measuring what the manager does after a loss, how much exposure overlaps, how long losers are held relative to winners — catches them.

The audit pushed back just as hard in the other direction, and was right to.

  • The gates were not forward-validated. Eligible managers’ forward expectancy was worse than eliminated managers’ at three of four cutoffs. Risk-gated forward drawdown was better at the first two cutoffs and worse at the last two. The honest conclusion: "the continuous score has moderate forward rank signal" is supported by the evidence; "the eligibility and risk gates are forward-validated" is not. We report the second claim as unproven.
  • The bootstrap assumption was too strong. Resampling individual trades independently assumes they are independent. Grid baskets, regime clusters and batched copied orders are serially dependent, so an independence-based lower bound can be optimistically biased.
  • The five score components are not five independent confirmations. They are derived from the same profit-and-loss series, so agreement between them is partly structural.
  • A size proxy is not size. On the complete-history venue, trade volume is not published at all — for all 50,855 trades. The first pipeline inferred a within-manager size proxy from the ratio of profit to price movement, which is proportional to position size for a given instrument. That is genuinely useful for detecting dramatic recovery sizing within one manager’s own record. It is not volume, and must never be described as volume or used to compare across managers.

How the disagreements were resolved. Not by averaging the two rankings, which would have been meaningless. The resolution rule was: a manager must survive both the ranking logic and the behavioural hazard tests. Rank tells you where to look; hazard detection tells you what to exclude. Where they conflicted, exclusion won.

10. The availability correction

One episode deserves its own section, because the mistake is easy to make and expensive.

Manager Beta — a gold intraday specialist with the strongest record on the windowed venue — was the top-ranked manager in the first pipeline and the largest single allocation in its proposed portfolio. In a later evidence refresh, the manager’s record came back as unavailable. The audit concluded, reasonably on that evidence, that the manager was not established as currently copyable, and rejected the portfolio that depended on it.

A follow-up check found the manager present, active and copyable. The earlier failure had been a temporary evidence-availability problem, not a signal about the strategy.

The corrective analysis then did the right thing: instead of simply reinstating the manager, it re-examined what the retained evidence actually said.

  • The visible window was exactly the previously retained window — 200 closed trades, spanning from mid-2025 to mid-July of the study window. Stored and current trade IDs matched 200 of 200, with zero new IDs, zero missing, and zero field conflicts.
  • Current open positions: zero. No hidden basket had accumulated.
  • Window profit factor 2.73, win rate 65%, over 378 days of retained evidence.
  • Investor count had declined modestly, from 129 to 123, and invested capital by roughly 2%.

That no newer closed trades were visible at the final observation point is freshness uncertainty, not evidence loss, and a modest decline in investors is not evidence of strategy deterioration. There was no behaviour change, no new open basket, and no deterioration in the closed-trade record.

The general principle: availability and freshness are separate axes from quality and deterioration. An evaluation should record "I currently cannot see this" as its own state, and must not silently convert it into "this has got worse." The two demand completely different responses — one is a reason to look again, the other is a reason to stop.

There is a second lesson embedded in the correction. When Manager Beta was reinstated, its corrected score placed it 13th on its venue and 22nd cross-platform — yet it was still selected for the shortlist over managers ranked above it. That is not inconsistent. Rank and portfolio construction answer different questions. Rank asks "how strong is this record?" Portfolio construction asks "which combination can this capital actually fund, with acceptable correlation and acceptable hazard?" Several higher-ranked managers had far shorter evidence, or behaviour incompatible with a small account, or minimums that would consume the entire allocation.

By comparison, Manager Delta — which the audit had proposed as Manager Beta’s replacement — carried a very high observed profit factor of 7.55 but only 61 trades over 69 days and a reported drawdown of 18.9%. Substituting a short, high-drawdown record for a 378-day one, on the strength of a single failed availability check, was the wrong trade. The corrected analysis reversed it.

11. The final shortlist and how allocation was reasoned

The final illustrative shortlist was three managers: Alpha, Beta and Gamma. Alpha was the strongest final research case and the highest-confidence candidate for controlled demo observation. Beta was restored after the availability and freshness correction described above. Gamma was capped at the minimum because of measured overlap, add-on and deferred-exit hazards.

Three portfolios were proposed across the study. Showing all three, and why they changed, is more useful than showing only the last — because nothing about the underlying records changed between them.

Manager Alpha Manager Beta Manager Gamma Manager Delta Manager Epsilon Reserve
Initial $402 $372 $225 $1
Post-audit revision $500 $250 $100 $150
Corrected final $450 $350 $100 $100

Illustrative research allocation on a nominal $1,000, for controlled demo observation only. This is a worked example of allocation reasoning, not a recommendation or endorsement of any manager, and not a suggestion that any reader should copy anyone.

Manager Alpha — the anchor allocation, on evidence rather than returns. 2,358 trades over approximately 1,109 days, profit factor near 2.0, roughly 80% winners, no open positions at final observation, and — uniquely in this study — the one manager independently accepted by both pipelines. That cross-pipeline agreement is what makes it the strongest final research case: two methods that disagreed about almost everything else at the top of the ranking both kept it.

Note what it did not win on. Its per-trade edge is the lowest of the shortlist. It receives the anchor weight because its evidence is the deepest and most reconcilable, not because its numbers are the prettiest. That is the whole thesis of the paper applied to a single decision.

It is also not low risk, and the study says so plainly: losing trades are held about 6.1× longer than winners, the worst single loss is worth about 24 average wins, and overlapping exposure reaches 13 simultaneous positions. Any copier would need an account-level equity stop.

Manager Beta — restored to the shortlist. 378 days of retained evidence, profit factor 2.73, 65% win rate, 7.9% reported drawdown, zero open exposure at final observation. It sits below Alpha because the windowed evidence regime caps confidence structurally — not because of anything in its record.

Manager Gamma — capped at the minimum, deliberately. This manager has the longest record in the observed sample — 655 trades across 1,441 days of fully reconciling history, profit factor 2.09. On evidence quality alone it would deserve more. It is capped because its behaviour is the most hazardous of the three: 62% of entries overlap, 60% are same-side add-ons, and maximum concurrency reaches 19. It also carried two small legacy positions, one open for well over a year, aggregating to about −0.38 against more than 6,170 in observed closed profit.

That last figure matters for calibration in both directions. The stale positions are not a material hidden loss — an earlier assessment treated them as sufficient to reject the entire portfolio, and that was an overreaction. But they do directly demonstrate deferred-exit behaviour. The proportionate response is neither rejection nor indifference: it is a minimum-sized allocation — participate, cap the exposure.

Why the weights changed at all. Between the three versions, no manager’s underlying record changed. What changed was the interpretation: an availability state was corrected; a measured hazard was priced as a cap rather than used as a veto; and an operational reserve was introduced. The final allocation holds $100 unallocated — not risk capital, but a buffer for platform friction, minimum-investment changes and freshness checks. The initial portfolio’s $1 reserve was, in hindsight, an optimisation artefact rather than a decision.

On correlation. Standardised daily profit-and-loss correlations across the shortlist were approximately 0.02 (Alpha–Beta, over 50 shared active days), 0.08 (Alpha–Gamma, over 87 shared days) and 0.04 (Beta–Gamma, over only 14 shared days). The first two are useful evidence of low realised dependence. The third is not — 14 shared days is too thin, below our own 20-day threshold, and we flag it rather than quoting it as diversification. No correlation measured on calm days predicts correlation during a liquidation event.

To restate the framing, because it matters: this shortlist is an illustrative due-diligence outcome on an observed sample, produced to test a method. It is not advice, not an endorsement, and not a list anyone should act on. What is intended to transfer is the reasoning: deepest evidence earns the anchor weight, structurally limited evidence earns less, measured hazard earns a cap rather than a veto, and thin correlation estimates get flagged rather than trusted.

12. The risk features that actually mattered

Ranked by how often they changed a conclusion in this study.

Deferred-loss accumulation. The dominant pattern: closing winners while holding losers open. The closed record looks near-perfect; the equity-based drawdown is the only tell. Seventy managers showed the signature. Detect it by comparing win rate against reported equity drawdown, and by checking whether open positions carry material floating loss.

Post-loss size escalation. Position size rising after losing trades — recovery or martingale behaviour. It produces smooth equity curves until it does not. Two managers rejected in this study showed roughly 3.0× and 1.63× median size increases following losses. Detect it by comparing typical size after a loss against typical size after a win. Above about 1.4× warrants serious scepticism.

Overlapping entries and same-side add-ons. Multiple positions open simultaneously in the same direction. This is not automatically bad — our own research found bounded overlap genuinely improved risk-adjusted returns — but it multiplies exposure to a single adverse move. The distinction that matters is whether add-ons are bounded and conditional (a limited number, each requiring a stronger signal than the last) or unbounded averaging (adding indefinitely to a losing position). Shortlisted managers ranged from 29.6% to 60.2% same-side add-ons, with maximum concurrency from 13 to 19.

Holding-time asymmetry. The ratio of losing-trade duration to winning-trade duration. A manager holding losers 6× longer than winners is, mechanically, converting a would-be loss into a lengthy open position. This single ratio explains a great deal of what makes high win rates possible.

Loss tails. The worst single loss expressed in units of average win. A manager whose worst loss equals 24 average wins needs a long run of wins to recover from one bad trade. Average metrics hide this entirely; it is a property of the tail.

Maximum concurrency. The largest number of positions open at once. It sets the worst-case simultaneous exposure you inherit.

Manager age and trade count. The two cheapest and most reliable filters available. A 61-trade, 69-day record cannot support a confident conclusion regardless of how good the ratios look. Our gates required at least 40 trades and 30 days as a floor, and in practice we weighted multi-year records far more heavily.

Stale open positions. Positions left open for extended periods. Individually they may be immaterial; as a behavioural signal they reveal how a manager treats losers.

Active-day overlap and correlation. Discussed in §6 and §11. The key discipline is refusing to quote a correlation computed on too few shared days.

Advertisement

13. From observed behaviour to a testable hypothesis

The second half of the project asked a different question. If observable trade behaviour reveals something about how a manager operates, is that something independently valuable — or does it only work in the manager’s hands, with their unobserved judgement and infrastructure?

We tested this on Manager Alpha — a single-instrument specialist trading one FX cross — because it had by far the deepest and most reliable evidence in the observed sample.

What was independently established. Comparing the manager’s entry timestamps against matched no-trade control periods — a leak-free design, controlling for time of day so that "entries happen during active hours" cannot masquerade as an edge — produced a clear and statistically overwhelming result:

  • roughly 77.6% of buy entries occurred below the 20-period moving average, and 71.5% of sells above it;
  • median signed 60-minute price displacement at entries was about −9.6 pips, against −0.1 pips for matched controls;
  • median signed 60-minute z-score was −1.919 at entries, against 0.005 for controls, with about 94.3% of signed z-scores negative;
  • every tested displacement horizon from 5 to 120 minutes was overwhelmingly significant.

The conclusion is narrow and we state it narrowly:

On this instrument, this manager usually enters against a statistically extreme recent price displacement.

That is a behavioural characteristic inferred from observable evidence. It is emphatically not the manager’s strategy. It says nothing about their instrument selection, entry timing within the extension, position sizing, basket management, or exit discretion. Two independent analyses, using different price data at different resolutions, agreed on this one fact and agreed that it was not enough to reproduce anything.

What the analysis also revealed — and this turned out to be the important part. The manager’s ~80% win rate is not produced by entry selectivity. Win rate is essentially flat across entry extension buckets, from 0.78 to 0.835 regardless of how stretched the market was. The win rate comes from the exit: holding through adverse movement until price reverts, rather than stopping out.

That is a profound problem for anyone hoping to systematise it. The behaviour that produces the attractive statistic is precisely the behaviour that creates unbounded tail risk. Any hard stop lowers the win rate. A wide stop preserves most of it — but exposes the fat left tail that the manager’s own record shows in its worst-loss figure.

Bracket structure testing made this concrete. Applying fixed take-profit/stop-loss brackets to the manager’s own entry points, measured in units of average true range:

Take-profit / stop-loss Hit rate Expected return per unit risk
1.0 / 1.0 51.3% +0.03
2.0 / 1.0 32.8% −0.02
1.4 / 2.0 63.7% +0.17
1.0 / 2.0 72.5% +0.17
0.75 / 2.5 85.5% +0.28

The edge lives in a small take-profit with a wide stop — the same asymmetric structure the manager’s own behaviour implies. Larger take-profits weaken it; letting winners run makes it negative. Entry timing barely mattered (shifting entries by up to three bars moved the 1:1 hit rate only between 0.51 and 0.52); entry structure mattered enormously.

14. Two independent verdicts on the strategy question

Both research tracks then tried to turn behavioural understanding into a validated, executable strategy. Both failed — in different ways, for different reasons, and the pair of failures is more instructive than either alone.

14.1 The vendor-data problem

The first track built a regime-gated mean-reversion strategy on the structure identified above, with a genuinely sealed holdout: the final evaluation period was hash-locked and opened exactly once.

On smoothed vendor price data, results were positive and survived stress testing:

Test Result
Training period +3.54%, PF 1.18, 75% win rate, 3.1% drawdown
Sealed holdout (single run) +1.67%, PF 1.26, 77% win rate, 2.1% drawdown
Holdout with +3 pip spread still positive
Holdout with two-bar entry delay still positive

On real bid/ask data from a tick-level provider, the same strategy was negative at every stop width tested, with an observed win rate of 62–71% against a break-even requirement of about 77% for the 0.75/2.5 structure.

The explanation is mechanical, and it is the most transferable technical lesson in this paper. Smoothed vendor OHLC data reports a bar’s high and low as summary values. It does not reproduce the intrabar path. A wide stop placed 2.5 ATR away is only triggered by a genuine excursion — and smoothed data systematically under-represents those excursions. The strategy therefore appeared to survive adverse moves that, in reality, would have stopped it out. That inflated the win rate from a real ~64% to a fictitious ~77% — which is precisely the difference between losing and winning for this structure.

If your strategy depends on a wide stop not being hit, you cannot validate it on smoothed OHLC data. The data is not wrong; it is answering a different question than the one your backtest is asking.

A later architectural refinement — allowing a bounded second position on a deeper extension, which the manager’s own 98% same-side overlap suggested — improved the sealed holdout materially (+4.72%, PF 1.45, Sharpe 1.65, versus +1.67% and PF 1.26 for a single position). It did not close the vendor gap. Real-tick win rate stayed at 65.9%. Architecture improved the strategy; it did not change the verdict.

The second track ran a deliberately broader search, on a different instrument, with an execution model built to broker specifications: minute-level bid/ask data, historical spreads, explicit commission and slippage, real volume minimums and steps, mark-to-market drawdown, and a candidate set fingerprinted and committed before the final holdout was opened once.

It explored six strategy families — impulse exhaustion/fade, range breakout, liquidity sweep, volatility compression/expansion, session range, and a supervised model trained on manager entries — across 1,584 screened candidates, 48 finalists, 24 state ensembles and 38 parameter neighbours.

Family Development Combined walk-forward Holdout
Impulse fade +1.74% −4.75% −0.11%
Range breakout −15.64% −20.29% +2.01%
Liquidity sweep −20.27% −23.28% −2.59%
Compression/expansion −0.36% −4.98% +1.18%
Session range −33.32% −32.65% +2.37%
Manager-supervised model −50.66% −30.68% −8.28%
State ensemble −4.72% −19.76% −1.13%

No candidate had all-positive folds. Every one lost under doubled costs. Zero of the 38 parameter neighbours was profitable. The two families showing holdout profit had already failed development and walk-forward before the holdout was opened — so those figures are not validation, they are noise in a period that happened to be favourable. One of them flipped from +2.37% to −2.05% and −4.89% when the bar boundaries were shifted by five and ten minutes.

The conclusion was to build no executable strategy at all from that research, on the correct principle that producing an implementation after a failed validation gate adds deployment risk without validated edge.

14.3 What the audit found in the first track’s own work

The independent audit’s review of the first track’s strategy work found real defects, and they are worth listing because they are common failure modes:

  • Holdout contamination. The final 20% of data had been used in an earlier stress test; the strategy was then redesigned in light of that result, and the same interval was reused and described as never-seen. The headline result reproduced numerically but was not an unbiased holdout. (The later, genuinely sealed holdout described in §14.1 was built specifically to fix this.)
  • Look-ahead in the entry-context study. Entry-context calculations used the closing value of an hourly bar that was still open at many entry timestamps — a subtle leak that inflates apparent entry quality.
  • A reporting discrepancy. A stress result was stated as +9.4% in prose while the committed data showed +4.98%.
  • A non-monotonic model artefact. A sizing defect made higher spread capable of increasing returns in one configuration. That is not robustness; it is a bug, and it is the kind of result that should immediately halt a backtest review.
  • Parity gaps between backtest and live code — differing cooldown handling, holding-time semantics, volume handling, and position ownership. A strategy validated in a backtester that behaves differently from the deployed code has not been validated.
  • Compilation never established. The strategy implementation had never been compiled against the real platform API. When it finally was, it failed to compile — two genuine errors against the actual API surface. A strategy whose implementation has never been built is not a strategy that is ready for anything.

15. The forward-demo probe

One question survived all of this in a form that historical data cannot settle:

Do real broker fills behave like the optimistic smoothed model, or like the pessimistic bid/ask model?

That is not answerable from any vendor’s history. It requires observing actual fills at an actual broker. So the final deliverable is not a trading system — it is a measurement instrument.

What it is. An experimental automated strategy running on a demo account only, on a single instrument and timeframe, implementing the bounded two-position mean-reversion structure the research identified. It exists to record execution reality with enough fidelity to settle the question.

What it is explicitly not. It is not validated as profitable. It is not approved for live capital. It does not replicate any strategy manager’s method — it tests whether a generic, publicly-known structural pattern (fading statistical extension with a small target and wide stop) survives retail execution costs. The manager research told us where to look; it did not supply a strategy to copy.

What it measures. For every trade: the intended reference price against the actual fill, and the difference in pips; the spread at the moment of the decision; commission and swap actually charged; the exit reason as reported by the platform rather than inferred; maximum adverse and favourable excursion; and holding time. For every evaluated bar, including ones where no trade was taken: the full indicator state and the specific reason the signal was rejected. Rejected signals matter as much as taken ones — without them you cannot distinguish "the edge failed" from "the risk controls never let it trade."

The decision rule was fixed in advance, which is the entire point of pre-registration:

  • Realised win rate at or above ~0.74 and positive net profit after real commission and swap → the optimistic model was right; promote to provisional.
  • Realised win rate near ~0.65, or negative net → the pessimistic model was right; abandon the strategy family.

Engineering notes worth generalising. A deployment audit of the probe against the actual platform API — rather than against its own comments — found fourteen material defects, two of them critical. Both are instructive:

  1. The persistence design would have produced no data at all in the platform’s default cloud execution environment, because it wrote to local files a cloud instance cannot access. The strategy would have started, reported healthy, traded, and silently recorded nothing. The runtime choice alone would have determined whether the experiment produced any evidence.
  2. A state-loss interaction changed trading behaviour, not just record-keeping. After a restart with missing state, adopted positions were stored with an entry-extension value of zero. The add-on gate compared each new signal against that zero baseline, found it "deeper," and permitted additional positions at any depth — silently converting a bounded two-position architecture into the unbounded averaging it was designed to exclude. Now, a position whose entry extension cannot be verified blocks add-ons entirely: lost state makes the system more conservative, never less.

The generalisable lesson: an automated strategy’s failure modes are not only in its signal logic. They are in what happens when it restarts, when its environment differs from your assumption, and when its record-keeping fails. Those paths deserve the same scrutiny as the entry rule.

16. Limitations and assumptions

Stated plainly, because a due diligence study that hides its own weaknesses is not one.

On the manager evidence:

  • Visible data only. We studied what platforms display to prospective investors. Managers’ actual account sizes, leverage, internal risk rules and intentions are not observable.
  • Trade volume is not published on one venue at all — for all 50,855 trades. The size proxy used is proportional to size within one manager’s record on one instrument, and is not comparable across managers or instruments.
  • The windowed venue’s moving 200-trade limit bounds how far back any analysis can see. Older behaviour is only partially recoverable, and trades that closed before observation began are permanently unavailable.
  • The sample is what was observable, not a census. The 430 managers are those visible to the study during its observation period. Managers who failed and disappeared beforehand are absent entirely, so the observed sample is by construction more favourable than the full historical population — the classic survivorship problem. Managers who joined, paused, restricted their visibility or changed behaviour after observation are likewise unrepresented. No proportion in this paper should be read as a rate for either platform’s current directory.
  • Balance drawdown excludes floating losses by construction on one venue, so it understates the risk of grid and averaging strategies. Our effective-drawdown floor mitigates but does not eliminate this.
  • The risk gates are not forward-validated, for the structural reason that forward realised return cannot see losses deferred into open positions — the very blind spot the gates exist to cover.
  • Correlation estimates are weak. They are computed on shared active days, in calm conditions, over short overlaps. Tail correlation during a liquidation event is not estimable from this data.
  • Manager behaviour can change at any time without notice, and a strategy manager is not contractually bound to keep doing what their history shows.

On the strategy research:

  • All backtest costs are modelled, not observed. Spread, slippage, commission and swap are assumptions until measured at a real broker — which is exactly why the demo probe exists.
  • Single-instrument concentration and regime dependence both apply; the observed edge was stronger in recent periods, which may not persist.
  • Demo is not live. Demo fills, spreads and rejection behaviour can differ from live. A demo result is necessary evidence, not sufficient evidence.
  • Copy execution differs from direct execution. When you copy a manager, your fills, spreads, fees and timing are yours, not theirs. Copy slippage, minimum volume rounding and fee structures all degrade the copied return relative to the displayed one.
  • Small accounts silently skip signals, which biases a forward test. Position size is (equity × risk fraction) / (stop distance × pip value), and a broker enforces a minimum volume and step. When the computed size falls below that minimum, the trade simply does not happen. The wider the stop, the smaller the computed size — so the signals dropped are systematically the high-volatility ones. Our sizing analysis found that at a nominal $1,000 with 0.5% risk per trade, roughly 16% of this strategy’s signals would fall below a typical minimum and be skipped, against full coverage at around $3,000. Any forward test run on a small account is therefore measuring a calmer subset of its own strategy, and should say so. The running experiment records every such rejection in its own telemetry so the effect can be quantified rather than assumed.

Above all: past results do not guarantee future returns. Every figure in this paper describes what already happened.

17. A practical framework you can apply

Distilled from everything above, for a reader assessing a strategy manager on their own.

Step 1 — Establish what kind of evidence you have. Before any metric, ask: is this a complete history or a recent window? Does the platform tell you the lifetime trade count? If so, divide visible trades by lifetime trades — that ratio is your evidence coverage. A window needs to clear a higher bar than a complete record, not the same bar.

Step 2 — Cross-check the closed record against the equity record. This single step catches the most common hazard. Compare the win rate against the reported equity or balance drawdown. A very high win rate alongside a non-trivial drawdown is the signature of deferred losses. Then look at the open positions list: if it exists, check whether it carries material floating loss.

Step 3 — Challenge a low drawdown rather than admiring it. Ask what it is computed on. If it is realised balance, it excludes open positions by construction, and for a grid or averaging strategy it can be almost meaningless. A 1.4% drawdown on a strategy that holds losing baskets is not a 1.4% risk.

Step 4 — Look at behaviour, not just outcomes. The questions that matter most:

  • How long are losers held compared to winners? A ratio above ~3× is a red flag.
  • Does position size increase after losses? Above ~1.4× typical size after a loss versus after a win suggests recovery sizing.
  • How many positions are open at once, and how often are they in the same direction? Bounded, conditional add-ons are defensible; unbounded averaging into a loser is not.
  • What is the worst single loss worth in average wins? If it is 20× or more, one bad trade undoes a long good run.

Step 5 — Apply hard minimums before anything else. Trade count and record length are the cheapest filters that exist and among the most predictive. A 60-trade, two-month record cannot support a confident conclusion no matter how good its ratios look. Extreme returns on young accounts are leverage, not edge.

Step 6 — Combine win rate with payoff ratio, never either alone. Compute the break-even win rate as 1 / (1 + payoff ratio) and compare. If the actual win rate is not comfortably above it, there is no edge — before costs, let alone after them.

Step 7 — Treat "cannot see it right now" as its own state. If a record becomes unavailable, that is a reason to check again, not a reason to conclude deterioration. Equally, an unrefreshed record is not a fresh one: note when you last saw new activity.

Step 8 — Do not trust a correlation computed on a thin overlap. If two managers share fewer than about 20 active days, you do not have a diversification estimate. And no correlation measured in calm conditions tells you what happens in a crisis.

Step 9 — Observe in demo before funding. Watch the copy relationship in demo for a defined period, with the acceptance criteria written down in advance. Record fills, spread, fees, slippage, maximum floating loss and any change in manager behaviour. Then decide against your pre-written criteria, not against how the period happened to feel.

Step 10 — Write down what would make you stop. Before allocating, define removal triggers relative to the manager’s own demonstrated history rather than round numbers: drawdown exceeding ~1.5× their historical maximum; a losing streak beyond their historical maximum plus three; size after losses rising above ~1.4× size after wins; losers suddenly held more than 3× as long as winners; floating loss exceeding half the profit earned since you started; or trading migrating to instruments absent from the historical record.

18. Frequently asked questions

What is copy trading due diligence?

It is the process of independently assessing a strategy manager’s observable evidence — trade history, platform-disclosed statistics, behavioural patterns and risk characteristics — before allocating capital to copy them. It goes beyond the displayed return and win rate to test whether the record is complete, whether the risk figures measure what they appear to, and whether the manager’s behaviour contains hazards the headline statistics do not show.

Why is a high win rate a warning sign in copy trading?

Because it is trivially easy to manufacture by never closing losing trades. A manager who closes winners and holds losers open produces a near-perfect closed-trade record while accumulating unrealised losses that appear only in the equity drawdown figure. In our study, the single largest category of critical risk was exactly this pattern. Always read win rate alongside payoff ratio and equity-based drawdown.

What is the difference between profit factor and win rate?

Win rate is the proportion of trades that made money. Profit factor is gross profit divided by gross loss. A strategy can have a 90% win rate and a profit factor below 1 if its rare losses are large enough. Profit factor is the more informative of the two, but it can still be inflated when losing positions remain open and therefore sit outside the calculation.

Is a low maximum drawdown always good in copy trading?

No, and it should be challenged rather than admired. Check what the drawdown is computed on. If it uses realised account balance, unrealised losses on open positions are excluded by construction, so a strategy holding losing baskets can display a very low figure while carrying substantial real exposure. Mark-to-market equity drawdown is far more informative.

How many trades does a strategy manager need before their record is meaningful?

Our gates required a minimum of 40 closed trades and 30 days, but that is a floor, not a target. In practice we weighted multi-year records with thousands of trades far more heavily than short records with excellent ratios. A 61-trade, 69-day record with a profit factor of 7.5 is not stronger evidence than a 2,358-trade, three-year record with a profit factor of 2.

Can you tell what strategy a copy trading manager is running?

You can infer characteristics from observable behaviour — typical holding time, whether entries cluster after price extensions, how exposure overlaps, how size changes after losses. You cannot determine their actual method. In our study, two independent analyses agreed that one manager tends to enter against statistically extreme recent price moves, and both agreed that this single fact was nowhere near sufficient to reproduce their selection, timing, sizing or exits.

Why did a strategy that backtested profitably fail on real bid/ask data?

Because smoothed vendor OHLC data reports summary highs and lows and does not reproduce the path price took within each bar. A strategy relying on a wide stop not being hit will appear to survive adverse moves that would really have stopped it out. In our test this inflated the win rate from roughly 64% to roughly 77% — the difference between a losing and a winning system for that structure.

How should I evaluate copy trading managers on platforms that only show recent trades?

Treat the window as a biased sample rather than a record, because the manager’s own closing decisions determine what appears in it. Compare the visible window’s win rate and worst loss against the platform’s lifetime figures: if the window looks materially better, it is window-biased. Require more evidence from windowed records than from complete ones, and check whether the platform reports a lifetime trade count you can use to estimate coverage.

Should I diversify across multiple copy trading strategy managers?

Diversification only helps if the managers are genuinely independent, and that is harder to establish than it appears. Correlations computed on a small number of shared trading days are unreliable, and correlation in calm markets says little about behaviour during a liquidation event. Diversifying across different mechanisms and instruments is more defensible than diversifying across managers who merely have different names.

Is demo testing worthwhile before copying a strategy manager?

Yes, provided the acceptance criteria are written down before you start. Demo observation reveals copy slippage, actual fees, real spreads and any divergence between the manager’s displayed activity and what actually reaches your account. It cannot prove profitability, and demo conditions can differ from live, but it converts several assumptions into measurements.

What does it mean when a study says a strategy manager was "eliminated"?

In this study, eliminated means removed from the research shortlist under the study’s own disqualifying gates, on the evidence visible at the time. It is not a judgement that a manager is objectively bad, unsuitable for every investor, or doing anything improper. Many eliminated records belong to strategies working exactly as designed whose risk profile simply did not match what this study was screening for. The two largest elimination categories were about insufficient evidence rather than bad behaviour, and a different set of gates would produce a different shortlist.

How large does a demo account need to be to test an automated strategy fairly?

Large enough that the broker’s minimum volume never truncates a signal. Position size is account equity times the risk fraction, divided by the stop distance times pip value — so a wider stop produces a smaller size. When that size falls below the broker minimum, the trade simply does not happen, and the signals dropped are systematically the high-volatility ones. A forward test on an account that is too small is therefore measuring a calmer subset of its own strategy. Calculate the widest stop your account can size before you start, and compare it against your strategy’s typical stop distance.


About this research

This study was conducted independently, using only evidence that copy trading platforms make visible to prospective investors. It involved no privileged access, no confidential data, and no attempt to obtain any strategy manager’s proprietary methodology. All strategy managers are anonymised.

The analysis was deliberately run as two independent pipelines that did not see each other’s work until both had committed their results, and the disagreements between them are reported alongside the agreements. Where the two reached different conclusions, this paper explains the disagreement and how it was resolved rather than presenting a tidier consensus that did not exist.

Scope reminder. Every figure describes an observed sample of 430 managers during the study period, assessed under this study’s own definitions, at the time of observation. It is not a survey of every manager available on either platform, and it is not a current statement about any manager’s present behaviour.

Disclaimer. This article is for informational and educational purposes only. It is not investment advice, financial advice, or a recommendation to copy any strategy manager, use any platform, or enter any trade. Copy trading and leveraged trading carry a substantial risk of loss and are not suitable for all investors. Past performance is not indicative of future results. No strategy manager discussed here is identified, endorsed, or criticised as an individual, and "eliminated" throughout means excluded from this study’s research shortlist under this study’s gate definitions — not a judgement that any manager is objectively unsuitable for any investor. Any automated strategy described is experimental, runs on a demo account, and is not validated as profitable. You are solely responsible for your own investment decisions, and you should consider seeking advice from a licensed professional in your jurisdiction.

Trademarks and non-affiliation. cTrader, cTrader Copy, XM, and related names are trademarks or brands of their respective owners. This study is independent and is not affiliated with, endorsed by, or sponsored by any platform or broker mentioned.