Methodology · 7 min read

How we score traders: the 5 factors behind the composite score

A leaderboard sorted by return ranks outcomes. Allocating capital requires ranking process. This is what we measure, why each factor is in the score, and what happens to a leader whose score decays.

In short

HyperMirror ranks candidate leaders on a composite score built from five factors — realized PnL consistency, win rate, profit factor, position discipline and account survivability — rather than on raw PnL or ROI. Capital is then allocated in proportion to that score, each leader is isolated in its own sub-account so per-leader performance stays measurable, and score decay moves a leader to probation and reduced allocation while a hard risk breach triggers immediate replacement.

Why PnL and ROI rankings mislead

Every public leaderboard on Hyperliquid answers the same question: who made the most money recently. That is a fact about the past and almost nothing about the process that produced it. Two accounts can post an identical 30-day return while running completely different risk: one held modest size with room to spare on margin, the other ran near-maximum leverage in a thin alt perp and happened to be right.

A return ranking cannot distinguish them, and it systematically prefers the second. Sorting by outcome is mechanically a filter for whoever took the most risk in the direction that happened to work. Over any given month, the top of that list is populated partly by skill and partly by the surviving tail of accounts that sized aggressively — and the accounts that sized the same way and were liquidated are not on the list to warn you.

There are three further distortions specific to perpetual futures. Returns are leverage-dependent, so a percentage figure describes the position size as much as the decision quality. Returns are regime-dependent, so a strategy that only works while funding stays persistently positive looks indistinguishable from a durable edge for as long as that regime lasts. And returns are sample-dependent: one outsized winning trade can carry an entire curve, which reads as an edge but is not repeatable.

If capital is going to follow a leader, the ranking has to describe how the result was produced. That is what the composite score exists to do.

What the composite score is

The composite score is a single comparable number derived from several independent measures of an account's public on-chain history. It exists for two reasons: to make heterogeneous accounts comparable, and to make it hard for one impressive dimension to hide a weak one.

That second property matters more than the first. If a score were simply an average, an exceptional return could compensate for reckless sizing and the account would still rank highly — exactly the failure mode a leaderboard already has. Instead, each factor carries a floor. A candidate that fails a floor does not enter the basket regardless of how strong it looks elsewhere; clearing every floor is a precondition, and the score then determines how much capital a qualified leader receives.

The five factors are deliberately not five views of the same thing. Consistency measures repeatability, win rate measures hit frequency, profit factor measures the economics of wins against losses, position discipline measures behaviour, and survivability measures how close the account has run to failure. A strategy can look excellent on two of these and be uninvestable on a third.

We publish the factor names and what each is intended to capture. We do not publish exact formulas, weights or thresholds, for a straightforward reason: a fully published scoring function on a public leaderboard is a specification for gaming it. Wash-like turnover, cosmetic trade splitting and short-horizon curve shaping are all cheap for an account that knows precisely what is being measured.

Factor 1 — Realized PnL consistency

Consistency asks whether realized profit recurs across sessions and conditions, or whether it is concentrated in a handful of trades. A curve where a large share of total profit comes from one or two positions is not evidence of a repeatable process; it is one correct call, which is not something you can allocate against.

This factor is deliberately built on realized PnL rather than mark-to-market equity. Unrealized profit on an open perp position is a price quote, not a result. An account can carry a large paper gain for weeks and close it flat or negative, and an equity-based measure would have credited it the whole time.

Consistency also captures something a return figure cannot: whether the account produced results in more than one market state. An account whose entire history sits inside a single trending quarter has one observation, not a track record.

  • Distinguishes a repeatable process from one lucky position.
  • Uses closed results, so open-position marks cannot inflate it.
  • Rewards accounts that produced profit across more than one regime.

Factor 2 — Win rate

Win rate is the share of closed trades that ended profitable. On its own it is close to meaningless, and it is included precisely because its interaction with the other factors is informative.

A low win rate is entirely normal for trend and breakout strategies, which accept many small losses to catch a few large moves. A high win rate is normal for mean-reversion and scalping strategies. Neither is inherently better. What matters is whether the win rate is coherent with the payoff profile: a high win rate paired with occasional very large losses is the signature of an account that lets losers run, and a low win rate paired with small average wins is a strategy with no arithmetic path to profitability.

Win rate is also the factor most sensitive to how trades are counted. An account that splits one position into many partial fills and closes can inflate its apparent hit rate, which is one of several reasons the score is evaluated across factors rather than any single one.

Factor 3 — Profit factor

Profit factor compares gross profit to gross loss over the measured history. It is the most direct answer to the question a win rate cannot answer: when this account is right, how much does it make relative to what it gives back when it is wrong.

This is where the recurring failure mode of high-win-rate perp strategies becomes visible. Selling into every move, averaging down and closing small winners produces a flattering win rate and a profit factor barely above break-even — a strategy whose entire history is a series of small gains funding one eventual large loss. Profit factor catches that; win rate alone does not.

It has a known weakness. Over a short history, or an account with few closed trades, profit factor is unstable — a single outsized trade moves it substantially in either direction. That is why it is read together with consistency and sample depth rather than treated as decisive on its own.

  • Exposes strategies that survive on many small wins and one large loss.
  • Read alongside win rate, since neither is interpretable alone.
  • Unstable on small samples, so sample depth is part of the assessment.

Factor 4 — Position discipline

Discipline is the behavioural factor. It looks at how the account trades rather than what it earned, and it is the factor that most often disqualifies an otherwise impressive candidate.

The observable signals on Hyperliquid are concrete. Is position size stable relative to account equity, or does it jump erratically? Does effective leverage stay inside a band, or does it spike to near-maximum on individual entries? Does the account add to losing positions as they move against it? Does it concentrate size in thin alt perps where slippage and liquidation cascades are materially worse than in BTC or ETH? Does it hold large one-sided exposure through funding resets, paying a recurring cost that has nothing to do with the directional thesis?

Discipline matters more for a copied leader than for a trader operating their own capital. A mirrored account is executing against your balance and your risk limits, on markets whose liquidity you did not choose. An erratic sizing profile that a leader can absorb across their own book becomes concentrated risk once it is mirrored — and it makes the mirrored position harder to fill at a comparable price.

Factor 5 — Account survivability

Survivability asks how close the account has come to not existing. Two accounts with the same return are not equivalent if one of them repeatedly traded within a few percent of liquidation.

The signals here are margin headroom maintained during adverse moves, the depth and shape of historical drawdowns, whether recovery came from disciplined re-entry or from doubling down, and continuity — how long the account has traded and whether it kept operating through at least one genuinely adverse period rather than stopping and restarting after each loss.

This factor exists because liquidation is not a drawdown. A drawdown is something a strategy trades out of; liquidation is a permanent, realised end to the position and the capital that would have compounded from it. On leveraged perpetual futures, an account with a strong return and thin margin headroom is not a strong candidate that occasionally cuts it fine. It is a strategy whose historical returns were produced by accepting a real probability of total loss, and that probability does not disappear because it has not yet been realised.

  • Margin headroom held through adverse moves, not just on average.
  • Drawdown recovery behaviour: disciplined re-entry versus doubling down.
  • Continuity through at least one adverse regime, not a short clean streak.

From score to allocation

Once a candidate clears every floor and receives a composite score, the score determines its share of capital. Allocation is proportional to score rather than split evenly across the basket.

Equal weighting sounds neutral but makes an implicit claim: that the weakest qualifying leader deserves the same conviction as the strongest. Score-weighting says the opposite — exposure scales with evidence, which limits what any single decayed edge costs while still keeping the basket diversified. Per-leader ceilings sit on top of that, so no leader dominates the book even at a high score.

This only works because each mirrored leader trades in its own Hyperliquid sub-account. Isolation is not only about preventing a long and a short in the same market from netting to nothing; it is what keeps per-leader PnL measurable after allocation. If leaders shared one account, their positions would merge into a single net figure and there would be no way to re-score them on live results — the scoring system would be blind exactly where it needs to see.

Allocation and scoring therefore run as one loop: score determines capital, isolation makes the outcome per leader observable, and the observed outcome feeds the next scoring pass.

Score decay, probation and replacement

Scores are not set once at selection. They are refreshed continuously against cached snapshots of public trading history, so a leader's standing reflects recent behaviour rather than the state that qualified it.

Decay and breach are handled on deliberately different clocks. A gradual score decline — consistency weakening, profit factor compressing, discipline drifting — moves a leader to probation: its allocation is reduced and it is queued for replacement at the next rebalance rather than cut instantly. A hard breach, such as a risk-limit violation, does not wait for a rebalance and triggers replacement immediately.

The two-speed design exists because both errors are expensive. Cutting a leader on ordinary variance realises a drawdown and rotates capital into whoever is currently at the top of the leaderboard — typically an account at the peak of its own regime. Holding a genuinely broken edge because the drawdown might be noise is the other half of the same mistake. Running several scored leaders in parallel is what makes the distinction tractable: a leader falling behind while the rest of the basket holds up is information about that leader, not about the market.

None of this is predictive. Scoring describes what an account has already done and how it did it. It reduces the chance of allocating to a fragile strategy; it cannot rule out a leader failing immediately after qualifying, and there is always lag between a strategy breaking and the data showing it.

  • Slow decay: probation, reduced allocation, replacement at the next rebalance.
  • Hard breach: immediate replacement, no waiting period.
  • Parallel leaders provide the baseline that makes decay distinguishable from noise.

Why transparent scoring matters

A scoring method you cannot inspect is indistinguishable from a discretionary decision. If the only thing published is a ranking, there is no way to tell whether a leader was included because it cleared a standard or because its recent return looked good in a screenshot.

Publishing which factors are measured, what each is intended to capture, and what happens when a score decays makes the system accountable in a specific way: the rules are stated before the drawdown, not explained after it. That is the part worth verifying — not the numbers, which will change, but whether the process that produces them is written down.

Perpetual futures remain leveraged instruments, and a rigorously scored basket can still lose money. What scoring changes is the basis on which capital is committed and the conditions under which it is withdrawn.

Side by side

The five composite-score factors and what each is intended to prevent
FactorWhat it measuresFailure mode it catches
Realized PnL consistencyRepeatability of closed profit across sessions and conditionsA curve carried by one or two outlier trades
Win rateShare of closed trades that ended profitableHit rate incoherent with the payoff profile
Profit factorGross profit against gross lossMany small wins funding one eventual large loss
Position disciplineSizing stability, leverage behaviour, market and funding exposureErratic size and near-maximum leverage entries
Account survivabilityMargin headroom, drawdown recovery, continuity through adverse regimesReturns produced by accepting a real chance of liquidation

Past performance is not indicative of future results. Perpetual futures are leveraged instruments and carry a substantial risk of loss, including the loss of your entire position.

Diversified copy trading. On autopilot.

Score-weighted allocation across up to 10 elite Hyperliquid traders, each isolated in its own sub-account. Your funds never leave your account.

Non-custodial · Agent cannot withdraw · Cancel delegation anytime