Selection · 6 min read

Why raw Hyperliquid leaderboards are misleading

A leaderboard sorted by profit is not a ranking of skill. It is a ranking of outcomes, produced by a sample that has already had its failures removed, over a window short enough that luck can win it.

In short

Raw Hyperliquid leaderboards rank by realized PnL or ROI over a fixed window, which introduces four systematic distortions: survivorship bias, because liquidated and abandoned accounts leave the ranking; short-window luck, because a few leveraged trades can outrank hundreds of disciplined ones; risk blindness, because two accounts with the same return can have taken wildly different leverage and drawdown to get it; and no measurement of position discipline or account longevity. Ranking by outcome selects for variance, so the top of a PnL sort is systematically enriched with accounts that took the most risk and happened to be right. A robust evaluation measures return per unit of risk, consistency across many trades, drawdown depth and recovery, sizing and leverage behaviour, and survival across different market regimes — recomputed continuously rather than read once as a snapshot.

How Hyperliquid traders actually get ranked today

Hyperliquid is unusually transparent. Positions, fills and account equity are on-chain, so anyone can build a leaderboard, and many have. The standard form is the same everywhere: pick a window — 24 hours, 7 days, 30 days, all-time — sort accounts by realized PnL or ROI over that window, and show the top of the list. It is the first screen most people look at before deciding whose trades to follow.

The ranking is not wrong about what it measures. Those numbers are real and verifiable on-chain, without the self-reporting problems that affect performance claims elsewhere. The issue is narrower: a PnL sort answers "who made the most money over this window?" while the reader is asking "whose future trades should I mirror?" Those questions have different answers, and the gap between them is structural rather than noise.

Survivorship bias: the sample is already filtered

A leaderboard shows the accounts still standing. Hyperliquid liquidates positions when maintenance margin is breached, so an account that ran high leverage on a directional view and was wrong does not appear with a large negative number next to it — it stops trading, sits near zero equity, or the operator moves to a fresh address.

That changes the population you are choosing from. If many accounts run the same aggressive approach, some fraction are right for a while and post extraordinary returns while the rest are removed by liquidation. Among the survivors the approach looks like it works, because the evidence of its failure rate was deleted by the same process that produced the ranking. Restarts compound it: a trader on their fourth address shows a clean curve, since the first three are linked to nothing.

Short-term luck versus consistency

The second problem is sample size. Over a week of Hyperliquid perps trading, a handful of leveraged trades can produce a return that outranks hundreds of carefully sized ones. The leaderboard cannot tell them apart, because it sums outcomes without examining how they were produced.

Two accounts can end a month with identical realized PnL via completely different routes: one from three trades where a directional bet paid, the other from a few hundred trades with a modest positive expectancy each. Only the second has enough independent observations to say anything about the process. Shorter windows amplify this to dominance — a 24-hour ranking is close to a pure variance sort, where whoever held the largest position in the direction the market moved sits on top. Trade count, not calendar time, is what makes a result repeatable.

Drawdowns and the risk that produced the return

A PnL or ROI figure is a numerator without a denominator: it says what an account earned and nothing about what it risked. On a venue offering high leverage on perpetuals, that omission is severe. Two accounts with the same 30-day return are not comparable if one never saw equity fall far below its start and the other got there after a drawdown that consumed most of the account. On the leaderboard they are near-identical; as candidates to mirror they are not, because the second operates close enough to the liquidation boundary that an ordinary adverse move can end it.

The dimensions a return figure hides are all measurable from the same on-chain data:

  • Peak-to-trough drawdown and recovery time — how much of the account a bad sequence consumes, and whether the process repairs itself.
  • Leverage actually used rather than the maximum available, which largely determines how survivable the next mistake is.
  • Margin utilisation and proximity to liquidation. A thin maintenance buffer is a risk that shows up in returns only the one time it matters.
  • Adverse excursion on winning trades: positions deeply underwater before closing green indicate a tolerance for being wrong that changes what the win rate means.
  • Return per unit of risk. The same PnL earned with half the volatility is a materially different result.

Position discipline and account longevity

Beyond risk metrics there is a category the leaderboard has no column for: how the trader behaves. Discipline appears as consistency in position sizing rather than escalation, stable holding periods, and losses closed at similar magnitudes rather than widened in hope of recovery. Its absence has a recognisable signature — size increasing after a loss, holding periods stretching when trades go against the position, an occasional trade several times larger than normal. That last pattern is over-represented at the top of short-window rankings, because it is exactly what produces an outlier return when it works.

Longevity is the other missing dimension. Hyperliquid has traded through directional trends, sharp deleveraging events, funding regime changes and long chop. An account that has only traded a trending market has demonstrated fit to the regime it was born in, not that its approach survives a different one. Time in market matters as exposure to conditions, not seniority. A PnL sort penalises none of this; if anything, having no history and no sizing constraints makes an extreme return easier to achieve.

Why high returns alone are a poor selection filter

Together these effects point in one direction: ranking by outcome selects for variance. Among accounts of similar skill, the ones that took the most risk occupy the top rows, because high variance is the only way to produce an extreme result in a short window. The list is not sorted by edge; it is sorted by realised luck with edge as a minor contributor.

That is why mirroring the top of a leaderboard tends to disappoint in a specific way. You are not buying the process behind the number — you are buying the tail of a distribution at its most extreme, which is also the point least likely to repeat. Edge decay adds to it: a visibly successful public wallet on a transparent venue attracts followers, and crowded execution degrades the fills that produced the edge. By the time an account is high enough to be widely noticed, the conditions behind its numbers have already started to change.

What a more robust evaluation has to measure

The fix is not a cleverer sort of the same column. A serious evaluation of a Hyperliquid account needs return measured per unit of risk; consistency across enough trades that the result is not a small sample of a fat tail; drawdown depth and recovery as a direct measure of survivability; leverage and margin behaviour over time; sizing and exit discipline as a measure of process rather than outcome; and performance across distinct regimes rather than within one.

Two properties of the measurement matter as much as the metrics. It has to be continuous, because a snapshot only describes the moment it was taken and edge decays. And the metrics have to be combined into something comparable, because individually they conflict — the highest return rarely has the best drawdown, and the most consistent account rarely has the highest ROI. Without an explicit combination rule you are back to one column and its bias.

How composite scoring addresses the weaknesses

A composite score is that combination rule. Instead of ranking one outcome, it evaluates several independent properties of an account and reduces them to a single comparable number. Ours uses five factors, all computed from public on-chain data: realized PnL consistency, win rate, profit factor, position discipline and account survivability.

The mapping to the failure modes is direct. Consistency addresses small-sample luck by rewarding a repeated pattern rather than a large total. Profit factor sets gross gains against gross losses, so a return achieved through severe losses scores worse than the same return achieved cleanly. Position discipline reads sizing and exit behaviour, where escalation after losses becomes visible. Survivability covers both survivorship bias and longevity: thin margin buffers and few regimes score lower, and recent return does not compensate. Win rate, the most misleading of the five in isolation, becomes informative next to profit factor, because together they describe the shape of the distribution rather than only its centre.

Two design choices matter more than the factor list. The score is recomputed continuously rather than fixed at selection, so decay changes the assessment while the PnL curve still looks acceptable. And it sets allocation weight, not just membership — capital is distributed in proportion to the strength of the evidence, so weaker support receives less of it. That is a different response to uncertainty than picking the top row and committing fully.

This does not make selection reliable in the sense of guaranteeing outcomes. A composite score is an estimate from historical data, it can be wrong about a specific account, and a trader whose behaviour changes is recognised after the change rather than before. The claim is narrower: it removes the distortions that make a raw PnL sort systematically misleading, and it applies the same rule to every account.

Better ways to evaluate traders

When you look at a Hyperliquid leaderboard, ask what the number you sorted by cannot see. It cannot see the accounts that failed, how much risk produced the result, whether that result came from three decisions or three hundred, whether the trader escalates after losses, or whether they have traded a different regime. Every one of those is measurable from the same on-chain data — a single-column sort simply cannot express them.

If you want to see how a scored selection system reaches a different ordering from the same public data, the note on the five factors behind our composite score walks through each input, and the note on score-weighting explains why the score also decides how much capital each leader receives.

Side by side

What a PnL sort measures versus what selection requires
QuestionRaw PnL / ROI leaderboardComposite evaluation
Who made the most money?Answered directlyConsidered as one input
How much risk produced it?Not measuredLeverage, margin buffer, drawdown
Is the result repeatable?Not measuredTrade count and consistency
Do failures appear?No — removed by liquidation or restartSurvivability penalises thin buffers
Is the process disciplined?Not measuredSizing and exit behaviour
Has it survived other regimes?Not measuredAccount longevity across conditions
Updated over time?Snapshot per windowRecomputed continuously

Past performance is not indicative of future results. Perpetual futures are leveraged instruments and carry a substantial risk of loss, including the loss of your entire position.

Diversified copy trading. On autopilot.

Score-weighted allocation across up to 10 elite Hyperliquid traders, each isolated in its own sub-account. Your funds never leave your account.

Non-custodial · Agent cannot withdraw · Cancel delegation anytime