Wash trading and volume figures that cannot be real
Self-dealing volume costs almost nothing to manufacture and leaves statistical fingerprints; the cross-checks that expose it are simple arithmetic.
Wash trading is trading with yourself, or with a coordinated counterparty, so that activity is recorded without any change in beneficial ownership. It inflates reported volume at close to zero cost when a venue charges nothing or rebates the passive side, and reported volume feeds rankings, listing decisions and the perception that an asset can be traded. The result is a figure that looks like a measurement and is frequently a marketing output.
Why the incentive is structural
Volume is the industry's default ranking variable. Data aggregators order venues by it, listing committees use it as evidence of demand, index providers use it in eligibility screens, and token issuers use it to argue their asset is liquid. Anything used for ranking gets optimized, and volume is uniquely easy to optimize because a trade is simply two matched orders from accounts under common control. Fee structures reduce the cost further: zero-fee promotions, negative maker fees and rebate tiers can make a round trip free or slightly profitable. On-chain, the equivalent behavior appears in NFT markets, where an asset is sold back and forth between wallets to manufacture a price history, and in token markets where fabricated activity is used to qualify for liquidity mining rewards.
It is worth separating this from the equity term it resembles. A wash sale is a tax concept: selling at a loss and repurchasing a substantially identical asset within a defined window, which in some jurisdictions disallows the loss deduction. Wash trading is a market-integrity concept about fabricating activity. The words share a root and describe different things, and treatments differ by jurisdiction and by whether the asset is classified as a security or a commodity.
How the volume is manufactured
The simplest method is self-matching, where one account posts a resting order and the same operator sends the matching aggressive order. Many venues block identical-account matches, so the common form uses two or more accounts under common control, sometimes with a short random delay and varied sizes to look organic. Where the venue pays a rebate to the passive side and charges the aggressor less than the rebate, the round trip is free or slightly profitable, which removes the only natural brake. A related technique alternates aggressor roles between the accounts so that neither builds a lopsided history. On-chain, the pattern is different because every trade is public: an operator moves an asset between wallets they control, paying only network fees and pool fees, and the fees paid can themselves be recovered if the trades qualify for an incentive program. Tokens can also be cycled through a liquidity pool the operator owns, in which case the pool fee is paid to themselves and the true cost approaches zero.
Detection is not free of error in either direction. Institutional execution algorithms deliberately slice a large parent order into small, evenly spaced child orders, which resembles scripted activity under a timing test. A venue serving mostly automated participants will show a tidier size distribution than one serving individuals. So a single failed test is a reason to look further rather than a verdict, and the persuasive cases combine several independent tells with an incentive that explains why the behavior would occur at all.
The statistical fingerprints
Genuine trading is messy in specific ways, and fabricated trading is usually too tidy. Several tests are well established in the academic literature on market manipulation and require nothing more than a trade tape.
- Trade size distribution. Real markets produce a heavy-tailed mix of sizes with strong clustering at round human numbers. Bot-generated wash volume often shows either near-uniform sizes or an implausibly smooth distribution.
- Leading digits. The first digits of genuine trade sizes approximate the Benford distribution. Fabricated series frequently do not, and the deviation is measurable with a simple goodness-of-fit test.
- Inter-trade timing. Real arrivals cluster into bursts. Trades spaced at near-constant intervals across a whole day indicate a script, not a market.
- Price impact that never appears. Large recorded volume with no measurable price impact and no change in the spread is the strongest single tell. Real size moves prices; fabricated size does not, because both sides are the same participant.
The arithmetic cross-checks anyone can run
| Reported figure | Cross-check against | What an inconsistency implies |
|---|---|---|
| Very high 24-hour volume | Visible depth within one percent of mid | Volume many times the entire book with an unchanged spread is arithmetically implausible |
| Volume on a token traded mainly on-chain | DEX volume and pool TVL | Trading far above what the pools could absorb without moving price points to off-chain fabrication |
| Exchange volume for a chain asset | on-chain settled volume | These legitimately differ, but an extreme and stable ratio is worth examining |
| Turnover | 24-hour turnover against sector peers | Turnover far above comparable assets usually reflects a small float, fabricated volume, or both |
| Volume concentrated on one venue | Venue mix and fee schedule | Dominant share on a venue with zero fees and no institutional presence is the classic pattern |
A related and less discussed distortion is turnover computed from a denominator that is itself unreliable. If circulating supply is overstated, turnover looks low and the asset appears calm; if understated, turnover looks extreme. Two errors in the same figure can offset or compound, and the direction is not knowable from the ratio alone.
What responsible data handling looks like
Serious data providers do not simply sum every venue. They apply venue eligibility criteria, weight or exclude venues that fail integrity checks, publish which venues are included, and version their methodology so historical figures can be reproduced. This is why two reputable sources report different volume for the same asset on the same day, and the difference is informative rather than embarrassing: it is the size of the judgment being applied. A source that reports a single number with no venue list is asking to be trusted on the part that matters most.
The practical stance is not that all volume is fake, which would be as wrong as taking it at face value. It is that volume is a claim by a venue about its own activity, and claims of that kind require corroboration from something the venue does not control, such as on-chain settlement, observable depth or price behavior. The methodology page sets out venue inclusion rules used here, data sources lists what feeds each field, and the screener allows volume to be read against depth and turnover rather than alone.