Sizing a crypto economy
A research consultancy needed to say how big Nigeria's crypto market is, who uses it and where it is going. I took one exchange's transaction data and a survey, and turned a long list of research questions into a market size, behavioural segments, a dashboard and a two-year forecast.
Situation
A research consultancy was preparing a study of Nigeria's crypto market for a client: how big it is, who uses it and where it is going.
Task
As data lead on a two-month engagement, turn the client's open research questions into a market size, behavioural segments, a dashboard and a two-year forecast.
Action
Scaled one exchange's active users by its survey market share, converted platform accounts back into people, segmented users by behaviour, and forecast volume with trend and seasonality.
Result
Delivered the dashboard in Looker Studio and presented the findings to the client's management. The client published the study and hosted a launch session for it.
From research questions to a dashboard spec
The client arrived with questions, not metrics: "how many Nigerians use crypto?", "how does money come in and go out?", "where do stablecoins fit?". My first job was to turn each one into something the data could answer, with a definition, a calculation and a chart type. Agreeing that up front stopped the dashboard turning into a pile of charts nobody asked for.
| Research question | Metric | Visual |
|---|---|---|
| How big is the market? | Estimated unique crypto users (method below) | Headline number with trend |
| How active are users? | DAU / WAU / MAU: distinct users active in a day, week or month | Line chart over time |
| How much, how often? | Average transaction size and trades per user, by type | Bars and trend lines |
| What is crypto used for? | Mix of deposits, trades, withdrawals and transfers | Stacked bars |
| Who and where? | Active users by age group, gender and state | Map and breakdowns |
| How does money flow? | Inflows vs outflows by currency; the stablecoin journey from deposit to exit | Dual series and a flow diagram |
| What is traded? | Top trading pairs by value and count; pair × segment | Ranked bars, heatmaps |
| Who drives adoption? | Business vs retail accounts by users and value | Grouped bars |
| Where is it going? | Two-to-three-year volume forecast | Forecast with uncertainty band |
Every slice also needed filters for time, demographics, pair, transaction type and device, plus drill-down from a state or pair to its transactions. Charts use illustrative data rebuilt from the shape of the real results: values are indexed and lightly perturbed, and the real figures stay with the company. The sizing calculator is the exception: its inputs are pure simulation.
Sign-ups are not adoption
The first thing the activity data showed was that account openings are a poor guide to how many people actually use crypto. Openings came in sharp surges. The largest one, in the third year, moved sign-ups far more than it moved monthly actives, and it converted into a first trade at a fraction of the usual rate. Monthly actives told a different story: a climb to an early peak at the end of the first year, a fall to less than half within a couple of months, a long flat stretch, and a recovery that only regained the old peak in the fifth year.
That is why the dashboard led with monthly actives and first-transacting accounts, not sign-ups, and why the market-sizing method below starts from active users. The conversion view also shows the users who stayed becoming busier: transactions per active user rose well above their early level over the period.
Sizing the market: accounts are not people
One exchange knows its own users. The survey tells you what share of crypto platform use that exchange holds. Scaling up gives the number of platform accounts in the market. But many people use more than one platform, so accounts are not people. The survey also gives the average number of platforms per crypto user, and that is what converts accounts back into people.
N is the exchange's active users, s its share of platform use in the survey, and k the average platforms per user. Dividing by k matters: without it, everyone who uses two or three platforms would be counted two or three times. Move the inputs below to see how sensitive the estimate is.
Estimated unique users across the range of plausible market shares (log scale). The shaded band shows k ± 0.5 platforms per user, because both s and k come from a survey and carry sampling error. The dot is your current setting. Inputs are illustrative, not the client's.
The band matters as much as the line. At low market shares a single point of survey error moves the estimate a lot, so I reported a range, not a single headline number, and said which inputs drove it.
Three behavioural segments
Demographics were patchy (more on that below), so I segmented users by what they did instead. The rules had to be simple enough for the client to explain in a report:
- Trader: makes more trades than the average user.
- Saver: trades below average and withdraws less value than they deposit, so value stays on the platform.
- Cash-out: trades below average and withdraws more than they deposit, using the exchange as an exit route, for example receiving crypto and converting it to naira.
"Average" hides a judgement call. Trade counts are extremely skewed: a few accounts trade constantly and most trade a handful of times. The mean sits many times above the median, so the two give very different thresholds and very different segment sizes. The figure lets you see how sensitive the split is.
Scatter: each dot is an illustrative user, by number of trades and the ratio of value withdrawn to value deposited (both log scale). The population is generated so that, under the agreed rule (mean, 1×, no balanced band), the segment mix matches the shape of the real split. Bars: each segment's share of users, of trades and of deposited value. Threshold: trades. A balanced band above 0% pulls users whose flows roughly net out into their own group instead of forcing them into Saver or Cash-out.
With the agreed definition (above the mean), traders were a small minority of users, but they made the large majority of trades and held most of the deposited value. Savers were the largest group by count, and cash-out users were around a quarter: they deposited little and withdrew far more, which fits people receiving crypto from abroad or from other wallets and converting it to naira. Switch the rule to the median and traders jump to about half of all users, which is why I wrote the threshold into the definition rather than leaving "average" open.
Trading was concentrated too. Two pairs, stablecoin-to-naira and bitcoin-to-naira, carried well over two-thirds of both value and trade count, which is consistent with crypto being used partly as a dollar substitute and a payment rail, not only as an investment. Most users were domestic, with a diaspora tail in the US, UK and Ghana.
Share of value and of trade count across the most-traded pairs, lightly perturbed. Pair names are public market tickers. Blue bars are naira-quoted pairs; the others are quoted in dollars or stablecoins. Dollar-quoted pairs such as BTC/USDT take a bigger slice of value than of trades: fewer, larger tickets. Naira pairs dominate the count.
Forecasting two years ahead
The client wanted a two-to-three-year view. I used Prophet, which models a series as trend plus seasonal patterns plus noise. Underneath, its seasonality is a set of Fourier terms: sine and cosine waves at the yearly frequency and its multiples. More terms let the yearly shape bend more, at the risk of fitting noise. I also tested lagged volume and macro regressors (inflation and the naira exchange rate), since crypto use in Nigeria tracks currency pressure.
The figure rebuilds the core of that model in the browser: a linear trend, day-of-week effects and K pairs of yearly Fourier terms, fitted by ordinary least squares on two and a half years of daily volume, then summed to weeks. The volume series is rebuilt from the original model's own trend and yearly components (indexed, with noise added), and the day-of-week profile from the shape of daily active users. The last six months are held out first to check how well each K predicts data it has not seen.
Top: weekly volume (index, 100 = first-year average), the model's fit, and a two-year forecast with an 80% band. As in the original model, the band widens only slightly: most of the uncertainty is week-to-week noise, not the trend. Bottom left: the yearly shape from the original model against the shape your K terms can draw. It peaks in late January and again in November, sags through a long lull from late May to mid-July, and dips over the Christmas weeks. Bottom right: the fitted day-of-week effect, busiest on Fridays and quietest on Sundays. MAPE is mean absolute percentage error.
K = 1 is too stiff: one smooth wave cannot draw both the January peak and the November peak, so it flattens the mid-year lull too. By about four terms the yearly shape is mostly there, and beyond that the held-out error barely moves while the parameter count keeps climbing. That is the same trade-off I tuned in Prophet. I presented the forecast as a range and named the assumptions that would break it, chiefly a regulatory change or a sharp currency move.
Data-quality judgements
What I excluded or flagged, and why
- One state had implausibly many users relative to Lagos and the capital. That pattern usually means a default location or an IP-geolocation artefact, not real users. I excluded it from the geography read and said so.
- Age was missing for most transactions. Age and gender cuts were still shown, but flagged as low-confidence and not used for headline claims.
- Value fields were empty in the early months. I started value trends where the data became complete, rather than showing a fake growth curve from zero.
None of these is dramatic on its own. Together they decide whether a client's report survives the first sceptical reader.
Caveats and what I'd do differently
- One exchange is not the market. Its users may be younger, more urban or more trading-heavy than the average crypto user. The survey share corrects for size, not for mix.
- Survey shares are self-reported. People forget platforms and over-report popular ones. I would bootstrap the survey to put a proper interval on s and k, rather than a fixed band.
- Peer-to-peer and offshore activity is invisible in exchange data, and it is a large part of crypto use in Nigeria.
- Forecasts with macro regressors need forecasts of the regressors. With more time I would run scenarios (stable naira, further devaluation) instead of one central line.