1 of 61

Issue 2

bitfinex.com

2 of 61

Bitfinex Alpha | ISSUE 2

Thought leadership | Research | Market Analysis

Welcome to issue 2 of Bitfinex Alpha on HFT price prediction models!

Price movement and prediction are some of the most intriguing topics of conversation within the crypto community and the financial world in general.

Short-term oriented (day-) traders and longer term oriented investors focus on fundamental and technical analysis for their decisions. Market makers and other high frequency traders (HFT) face the additional challenge to determine the β€œmarket price” of an asset at any given moment. Even most professional traders simply rely on the ticker price, that is, the last executed price. But is the ticker price the best ”market price”?

In this issue, we will explore several price prediction methods in a high frequency environment and assess their accuracy compared with the commonly-used ticker price.

3 of 61

In the current finance and cryptocurrency exchanges system, the price of the last transaction is issued as the authoritative price that is used for all kinds of market statistics. Yet, OTC traders often use the mid-point of the spread between highest bid and lowest ask doubts as a reference point.

What is the better source for market prices: trade history or the order book?

With every issue of Bitfinex Alpha, we aim to give a deeper look into various topics and analysis from experts and thought leaders in the industry that will provide insights and knowledge to our users.

If crypto price movement and market analysis intrigue you, we’re excited for you to delve into Bitfinex Alpha’s issue 2 that features research on this question conducted by independent Blockchain Research Lab (BRL). The BRL is a non-profit research institution dedicated to independent science and research on blockchain technology for the benefit of society.

The views and opinions expressed in the note belong solely to the author and do not reflect those of Bitfinex.

4 of 61

Tl;dr: Key takeaways

Ticker price is good enough an indicator for most purposes, but not for high frequency trading:

  • Market might have moved without a trade (in volatile times or in illiquid markets)
  • Tiny orders can be used to manipulate the ticker price
  • Execution price of an order depends on the spread and thus the order book, which is not reflected in the ticker.

There are various other price metrics based on order book and/or trade history information. For example, Uniswap uses time weighted averages prices (TWAP) of past trades.

The study by the Blockchain Research Lab (BRL) found the volume-limited clearing price to be the best predictor of the price of subsequent trades. It contains only order book data, but no trade history data.

According to the study, a close second is the very simple spread mid-price, which is the average between the highest bid and the lowest ask order. For most traders it might thus be best to use the mid of the spread as reference price.

5 of 61

About the Blockchain Research Lab (BRL)

The non-profit organization Blockchain Research Lab (BRL) is dedicated to independent, interdisciplinary research on blockchain technology and to publishing the results for the benefit of society.

The BRL conducts peer-reviewed academic research in the field of blockchain, with an emphasis on cryptocurrency markets. The research follows a data-driven empirical methodology and sound economic models. Past investigations focused covered topics such as:

  • Cryptocurrency adoption
  • Market reactions to large Bitcoin transfers
  • Efficiency of crypto and stablecoin markets
  • Exclusive mining of Bitcoin transactions
  • Security token offerings (STOs)
  • Initial coin offerings (ICOs)

Articles by the BRL have been published in prestigious academic journals such as Finance Research Letters, Telematics & Informatics, Journal of Grid Computing, Decisions in Economics and Finance, Digital Finance, Scientometrics or Technological Forecasting and Social Change. All publications can be found on the publication site.

6 of 61

Discovering market prices

Which price formation model best predicts the next trade?

By AndrΓ© Meyer and Ingo Fiedler

28 May 2021

7 of 61

Abstract

For most purposes of technical analysis, valuation metrics and many other relevant financial methods, the price of the last transaction is considered representative of the market price. The straightforward argument is that at this price, supply and demand have last met. However, on closer examination, the question arises as to why a past event should be relevant to the future, and why other, potentially more recent information should not be used to discover a future price. Building on this question, we apply a range of new price formation models to current data available on cryptocurrency exchanges that depict level II market data, and compare their short-term forecast accuracy against the common-used ticker price and mid-price. Data on cryptocurrencies is used as the closest example to free markets, since cryptocurrency trading is continuous, markets never close, and interferences through oversight is extremely rare. We find that two of the five price formation models investigated outperform the widely used ticker as a price indicator for the next trade. We conclude that the volume-limited clearing price best predicts the price of subsequent trades. Its usage can thus enhance the explanatory power of various financial analyses.

8 of 61

Introduction

The concept of price has existed ever since man began to trade.

From very early on, economists such as Smith (1776), Ricardo (1817) and Stackelberg (1934) have examined the significance of prices and their origin. While most economists agree that a price marks the equilibrium between supply and demand, there are different definitions of prices and different forms of markets that influence the discovery of prices.

The interplay of supply and demand is perhaps best illustrated by exchanges with the characteristics of an order-driven market and a situation of perfect competition, where the order books reflect the supply and demand curves. For this reason the Limit Order Book (LOB) is an important field of research and has been studied in various ways (Glosten 1994; Biais et al. 1995; Foucault et al. 2005; Roşu 2009; Gould et al. 2013)

9 of 61

If sales and purchase orders meet or if a market order is executed, a trade is completed. The price at which such a trade takes place is displayed as the so called "ticker" and represents the current value of a share or a good.

This suggests that the ticker is a relevant reference point for the present and the near future.

The more recent the trade and thus the ticker, the more convincing this assumption.

A widely used alternative to the last price is the mid-price (or midpoint price), the midpoint between the best ask and the best bid offer (Laruelle et al. 2013; O’Hara 2015; Cont and Kukanov 2017; Ntakaris et al. 2018).

However, this indicator is associated with a number of weaknesses, as it fails to consider not only the volume of the last trade but also other important factors that can have a significant influence on market prices, such as the price depth of the bid and ask side (Kempf and Korn 1999; Ahn et al. 2001) or the tick size (Darley et al. 2000). Therefore, other measures have already been developed and discussed in the literature to discover more representative prices.

10 of 61

Besides asset pricing models (Sharpe 1963; Merton 1973; Bollerslev et al. 1988) and models that use fundamental and technical methods (Graham and Dodd 1940; Damodaran 2012; Taylor and Allen 1992; Damodaran 2012) to calculate future prices, many models have been created for the increasingly important high-frequency trading (Cont 2011; Agarwal 2012; O’Hara 2015; Avellaneda and Stoikov 2008). High- 2 frequency trading poses a challenge especially for market makers. They not only have to fulfil their primary task of providing liquidity but must also defend themselves against possibly better-informed traders (Menkveld 2013). Taking into account the often discussed diffusion of the price (TΓ³th et al. 2011; Mastromatteo et al. 2014), price criteria such as order book imbalance (OBI) or order flow imbalance (OFI) have been established in the literature in order to provide market makers with an important indication of price discovery (Eisler et al. 2012; Cont et al. 2014).

11 of 61

On the other hand, the literature is investigating strategies to liquidate large positions without a significant price impact (Easley and O'hara 1987; Lin et al. 1995). To measure the success of such strategies, the volume-weighted average price (VWAP) was introduced as a benchmark (Konishi 2002; Madhavan 2002; GuΓ©ant and Royer 2014; Frei and Westray 2015).

This measure already takes into account the volume, albeit ex post.

The weighted mid-price in turn integrates volume in the form of order book imbalance into the discovery of a future representative price. However, this method is of limited use for high frequency trading as it only takes into account the best bid and ask prices and their respective volumes, which are susceptible to constant cancellations in fractions of a second, as is common in HFT (Gatheral and omen 2010; Robert and Rosenbaum 2012).

12 of 61

Therefore, a number of new models based on the weighted mid-price were developed. Bonart and Lillo (2016) adapted the Madhavan et al. (1997) price formation model, taking into account quote discretization and liquidity rebates to introduce price definitions for large tick stocks. The approach by Jaisson (2015) incorporates conditional expectations and was adopted by Lehalle and Mounjid (2017) to form the so-called micro-price as the expected future midprice under the condition of the current mid-price and the degree of order book imbalance. Finally, Stoikov (2017) tests mid-prices, weighted mid-prices and microprices for their informative value and short-term forecast accuracy. He finds that the micro-price yields the most accurate results.

Building on the methods described above, we define new price formation models and test their accuracy in short-term price prediction in relation to the commonly used ticker price. The null hypothesis of this paper is therefore that none of our price formation models is superior to the ticker price as an indicator of the next trade.

13 of 61

We do not only use current data provided in every level II order book, we also test a model that uses the trade history as an indicator for pricing. The use of the trade history is inspired by the ticker price, which is itself a past-related value.

We test the forecast accuracy of our models using cryptocurrency market data.

Ghysels and Nguyen (2018) already use data from cryptocurrency exchanges in their work for new insights into price discovery. Our choice of crypto market data is primarily motivated by the non-stop 24/7 trading and the absence of distorting stabilizing price mechanisms or regulatory intervention.

14 of 61

For example, the auctions that are used on traditional stock exchanges after trade disruption, for pre- or post trade, but also in case of excessive price volatility, have a significant influence on price discovery. Madhavan and Panchapagesan (2000) find that while the opening auction of the NYSE increases price efficiency, it leads to stagnation of prices and thus reduces their flexibility. By contrast, the market for cryptocurrencies is cleaner in this regard and can thus be seen as prototype closer model of a free market. Furthermore, its data richness makes it the perfect environment to analyse price discovery.

Yet crypto markets also have disadvantages, specifically their lack of regulation entails a risk of manipulation. Many crypto exchanges are suspected of flaunting wrong volumes, making it difficult for researchers to obtain unspoiled results.

We therefore use data from exchanges that according to the Blockchain Transparency Institute (2018) report unspoiled volumes (Fusaro and Hougan 2019).

15 of 61

Price formation models

In the following, we will test five different price formation models, two of which are tested with three different parameter weights each, yielding a total of nine separate calculations. The results will be compared to the ticker price. The first price formation model we study is the midprice.

As already mentioned, the mid-price is often used both in research and in practice for short-term price predictions.

16 of 61

In this simple calculation method, the reference price is obtained as the mid-point between the highest buy and the lowest sell offer:

where π‘šπ‘ is the mid-price, 𝑝0 𝑏 is the highest bid and 𝑝0 π‘Ž is the lowest ask offer in the LOB. These prices are always in the first position of the respective order book side. Their distance to the best order of each order book side is therefore zero, as indicated by the index 0.

17 of 61

π‘€π‘šπ‘ is calculated using the imbalance (𝑖𝑏), which depends on the volume of the best bid (π‘ž0 𝑏 ) and ask (π‘ž0 π‘Ž ) offer:

Both the mid-price and the weighted mid-price are susceptible to frequent order changes or low volume orders, which merely serve to price discovery, especially in less liquid markets.

18 of 61

The clearing price is often defined as the price at which the market settles a commodity or security, i.e. 3 where the quantity delivered equals the quantity demanded. In practice, the ticker price is often assumed to be the clearing price because it is the price at which the last units of an asset were traded.

However, as the above-mentioned studies on the liquidation of large positions show (e.g. Cartea and Jaimungal 2016), this is not always consistent with the required volume. The following price formation model calculates the midprice for a given volume based on the original definition of the clearing price.

We call it the volume-limited clearing price (𝑣𝑙𝑐𝑝) and define it as follows:

where the total price of a market buy order for a fixed volume (volume limit) 𝑣𝑙 is given by the function 𝑠𝑓, which is defined as:

19 of 61

The total price of a market sell order for a fixed volume 𝑣𝑙 is denoted as 𝑠𝑓 and defined analogously using the bid-side offers. The volume-weighted average bid and ask prices are then formed by dividing the total prices of each side by the given volume. The average of these values then yields the volume-limited clearing price as the middle between the volume-weighted average prices of each side of the LOB.

In calculating the market price, this model includes the volume up to a fixed amount, thus excluding nonrepresentative orders located lower in the LOB that are often placed on crypto markets by speculators in the hope of a so-called fat-finger error, where an order is placed of a far greater size or price than intended, or in the wrong currency.

20 of 61

The model adapts to market conditions. If the specified volume (𝑣𝑙) does not exceed the volume of the best bid and ask offer, the results of the model are equal to the mid-price. Conversely, if 𝑣𝑙 does exceed the volume of the best bid or ask orders, the order book imbalance is also indirectly included in the calculation, not directly by the volume itself, but by the volume-weighting of prices.

In markets with low liquidity near the spread or in markets with frequent placement and cancellation of low volume orders at the top of the order book, the model produces more stable and consistent results, while it delivers the same results as the mid-price in liquid markets with high volumes close to the spread.

21 of 61

Our next model is related to the 𝑣𝑙𝑐𝑝 but uses a price limit to calculate the price.

We therefore call it the price-limited clearing price (𝑝𝑙𝑐𝑝). The prices are weighted by volume and the distance to a reference price. The further away an order price is from the reference price, the lower its weight. The 𝑝𝑙𝑐𝑝 is defined as:

where π‘Ÿπ‘ is the reference price, 𝑝𝑙 is the price limit, and 𝑑𝑀𝑖 π‘Ž and 𝑑𝑀𝑖 𝑏 are the distance weights of the ask and the buy side, respectively. We define the price limit (𝑝𝑙) as:

22 of 61

Starting from the reference price (π‘Ÿπ‘), depending on the order book side, the price limit (𝑝𝑙) is augmented (ask side) or reduced (buy side) by a distance that is calculated using percentage depth (𝑝𝑑). 𝑝𝑑 is an external parameter that we assumed to be either 1, 2 or 3 in our sample. The parameter indicates up to which price level of the respective order book page orders are included in the calculation. The larger the value the greater the price range within the orders will be included in the calculation. While this can mean to incorporate more information in form of more orders, it also increases the likelihood of including a bias from noise, for example, in the form of an order book that is asymmetric by chance. Such risk is especially likely for illiquid order books. We use the 𝑣𝑙𝑐𝑝 as the reference price, but the mid-price or any other reference value serve just as well. To calculate 𝑝𝑙𝑐𝑝, we must furthermore define the distance weights (𝑑𝑀):

23 of 61

The distance weights are calculated by discounting the volume of an order (π‘žπ‘– ) using the distance function (𝑑𝑓(𝑝𝑖 , π‘Ÿπ‘)) and the distance exponent (𝑑𝑒). We set the external parameter 𝑑𝑒 to 0.75 in our sample. The distance function is defined as follows:

If the price of an order is equal to the reference price, the function assumes the value 0, which would result in an error in the calculation of the weights. Therefore, the reference price must be chosen so as to make this impossible. Therefore, we use the 𝑣𝑙𝑐𝑝 as the reference price (π‘Ÿπ‘) in our sample. 𝑣𝑙𝑐𝑝 cannot be reached by any buy or ask price, since its value is within the spread. Hence, it is impossible that the condition 𝑝𝑖 = π‘Ÿπ‘ occurs.

24 of 61

It should be noted that even if the tick size does not allow a lower price scaling, this does not affect the hypothetical value of the 𝑣𝑙𝑐𝑝, since it is infinitely scalable. For ask side prices, the function assumes the value since here the order prices exceed the reference price. For the bid side, applies, as the order prices are below the reference price. An advantage of this model is that it takes into account that orders located towards the bottom of the order book are less relevant than those at the top. In addition, it allows users to specify to which price depth the LOB is taken into the calculation. The model also takes into account the OBI since the volumes are considered but discounted according to their distance to the reference price.

By incorporating these discounted volumes, the model also considers any imbalance of the order book, which may provide a valuable signal of excess demand or supply. Due to the weighting function of the model, imbalances closer to the reference price are weighted more heavily than more distant ones. However, there is a risk that large orders will be placed within reach of the reference price in order to manipulate 𝑝𝑙𝑐𝑝 and affect trading strategies based on it.

25 of 61

The next model uses the 𝑣𝑙𝑐𝑝 as the reference price as well and is furthermore using the price limit (𝑝𝑙) to adjust the reference price by the factor (1 + 𝑐𝑓(π‘Žπ‘“,ma)). This model, which we call the adjusted reference price, is defined as follows:

𝑐𝑓 is the cap function that limits the results of the adjustment function (π‘Žπ‘“) to a maximum adjustment (ma):

The adjustment function (π‘Žπ‘“) is determined by the adjusted weights (π‘Žπ‘€) and defined as follows:

26 of 61

Within π‘Žπ‘“ the adjusted weights (π‘Žπ‘€) are taken into account up to the price limit (𝑝𝑙) given by equation (7) and are defined as:

Where π‘žπ‘– is the volume of order 𝑖 and 𝑏 is an external parameter whose exponent is 100 times the distance function (9). This model uses the different order volumes up to the predefined price limit of each order book side to create an imbalance correction factor by which the reference price is adjusted.

The adjustment is limited by the external parameter ma. For our sample, we assumed 0.003 as the value for ma. While the price limit for the 𝑝𝑙𝑐𝑝 is determined by the prices of the respective orders, in the π‘Žπ‘Ÿπ‘, the price limit is determined by the volume of the respective orders.

27 of 61

The last price formation model we shall test is completely different from

the previous ones. It does not use current information but calculates a price based on historical transactions. The model thus builds on the ticker itself. While the ticker only uses the price of the last trade and thus has the disadvantage of being quite volatile, our model uses the prices and volumes all past transactions up to a certain age, though with decreasing weights. We call this price formation model the trade history model (π‘‘β„Ž) and define it as follows:

where the time weights (𝑑𝑀) are defined as:

28 of 61

The age limit allows us to determine how long a period should be considered relevant for the representative price. While the weight of older trades would eventually be discounted to zero, the age limit allows to cut off trades with a very low weight that do not add much information but, when included, would cost computing power and thus delay the result.

The reduced weights of older trades are in line with the common sense notion that events in the more distant past should have less bearing on the future. Table 1 lists the values we selected for the external parameters.

Table 1: Parameters used in our price formation models

29 of 61

Methodology

To assess the performance of our price formation models, we recorded their results in 44,640 minute-data points from 01.12.2018 00:00 to 31.12.2018 23:59 for the prices of 3 cryptocurrencies – Bitcoin (BTC), Litecoin (LTC), and Ethereum (ETH), each expressed in USD. These are the three largest-cap cryptocurrencies that use the proof-of-work mechanism which implies a natural price, as production of the cryptocurrencies entails hardware and electricity costs.

These mining costs make such crypto currencies comparable to commodities that must be extracted before they can be traded and that therefore also have a natural price.

Dataset

30 of 61

*The listed variables were recorded in the period from 01.12.2018 00:00 until 31.12.2018 23:59 using the application programming interface (API) of the respective crypto exchanges.

Table 2: Summary statistics

31 of 61

We continuously calculated and recorded all results of the price formation and discarded any raw data due to data storage limitations. To test the forecast accuracy of the models, we recorded all the trades that took place during this period and formed a minute-by-minute volume-weighted average price (ap). The aim of the models is to predict the next minute’s ap.

The data was recorded using the application programming interface (API) of the crypto exchanges, which we will refer as exchange A and exchange B, which we chose because they support trading against US-Dollar pairs rather than against a digital tokens tethered to US-Dollar, such as Tether tokens (USDt), as, in practice, their market price usually diverge by at least a few basis points. Exchange B features greater trading volumes than exchange A. For example, the average cumulative trading volume per minute during the sample period for the BTC/USD pair was 10.95885 BTC for exchange B and 8.452792 BTC for exchange A.

32 of 61

Depending on the cryptocurrency and the exchange, the data set contained between 7 (0.02%) and 7,653 (17.43%) gaps. These gaps in the reporting can be due either to server problems or to the fact that no trading took place and therefore no prices and volumes could be recorded. Since the results of the mid-price, 𝑣𝑙𝑐𝑝, 𝑝𝑙𝑐𝑝 and π‘Žπ‘Ÿπ‘ models are based on order book data, the gaps of these models are exclusively due to server connectivity problems. In this case the order book was not accessible via web sockets and therefore no data could be recorded.

The number of recording gaps due to server connectivity problems is between 745 (1.67%) and 753 (1.69%). In addition to these gaps, the trade history model (π‘‘β„Ž) also has gaps that result from a lack of trading over a period that exceeded the external parameter π‘Žπ‘™, so that no result could be determined. Therefore, this model has between 751 (1.69%) and 3,219 (7.21%) gaps, depending on the currency pair and the exchange.

At 6,500 (14.56%) to 7,653 (17.43%), the currency pairs ETH/USD and LTC/USD as traded on exchange A featured the most gaps.

33 of 61

So as not to distort the forecast performance of the models, the gaps were not filled by interpolation or similar methods. Rather, it was assumed that no trade would have been possible even at other prices for the gaps. Depending on the type of accuracy measure and the cryptocurrency and exchange, varying numbers of observations were used to calculate the measures of forecast accuracy. As a result, some of the forecast measures are more meaningful than others, which is why we provide the number of observations used for the calculation for each forecast measure, cryptocurrency and exchange.

For example, only 6,947 observations were available to calculate the Mean Directional Accuracy (MDA) for LTC/USD on exchange A, as compared to 43,472 observation for other measurements regarding BTC/USD on exchange B.

34 of 61

During the observation period (December 2018), the crypto market experienced a phase of sideways movement and relative price stability. During that month, the Bitcoin price on exchange A fell from 3,973.253 USD to 3,690.607 USD. The prices on exchange B and the prices of LTC/USD and ETH/USD show a similar pattern, though ETH/USD rose slightly over the period. The lowest price for a Bitcoin traded on exchange A was 3,124.043 USD while the highest price was 4,258.983 USD, a range of 1,134.94 USD. Despite this large range, the standard deviation around the mean of 3,672.419 USD was only 284.5603 USD, or 7.75% of the mean.

For ETH/USD and LTC/USD, the standard deviation amounted to 17.26% and 12.07% or the respective means. Thus, the price of Bitcoin was more stable than the other two currencies, yet all three crypto assets were more volatile than most stock prices. Table 2 provides an overview of the dataset.

35 of 61

To assess the forecast quality of the price formation models, we first looked at the mean errors(𝑀𝐸) between the weighted average price at minute 𝑑 and the prediction of each price formation model at time 𝑑 βˆ’1. The mean error thus expresses how well the price formation model (π‘₯π‘‘βˆ’1 ) can determine the price of the next minute (π‘Žπ‘π‘‘ ).

The mean error itself gives little information about the quality of a forecast, which is why it should be interpreted in conjunction with other measures. But it can give an indication of potential systematic distortions, which can be confirmed by the distribution of the forecast errors. The mean absolute error (𝑀𝐴𝐸), on the other hand, uses the absolute values of the forecast errors and thus provides information about the quality of the price formation models, specifically the average absolute difference between the realized volume-weighted average price and the forecast result of the respective price formation model.

Measures of forecast quality

36 of 61

The smaller the 𝑀𝐴𝐸 the better. It is defined as:

Next, the root-mean-square error (𝑅𝑀𝑆𝐸) is often mistakenly interpreted as 𝑀𝐴𝐸 but it is the square root of the average quadratic error, so any error has a squared impact on the 𝑅𝑀𝑆𝐸. Therefore, larger errors have a stronger effect on the 𝑅𝑀𝑆𝐸 than on the 𝑀𝐴𝐸. In conjunction with the 𝑀𝐴𝐸, the 𝑅𝑀𝑆𝐸 can provide information on the size and frequency of outliers of the price formation models. The 𝑅𝑀𝑆𝐸 is:

All of the above forecast measures are scale dependent, so their results are not comparable across different scales and thus different cryptocurrency pairs. We therefore also draw upon the mean absolute percentage error (𝑀𝐴𝑃𝐸)which is defined as:

37 of 61

Each deviation between the actual volume weighted price at time 𝑑 (π‘Žπ‘π‘‘ ) and the price calculated by a price formation model at time 𝑑 βˆ’ 1 (π‘₯π‘‘βˆ’1 ) is scaled by π‘Žπ‘π‘‘ .

The resulting absolute percentage errors are added up and divided by their number.

The result is multiplied by 100 for an easily interpretable percentage.

Another scale-independent measure of forecast quality is the Mean Directional Accuracy (𝑀𝐷𝐴). It compares the direction of movement of the actual volume-weighted price between times 𝑑 βˆ’1 and 𝑑 to the corresponding price movements predicted by each price formation model and counts the number of matches. 𝑀𝐷𝐴 is defined as:

The signum function 𝑠𝑖𝑔𝑛(π‘Žπ‘π‘‘ βˆ’ π‘Žπ‘π‘‘βˆ’1 ) extracts the sign of the result and is an indicator function which delivers the value 1 if the Boolean expression is true and the value 0 if it is false. In the following, these five forecast measures will be used to assess the accuracy of the price formation models introduced above.

38 of 61

Results

Table 3 provides an overview of the distribution of forecast errors. While the number of forecast errors on exchange B is between 42,446 and 43,472 for all three cryptocurrency pairs, the number of observations on exchange A varies considerably depending on the cryptocurrency pair. For example, the number of forecast errors on exchange A ranges from 11,872 (LTC) to 41,400 (BTC), depending on the currency pair.

This is due to the lower number of trades. Increased trading and higher volumes could have a stabilizing effect on the forecasting power of the price formation models. This is also reflected in the standard deviation of the forecast errors, which is lower on exchange B for all three currency pairs than on exchange A. Furthermore, the forecast errors of the mid-price and the 𝑣𝑙𝑐𝑝 have the lowest standard deviation across all currency pairs and exchanges.

Descriptive results

39 of 61

The mean error (𝑀𝐸) is closest to the ideal value of 0 for the price formation model th, while we often find the greatest distance for 𝑝𝑙𝑐𝑝 with a percentage depth (𝑝𝑑) of 3. The forecast errors can cancel each other out, which is why this information only permits conclusions on any systematic bias in the forecast. Therefore, it makes sense to also interpret skewness and kurtosis. The ideal distribution of the prediction errors should have the bulk of its mass around 0 and be symmetrical, i.e. free of systematic bias. Therefore, a skewness of 0 is desirable.

None of the price formation models we tested conform with these expectations for the present data set; each model is skewed either left or right, depending on the exchange and currency pair. However, the data set does not allow a definitive conclusion on a systematic bias of the models, as the sample size is insufficient.

40 of 61

Furthermore, factors such as trading volume and trend can influence the skewness of the models. When considering the kurtosis, all models have a leptokurtic distribution, meaning that outlier forecast errors cause a higher kurtosis compared to a normal distribution. It should be noted that the crypto market is highly volatile and numerous and large outliers are to be expected.

The range of forecast errors is depicted by the range between the minimum and maximum values of the distributions. Based on the mean of the volume weighted average prices per minute, 𝑝𝑙𝑐𝑝 (pd=2) produces the largest outlier, at -11.65193 USD (-10.79%) for ETH/USD on exchange A, which gives an impression of in what percentage range the maximum forecast error of this sample moves. Taking this approximation into account, the 𝑣𝑙𝑐𝑝 features the lowest negative and positive outliers for the currency pair LTC/USD on exchange B, at -0.4422455 USD (-1.54% off mean ap) and 0.4973602 USD (1.73% off mean ap).

41 of 61

Tables 4 and 5 summarise the performance of the price forecast models according to the accuracy measures. We find that only the mid-price and the 𝑣𝑙𝑐𝑝 outperform the ticker. The 𝑣𝑙𝑐𝑝 delivers the best results in all settings except BTC/USD on exchange B, where the mid-price fares best in terms of MAE, RMSE and MAPE. The π‘Žπ‘Ÿπ‘ outperforms the ticker regarding MAE, RMSE and MAPE only for the currency pairs ETH/USD (except pd=1) and LTC/USD on exchange A.

In terms of the MDA, the π‘Žπ‘Ÿπ‘ outperforms the ticker for all currency pairs on exchange A. This also applies to LTC/USD on exchange B. All of the models perform well in terms of MDA Only the ticker and the π‘‘β„Ž predicted the right direction for LTC/USD on exchange A in less than 50% of the cases. In all other cases, the MDA was higher than 50%, reaching almost 80% using the 𝑣𝑙𝑐𝑝 for the BTC/USD currency pair on exchange B. Therefore, the MDA is also the only measure in which the 𝑝𝑙𝑐𝑝 performs better than the ticker, but only on the exchange A.

Forecast accuracy results

42 of 61

If you take the other measures into account in addition to the MDA, one of the three variants of the 𝑝𝑙𝑐𝑝, with the exception of the BTC/USD currency pair on exchange B, always yields the worst results. The π‘‘β„Ž, with one exception, has the worst MDA scores.

π‘‡β„Ž also has high error values in MAE, RMSE and MAPE that make this price formation model the most inaccurate one regarding BTC/USD on exchange B. Percentage depth (𝑝𝑑) affects the two price formation models in a different way. While the results of the 𝑝𝑙𝑐𝑝 get worse with increasing pd except one case, the results of the π‘Žπ‘Ÿπ‘ are very different for each currency pair and crypto exchange. The π‘Žπ‘Ÿπ‘ delivers the same MAE, RMSE and MAPE for the pd of 2 and 3. This is not the case for the MDA values where all 3 variants always deliver different results.

43 of 61

*The listed variables were recorded in the period from 01.12.2018 00:00 until 31.12.2018 23:59 using the application programming interface (API) of the respective crypto exchanges.

Table 3: Descriptive results of forecast errors

44 of 61

For the purpose of short-term price prediction, we tested various price formation models to determine the best possible reference price for next minute transactions. Our point of comparison was the widely used reference price, the price of the last trade, or ticker. To evaluate the results, we discussed, compared and applied five different measures of forecast accuracy.

Especially the 𝑣𝑙𝑐𝑝 produced satisfactory results in terms of MAE, RMSE, MAPE and MDA. With few exceptions, the 𝑣𝑙𝑐𝑝 yielded the smallest forecast errors. However, the mid-price, one of the simplest methods, performed similarly well, while the performance of the more complex models such as 𝑝𝑙𝑐𝑝 and π‘Žπ‘Ÿπ‘ fell short of the ticker.

Discussion

45 of 61

The model π‘‘β„Ž, while comparable to the ticker in its use of past data, also scored worse than the ticker, in particular in terms of the MDA. As mentioned before, the RMSE weighs outliers more heavily than the MAE. Joint consideration of these two errors can therefore also provide information on how often a price formation model delivers outliers.

However, the results presented in Table 4 do not provide any indication that a price formation model has extreme outliers as there are no significant differences between the RMSE and the MAE.

*Lowest value (highest degree of accuracy)

**Highest value (lowest degree of accuracy)

Table 4: Prediction accuracy statistics

46 of 61

All 4 error measures can be compared across the exchanges for a given currency pair. A uniform picture for all measures and currency pairs emerges; on exchange B, which has the highest trading volume for all 3 currency pairs, all price formation models perform better than they do on the less busy exchange A. The MAPE allows a comparison of the price formation models beyond currency pairs and crypto exchanges. This confirms the influence of trading volume, as described above.

The greater the volume, the smaller the MAPEs, which supports the argument that liquidity fosters market efficiency. Conversely, this finding enables the detection of wash trading, that is, trades in which the buyer and the seller are the same entity. Such trades do not add to price discovery and are considered manipulative. As exchange B has more trading volume than exchange A in all currency pairs, the MAPEs are also lower in the former. However, considering the BTC/USD pair on exchange A and ETH/USD on exchange B, it appears that regardless of the crypto exchange, the high-volume currency pair (BTC/USD) leads to lower MAPEs.

47 of 61

The ME of the price formation models points to systematic bias in the forecasts. The mid-price and 𝑣𝑙𝑐𝑝 are positive for some currency pairs and exchanges while they are negative for others. It is therefore unclear whether these two models are biased. For the remaining models the MEs are negative, which could be an indication of a general overestimation. However, there is also the possibility that individual extreme overestimates of the models may overcompensate those of the underestimates and thus merely give the impression of systemic distortion.

*Lowest value (worst performance)

**Highest value (best performance)

48 of 61

The examination of the different weights of the percentage depth parameter used in the 𝑝𝑙𝑐𝑝 and π‘Žπ‘Ÿπ‘ models has shown that more collected data in terms of higher parameter weighting does not enhance the predictive power for this data set. This could mean that the information value contained in orders that are placed deep in the book but intended to be executed is diluted and outweighed by the misleading information contained in orders that are placed solely to move the market.

The measures of forecast quality suggest that the 𝑝𝑙𝑐𝑝 model fares worse with increasing percentage depths, so the benefit of recording the data at higher percentage depths may be doubted. However, the opposite often applies to the π‘Žπ‘Ÿπ‘ model. Although the π‘Žπ‘Ÿπ‘ model delivered roughly the same results for percentage depths of 2 and 3, raising the parameter from 1 to 2 improved the forecast quality, which justifies additional data recording effort.

49 of 61

The analyses allow us to reject our null hypothesis, which held that none of the price formation models we examined is superior to the ticker. The models 𝑝𝑙𝑐𝑝 and mid-price in fact delivered superior forecast quality. This holds implications for science and practice. Other prediction models such as the ARIMA model could use the 𝑣𝑙𝑐𝑝 instead of historical price data and possibly produce more accurate results. The 𝑣𝑙𝑐𝑝 may indeed become a more relevant reference price than the ticker. This model should thus be taken into account in future research.

Our results indicate that the 𝑣𝑙𝑐𝑝 is a good short term prediction model and may therefore be used by investors to estimate the value of future transactions. Used as a signal for trading algorithms and deep learning, it may facilitate more accurate results. Market makers can base their price discovery strategy on it or combine currently used methods with the results of 𝑣𝑙𝑐𝑝. The results suggest that it might be useful for crypto and traditional exchanges to display the 𝑣𝑙𝑐𝑝 n next to the ticker and the order book to provide traders with additional information

50 of 61

Yet the present study also has a number of limitations which we will discuss briefly in the following. Firstly, it suffers from limited data quality and availability. Since the interfaces from which the data were obtained were not always accessible, the record has gaps, which reduce the representativeness of the results and necessitate further investigations on other 11 datasets. Furthermore, it would be worthwhile to check whether the results we obtained from minute data also apply to shorter or longer intervals.

The crypto pairs analysed in this study were selected on the basis of their large market capitalisation; the crypto exchanges were picked because they list USD currency pairs. It remains unclear whether our results would hold when applied to other exchanges with their own sets of procedures and rules, and to other cryptocurrencies, or to altogether different classes of assets.

51 of 61

Some of the price formation models use externally specified parameters. However, with the exception of three different depth values we did not test different values for each parameter, so we cannot state the optimal values for each parameter in a given situation.

The 𝑣𝑙𝑐𝑝 has emerged as the most accurate model, so further research in this area appears to be warranted. Determining the right volume limit is an interesting challenge for future research. Yet in the less successful models, too, different external parameters could be tested to improve prediction performance. Furthermore, the robustness of the model parameters should be examined more closely. Finally, we did not test a number of existing price formation models such as the micro-price.

The implications for further research immediately arise from these limitations. The robustness of the models should be checked for other cryptocurrency pairs and crypto exchanges. An analysis of the price formation models over different time series by means of the Diebold-Mariano test for differences in the mean square error or an encompassing test such as the FairShiller test are promising tasks for further investigation in this area (Diebold and Mariano 1995; Mizon 1984; Mizon and Richard 1986; Fair and Shiller 1988).

52 of 61

The comparison of traditional financial data and crypto data is another interesting subject for further research. It would clarify whether the price formation models perform equally well in both markets. Another possible field of research would be the comparison with traditional financial market forecast models.

It remains to be determined whether the price formation models, which are based primarily on current data, are superior to the traditional methods of technical analysis, which are mostly based on historical data or whether there can be a meaningful combination of both approaches.

Finally, as already mentioned, it would be worthwhile to test how well the price formation models perform when fed not with minute data but with longer or shorter time periods. More specifically, it would be interesting to see whether the 𝑣𝑙𝑐𝑝 remains the most accurate model when applied to different time periods.

53 of 61

Conclusion

In the current system of finance and cryptocurrency exchanges, the price of the last transaction is issued as the authoritative price. Yet doubts arise as to whether the ticker, looking solely at the past and ignoring the information contained in trading volume, can be considered representative of future transactions.

We thus proposed a number of alternative market price formation models and assessed their ability to predict the price of the next trade. Various tests were conducted to determine whether any of these computational models could beat the ticker and can therefore be considered more representative of future trade.

54 of 61

In a first step, different measures of forecast quality were created and discussed. Additionally, the price formation models were checked for systematic over- or underestimations using their mean errors. It turned out that, depending on the currency pair and crypto market, the bias produced by a price formation model may shift. An analysis of the distribution of the forecast errors revealed a leptokurtosis for all price formation models. We found that the mid-price and the volume-limited clearing price (𝑣𝑙𝑐𝑝) provided more accurate forecasts than the other models and the ticker.

In conclusion, this work has revealed a number of new findings that may prove very valuable in financial market practice and that furthermore provide a solid foundation for further research. The simple mid-price and 𝑣𝑙𝑐𝑝 models produced the most accurate forecast results and are therefore most representative of future transactions. The null hypothesis of this work that none of the examined price formation models is superior to the ticker can thus be rejected.

55 of 61

References

Agarwal, Anuj (2012): High frequency trading. Evolution and the future. In Capgemini, London, UK, p. 20.

Ahn, Hee-Joon; Bae, Kee-Hong; Chan, Kalok (2001): Limit Orders, Depth, and Volatility. Evidence from the Stock Exchange of Hong Kong. In The journal of finance 56 (2), pp. 767–788. DOI: 10.1111/0022-1082.00345.

Avellaneda, Marco; Stoikov, Sasha (2008): High-frequency trading in a limit order book. In Quantitative Finance 8 (3), pp. 217–224. Biais, Bruno; Hillion, Pierre; Spatt, Chester (1995): An empirical analysis of the limit order book and the order flow in the Paris Bourse. In The journal of finance 50 (5), pp. 1655–1689.

Blockchain Transparency Institute (2018): Exchange Volumes Report. December 2018. Available online at https://www.blockchaintransparency.org/, updated on 3/29/2019, checked on 3/29/2019.

Bollerslev, Tim; Engle, Robert F.; Wooldridge, Jeffrey M. (1988): A capital asset pricing model with time-varying covariances. In Journal of political Economy 96 (1), pp. 116–131.

Bonart, Julius; Lillo, Fabrizio (2016): A Continuous and Efficient Fundamental Price on the Discrete Order Book Grid. In SSRN Journal. DOI: 10.2139/ssrn.2817279

56 of 61

Cartea, Alvaro; Jaimungal, Sebastian (2016): A closed-form execution strategy to target volume weighted average price. In SIAM Journal on Financial Mathematics 7 (1), pp. 760–785.

Cont, Rama (2011): Statistical modeling of high-frequency financial data. In IEEE Signal Processing Magazine 28 (5), pp. 16–25.

Cont, Rama; Kukanov, Arseniy (2017): Optimal order placement in limit order markets. In Quantitative Finance 17 (1), pp. 21–39.

Cont, Rama; Kukanov, Arseniy; Stoikov, Sasha (2014): The price impact of order book events. In Journal of financial econometrics 12 (1), pp. 47–88.

Damodaran, Aswath (2012): Investment valuation. Tools and techniques for determining the value of any asset: John Wiley & Sons (666).

Darley, Vince; Outkin, Alexander; Plate, Tony; Gao, Frank (Eds.) (2000): Sixteenths or pennies? Observations from a simulation of the NASDAQ stock market: IEEE.

Diebold, Francis X.; Mariano, Roberto S. (1995): Comparing Predictive Accuracy. In Journal of Business & Economic Statistics, pp. 253–263.

Easley, David; O'hara, Maureen (1987): Price, trade size, and information in securities markets. In Journal of financial economics 19 (1), pp. 69–90.

Eisler, Zoltan; Bouchaud, Jean-Philippe; Kockelkoren, Julien (2012): The price impact of order book events. Market orders, limit orders and cancellations. In Quantitative Finance 12 (9), pp. 1395–1419.

Fair, Ray; Shiller, Robert (1988): The Informational Content of Ex Ante Forecasts. Cambridge, MA: National Bureau of Economic Research.

57 of 61

Foucault, Thierry; Kadan, Ohad; Kandel, Eugene (2005): Limit order book as a market for liquidity. In The Review of Financial Studies 18 (4), pp. 1171–1217.

Frei, Christoph; Westray, Nicholas (2015): Optimal execution of a VWAP order. A stochastic control approach. In Mathematical Finance 25 (3), pp. 612–639.

Fusaro, Teddy; Hougan, Matt (2019): Memorandum: Bitwise Asset Management Presentation to the U.S. Securities and Exchange Commission. Available online at www.sec.gov/comments/sr-nysearca-2019- 01/srnysearca201901-5164833-183434.pdf, checked on 3/29/2019.

Gatheral, Jim; Oomen, Roel C. A. (2010): Zero-intelligence realized variance estimation. In Finance Stoch 14 (2), pp. 249–283.

Ghysels, Eric; Nguyen, Giang (2018): Price Discovery of a Speculative Asset. Evidence from a Bitcoin Exchange. In SSRN Journal. DOI: 10.2139/ssrn.3258508.

Glosten, Lawrence R. (1994): Is the electronic open limit order book inevitable? In The journal of finance 49 (4), pp. 1127–1161.

Gould, Martin D.; Porter, Mason A.; Williams, Stacy; McDonald, Mark; Fenn, Daniel J.; Howison, Sam D. (2013): Limit order books. In Quantitative Finance 13 (11), pp. 1709–1742

Graham, Benjamin; Dodd, David L. (1940): Security analysis. Principles and techniques: McGraw-Hill Book Company.

GuΓ©ant, Olivier; Royer, Guillaume (2014): VWAP execution and guaranteed VWAP. In SIAM Journal on Financial Mathematics 5 (1), pp. 445–471.

Jaisson, Thibault (2015): Liquidity and impact in fair markets. In Market Microstructure and Liquidity 1 (02), p. 1550010. .

58 of 61

Kempf, Alexander; Korn, Olaf (1999): Market depth and order size. In Journal of Financial Markets 2 (1), pp. 29– 48. DOI: 10.1016/S1386-4181(98)00007-X.

Konishi, Hizuru (2002): Optimal slice of a VWAP trade. In Journal of Financial Markets 5 (2), pp. 197–221.

Laruelle, Sophie; Lehalle, Charles-Albert; PagΓ¨s, Gilles (2013): Optimal posting price of limit orders. Learning by trading. In Math Finan Econ 7 (3), pp. 359–403. DOI: 10.1007/s11579-013-0096-7.

Lehalle, Charles-Albert; Mounjid, Othmane (2017): Limit order strategic placement with adverse selection risk and the role of latency. In Market Microstructure and Liquidity 3 (01), p. 1750009.

Lin, Ji-Chai; Sanger, Gary C.; Booth, G. Geoffrey (1995): Trade size and components of the bid-ask spread. In The Review of Financial Studies 8 (4), pp. 1153–1183.

Madhavan, Ananth (2002): VWAP strategies. In Trading 1, pp. 32–39

Madhavan, Ananth; Panchapagesan, Venkatesh (2000): Price discovery in auction markets. A look inside the black box. In The Review of Financial Studies 13 (3), pp. 627– 658.

Madhavan, Ananth; Richardson, Matthew; Roomans, Mark (1997): Why do security prices change? A transaction level analysis of NYSE stocks. In The Review of Financial Studies 10 (4), pp. 1035–1064.

59 of 61

Mastromatteo, Iacopo; Toth, Bence; Bouchaud, Jean-Philippe (2014): Agent-based models for latent liquidity and concave price impact. In Physical Review E 89 (4), p. 42805.

Menkveld, Albert J. (2013): High frequency trading and the new market makers. In Journal of Financial Markets 16 (4), pp. 712–740.

Merton, Robert C. (1973): An intertemporal capital asset pricing model. In Econometrica 41 (5), pp. 867–887.

Mizon, Grayham E. (1984): The encompassing approach in econometrics: Australian National University, Faculty of Economics and Research School of …

Mizon, Grayham E.; Richard, Jean-Francois (1986): The Encompassing Principle and its Application to Testing Non-Nested Hypotheses. In Econometrica 54 (3), p. 657. DOI: 10.2307/1911313.

Ntakaris, Adamantios; Magris, Martin; Kanniainen, Juho; Gabbouj, Moncef; Iosifidis, Alexandros (2018): Benchmark dataset for mid-price forecasting of limit order book data with machine learning methods. In Journal of Forecasting 37 (8), pp. 852–866. DOI: 10.1002/for.2543.

O’Hara, Maureen (2015): High frequency market microstructure. In Journal of financial economics 116 (2), pp. 257–270. DOI: 10.1016/j.jfineco.2015.01.003.

Ricardo, David (1817): On the principles of political economy and taxation: London: John Murray.

Robert, Christian Yann; Rosenbaum, Mathieu (2012): Volatility and covariation estimation when microstructure noise and trading times are endogenous. In Mathematical Finance 22 (1), pp. 133–164.

60 of 61

Roşu, Ioanid (2009): A dynamic model of the limit order book. In The Review of Financial Studies 22 (11), pp. 4601–4641.

Sharpe, William F. (1963): A simplified model for portfolio analysis. In Management Science 9 (2), pp. 277–293.

Smith, Adam (1776): An inquiry into the nature and causes of the wealth of nations. Volume One. In : London: printed for W. Strahan; and T. Cadell, 1776.

Stackelberg, Heinrich von (1934): Marktform und Gleichgewicht: J. springer.

Stoikov, Sasha (2017): The Micro-Price. In SSRN Journal. DOI: 10.2139/ssrn.2970694. Taylor, Mark P.;

Allen, Helen (1992): The use of technical analysis in the foreign exchange market. In Journal of international Money and Finance 11 (3), pp. 304–314.

Tóth, B.; Lempérière, Y.; Deremble, C.; Lataillade, J. de; Kockelkoren, J.; Bouchaud, J.-P. (2011): Anomalous Price Impact and the Critical Nature of Liquidity in Financial Markets. In Phys. Rev. X 1 (2), p. 383. DOI: 10.1103/PhysRevX.1.021006.

61 of 61

Bitfinex Alpha | ISSUE 2

Thought leadership | Research | Market Analysis

If you would like to speak to the Bitfinex team you can contact - OTC@Bitfinex.com