Files
quanxiel/course/Full Algorithmic Trading Using Python/17-bitcoin-ml-bot.md
T

24 KiB
Raw Blame History

17 · Bitcoin ML Bot(用神经网络做比特币)

  • 系列:Full Algorithmic Trading Using Python
  • 频道:TradeOptionsWithMe | 本集:Bitcoin ML Bot(用神经网络做比特币)
  • 时长:31 分 7 秒 | 原视频:https://youtu.be/waiBgdalmSE
  • 本地视频:videos/17-bitcoin-ml-bot.mp4

🎯 本集要点(中文导读)

用机器学习(神经网络)做比特币交易(思路同样适用于其它标)。

  • 思路:用神经网络输入"过去 30 天的价格+成交量",预测"下一根收盘价会高于还是低于当前收盘价",据此做多或做空。
  • 神经网络简介:输入层 → 若干隐藏层(每个神经元做加权计算 + 激活函数)→ 输出层;通过训练数据调整权重来学习。
  • 工具:Python 的 TensorFlow + Keras 搭建网络。
  • ⚠️ 作者强调的关键点:
    • 用 AI/神经网络时,最好把它当辅助工具(优化现有流程、作为策略补充),不要只依赖网络预测。
    • 严防过拟合:本集用 2020–2022 训练,在这一区间回测"必然好看"(模型见过),不能代表真实表现。
    • 神经网络还可用于预测:波动率、长期趋势、成交量、与其他资产的相关性、指标值等。

📝 完整文稿(英文 · 自动转写)

由本地 ASR(faster-whisper)从视频音轨转录,未人工校对,供检索/精读使用。原始讲解以上方视频为准。

Welcome to another video. In this video you will learn how to create a cryptocurrency trading algorithm that uses machine learning techniques to make trading decisions. Even though I will use Bitcoin as the underlying asset in this video, note that you can also create such a bot for other cryptocurrencies, stocks, forex or whatever you want to trade. We will code this trading bot using Python and the QuantConnect platform. Note that you can follow along and do everything that I do in this video with a free QuantConnect account. If you haven't already, you can create your free account using the link in the description box below. If you are new to QuantConnect, don't worry. I will go over everything step by step in this video. However, if you want to become an algorithmic trader yourself and learn how to develop your own trading bots, you could check out my free algorithmic trading course on my channel. Before we will head over to QuantConnect and actually start coding, let me quickly outline what we will do in this video on a conceptual level. Like I already mentioned, the strategy that we will be implementing will make trading decisions based on artificial intelligence techniques. More specifically, we will create a neural network that takes the price and volume information of the last 30 days of Bitcoin as the input and then tries to predict whether the next closing price will be above or below the last closing price. Based on this prediction, our trading bot will then either buy or set a Bitcoin coin. If you're not familiar with neural networks, three blue one brown has a great introductory video on them. But in very simple terms, a neural network is a weighted graph that consists of a bunch of nodes that are called neurons. The first layer of a neural network is the input layer. And this is where you can pass what you want the network to use as an input. In our case, this would be where we pass the 30 day price and volume data. The last layer is the output layer, and we will use the value output here as our prediction for the direction. In between the input and output layer, they usually are a variable number of hidden layers. Each neuron applies a certain function to the input it receives, and then outputs their respective results based on a so-called activation function. Depending on where this neuron is, that output will then be used as the input of another neuron or as the final output. The connection between the neurons are labeled with certain weights that are used for the calculations in each neuron. Through certain techniques, you can train a neural network by feeding it a bunch of training data and then adjusting the weights based on the accuracy of its predictions. Depending on the problem, the predictions thereby become more and more accurate and sophisticated. For now, you can just think of this as a black box that spits out a certain output based on what you give it as an input. We will be creating our own neural network using the TensorFlow and Keras modules for Python. With that being said, let us now head over to QuantConnect and start actually writing some code. Note that for the best learning effect, I highly recommend coding along. However, I will also post a link in the description box below that allows you to clone the entire code from this video for you to play with after watching this video. Inside of QuantConnect, we will go to the lab tab and create a new algorithm from there. This will bring you to a screen like this with a bare-bones template for a new trading algorithm. Before we code out the actual algorithm though, we will first create our neural network in QuantConnect's research environment which uses Jupyter notebooks. The first thing we do in this research notebook is import all the required dependencies that we will use. This mainly includes some things from TensorFlow and Keras. You can run the code in a given code cell by clicking Shift and Enter. After we are done with our imports, let's save the start and end dates that we will use for the training and testing data for the neural network that we will create. For this video, I will use two years from 2020 until 2022 but feel free to try out different time periods. Now we will initialize an instance of the QuantBook which is QuantConnect's research environment class that gives you access to all of QuantConnect's research API methods. The first of these methods that we will use is AddCrypto which will add the data for Bitcoin over the chosen time period. To actually access the specific historical price data, we use the history method which will return a data frame consisting of Bitcoin's price data over the last two years. For our neural network, we will not use the raw price data though. Instead, we will use the percentage change of the open, high, low, close and volume data of every day since we just care about the changes in this data and not the actual absolute values. Before recording this video, I took a look at this data and saw that there are some problems in the volume column. Some of the values got assigned infinity as a value. Since we can't use this for any calculations, we will place the occurrence of infinity with a max value in the volume column. Now that we've cleaned up the data frame, we have to bring the data into the correct format so that we can actually use it to train a neural network. For this, we create some variables. The Features list is basically the input list and the labels list is the desired output. When we later actually use this neural network, we won't yet know what the desired output would be since we are trying to predict it. But when training it, we will use the labels list to improve its accuracy. So each element in the Features list should consist of the open, high, low, close and volume values for each of the last 30 days. An element from the labels list on the other hand will be one if the price was up on that date and zero if it was down. Before we can now actually use these lists for the neural network, we need to turn them into non-pay arrays since that is a required input format for the tensorflow method that we will use. Next up, let's divide this data into training and testing data. We will use this training data to actually train and improve the neural network and the testing data to test it. Here it is very important to not mix the two up. It is very easy to create a model that performs well on training data since it can just remember the desired output and overfit itself to that data. However, in that case, the model would not perform very well on new data that it has not seen before. And that's exactly what we will use the test data for, to make sure that it is not too overfit to the training data. So let's use 70% of the two-year data for training and the rest for testing. A common pitfall here is that the output of the training data is not evenly distributed. If for example, 80% of the days that we use to train the model with are up days, an easy way for the model to achieve an 80% accuracy is by just always predicting that a day will be an update. However, that is not a very intelligent guess and that won't work on other data. That's why we check the ratio of up to down days here. As you can see in the training data, almost 56% of the days are up days. Ideally, this would be as close to 50% as possible. So let's check if this improves if we consider the last part of the two-year time frame instead of the first part. As you can see here, this actually leads to a slightly more even distribution of about 53% updates. This is still not optimal, but I leave it for now. Now we can finally start building the actual neural network. For this, we will use Keras sequential model, which basically allows us to build a network layer by layer. We will start with a density-connected layer as the input layer for our neural network. For this layer, we will specify the shape of the training data as the input shape so that the network knows what input to expect. As I briefly mentioned earlier, each neuron in a neural network has an activation function that decides what a neuron should output. If you want the output to be binary, like in the case of a classification problem, a common activation function is a step or sigmoid function since it transforms the input to a value between 0 and 1. Another very common activation function is the real-you activation function which we will use here. Next up, we will create another dense layer with the same function. After that, we want to output our result in one last neuron, but for this, we first need to add a flat layer that flattens the input data. Then we can create the last layer with only one output. Since we here want a value between 0 and 1 to encode the predicted direction, we will use the sigmoid function. If you want a full tutorial on how to apply Python, TensorFlow and Keras for deep learning, you could check out a great tutorial series by CentDex that I linked in the description box below. Note that finding the optimal number and types of layers and connections for a specific neural network depends on the problem and is still a heavily researched problem. So here it often helps to play around a little to see what yields good results. After defining the neural net, we can now configure the model for training. Here we define a loss function, an optimizer, as well as the metric we want to monitor when training the network. The binary cross entropy loss function is a common choice for binary classification problems like this one. And the addm optimizer is probably one of the most used optimizers as of right now. Remember that when optimizing a model to always pay attention to what metric you're optimizing for, since otherwise your model might be trying to achieve something that wasn't actually a intended goal. Now we can finally fit our training data to this neural network to actually train it. Usually you go through the entire training data multiple times when fitting neural networks. Each of these cycles is called an epoch. The more epochs you run, the more the neural network will adjust itself to the training data. However, here it is important not to overfit the model by letting it run for too long. We will let it run for five epochs for now. While it runs, we can actually see how the accuracy and whatever other metrics we configured above change for each training cycle. As you can see here, the prediction accuracy for the training data seems to be steadily increasing for each epoch we let it run. Note that these values will be different every time you run the fitting since there are some probabilistic processes involved. So if you're coding along, keep in mind that the neural network you will get probably will not be the same as the one in this video. Now that the model is trained, we want to test its performance using the testing data that it has not yet seen. For this, we let it run an X test and then format the data so that we can plot it on a graph. On this graph, the blue line represents whether a given day actually was an update or not, while the orange line symbolizes the prediction of the network. If the orange line is above 0.5, we consider it to have predicted an update while below 0.5 means it thinks the day will end in the red. If we were to increase the number of epochs that we use to train our model, we would see the orange plot edge and closer and closer to the blue plot. But this would likely not improve the model's accuracy on the general data but just overfitted to the training data. Even though this graph gives us a good general picture, we can't actually use it to analyze its accuracy very well. So let's instead output the model's accuracy and error first on the training data and then on the testing data so that we can compare the values with each other. As you can see, the accuracy for the test data is lower than the accuracy for the training data. However, the difference isn't that big so I will just leave the neural network as it is. But as you can see, even with 30 days of price and volume data, this network does not manage to accurately predict Bitcoin's daily price direction much more than 50% of the time, which just shows how hard it is to do this. Like I said, I will leave it as it is right now for the rest of this video, but I highly recommend for you to code along or clone this notebook using the link below and play around with the network yourself to see if you can manage to improve it without overfitting it. Now we actually want to use this model that we just created inside of a QuantConnect algorithm to make trading decisions based on the predictions. However, to do so, we first save the model as an object into QuantConnect so called object store, which allows you to save models such as this one so that you can access and use them from wherever you want to. To accomplish this, we turn the model into a JSON object and then save it into the object store using the key Bitcoin price predictor. Let me now show you how we can access this model directly from the object store. I will show you how to do this in the research notebook, but we can then basically use the exact same code inside of the trading algorithm that we will create in a few minutes. To get the model from the object store, we first check where that exists. If it does, we read it in as a string by importing it using dot read. Thereafter, we turn it into a JSON object and then turn it into the neural network we want using Keras sequential class. To test if everything is working as intended, let us get the prediction for today's Bitcoin price direction. For that, we access the past 40 days of price history once again using the history method. To actually use this data to get the prediction from the model, we now just have to turn it into the right input format. So let's focus on the open, high, low, close and volume columns and get their percentage change values. Then we take the last 30 rows, add them to a list and wrap an umpire array around it. With that, we can now use this as the input data for our neural network and check whether the rounded output equals zero or one. Note that the output of the prediction is a double nested array, which is why we first have to index it. With these few lines of code, we can now get the prediction for any date that we want to using the model that we created. So let us now copy this code and use it to create an actual trading bot. For this, we go back to the main.py file. Here, we once again need to import the tensorflow.keros sequential object as well as the JSON library. The first thing we do after that is implement the initialize method, which is used to initialize our trading algorithm class. Here, we for instance set the backtest time frame for this algorithm. I will go with 2018 until 2020. Here it is important to keep in mind that if you use the year 2021 to backtest this algorithm, you might get quite biased results since this is the time frame we used to train the neural network with. After setting the backtest time frame, we can copy paste the code from the research notebook to import the actual model. Note that the quantbook class is a wrapper class around the quant connect algorithm class, which is why it is important to use self instead of the QB quantbook identifier for some of the methods. In the next step, we will set the brokerage model so that we can buy and sell Bitcoin however we want to. For this, we choose bid finics as the exchange and a margin account as the account type. Then we set the starting cash balance for this backtest. Here I will go with 100,000 USD but feel free to enter a different value here. Next up, we need to add the actual cryptocurrency to this algorithm so that we can trade it. We do this with add crypto. Quant connect supports data in many different resolutions, even including minutely resolution, but since we are only interested in the daily data, we will go with daily resolution. Last but not least, we set bitcoin to be the benchmark for this algorithm. This will allow us to later compare this strategy's performance versus just investing in Bitcoin. Now we are done with the initialize method. Next up we will implement a method that will allow us to get the prediction for the next day using the model we trained in the previous part of this video. For this, we can basically just copy the code from the research notebook that transforms the data into the right shape and once again replace the QB identifier with the Python self class identifier. So here the history method gets the past 40 days of data and then it is transformed into the right input format for the model. Then we check whether the model's output for that input is closer to one or zero and then return it down or up respectively. Note that since we will call this method every day in this algorithm, this is actually not a very efficient way to implement this. The history request will always have to fetch 40 days of data every day, even though we already fetched 39 of the same days on the day before. A more efficient way to accomplish this would be to use a rolling window that stores the intermediate data, but I will not be doing that here. If you want to learn how to do this and all the details of developing algorithms in Quant connect, you should check out my free algorithmic trading course. A possible variation of this algorithm would be to use the actual output values of the neural network instead of the rounded counterpart. You could then interpret values closer to one or closer to zero as more confident predictions than values closer to 0.5. You could then, for instance, reflect this confidence level by adjusting the position size, but we will just stick with the rounded values for now. All that's left to do now is implement the on data method, which is called every time the algorithm receives new data, which should be daily since we added data at a daily resolution. So every time this is called, we want to get the prediction of the model using the method that we just implemented. If that returns up, we want to set our holdings in Bitcoin to 100%, which is equivalent to buying as much of it as we can. Otherwise, if the model predicts down, we want a short Bitcoin. But due to the high risk nature of shorting and volatility in crypto, we only set this to 50% of our available capital. Alternatively, you could also decide to just go flat in this case. With that, we can now click on Build and Backtest to find out how this algorithm performs over the given time frame. This will start the backtest, which should be done in a few seconds to a minute. When the backtest is done, a performance report like this one will be generated that allows you to analyze the strategy's performance. If you scroll down, you can also find more performance metrics, as well as the option to view every order that was sent by the algorithms separately. This is great to check if the algorithm is actually doing what it's supposed to do. If we look at the equity chart at the top, we can see that this bot did not seem to perform that well over the chosen backtest time frame. However, if we look at the benchmark chart of Bitcoin's price, we can also see that Bitcoin itself didn't exactly perform that well during 2018 until 2020. Nonetheless, our algorithm's performance definitely could use some improvements. However, the point of this video was not to create the best possible trading strategy, but instead to show you how easy it is to develop your own trading algorithms, even if you want to use more complex machine learning and AI models. Once you've developed and thoroughly tested an algorithm that you are satisfied with, QuantConnect allows you to connect the algorithm to your broker platform so that it can actually trade with real money. Besides applying a neural network such as the one presented in this video to Bitcoin, QuantConnect also allows you to do the exact same for almost all assets. In fact, you could easily clone the code from this video and simply use add equity instead of add crypto in the research notebook, as well as the algorithm itself to try if this works better for some stock or market ETF. Just now that you should also just remove the brokerage model line in that case. In general, a neural network probably works best for more predictable stable assets, which are highly volatile cryptocurrency probably isn't. Here it is best to find a period that is as representative of that asset's general price behavior as possible. Finding such a period is a challenge in itself though. That being said, short term price changes such as daily changes are notoriously hard to predict across all asset classes since the data is very noisy. Therefore, it generally makes more sense to focus on more stable variables when using AI techniques such as neural networks. Besides using a neural network to predict the direction, there are many more use cases where you could apply such a technique. Possible examples include predicting the upcoming volatility, for instance by using the percentage change as the output, trying to predict the longer term trends, the volume, correlation to other assets, indicator values and more. But in general, it is a good idea to think about AI techniques such as neural networks more as a supplementary helping tool in your strategies instead of solely relying on the predictions of a neural network as your strategy. You can use them to improve and optimize existing processes and as additions to your strategy. Also, always make sure to pay attention not to overfit your models or to draw too many conclusions from the results that stem from the training data since the model already has seen this data. For example, in this video, we used the time frame 2020 up to 2022 to train the model. So the testing of our trading algorithm on this time frame will not necessarily yield representative results since the model knows this data already. That said, I really hope you enjoyed this video and learned a lot. To learn more about creating your own trading algorithms, check out my free algorithmic trading course in which you'll learn everything you need to know on how to develop your own trading algorithms including dozens of example trading bots and much more. If you enjoyed this video, make sure to smash the like button, subscribe and turn on the notification bell. Thanks for watching.