Files
quanxiel/course/Full Algorithmic Trading Using Python/16-data-driven-research.md
T

18 KiB
Raw Blame History

16 · Data-Driven Research(研究环境 · QuantBook)

  • 系列:Full Algorithmic Trading Using Python
  • 频道:TradeOptionsWithMe | 本集:Data-Driven Research(研究环境 · QuantBook)
  • 时长:21 分 52 秒 | 原视频:https://youtu.be/XUib6Y3eePs
  • 本地视频:videos/16-data-driven-research.mp4

🎯 本集要点(中文导读)

本集讲 Data-Driven Research(研究环境)——写策略前的"研究/探索"环节。

  • Research Environment(研究环境):QuantConnect 内建 Jupyter Notebook,用于交互式探索数据、可视化、找相关性、建模型。
  • QuantBook 类:研究环境里不能用 QCAlgorithm,改用 QuantBook(是 QCAlgorithm 的包装)——可访问历史数据、consolidate、画图、建指标等;但不能用事件驱动方法(OnData、universe 事件处理器)。
  • 数据:以 Pandas DataFrame 组织(带标签的表格);可访问股票/期货/期权/外汇/加密的历史与另类数据。
  • 典型用法:探索数据 → 找相关/因子 → 建 机器学习/统计模型(本集演示了一个线性回归示例)。
  • 把研究转成算法:把 QuantBook 代码直接导入/复制到 QCAlgorithm(很多方法通用,通常把 qb 换成 self)。
  • ⚠️ 提醒:线性回归只能捕捉线性关系;示例数据量小,不代表样本多就一定更好。

📝 完整文稿(英文 · 自动转写)

由本地 ASR(faster-whisper)从视频音轨转录,未人工校对,供检索/精读使用。原始讲解以上方视频为准。

Welcome back to the next video of this Python algorithmic trading course. So far the focus of this video course was on the implementation side of the algorithmic development process. However, often you would want to perform certain research before actually coding out to strategy and backtesting it. During this research you could explore the available data, visualize it, analyze it for any potential correlations and even create advanced models to get a good idea of what kind of strategy you could use in the given sector. For this, QuantConnect has a built-in research environment that uses Jupyter Notebooks so that you can analyze data in an interactive manner. In this video we will create a new QuantConnect project and explore this research environment by going through some of the available features in an example research notebook. Typically the data inside of the research environment will be structured into Pandas Data frames. Pandas is a data analysis library for Python that basically allows you to structure data into tables and easily access and manipulate that data. These tables are so-called Pandas Data frames and you can think of them as labeled spread sheets. If you have never used Pandas before it might be a good idea to watch some introductory videos on how to use it since I will not cover all the Pandas basics in this video. That being said, before we head over to QuantConnect and look at the research environment in action let me quickly outline some of the theory relevant to the research notebooks. First and foremost, since you aren't coding actual trading algorithms in the research environment you can't use the QC algorithm class. Instead you have to use the so-called Quantbook class which is a wrapper around the QC algorithm class. This means that the Quantbook instance will allow you to access methods of the QC algorithm class inside of the research environment plus more. However, note that since you aren't actually back testing anything per se you can't use event driven methods such as on data or universe selection event handlers. But you can access all the historical data, consolidate it, visualize it with charts, create indicators and much more. QuantConnect allows you to access and analyze historical and alternative data for equities, functions, futures, forex and crypto. Another use case of the research environment is to develop machine learning or other advanced data science models that you can then use for your actual trading algorithms. After sufficiently analyzing the data in the research environment you would want to translate the Quantbook based research code into QC algorithm based code that can actually trade for you. For this you can actually import certain classes, models and methods from the research notebook directly into your QC algorithm. Furthermore, many methods can be directly copied into the algorithm since the Quantbook simply is a wrapper class. That being said, let us now actually go over to QuantConnect and create our first research notebook. To do so we have to create a new algorithm project like we always do. You can do this from the lab tab. However, instead of actually editing the main file of the project, we now click on the research notebook file in the sidebar. This will open the research environment including a few sample lines of code that we will remove for now. Note that it can be a good idea to code along and play around with the research environment yourself to better understand the concepts presented in this video. However, you can also clone the finished notebook of this video using a link that I'll put in the description box below. That said, the first thing we do here is import two libraries, namely NumPy S&P and Matplotlib.pyplot as PLT. We will use NumPy for calculation purposes and Pyplot for plotting. To execute a code block you can hold shift and press enter. If you're not yet familiar with Jupiter notebooks, you can click on the help button at the top for more useful shortcuts. Next up, we create an instance of the Quantbook class and save it to the variable QB. Since we now have a Quantbook instance, we can start adding data for the securities that we want to analyze. For this example notebook, we will focus on seven financial stocks, namely JP Morgan, Bank of America, Morgan Stanley, Schwab, Goldman Sachs, American Express, and Citigroup. To add these stocks to the Quantbook, we can use the usual add equity method. The same goes for other asset classes. We save the symbol object of each of these securities into the symbols list. Now we're interested in looking at the price to earnings ratio of the just added securities and its relationship to the price. So firstly, let's visualize how the PE ratio of these stocks has changed over the course of 2021. For this, we can use get fundamental and pass the symbols list, PE ratio, and a time frame as arguments to get the data. This method will return a Pandas data frame with a PE ratio over the specified time frame for each of our symbols. Note that in this data frame, you might notice that the symbols of some of the stocks deviate from the ticker that we specified earlier. For instance, the Bank of America symbol starts with NB. In contrast to tickers, QuantConnect symbols are constant and always refer to the underlying entity regardless of rebranding or name changes. However, to better understand the data, we will rename the columns to the actual current company names. Note that here it is important to respect the order of the columns in the data frame. To actually plot the PE ratios, we can use matplotlib. Before plotting it, however, I will quickly add some appearance parameters to increase the size of the chart and add some labels. As you can see, most PE ratios stayed relatively constant throughout the year. The biggest change was in America Express' PE ratio, which dropped from over 30 to under 20. Next, let's look at and compare the average PE ratio of these stocks throughout 2021. For this, we simply take the mean and then sort the values. Here you can see that Goldman Sachs had the lowest PE ratio closely followed by Citigroup. Schwab was the stock with the highest mean PE ratio of almost 30. Another interesting aspect to consider is the relationship between the returns of each stock in 2021 and the respective PE ratio. However, before we can look at that, let's first look at the returns that each of these stocks generated in 2021. To calculate these returns, we first need to look at the price data for each of these stocks. We can access the price history using the history method and passing the symbols list, time frame and resolution. This will return a multi-index pandas data frame with the open, close, high, low and volume history for each of these stocks. But since we are only interested in the close price, we only save that information into the history variable. We once again rename the columns for better readability and then output the first five columns of this new data frame using the head method. With this data frame, we can now easily plot the returns of these stocks using the percent change method. Since we always need two values to calculate the percentage change, we have to start at the second row of the data frame. Once again, we use matplotlib for the plotting. As you can see, there definitely seems to be some form of correlation between these stocks which shouldn't be too surprising since they're all in the same sector. Nonetheless, there is some deviation across their returns with Schwab being the best returning stock and Citigroup the worst one in 2021. Now we want to analyze the correlation between these 2021 yearly returns and the average PE ratio of these stocks. To calculate the correlation coefficient, we can use one of NumPy's methods and pass the last row of the returns data frame and the average PE ratio list as arguments. Here we can see that there seems to be a slight positive correlation between these two. In other words, according to this, granted very limited data, a higher average PE ratio led to higher returns for these stocks in 2021. A high PE ratio means that the company shares cost a lot compared to their earnings. There could be multiple reasons why this might be the case. We can also visualize this relationship with a scatterplot which is exactly what I will do next. Note that this is just to give you some examples of how you can use the research environment. These aren't very meaningful findings, especially since we are only considering a very insignificant number of stocks over one arbitrary time period. In the next step, let's take a look at how we can add data for more complex securities such as options. I will quickly demonstrate this using Bank of America as the underlying security for the options. Just like in standard QC algorithms, we can use Add Option and Set Filter to Add Options to the Quantbook. The Set Filter method takes the number of strikes that you want to consider above and below the current price as well as the time frame until expiration as arguments. Here I will go with a simple filter of a 10 point wide strike range and 20 to 50 days until expiration. To actually get historical data for the options over a certain time period, you'll have to use the GetOptionHistory method. You then have to pass the Request symbol as well as the time frame that you're interested in. Here it is important to understand that this will actually not directly return a data frame. Instead, it will return an Option History object. This Option History object has a bunch of helper methods that you can use to filter the option chains. For instance, you can use Get Strikes and Get Expiry dates to access the available strikes and expiration dates for the requested time frame. To access the actual historical data, you'll have to use the GetAllData function. This will then return a data frame containing information such as open, close, low, high and volume among others for all the available options. Now you can visualize and analyze this data to whatever extent you want to. However, instead of doing that now, let me rather show you some other useful aspects of the research environment. So let's instead look at how you can add, use and visualize indicators here. I will demonstrate this by plotting a BollingerBand indicator on Bank of America stock. But this works similar for other indicators. First, we need to create an instance of a BollingerBand object. We can do this with this constructor. As arguments, we pass the number of days that the moving average of the BollingerBands should take into account as well as how many standard deviations the outer bands should deviate from the mean price. Here I will go with a 30-day moving average and two standard deviations. Next, we get historical data for this indicator by applying it to Bank of America over the past 360 trading days. Note that here we specify that we want the BollingerBand to be created from the open price instead of the default closing price option. Last but not least, we use Matplotlib to plot this indicator at all its related data. Here you can see that besides the price and the upper, lower and middle band, we also plotted other metrics such as the band width. But since we only are interested in the actual indicator plots, we want to drop the columns referring to these other metrics. Thereafter, we create the plot again and get a less cluttered view of the actual indicator bands. Now you could further analyze the indicator data by for example putting it into relation with the underlying price. But once again, I will not do that here. Besides the rather simple analysis tool shown so far, the research environment also supports countless advanced Python data science and machine learning libraries. So it's very much possible to develop complex machine learning models and then import and use these in your actual algorithms. Building such a model is outside of the scope of this video. Nonetheless I want to give you a taste of the potential. That's why we will create a very simple linear regression model that attempts to predict stock prices based on the BollingerBand SMA. For this, we will use the last 60 days of 2021 of Bank America's Price and BollingerBand data to train and test this regression model. Furthermore, we will use scikit-learn which is a Python statistics library to create the linear regression model. To do this, we have to add the historical price data for Bank of America which we can once again do with the history method. Since we are only interested in the closing price, we can refactor this data frame and only save the daily closing prices into a list. The idea behind this very simple model will be that it will get the value of the middle band of the BollingerBand indicators the input and it will then try to predict the closing price purely based on that input. Note that the middle band basically just is a 30 day moving average of the daily opening prices. But to do this, we first need to divide our 60 days of data into training and test data. We will use the first 30 days as training data and the last 30 days to test our model. We can train a linear regression model using scikit's fit method. However, before we can do that, we need to bring the data into the right format which is exactly what we are doing here. After fitting the model to the data, we want to use the model to predict the last 30 days of the closing prices so that we can evaluate the model's accuracy. We can do this using the dot predict method. However, before that, we once again have to bring the data into the right format. To actually evaluate the accuracy, we will plot the training and testing data on a sketch plot. Since we want to be able to differentiate between the training and testing data, we will call the training data blue and the testing data green. Furthermore, we will plot the linear regression line in red. The x axis represents the middle band of the BollingerBand indicator while the y axis represents the closing price. Let me now quickly explain what this plot shows since you might not be very familiar with linear regression models. A linear regression model tries to find a linear relationship between two variables. In this case, the variables are the indicator value and the closing price. The red line represents this model. So what the model asserts is that the higher the middle band's value is, the higher the closing price of the stock will be for that day. To give you an example, it says that if the value of the middle band is 44, it predicts the closing price to be 46. If we look at the blue dots, we can actually see that they are somewhat aligned with the red line, which is because the model used these data points to create that line. However, if we look at the green dots that represent the testing data, we can see that the red line does a very bad job predicting these closing prices. This should not be too surprising since one simple moving average value is not enough to accurately predict the closing price of a stock. A stock can even have the exact same indicator value in the morning on different days and still close at very different prices. So unsurprisingly, that does not seem to be a significant linear relationship between the daily 30 day moving average values and daily closing prices. Also note that this is a very simple example. We only looked at 30 days of data for this model, which isn't exactly a lot. Nevertheless, that does not mean that using more data would lead to a more useful model in this specific case. In general, it is also important to understand that linear regression models can only be used to find linear relationships between variables. It cannot be used to find more complex relationships. Nevertheless, such models can be very useful in some scenarios. I hope that this brief example gave you a good overview of the potential that the research environment offers for building your own models and all the data that you can use to build such models. Besides the data presented in this sample notebook, you can also add data for a bunch of other asset types, alternative data and much more. For more details on this, I recommend checking out the docs. If you want to copy the sample notebook that I created in this video, you can clone this code by using a link in the description box below. For even more sample notebooks, make sure to check out the documentation pages as well as the community forum. Once you're done with your research, you can also easily convert the code from a research notebook into code supported by the QC algorithm. Some functions you can even directly copy or import into your algorithms. For the helper methods, you can often just replace the QB quantbook identifier with a self-pilot Python class identifier. With that being said, I really hope you enjoyed this video and learned a lot. If you did, make sure to smash the like button and subscribe and turn on the notification bell. Thanks for watching.