25 lines
25 KiB
Markdown
25 lines
25 KiB
Markdown
# 10 · Backtesting & Performance Analysis(回测与绩效评估)
|
||
|
||
- **系列**:Full Algorithmic Trading Using Python
|
||
- **频道**:TradeOptionsWithMe | **本集**:Backtesting & Performance Analysis(回测与绩效评估)
|
||
- **时长**:25 分 45 秒 | **原视频**:https://youtu.be/oFvYbfDOJ5c
|
||
- **本地视频**:[videos/10-backtesting-performance.mp4](videos/10-backtesting-performance.mp4)
|
||
|
||
## 🎯 本集要点(中文导读)
|
||
|
||
本集讲 **策略评估** 与 **参数优化**——非常关键的一集。
|
||
|
||
- ❌ **只看总收益是错的**:相同收益下,波动/回撤小的策略更好。要看 **回撤、波动率、稳定性**。
|
||
- ✅ **风险调整收益**:最常用 **夏普比率(Sharpe Ratio)** = (策略收益 − 无风险利率) / 超额收益标准差。**长期 >1 算不错**。还有 Sortino、Calmar、Treynor 等。
|
||
- **其他统计**:胜率、平均盈利/亏损——**要结合起来看**(例:胜率 10% 但平均盈利是平均亏损的 20 倍,仍可能很有 edge)。
|
||
- **始终用基准对照**(相关市场指数),别只比"是否盈利"。
|
||
- **优化(Optimization)**:在 QuantConnect 里可批量回测不同参数(如 SMA 天数),按 Sharpe 等排序找相对最优。
|
||
- ⚠️ **极易过拟合**:参数要**有逻辑意义**;优化本身会引入**前视偏差**(用这批数据挑最优参数)。务必留出样本外检验。
|
||
- 建议关注指标:Sharpe、回撤、年化收益、交易次数等。
|
||
|
||
## 📝 完整文稿(英文 · 自动转写)
|
||
|
||
> 由本地 ASR(faster-whisper)从视频音轨转录,未人工校对,供检索/精读使用。原始讲解以上方视频为准。
|
||
|
||
Welcome to the 10th video lesson of my algorithmic trading course using Python and the QuantConnect platform. In this video lesson, we will discuss how you can analyze and evaluate the results of your trading algorithms. After that, I will also show you an optimization procedure that you can use inside of QuantConnect. As always, I will present this with the help of an example bot after going over it on a conceptual level. First and foremost, let me start by presenting how you should not measure the performance of your trading strategy. Even though this is what most beginners do, you should not just look at the total returns that your strategy generated over a certain time period and only rely on this number alone for your evaluation. This is not a good measure for performance, since not all returns are created equally. Even if you are presented with two strategies that achieve the exact same return over a given time period, one might be a lot more desirable than the other. The reason for this is that one strategy might achieve these returns with much lower volatility and a smaller drawdown than the other strategy. Besides returns, you should therefore always consider other metrics, such as the drawdown, volatility and stability of this strategy. Since it can be quite hard to look at all these metrics individually, you can risk adjust your returns to account for all these factors with one simple number. The most common way to risk adjust your returns is by using the so-called Sharpe Ratio. The Sharpe Ratio is calculated by subtracting the risk-free rate from your strategy's returns and then dividing this figure by the standard deviation of the excess returns of your strategy. The higher this number is, the better the risk adjusted returns of your strategy are. A strategy with the Sharpe Ratio above one over a significantly long enough time frame is considered relatively good. Since the Sharpe Ratio accounts for the volatility of your returns, you can use it to better compare the performance of different strategies. Next to the Sharpe Ratio, there are many other metrics that you can use to risk adjust your returns. Other examples include the Tray Neuro, the Calima and the Sotino Ratio. In addition to risk adjusting your returns, it also makes sense to look at other stats about your strategy so that you can better understand its performance characteristics. For instance, it can be very useful to look at things such as the win rate as well as the average win and loss of your strategy. Note that here it is important to always consider these stats in conjunction with each other, by themselves that don't really tell you much. A strategy with a win rate of 10% for example might sound like an undesirable strategy. However, if this strategy has an average win 20 times the size of its average loss, this strategy seems to have an edge and should be very profitable given enough time. Besides that, you should not use profitability as a benchmark to compare the performance too. Just because a strategy managed to generate a profit does not mean that it actually is a good strategy. Instead, you should always rely on some relevant market index as a benchmark to compare the performance of your strategy too. This benchmark should have a similar level of risk as your strategy and be relevant in the same market sector. For example, if you mainly trading large-cab US stocks, you could use the S&P 500 index as a benchmark. So even if your strategy does not manage to generate a net profit over a given time period, it might still outperform the S&P 500 index and thus be a good investment compared to similar alternatives. In other words, it is important to consider the current market conditions when analysing a strategy's performance. Is the traded sector currently going through a bear market or is it having a historic bull run? These are some very relevant questions to ask yourself. Note that if your strategy always trades at a 2x leverage, it might be taking on more risk than a simple buy and hold on leverage S&P position. So besides the market sector, make sure to account for the risk. Will to try to consider the directional exposure of your strategy and its correlation to other markets and strategies. Usually you don't want your strategy to have a too high correlation with another strategy that you're trading. Furthermore, if you don't want more exposure to a given sector, you should make sure that your strategy is not highly correlated to that sector since you otherwise just basically are increasing your exposure to that sector. In general, it is important to consider personal preferences. Just because a given strategy is good for me does not mean that it is a good fit for you. Everyone has their own preferences that are determined by their risk tolerance, account size, other investments, goals and more. In algorithmic trading, you usually back test your strategies over a long enough time frame to get the best possible idea about the potential of a strategy. However, when back testing a strategy and evaluating the performance from a back test, there are some things to look out for. Before we move on to coding and evaluating the back test of an example bot, let me quickly outline what to look out for when back testing. First of all, it is essential to understand that the performance shown on a back test is not nearly as significant as actual real life trading results. This is the case because back testing makes a bunch of assumptions about the executions, fees, data, timing and availability of your trades. These assumptions might very well be unrealistic and thus lead to unrealistic results. So always make sure to take back testing results with a grain of salt. If a strategy achieves an incredible result in a back test, always be wary and try to look deeper into it. Besides unrealistic executions, costs, margin or trading models, you might also have full prey to some biases that could lead to unrealistic behavior. Some examples of such biases include overfitting, look ahead bias and survivorship bias. We've already covered these in previous videos, but let me quickly outdone a few ways to avoid overfitting since it's such a common problem. Overfitting is the act of fitting your strategy too closely to the back test data. At some point, your algorithm won't actually make meaningful trading decisions, but instead, just make certain decisions because it knows the back test data too well already. The most common mistake, especially for newcomers, is to constantly adjust their algorithms and then back test to check whether or not the performance now is better than before. This is the easiest way to end with a useless overfit algorithm. In general, when developing an algo, you should not try to back test it too often unless you're purely doing so for back testing and development reasons. One of the most common ways to try to avoid overfitting is by dividing the available data into two subsets. So if your data covering 10 years, you might want to divide it into two five year intervals. You can then use one of the subsets when developing your algorithm and the second subset for the testing. So first when you're completely done with coding and developing the algo, you should use the test data set. You can then compare the performance over the development and test time frame. If it is significantly worse over the test time frame, the algorithm might be overfit to the development data sets. Another good practice is to keep the number of parameters in your algorithm relatively low. Parameter's are fixed numbers that you can set that affect the outcome. For example, if you have an algorithm that always liquidates its holdings after 10% loss, this 10% number would be a parameter. The reason why too many parameters can be harmful is that often these parameters are set quite arbitrarily. In the just mentioned example, for instance, the likelihood is no logical reason as to why the strategy does not exit after 8% or 12% or some other percentage instead of 10%. Instead of using many fixed parameters, it can often be useful to use more dynamic models. For cutting losses, you could for example use the recent volatility of the traded asset to determine the point after which to close the position. More dynamic models also allow your algorithm to adapt to changing market conditions and thus can lead to more versatile strategies. If you are using fixed parameters however, try to always come up with a reason for why the parameters value is what it is. This reason should never just be because the back test performance with this value is better than the performance with some other value. Another good way to find out whether your strategy's back test performance is realistic or not is by paper trading it. If you let it trade live but without real money, the resulting performance can no longer be affected by the previously mentioned biases. So if the paper trading performance is similar to the back test performance, the back test results might actually be trustworthy. Note however that here it is important to let it paper trade for long enough. Performance from a few weeks isn't very significant and should not be used to make important decisions. Even if everything looks good, you should still always start slowly when going live with an algorithm. Always start with small amounts and then scale up from there. You don't want to deal with technical issues associated with deployment if the algorithm has access to half of your net worth. After deployment you should still actively monitor your algorithm and double check if it's doing what it's supposed to do. This however does not mean that you should constantly interfere with it just because you don't agree with the given trade. If you do that, you can't actually evaluate its performance since you are not giving it the chance to trade its program strategy. That said, let's now move on to the coding part of this video. For this, we will first call a simple example bot and then look at the back test performance report for this strategy together. But first, let me quickly outline the idea behind this trading spot strategy. The strategy that we will be implementing is a very simple long-term investing strategy. What we will do is look at SPY's 30-day moving average. If SPY's price is above this average, we say that there is an uptrend and in that case we want to allocate 80% of our capital to SPY and 20% to B and D which is a bond market ETF. If SPY's price is in a downtrend, so below its 30-day moving average, we want to flip this allocation around and put 80% to B and D and 20% to SPY. We do this since bonds and equities at least once used to have a somewhat negative correlation and bonds can be seen as a safer and less risk investment than equities. That's why we switch to them in case of a downtrend. But speaking of the long term, bonds have had historically lower returns than equities which is why we still want most of our portfolio in SPY if it isn't an uptrend. If SPY's price does not cross its 30-day simple moving average for a month and thus we don't adjust our positions, we want to rebalance our portfolio back to an 80-20 allocation. With that, I hope you now understand the strategy that we will be implementing. So let's now head over to QuantConnect and start coding. If you haven't already, you can use the link in the description box to create your free QuantConnect account. I highly recommend coding along for the best learning experience, but you can also use another link in the description box to clone the code from this video. So let's now head over to the lab tab and click on Create New Algorithm. As always, we will then start off with a template algo consisting of the initialise and on data methods. For the backtesting time frame, I will go with 2018 until 2021. As usual, I will leave the starting cash balance at $100,000. Next up, I will add the two securities that we will want to trade. We can do this with add equity and save the symbol into class variables. Since this is a long-term investing bot, I will go with daily resolution here. Thereafter, we will add the simple moving average for the SPY data. For this, we can use the self.sma helper method. Then all we have to specify are the symbol, number of bars, and resolution. We save this into the self.sma variable, which we can later use to access the moving average indicator values. Last but not least, we will add two helper variables. The first is self.rebalance time, which we will use to keep track of rebalancing and self.op trend, which we will use to save the state of the market. We initialise this to true. Next up, let's move on to the on-data method in which we will implement the actual trade logic. The first thing we do in the on-data method is check whether our sma indicator is ready and that there is data available for SPY and B and D. For this, we can use the ezready attribute of the indicator. If this is not the case, we simply return since we aren't ready yet. Now, we move on to the actual trading decision making. Here, we firstly check whether SPY's current price is greater or equal to the 30-day sma value. If this is the case, we know that SPY is in an up trend. However, since we do not want to adjust our holdings if we already are invested and does not get time to rebalance, we check if it's time to rebalance or we previously have not been in an up trend. If one of these conditions is fulfilled, we will adjust our holdings with two set holding statements. First, we use set holdings to allocate 80% of our capital to SPY and then we do the same for B and D and 20%. Furthermore, we set the up trend variable to true and adjust the rebalance time to plus 30 days. If the out-to-if statement is not fulfilled, we know that SPY's price is below its 30-day moving average, which means that it is in a down trend. Here, we need to perform the same check as before, so we check whether it's time to rebalance or if we have previously been in an up trend. In either case, we adjust the holdings by allocating 20% to SPY and 80% to B and D. Once again, we use set holdings for this. Next up, we set the up trend variable to false and set the rebalance time to 30 days from now. This ensures that we rebalance these holdings only if the trend changes or 30 days pass by without a change in trend. Now, we already done implementing the decision making of this trading bot. Nonetheless, there is one small thing I want to add. I will use safe.plot to add the SPY simple moving average to the benchmark chart that plots SPY's price. This will help us see during which periods we had which asset allocation. In general, it can be very useful to create custom plots such as this one when analysing your algorithm since this can significantly help with the understanding. That said, let's now build and back test this bot. When the back test is finished up, you will see a performance report like this one. At the top, there's a bar with some of the most commonly used performance metrics. The leftmost value is the probabilistic sharp ratio. This represents the theoretical probability that the sharp ratio of your strategy is above one over the specified time frame. This takes into account the distribution of daily returns of your strategy. Besides that, you can see other measures such as the unrealized and realized returns, fees, the value of all your holdings and more. Below that, you can look at one of the most interesting charts, namely the equity chart. This shows the performance of your strategy over the chosen back test time frame. Here we can see that this strategy actually performed quite well. At the bottom of the equity chart, we also see a bar chart showing the distribution of the daily returns. On the right hand side, we can select to also display the benchmark chart which is a chart of SPY. Since we plotted the SMA onto this chart, we can also see it here. Even though this chart is good for a quick comparison, it could also quickly adjust the algorithm to buy and hold SPY over the same time frame, so let me quickly do that so that we can better compare the performance. All we have to do for this is add a set holdings and return statement to the beginning of the on data method. After that, we can then build and back test this strategy again. This would then generate the same report for a SPY buy and hold strategy. As you can see, the SPY buy and hold strategy achieved a return of almost 50% over this time frame, which actually is about 10% more than the other strategy. However, if we look at the probabilistic sharp ratio or the equity chart, we actually see that this higher return was achieved with a much higher volatility. This means if we use the sharp ratio to risk adjust the returns, the buy and hold strategy actually underperformed our custom dynamic allocation strategy. That said, let me now go back to the report of our strategy. Here we can scroll down to see a bunch of other performance metrics and stats that can be used to learn more about its performance. In the top right corner, we can see that the sharp ratio for this strategy over this time frame is slightly above 1. Other interesting stats, such as the number of trades, winning percentage, average win and loss can also be found here. Since covering all these metrics in detail will be quite boring and too much for this video, I highly recommend googling some of these and reading how they are calculated and how they can be used. Really understanding these stats can help you develop a much better understanding of your strategy's performance. By skimming this back test report, we can conclude that this strategy does seem to reduce the overall volatility compared to simple S&P buy and hold strategy. However, this reduction in volatility does come at a cost which is a reduction in net returns. Depending on your appetite for risk, the added stability might very well be worth this trade off. But here it is important to keep in mind that this reduction in returns might be even more significant over longer time frames. However, note that this is just based on a quick look at this one back test. It would be a very good idea to look at some other time frames to see the broader picture. Especially looking at the performance during other market conditions, such as the 2008 crisis can be a very good idea and a way to stress test and better understand the performance of your bot. By looking at the performance during the early 2020 crash and comparing it to the benchmark chart, we can already see that this strategy does seem to better handle down trending markets, which isn't very surprising. However, instead of analyzing this strategy in even more detail, I will now briefly show you QuantConnect's parameter optimization feature. If you want to analyze this strategy in more detail, you can clone the code and test it yourself using the link in the description box below. For the parameter optimization, we first have to decide which parameter or parameters we want to optimize. Since this strategy actually does not have that many parameters, we will use the number of bars of the simple moving average for the optimization. Currently, this parameter is set to 30 bars, but since this is not based on anything, we could try to look at the effect of changing this parameter on the performance. To do so, we first have to define it as a parameter. This can be accomplished with self dot get parameter and then the name that you want to give this parameter. We will call this sma underscore length and save it into the length variable. Then we will expand sidebar and scroll down to the algorithm parameter section. Here, we click on add new parameter and add a parameter with the name sma underscore length. After that, we click on this parameter and enter 30 as the initial value for this parameter. Now the algorithm uses the get parameter helper to get the value of this parameter from here. This is useful since when you have already deployed an algorithm to live trading, you would usually not want to edit the source code anymore. This feature then still allows you to change parameters without touching the actual code. However, in some cases, the get parameter method might not be able to get a value. To account for this case, we need to check whether or not it returned the value and if it did not, we used 30 as default value. This ensures that the length variable always is well defined. So now we only have to specify this length variable to be the second argument of the sma helper method. To optimize this parameter, we can now click on the optimize button next to the back test button. This will bring you to a menu where you can select what you want to optimize for. Here, you can maximize or minimize things such as the sharp ratio and your return or drawdown. We will go with maximizing the sharp ratio. Next up, we have to select which parameters we want to optimize. Since we only defined one parameter, we will use the sma underscore length parameter here. As the minimum value, I will go with 15 and as the maximum I will specify 90. Under additional settings, I will set the step size to 2. This means that this optimization will back test our strategy for every second number between 15 and 90 for the sma underscore length parameter. Optionally, we can also add some constraints for the sharp ratio, drawdown or annual return, but we won't do that here. In the final step, you now have to select an optimization node as well as the max number of nodes that you want to use for this optimization. Note that this optimization feature is not free of charge. Depending on the number of back test and complexity, this can cost anything from a few cents to hundreds of dollars. But an optimization like this one should not cost more than a couple of cents. I will just deploy the max number of nodes for this example since this will take the least amount of time. Deploying with the smaller number of nodes, let's confirm the optimization is working before scaling up the cluster. When the optimization has finished, you will see a screen like this one. At the top, you can see the equity chart of all the back tests. As you can see, the performance of these back tests does vary quite significantly, which means that the value for the time span of the spysma is important. Below this top chart, you can see three heat maps for the sharp ratio, compile return and drawdown. Beneath these charts, there's a table in which you can see all the exact values for the different parameter values. Here we can sort the back tests by different metrics. For example, if we sort them by sharp ratio, we see that the back test with the best sharp ratio is that with a 39 day simple moving average. This back test has a sharp ratio of over 1.2. By back testing so many times, the optimization does include some look ahead bias however. Besides the columns that you can see here, you can also add more metrics by clicking on the columns button to the right on the table. Here you can add other stats such as the win rate, average size of win or loss, number of trades, and more. For more details on a specific back test, you can click on the back test name. This will bring you to a standard in-depth performance report of that back test. Note that even though this optimization feature can be very useful, you should still be very careful when using it since it can easily lead to overfitting. Like I mentioned earlier, the parameter value should still make logical sense. That said, I hope that this video gave you a good overview of what to look out for when analyzing and evaluating the performance of a trading strategy. I hope you now understand that you should not only look at the returns and profitability of a strategy. Instead, always use some benchmark for reference and consider how the returns were achieved. If you have any questions or comments, don't hesitate to let me know in the comments section below. Otherwise, make sure to smash the like button, subscribe and turn on the notification bell for more videos like this one. Thanks for watching.
|