Files
quanxiel/course/Full Algorithmic Trading Using Python/08-dynamic-universes.md
T

24 lines
24 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# 08 · Dynamic Universes(动态证券池与基本面选股)
- **系列**:Full Algorithmic Trading Using Python
- **频道**:TradeOptionsWithMe | **本集**:Dynamic Universes(动态证券池与基本面选股)
- **时长**:25 分 30 秒 | **原视频**:https://youtu.be/HfGHS-HeDK8
- **本地视频**:[videos/08-dynamic-universes.mp4](videos/08-dynamic-universes.mp4)
## 🎯 本集要点(中文导读)
本集讲 **动态 Universe(动态证券池)** 与 **基本面筛选**,重点解析 **幸存者偏差**。
- **Universe(universe 选择)**:让算法动态地、按市场条件决定"今天交易哪些标的",例如:只交易接近 52 周新高的股票、高波动股票、盈利稳健的公司、或财报公布前后的股票等。
- **幸存者偏差(Survivorship Bias)**:示例——"10 年前市值 >100 亿的股票买入持有"的回测,若只统计**今天还活着**的股票,结果会虚高(破产/被并购的公司被忽略)。需用"10 年前**当时存在**的全部股票"来筛选。
- **选择偏差(Selection Bias)**:即使只交易单一标的也有此偏差——如"只做苹果"要警惕:表现好可能只是**苹果本身强**,而非策略有效。要问:"如果苹果不再复制过去 10 年,策略还成立吗?"
- **QuantConnect 的 Universe 机制**能方便地规避这些偏差,并支持用**基本面数据(Fundamentals)**做筛选。
- 本集实现一个**基本面选股**示例(按市值等因子选若干只股票,定期调仓)。
- ⚠️ 作者提醒:示例里"取 10 只""只看流动性前 200"等**参数是随意的**,别当成好策略;这类练习的价值在于动手改进(换因子、加技术指标/风控)。
## 📝 完整文稿(英文 · 自动转写)
> 由本地 ASR(faster-whisper)从视频音轨转录,未人工校对,供检索/精读使用。原始讲解以上方视频为准。
Welcome to the 8th video of this algorithmic trading course. If you're new to this series, I highly recommend going back and starting with the first video of this series. So far, we've only created trading bots that trade in one specific security such as SPY. Besides that, we've learned how to add data for multiple different securities so that you could create algorithms that trade a handful of stocks or other securities. However, we've not yet learned how to create a dynamic universe of securities that changes based on the given market conditions. If for instance, you only want to trade stocks that are near 52-week high, stocks with high volatility, investing companies with solid earnings, only trade stocks around the earnings announcements or something else, you will need to set up a universe for this. That's why we'll cover how to set up a universe in this video. As always, we will first look at this on a conceptual level and thereafter we will end the video with an example bot demonstrating the learned concepts. Before we get into specifics however, let me first outline one of the main problems we look to address with Universe selection, namely survivorship bias. If you aren't using a dynamic universe of tradable instruments, you might be falling prey to this bias without even noticing it. In my opinion, survivorship bias is best understood with a brief example. Let's say you want to create a trading bot that buys and holds big US companies. To test if this strategy works, you look at all US stocks and you'll get their market cap 10 years ago. If the market cap was over 10 billion US dollars, you add the stock to your universe. To test if investing in large cap US stocks works, you now just have to test how this bot would have performed if it would have bought these stocks 10 years ago, right? Actually no. Testing this strategy like this would lead to a much better result than you would actually have achieved. The problem here is that in your selection process, you only consider those US stocks that still exist. So all the non-surviving stocks that for instance when bankrupt or merged aren't even considered. To fix this, you would have to look at all the US stocks that existed 10 years ago and then filter them by market cap. For some of you this might seem like an obvious mistake, but often things are a lot more subtle and harder to detect. The main takeaway here is to always be careful when you're confronted with the selection process. Make sure to think about if there are any invisible filters. Ask yourself whether new unknown data is introduced that at the time of the actual selection was not yet actually known. Even if you're creating a bot that only trades in one security, you can full-preat a selection bias since the choice of this security likely was not random. If for instance you decide to create a bot that only trades Apple, you have to keep in mind that Apple is one of the best performing stocks over the past decades. So here you have to ask yourself whether the performance generated from this bot can be attributed to its strategy or if it's just a result of choosing Apple as the security that is trades. Ask yourself would this strategy still work if Apple does not replicate its performance from the past 10 years? Luckily, QuantConnect's universe selection makes it quite easy to account for survivorship and selection bias, but especially if you're using custom data or a fixed universe you should definitely keep this in mind. Next up let's start looking at how you can implement a dynamic universe selection bot in QuantConnect. The idea here is that your algorithm receives a list of about 8000 daily stocks and then applies certain filters to this list and adds those that pass these filters to your universe. For this you would first start with the course filter that you can use to filter stocks by basic properties such as price, volume and the availability of data. You can then apply a fine filter to filter by other values such as fundamental company information, indicator values and more. Let's now look at the specific Python statements used to create a universe and the associated filters. To add a universe you can use the add universe method. This first takes a reference to your course filter method as the first argument and if you also want a fine filter you can optionally specify a reference to such a method as the second argument. In the next step you would want to implement your course and optionally the fine filter that is specified here. The main properties that you can filter by in the course filter are a stocks price, its dollar volume and whether or not it has fundamental data. Some tickers might have no or only limited fundamental data so if you're creating a fundamental investing bot you would want to filter out those stocks without fundamental data. Here is an example of a course filter that finds the 500 most liquid stocks that have a stock price above $10. The first line of this code sorts the security object inside of the course collection by dollar volume from highest to lowest. The next line saves the symbol for each of these stocks into a new list if the stock's price is above $10. Last but not least we will return the first 500 elements of this list. If you also implemented a fine filter this list would then be passed on to your fine filter where you could further filter out unwanted stocks. So let's look at an example of how we could implement a fine filter. As parameters a fine filter function takes a list of symbol objects that were passed on from the course filter. We can now filter this fine list for instance by looking at different fundamental stats of the associated companies. QuantConnect allows you to access hundreds of fundamental ratios, stats and other information that can be useful when analyzing stocks. For a reference table of all the supported fundamentals accessible through the QuantConnect API check out the link documentation page in the description box below. If for instance you want to find the 10 stocks out of the stocks in the fine list that have the lowest price to earnings ratio you could access this value with valuation ratios dot p e ratio. You can then sort the stocks by this ratio and return a list containing the symbols of the first 10 stocks. So after going through the just shown course and fine filters this algorithm would have a universe of 10 of the most liquid US stocks with a price over $10 and the lowest p e ratio. You can now iterate through this universe and potentially establish a position depending on the remaining implementation of your bot. You can access the securities currently in your universe with the self dot active securities. In comparison to the self dot securities array this collection only keeps track of the securities currently in your universe. If the securities removed from your active universe it will no longer be in this active securities array but it will still be accessible in the normal securities collection. This is important since you still might want to access this security for tracking purposes such as checking the fees spent or volume traded. You might remember that when adding individual securities you can change settings such as the data resolution, data mode, set leverage and more. Such things can of course also be customized when using a universe model. For this the firstly a global universe settings that you can change in the initialized method with self dot universe settings. For instance you can set the resolution for the added securities by setting self dot universe settings dot resolution to the desired resolution. Without specifying anything the default resolution is minutely. Other attributes such as leverage can be set in a similar way. Changing individual security properties is also possible. For this you need to set a security initializer which you should specify in the initialized method. Set security initializer takes a reference to a custom method as an argument. The method that you specify here will be called every time a new security is initialized which for example happens when it is added to your universe. So if you want the data normalization mode for your securities to be raw you could accomplish it like this. Inside of your custom security initializer you just set the data mode of the security that will be passed as an argument for the initialization. When a security enters or leaves the active universe a so called security changed event occurs. Like for all important events there is an event handler that is automatically called when this happens. This event handler is called onsecurities changed and its parameter is a security changed object which you can use to look at the added and removed securities. More specifically you can access the added securities with changes dot added securities and the removed ones with changes dot removed securities. One common use case of this onsecurities changed event handler is to initiate the exit process for the stocks that have left your active universe. Usually you only want to position in those securities that are currently in your universe. To close any potential positions in the removed securities you can iterate through this collection and liquid at any positions in these tickers. Before we move on to coding an example bot to demonstrate these features let's look at one last thing in the universe creation shortcuts. If you just want to use a simple dynamic universe without very specific requirements there are some shortcut helpers that QuantConnect provides. If for example you just want your universe to consist of the 50 US stocks with the highest dollar volume you can create such a universe in one line with self dot universe dot dollar volume dot top 50. For the bottom 50 you just have to replace top with bottom. Alternatively you can also filter by percentile. Here you can either pass a specific percentile or a range. This last line for instance would create a universe for stocks between the 70th and 80th dollar volume percentile. That said let me now present to you the main strategy idea behind the bot that we will implement in the remainder of this video. We base our strategy on a widely accepted and researched effect found in the markets, namely the so called size effect. Historically speaking small cap stocks have outperformed the large cap counterparts. This means that there is the potential to earn a premium without taking unpropotionally more risk when trading small cap stocks. There are many different explanations as to why this effect exists but there are two that are the most widely accepted. The first is that small cap stocks come with much worse liquidity than large cap stocks which makes them less suitable for institutional trading firms. The second reason is that smaller firms have much more space to grow. A smaller company could easily double triple or even 10 eggs its earnings within a few years. This is much much harder to do for a 100 billion dollar company. For a complete breakdown as well as links to various research papers about this size effect check out the Quantpedia entry that are linked below. We will try to implement a simple strategy that hopefully explores this size effect at least to some extent. However, we don't want our bot to invest in really bad micro cap junk stocks since these usually aren't very good investments. Instead, we want to focus on the 200 most liquid US stocks and then sort these by the market capitalization. We then invest in the 10 stocks with the lowest market cap out of these 200 stocks. For this we allocate an equal amount of our portfolio to each of these 10 securities. We rebalance the portfolio once a month. This means that every month we once again look at the most liquid US stocks sort them by market cap and if the lowest 10 have changed since last month we just our portfolio to account for this change. Note that this means that we aren't specifically investing in small cap stocks. Instead, we are investing in those stocks with the lowest market cap out of the 200 most liquid stocks ranked by the dollar trading volume. These likely still are very large cap stocks but they have lower market caps compared to other stocks on this top 200 most liquid stocks list. That said, let's now start writing some code. For this we will head over to the lab tab of the QuantConnect platform. If you haven't created your free account yet, let's link in the description box below. I highly recommend coding along. For now we won't be using the strategy builder. Instead, we just start with a blank template algorithm consisting only of the initialised and on-date methods. As always, we start by setting the start date and date and starting balance for our back test. Here, I just go with a two year time frame and a starting balance of $100,000 but feel free to try out different values. After that, I will initialise a couple of helper variables that we will need later for this algorithm. The first is Rebalance time which will help us implement the monthly rebalancing cycle. We initialise this to the earliest possible date since we want our bot to start trading right away. Then we create an empty set and save it to the active stocks variable. This will be used to keep track of the stocks in our universe. More on this later though. In the next step, we use the Add Universe method to create a dynamic universe. Here, we pass a reference to a course and fine filter which we still have to create. Before we implement these two filter methods, let me change the resolution of the securities that will be added to the universe. The default resolution is minutely but since this bot's average holding time is over a month, we won't need that lower resolution. Instead, we will use hourly resolution. This will also make back testing a lot faster. Last but not least, we create a variable called putForYourTarget which will be a list of putForYourTarget objects. With that, we are now done with the initialise method and can start implementing the course filter. As already mentioned earlier, in a course filter you can filter by price, volume and whether or not a stock has fundamental data. We can check this with dot has fundamental data which returns true if there was fundamental data available for the stock in the previous month. Note that this is not necessarily a guarantee that there will be fundamental data available now, so double checking if data is available for the specific fundamentals before using them can still be worthwhile. We will use this method to filter out the 200 most liquid stocks by dollar volume that are priced above $10. But before we implement the filtering, we first want to check if it's time to rebalance our portfolio. If it's not time to rebalance, it's pointless to even look for changes in the universe. To check if we should rebalance, we check if the current time is before the time saved in the rebalance time variable. If this is fulfilled, we return since we don't yet want to rebalance. If it's not fulfilled, we continue and set the rebalance time variable to one month from now so that we only rebalance once a month. After that, we can implement the filtering. After this, we first sort the securities in the course collection by their dollar volume. We order them in reverse since we're interested in those with the highest dollar volume. Thereafter, we return a list of symbols for the first 200 securities that are priced above $10 and that have fundamental data available. Note that the choice of 200 securities is quite arbitrary. You could easily bump this up to 300, 500 or even more. Just note that if you filter out too few securities, your algorithm might take a lot longer to backtest since it has to perform a lot of work. Generally, you should try to keep your active universe comparatively small. Creating a universe with thousands of securities will lead to a lot of performance problems and might require you to upgrade your server nodes RAM. Now the collection of securities that were returned from the course filter gets passed on to the fine filter which we will implement next. In here, we simply want to sort the equities by their market cap from lowest to highest. Out of these 200 most liquid stocks, we will then add the 10 stocks with the lowest market cap to our universe. Note that in the return statement, we filter out those stocks with the market cap of zero. We do this because some securities might not have an available value for their market cap. If that's the case, zero is the default replacement value but since that's obviously not the actual market cap, we don't want to include these stocks in our universe. Before we move on to the on data method, we will now implement the onsecurities changed event handler. This method is automatically called every time the universe of our algorithm gets updated. So in our case, this should happen once a month. The first thing we want to do in the onsecurities changed event handler is liquid at any open positions in those symbols that have been removed from our universe. To accomplish this, we can iterate through the changes.removesecurities collection. For each of these elements, we then call the liquidate function. Besides that, we want to remove them from the active stocks variable that keeps track of our universe. After that, we now iterate over the added securities which we can access through changes.addedsecurities. We then add each of these symbols to the active stocks variable. We should not try to send actual trade orders for the just added securities right in the onsecurities changed method. The reason for this is that the data for these securities first has to be added to the algorithm. This usually takes at least one iteration after the securities were added. You might be asking yourself why we are using our own self.active stocks variable and not just the self.active securities array since that should keep track of the active securities as well, right? The problem with the active securities array is that the removed security objects are not immediately removed from this collection. After liquidation, they are added to a pending removal list where they usually stay for at least one iteration. If we then open a position in that symbol before it is fully removed, it won't be removed from the active securities array. Therefore, we would somehow have to keep track whether this iteration has passed or not before opening a new position. Taking track of this would require more effort than just creating our own active stocks set. Next up, there's one last thing we do in the onsecurities changed method, namely create a list of portfolio target objects. For this algorithm, we'll go with an equal allocation to each of the equities in our universe. So we create a portfolio target object with a weight of 1 divided by the length of the active stock set which should be 10% for each symbol. Now we can move on to the on data method where we will use this list to set the portfolio holdings. The first thing we do in the on data method is check whether or not the universe has changed. We can do this by checking whether or not the portfolio target list is empty. If it's empty, nothing has changed and there's no reason to rebalance so we just return. Otherwise we want to rebalance our portfolio. However, we only want to do so if the algo has data for all the 10 securities in our active universe. This we can check by iterating through our active stock set and checking if each symbol is in the data parameter of this method. If this is not the case for one of the securities, we return since the data has not yet been added correctly. So in the remainder of this method, we now know that the data of all stocks has been added correctly and it's time to rebalance our portfolio. Now what that's left to do is pass the portfolio targets list to the set holdings method which will take care of the rebalancing for us. As a last thing, we now have to set the portfolio targets list to an empty list so that we don't try to rebalance the portfolio before the universe changes again. With that, we are now done and have successfully implemented our strategy. So let's now build and back to this bot. Depending on the specified time frame, this might take a few minutes. So let me just fast forward to when the back test is completed. As always, when the back test is finished up, you can see a performance report like this one. The first thing that we usually look at is the equity chart. Here we can see that the strategy seemed to perform quite well during 2019 but it had a big dip in the early 2020 crash. If we look at the benchmark chart and compare it to the equity chart, we can see that the two definite seem to be correlated. This should not be very surprising since this bot invested in 10 of the most liquid US stocks which likely are also holdings of the S&P 500 index. If we would increase the number of stocks, this correlation likely would increase even more. But be aware that over other time frames, the performance might deviate significantly. Note that I would not recommend actually funding such a bot with real money. One problem is that this bot still has quite a few totally arbitrary parameters. There's no good reason why we are investing in 10 instead of 12, 20, 30 or even 50 equities. Another example of an arbitrary parameter is that our course filter only looks at the 200 most liquid stocks. This strategy does not seem to be a very good choice if your goal is to explore the size factor. If your sum idea is about how to potentially improve this strategy or if you just want to play around with the code and take a closer look at this performance overview, you can clone the code from this video by using the link in the description box below. One possible extension is that you could look at other fundamentals instead of only looking at the market cap. Furthermore, you could also incorporate some technical indicators, risk management techniques and other possible improvements. Practicing by trying to code things like this yourself is probably the best way to get better at algorithmic trading and coding. I also have another video in which I implemented a more advanced fundamental investing bot. If you want to check that video out, there's a link in the description box below. Just know that there are some concepts in that video that have not yet been covered in this video series. One such example is the algorithm framework which is something we will cover in detail later in this series. So if you don't understand everything in that video yet, don't worry. In the next video, we will probably look at how to add external data such as your own custom data or data scraped from somewhere to QuantConnect. That said, I hope you enjoyed this video and learned a bunch of new things. If you are enjoying this series, make sure to smash the like button, subscribe and turn on the notification bell for more content like this. Thanks for watching.