# 09 · Twitter Trading Bot(自定义数据 · 推文情绪) - **系列**:Full Algorithmic Trading Using Python - **频道**:TradeOptionsWithMe | **本集**:Twitter Trading Bot(自定义数据 · 推文情绪) - **时长**:26 分 52 秒 | **原视频**:https://youtu.be/X7XwkHsE-4Y - **本地视频**:[videos/09-twitter-trading-bot.mp4](videos/09-twitter-trading-bot.mp4) ## 🎯 本集要点(中文导读) 本集讲 **自定义数据(Custom Data)**,示例:用**马斯克推文情绪**交易特斯拉。 - **为什么强大**:可引入社交媒体、天气、SEC 文件、新闻等任意外部数据。 - **添加**:`self.AddData(数据类型, ticker, resolution)`(在 `Initialize()` 中;与 `AddEquity` 类似,只是更通用)。 - **自定义数据类**:新建类继承 `PythonData`,**重写**两个方法: - `GetSource()`:返回 `SubscriptionDataSource`(数据源 URL / 本地文件 / 等;传输方式有 RemoteFile、LocalFile 等)。 - `Reader()`:解析你的数据格式(自定义,需手动定义结构)。 - **本集示例 bot**:读取马斯克推文 → 用 `NLTK` 的 **VADER 情感强度分析器**打分(polarity)→ 正/负情绪驱动做多/做空特斯拉;用 `self.Plot` 把情绪分数可视化。 - ⚠️ 局限:示例用的是**历史推文 CSV**;换时间框/实盘需要实时数据源并调整 Reader;且通用情感分析器对**推文语料**未优化,还有很大改进空间。 ## 📝 完整文稿(英文 · 自动转写) > 由本地 ASR(faster-whisper)从视频音轨转录,未人工校对,供检索/精读使用。原始讲解以上方视频为准。 Hello and welcome to the ninth video of this algorithmic trading video course. As always I recommend starting with the first video if you're new to this course since you otherwise might not fully understand some of the covered topics. In this video we will cover how you can add custom data to QuantConnect and use it for the decision making of your algorithms. The ability to add custom data is very powerful since it opens so many doors. You could for instance add social media data, weather data, SEC filings, news data or whatever data you can find. Even though QuantConnect already offers a huge variety of different data, adding your own data gives you even more possibilities. To demonstrate this feature I will create a trading bot that trades Tesla stock based on simple sentiment analysis of Elon Musk's tweets. But before we get into coding this trading bot let me start by presenting the whole spectrum of options that QuantConnect provides when it comes to adding custom data. To add custom data you can use the add data helper method. You usually would do this in the initialised method of your algorithm since you only want to do this once in the beginning. Add data takes three arguments. The first one is the type of data that you want to add. Usually you would have to create a custom class for this and use this class as the type. You will see a concrete example of this in a few minutes. The second argument is a string which represents the ticker for that data. And the third argument is a resolution which specifies how often the custom data should be pulled. If for instance you specify a minutely resolution the algorithm will check for the new data once a minute. You might be able to see that the add data method is very similar to the add equity method. In fact they both fundamentally do the same thing. The add data method is just a little more general. Like I just said the first argument of the add data method is the type of the data that you want to add. For this you should create your own custom class. This class should extend the Python data class and it should override the reader and get source methods of that class. The get source method is used to get the source of that data and the reader method is used to read your custom data. Since the data is custom and not standardized you need to specify the structure of that data in the reader method. Let me present how to implement these methods with an example of custom weather data. For this let me start with the get source method which is actually very simple to implement. All you have to do is return a subscription data source object. To create such an object you need to pass two arguments to its constructor. The first is the actual source of your data and the second one is a type of subscription transport medium. As the source you usually would specify a URL such as a dropbox URL where you are uploaded your custom data as a CSV file. For the subscription transport medium you have four options. Probably the most common one is remote file which means that the algorithm expects one file with all the data in your remote location such as your dropbox. Alternatively you could also use local file if the data is accessible locally. There is a nice remote and local file that also are options REST and Streaming. Rest is used if each line of the custom data should be pulled individually and comes from a REST call. This might be the case if you have a subscription to a data vendor that provides you with custom data in real time. For our example we have custom weather data in a dropbox. The URL to this data is specified as the source and for the subscription transport medium would choose remote file. In some cases you might want the get source method to return different data depending on whether you are back testing your algorithm or actually live trading it. If for example you have a real time data subscription you can't use that for back testing and might instead want to use a CSV file for the historical data when you are back testing. For this you can use the isLive flag which is true when you are live trading and falls if you are just back testing. Now the data specified in the CSV file will get sent to the reader method line by line. The goal of the reader method is now to standardize this data so that your algorithm understands what it looks like and it can actually use it. Before I show you how you can implement this reader method let me first quickly show you what the data actually looks like. As you can see this data has 11 entries divided into 4 columns. Column A specifies the data for each row. The first four digits specify the year, the next to the month and the last to specify the day. Column B represents the max temperature in degrees tells us for that day. Column B represents the daily average and column C is the minimum temperature. I know this data does not look very interesting and I actually don't even know what location this weather data was taken from. However in my opinion this example is a great and simple way to understand how to use custom data in QuantConnect. Theoretically you could use weather data such as this one for trading in the agricultural space since this is heavily weather dependent. Here it might also be useful to take precipitation, the weather forecast and other variables into account. That said you now hopefully have a good understanding of how this custom weather data is structured. Remember that this data now gets fed to the reader method line by line. The first thing we want to do in the reader method is check whether a meaningful line of data was passed to it. We do this by checking if the line of data is not empty and that the first value is a digit. This ensures that we don't try to use the data from the header index and don't continue if there is no data left. So if this is not fulfilled we just return an empty object. Otherwise, we split the line at its commas and save the generated list to the data variable. Remember that we can split at commas since the file is a CSV file which stands for comma separated values. Then we create an object of the weather class. Since this class extends the Python data class there are three important properties that we should set, namely the symbol, value and time property. For the symbol property we should always use config.symbol. For the time we need to correctly format the values from the time column in the CSV file. We can do so with this date time helper method in which you can specify the format of the date and time. Here it is very important to understand that data in QuantConnect is passed on at its end time and usually the data and time in a custom file is the start time. That's why it's essential to add a time delta to the start time. If you don't do this your algorithm might access this data before it actually would be known. This is called lookahead bias and it can lead to very unrealistic results. That's why we add a time delta of 20 hours to the time of the weather data. If we would not do this the algorithm could use the weather data for a given day at the beginning of the day before it actually would have been known. Next up we set the value attribute of the weather object. For this we will take the daily average temperature which was saved into the second column of the weather data file. Note that the value attribute always is a decimal object. Besides that we also create two more properties for the max and min temperature that were saved in the first and third column. Last but not least we will turn this weather object. Since custom data can sometimes lead to some unexpected errors it usually is a good idea to add some exception handling to the reader method. I will show you a simple example of this when we get to the coding part of this video. You can now access this custom weather data like you would access other price data. Inside of the on data method you can index the data slice object with the symbol of the custom data and then access values either with the get property helper or the dot value attribute. Note that if this still seems a little overwhelming to you don't worry too much. We will go over another example step by step when we implement the Twitter trading bot in a few minutes. Besides adding custom data that you want to access on the go line by line, QuantConnect also allows you to add static data. If for instance you just want to add a file containing a list of tradable symbols for your universe or an AI training model file you can accomplish this with the self dot download. As an argument you can for example once again just pass a drop box download URL. If the file is a CSV file you can then use pandas to convert this file to pandas data frame. Note that if you want to use custom data from a data vendor such as Quantal, Tingo or Intriniou QuantConnect does offer direct helper methods for this. Here you only have to provide your API token and you are good to go. For the specifics on how to do this I recommend checking out the documentation page. That said let me now move on to the Elon Musk Twitter Tesla trading bot that we will code in this video. As always I will first give a rough outline as to how this strategy works on a conceptual level before starting to actually code it. The idea behind this bot is that we will use the tweets from Elon Musk to perform sentiment analysis and based on the results of this analysis we will then make a trade decision for Tesla stock. For this we will only consider those tweets that mention Tesla in some way. For these tweets we then use a sentiment analyzer to find out whether this tweet has a positive or negative connotation to it. If it is positive we establish a long Tesla position for that day. If it is negative we short Tesla for that day. If it is mostly neutral or slightly positive or slightly negative we don't do anything. For the sentiment analysis we will be using Python's natural language toolkit NLTK package. Note that if you aren't familiar with this module don't worry we will just be using a pre-trained sentiment and density analyzer but if you are familiar with it you could try using your own custom models for better analysis. Since QuantConnect does not directly offer tweets as data we will have to find a dataset containing Elon Musk's tweets and then import this custom dataset so that we can use it for back testing this strategy. Luckily I found a dataset containing Elon Musk's tweet from 2012 until 2017 for free on Kaggle.com. This means we can use this dataset to back test the strategy at least from 2012 until 2017. However, before we actually import this dataset I firstly want to perform some pre-processing on it since we otherwise might run into some unwanted problems. For the pre-processing I will import this dataset into a Jupyter notebook where we will use Pandas to import it into a data frame. Note that if you want to code along when I code this trading bot you won't have to do these steps since you can access the pre-processed file through the link in the description box below. The unprocessed dataset contains 5 columns including values such as retweets, the row number and the count. But we are actually only interested in two of these columns namely the tweet itself and its date and time. So firstly we want to remove the three unwanted columns. Next up I want to change the order since we want the most recent tweets to be at the bottom and the least recent ones at the top. Besides that we will remove all URLs from the tweets since they can lead to some unexpected formatting problems and that don't really add any value to our sentiment analysis. After that I exported the pre-processed data to a new csv file and uploaded it to Dropbox. Through this Dropbox URL you can now import the tweet data into QuantConnect where we will use it to test our strategy. If you have not yet created your free QuantConnect account you can do so using the link in the description box below. In QuantConnect we will head over to the lab tab where we will create a new algorithm. As always so far we won't be using the strategy builder tool. Instead we start with the blank template containing only the initialised and on data methods. The first thing we do in the initialised method is set the back testing time frame to the time covered by the mask tweet dataset. So we set the start time to November 2012 and the end time to the beginning of 2017. The starting balance I will just leave at $100,000. Before we move on let me input sentiment and density analyser from NLTK sentiment. We will use this sentiment analyser later to analyse the sentiment of the tweets but more on that later. Thereafter we want to add the data that this algorithm will use. Firstly we add tesla data which we can do with the usual add equity helper method. Since the tweets can be sent pretty much at any time of the day and we have specific time stems for them we will set the resolution for the tesla data to minutely data. Then we want to add our custom data. We can do so with the add data method which takes three arguments. The first is the type that this data will have. I will pass mask tweet for this which is a class that we will implement in a few seconds. As the second argument you need to specify a ticker symbol that you can use to access this data with. For this you can choose whatever name you want to. Just try to pick one that isn't already occupied by some stock or other asset. I will just go with mask tweets. Last but not least you need to specify resolution. This resolution will determine how often your algorithm follows your custom data source for potential new data. Here I will also go with minutely data. Since the tweet data actually has timestamps with up to second accuracy we could also go with second data but this would slow back testing down and in my opinion isn't that important for now. Note that just like for equities and other assets we can use dot symbol to save the symbol for this data to a variable. We will do this so that we can later use this symbol to request the data. For this you could also use the specified ticker but that can come with unwanted ambiguity problems. For details on this go watch the third video of this series. Next up let's start implementing the mask tweet custom data class. Like I mentioned in the theory part of this video this class has to extend the python data class and it has to implement the get source and reader methods. Let's start with the get source method but before we do so make sure to note that this class no longer is part of the QC algorithm class so you won't be able to access the usual helper methods such as self dot debug in here. Before we move on let me save an instance of the NLTK sentiment intensity analyzer into a class variable named self dot sia. In the reader method we will use this for the sentiment analysis. That said let's implement the get source method now. Here we simply want to set a drop box URL as the data source. Note that here it is very important to have dl equals 1 at the end of the URL. dl stands for download and we want the algorithm to download this data. If you would want to view this dsv file online you could do so by setting dl equal to 0 and opening the URL in your browser. However for the algorithm you will always have to set dl to 1 while using drop box URLs. Then all we have to do is return a subscription data source object. For this object we will just use the saved URL as the source and specify this to be a remote file. If you would want to live trade such a bot you would have to connect this bot with some kind of twitter API that sends you real time tweet updates. In that case you would use the REST or streaming subscription transport medium option. To differentiate between back testing and live trading mode you could then use the is live parameter of this method. But since we will just use this for back testing we won't have to worry about this for now. Next up let's move on to implementing the reader method which handles bringing the data into the right format so that your algorithm can actually use it in a meaningful way. Remember the reader method receives data from your remote file on a line by line basis. So the first thing we need to check in here is whether the line actually contains relevant data. If for instance we receive an empty line or the first index row we won't do anything which is why we return none. Otherwise we split the line by commas since the data comes from a comma separated file. The data variable now is a list containing two elements. The first element is the date and time of the tweet and the second element is the tweets content itself. Next we create an instance of the mask tweet class and save it to a new variable named tweet. The next step is to set the symbol time and value of this mask tweet object. So the symbol we can just use config.symbol. Config is one of the parameters of the reader method. As for the time we have to specify the format of the time column in our custom data. The format is year, month and day separated by a dash and hour, minute and second separated by a colon. Note that we add a time delta of one minute to this time since this will be the end time of this data point. Since tweets should be available pretty much right away we could also use a time delta of one second but since we are using minutely data anyways this likely won't make a big difference. Be aware that when using price data or other data covering a certain period to always correctly adjust the time to the end time of the data. Otherwise the results can be very unrealistic. After setting the time we convert the tweet which is saved into the second place of the data list to lowercase and save it to the content variable. Now we want to perform some simple sentiment analysis on the content of each tweet. However, we are only interested in those tweets that reference Tesla. Since the tweet was converted to lowercase we can check whether it contains the ticker symbol of Tesla or the written out version. If it does we use our sentiment intensity analyzer to calculate a polarity score for the tweet. This polarity score will return a dictionary with scores for negative, neutral, positive and compounded sentiment. This tweet says that Tesla just had its first profitable quarter thanks to awesome customers and hard work by a super dedicated team. If we would use the polarity scores method on this tweet it will return the following dictionary. As you can see according to the sentiment intensity analyzer this is the most positive tweet which does make sense. Note that the compounded sentiment is not just an average of the others. An example of a tweet with a negative sentiment would be this one. Not all good news Virginia DMV commissioner just denies Tesla a dealer license. This is interpreted to be a neutral to negative statement. But note that often the sentiment analysis isn't as accurate as in these two cases. What we want to do is use the sentiment to either establish a long or short Tesla position. We will only use the compounded score for this though. That's what we save it as the value of the tweet data object. In the case that Tesla was not referenced in the tweet we just assign it a score of zero which would be equivalent to a completely neutral tweet. Next up I also want the content of the tweet to be accessible in the data bar. To accomplish this we can simply create a property for the tweet object with whatever name you want to and save the content to it. We will use tweet as the name for this property. Earlier in the video I mentioned that sometimes that can occur some unexpected problems inside of the reader method since we are dealing with custom data. That's why it makes sense to add some form of exception handling here. So let's put a try statement around this code and if a value error is thrown we simply return none. After that all that's left to do is return the tweet object and then we are done with the reader method. Now we're almost done. We just have to quickly implement the on data method and the exit logic. The first thing we do in on data is check whether the data slice object contains any data for the musk tweet. We check this with a symbol object of that data. If that's the case we can save the score of the current tweet with data.value. To access the content of the tweet we can use .tweet which is the name that we gave this property in the reader method of the musk tweet class. For this algorithm we will establish a long position if the compounded sentiment score is above 0.5 and we establish a short position if the score is below negative 0.5. We can check this with a simple if-l-f condition and use set holdings to go long or short. For demonstration purposes I will also log the score and content of the tweets that meet one of these thresholds. Now all that's left to do is add a method that liquidates any open position before the market closes since we don't want to hold Tesla overnight based on one tweet. For this we can schedule an exit positions method in the initialized method. We schedule it to be cooled 15 minutes before the close on every day that Tesla trades. In the exit positions method we then simply have to cool self.liquidate to close all open positions. With that we are now done implementing this Twitter sentiment analysis trading bot. So let's now build and back test it. As you can see this strategy seemed to perform quite ok over this time frame. However, here it is important to note that it isn't a very long time frame and Tesla had a huge bull run over the past few years. If we scroll down to the stats overview we can see that this bot made 165 trades over this time frame which is a solid amount. If we go to the logs tab we can see an overview of the scores and content of those tweets that the algorithm used to actually base trades on. Some of them make sense such as the tweet on the 4th of December 2012 that says that Musk is happy to report that Tesla was narrowly cash flow positive and expects continued improvement. This had a polarity score of almost 0.9 and led to a long position. However, some of the other scores and tweets aren't as clear. If you want to look at these results in detail you can use the link in the description box below to clone this algorithm. Note that to test this strategy over another time frame or live traded you would need to find a data source that gives you access to Elon Musk tweets over the desired time frame or in real time. This data will likely be structured slightly differently than the data from this video. That means you would probably have to adjust the reader method to some extent to use that data. Also note that the sentiment analysis here is very simple and the pre-trained sentiment intensity analyzer likely isn't optimized for Twitter text data since this data usually has a different structure than other text data. So there certainly are many areas to improve upon. That said, I hope that this video gave you a good overview of how to import custom data into QuantConnect. I also hope that it shows you why this is a very powerful feature since it gives you so many more options as to what you can do. For some other examples of algorithms that use custom data check out the algos reference in the documentation page. With that being said, if you're enjoying this series make sure to smash the like button, subscribe and turn on the notification bell. Thanks for watching.