如何使用Snscrape仅获取Twitter平台的原生推文?
筛选Snscrape采集的Twitter原生推文
要只获取原生Twitter推文,排除Instagram、Foursquare等第三方平台同步的内容,直接通过tweet.sourceLabel字段做过滤即可。原生推文的来源标签必然包含"Twitter"关键字(比如「Twitter Web App」「Twitter for iPhone」这类官方客户端标识),只需在采集时加个判断条件。
修改后的代码如下:
import snscrape.modules.twitter as sntwitter keyword = '(COVID OR Corona Vírus)' maxTweets = 30 tweets = [] for i,tweet in enumerate(sntwitter.TwitterSearchScraper(keyword + ' since:2020-01-01 lang:pt').get_items()) : if i >= maxTweets : break # 仅保留来源为Twitter官方渠道的推文 if "Twitter" in tweet.sourceLabel: tweets.append([tweet.date, tweet.id, tweet.content, tweet.user.username, tweet.sourceLabel])
补充说明
- 用
"Twitter" in tweet.sourceLabel的判断方式比较灵活,能覆盖所有Twitter官方客户端的推文,避免遗漏新出的客户端标识。 - 如果需要更精准的筛选,也可以定义一个官方来源列表做精确匹配,比如:
official_sources = ["Twitter Web App", "Twitter for iPhone", "Twitter for Android", "Twitter Lite"] if tweet.sourceLabel in official_sources: tweets.append([tweet.date, tweet.id, tweet.content, tweet.user.username, tweet.sourceLabel])
不过这种方式需要你定期更新列表,适配Twitter客户端名称的变化。
内容的提问来源于stack exchange,提问作者fygeng go
相关产品推荐
相关产品推荐

