You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过Tweepy筛选仅提及国家名称的美国地区Twitter推文?

问题:如何筛选仅提及国家名称的美国地区Twitter推文?

我想仅收集提及国家名称的美国地区Twitter推文。已知Twitter API的context_annotations字段可识别推文是否提及国家,其中国家对应的domain编号为160。目前我已经能获取美国地区的推文,但不会用Tweepy编写筛选逻辑,当前代码如下:

client = tweepy.Client(bearer_token=bearer_token)

# Specify Query
query = '"favorite country" place_country:US'                   
start_time = '2022-03-05T00:00:00Z' 
end_time = '2022-03-11T00:00:00Z' 

tweets = client.search_all_tweets(query=query, tweet_fields=['context_annotations', 'created_at', 'geo'], 
                                  
                                  place_fields = ['place_type','geo'], expansions='geo.place_id',
                                  start_time=start_time,
                                  end_time=end_time, max_results=10000)

# Prepare to write to csv file
f = open('tweetSheet.csv','w')
writer = csv.writer(f)

# Write to csv file
for tweet in tweets.data:
    print(tweet.text)
    print(tweet.created_at)
    writer.writerow(['0', tweet.id, tweet.created_at, tweet.text])

# Close csv file
f.close()

解决方案

要实现筛选仅提及国家名称的推文,核心是检查每条推文的context_annotations字段,判断是否包含domain.id为160的标注(对应国家类实体)。以下是修改后的完整代码:

import tweepy
import csv

client = tweepy.Client(bearer_token=bearer_token)

# 指定查询条件
query = '"favorite country" place_country:US'                   
start_time = '2022-03-05T00:00:00Z' 
end_time = '2022-03-11T00:00:00Z' 

tweets = client.search_all_tweets(query=query, tweet_fields=['context_annotations', 'created_at', 'geo'], 
                                  place_fields=['place_type','geo'], expansions='geo.place_id',
                                  start_time=start_time,
                                  end_time=end_time, max_results=10000)

# 准备写入CSV文件
with open('tweetSheet.csv', 'w', newline='') as f:
    writer = csv.writer(f)
    # 写入表头(可选)
    writer.writerow(['序号', '推文ID', '创建时间', '推文内容'])
    
    count = 1
    if tweets.data:
        for tweet in tweets.data:
            # 检查是否存在国家相关的标注
            has_country_reference = False
            if tweet.context_annotations:
                for annotation in tweet.context_annotations:
                    if annotation.domain.id == 160:
                        has_country_reference = True
                        break
            # 仅保留符合条件的推文
            if has_country_reference:
                print(tweet.text)
                print(tweet.created_at)
                writer.writerow([count, tweet.id, tweet.created_at, tweet.text])
                count += 1

关键修改说明

  • 新增筛选逻辑:遍历每条推文的context_annotations,判断是否存在domain.id为160的标注,仅保留符合条件的推文
  • 使用with语句管理文件,避免手动关闭时可能出现的资源泄漏问题
  • 新增表头和序号,让CSV文件结构更清晰
  • 增加空值判断:先检查tweets.data是否存在,避免无数据时触发报错

内容的提问来源于stack exchange,提问作者Rashid Abramov

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 14:55:19