You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Twitter学术研究V2 API批量获取推文遇KeyError问题求助

问题

使用Twitter Academic Research V2 API从用户列表批量抓取推文并存储到DataFrame时,单用户抓取正常,但多用户列表触发KeyError: 'users'错误。代码及报错信息如下:

import tweepy
from twitter_authentication import bearer_token
import time
import pandas as pd
import time
client = tweepy.Client(bearer_token, wait_on_rate_limit=True)

# list of twitter users
csu = ["Markus_Soeder", "DoroBaer", "andreasscheuer"]

csu_tweets = []
 
for politician in csu:
    for response in tweepy.Paginator(client.search_all_tweets, 
                                     query = f'from:{politician} -is:retweet lang:de',
                                     user_fields = ['username', 'public_metrics', 'description', 'location'],
                                     tweet_fields = ['created_at', 'geo', 'public_metrics', 'text'],
                                     expansions = 'author_id',
                                     start_time = '2022-12-01T00:00:00Z',
                                     end_time = '2022-12-03T00:00:00Z'):
        time.sleep(1)
        csu_tweets.append(response)

end = time.time()
print(f"Scraping of {csu} needed {(end - start)/60} minutes.")

result = []
user_dict = {}
# Loop through each response object
for response in csu_tweets:
    # Take all of the users, and put them into a dictionary of dictionaries with the info we want to keep
    for user in response.includes['users']:
        user_dict[user.id] = {'username': user.username, 
                              'followers': user.public_metrics['followers_count'],
                              'tweets': user.public_metrics['tweet_count'],
                              'description': user.description,
                              'location': user.location
                             }
    for tweet in response.data:
        # For each tweet, find the author's information
        author_info = user_dict[tweet.author_id]
        # Put all of the information we want to keep in a single dictionary for each tweet
        result.append({'author_id': tweet.author_id, 
                       'username': author_info['username'],
                       'author_followers': author_info['followers'],
                       'author_tweets': author_info['tweets'],
                       'author_description': author_info['description'],
                       'author_location': author_info['location'],
                       'text': tweet.text,
                       'created_at': tweet.created_at,
                       'quote_count': tweet.public_metrics['quote_count'],
                       'retweets': tweet.public_metrics['retweet_count'],
                       'replies': tweet.public_metrics['reply_count'],
                       'likes': tweet.public_metrics['like_count'],
                      })

# Change this list of dictionaries into a dataframe
df = pd.DataFrame(result)

报错回溯:

---------------------------------------------------------------------------
KeyError                                  Traceback (most recent call last)
~\AppData\Local\Temp/ipykernel_25716/2249018491.py in <module>
      4 for response in csu_tweets:
      5     # Take all of the users, and put them into a dictionary of dictionaries with the info we want to keep
----> 6     for user in response.includes['users']:
      7         user_dict[user.id] = {'username': user.username, 
      8                               'followers': user.public_metrics['followers_count'],

KeyError: 'users'

单用户抓取时无此错误,请问问题原因是什么?


原因分析

出现KeyError: 'users'的核心原因是列表中存在某个用户在指定时间范围内没有符合条件的推文:

  • 当用户无符合条件的推文时,API返回的响应对象中既没有tweet数据(response.data为空),也不会生成includes['users']字段
  • 单用户抓取时你选的用户刚好有推文,所以响应正常返回includes['users'];但多用户批量抓取时,至少有一个用户在2022-12-01至2022-12-03期间没有发布非转发的德语推文,代码直接通过response.includes['users']取值时就触发了KeyError

解决方案

修改代码,在处理用户数据和推文数据前增加空值判断:

result = []
user_dict = {}
# Loop through each response object
for response in csu_tweets:
    # 先判断是否存在用户数据,避免KeyError
    if 'users' in response.includes:
        for user in response.includes['users']:
            user_dict[user.id] = {'username': user.username, 
                                  'followers': user.public_metrics['followers_count'],
                                  'tweets': user.public_metrics['tweet_count'],
                                  'description': user.description,
                                  'location': user.location
                                 }
    # 判断是否存在推文数据,避免response.data为空时报错
    if response.data is not None:
        for tweet in response.data:
            author_info = user_dict[tweet.author_id]
            result.append({'author_id': tweet.author_id, 
                           'username': author_info['username'],
                           'author_followers': author_info['followers'],
                           'author_tweets': author_info['tweets'],
                           'author_description': author_info['description'],
                           'author_location': author_info['location'],
                           'text': tweet.text,
                           'created_at': tweet.created_at,
                           'quote_count': tweet.public_metrics['quote_count'],
                           'retweets': tweet.public_metrics['retweet_count'],
                           'replies': tweet.public_metrics['reply_count'],
                           'likes': tweet.public_metrics['like_count'],
                          })

也可以在抓取阶段增加日志,定位无数据的用户:

for politician in csu:
    print(f"开始抓取用户 {politician} 的推文...")
    response_count = 0
    for response in tweepy.Paginator(client.search_all_tweets, 
                                     query = f'from:{politician} -is:retweet lang:de',
                                     user_fields = ['username', 'public_metrics', 'description', 'location'],
                                     tweet_fields = ['created_at', 'geo', 'public_metrics', 'text'],
                                     expansions = 'author_id',
                                     start_time = '2022-12-01T00:00:00Z',
                                     end_time = '2022-12-03T00:00:00Z'):
        response_count +=1
        time.sleep(1)
        csu_tweets.append(response)
    if response_count ==0:
        print(f"警告:用户 {politician} 在指定时间范围内无符合条件的推文")

内容的提问来源于stack exchange,提问作者Maxl Gemeinderat

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 22:25:48