使用Twitter学术研究V2 API批量获取推文遇KeyError问题求助
问题
使用Twitter Academic Research V2 API从用户列表批量抓取推文并存储到DataFrame时,单用户抓取正常,但多用户列表触发KeyError: 'users'错误。代码及报错信息如下:
import tweepy from twitter_authentication import bearer_token import time import pandas as pd import time client = tweepy.Client(bearer_token, wait_on_rate_limit=True) # list of twitter users csu = ["Markus_Soeder", "DoroBaer", "andreasscheuer"] csu_tweets = [] for politician in csu: for response in tweepy.Paginator(client.search_all_tweets, query = f'from:{politician} -is:retweet lang:de', user_fields = ['username', 'public_metrics', 'description', 'location'], tweet_fields = ['created_at', 'geo', 'public_metrics', 'text'], expansions = 'author_id', start_time = '2022-12-01T00:00:00Z', end_time = '2022-12-03T00:00:00Z'): time.sleep(1) csu_tweets.append(response) end = time.time() print(f"Scraping of {csu} needed {(end - start)/60} minutes.") result = [] user_dict = {} # Loop through each response object for response in csu_tweets: # Take all of the users, and put them into a dictionary of dictionaries with the info we want to keep for user in response.includes['users']: user_dict[user.id] = {'username': user.username, 'followers': user.public_metrics['followers_count'], 'tweets': user.public_metrics['tweet_count'], 'description': user.description, 'location': user.location } for tweet in response.data: # For each tweet, find the author's information author_info = user_dict[tweet.author_id] # Put all of the information we want to keep in a single dictionary for each tweet result.append({'author_id': tweet.author_id, 'username': author_info['username'], 'author_followers': author_info['followers'], 'author_tweets': author_info['tweets'], 'author_description': author_info['description'], 'author_location': author_info['location'], 'text': tweet.text, 'created_at': tweet.created_at, 'quote_count': tweet.public_metrics['quote_count'], 'retweets': tweet.public_metrics['retweet_count'], 'replies': tweet.public_metrics['reply_count'], 'likes': tweet.public_metrics['like_count'], }) # Change this list of dictionaries into a dataframe df = pd.DataFrame(result)
报错回溯:
--------------------------------------------------------------------------- KeyError Traceback (most recent call last) ~\AppData\Local\Temp/ipykernel_25716/2249018491.py in <module> 4 for response in csu_tweets: 5 # Take all of the users, and put them into a dictionary of dictionaries with the info we want to keep ----> 6 for user in response.includes['users']: 7 user_dict[user.id] = {'username': user.username, 8 'followers': user.public_metrics['followers_count'], KeyError: 'users'
单用户抓取时无此错误,请问问题原因是什么?
原因分析
出现KeyError: 'users'的核心原因是列表中存在某个用户在指定时间范围内没有符合条件的推文:
- 当用户无符合条件的推文时,API返回的响应对象中既没有
tweet数据(response.data为空),也不会生成includes['users']字段 - 单用户抓取时你选的用户刚好有推文,所以响应正常返回
includes['users'];但多用户批量抓取时,至少有一个用户在2022-12-01至2022-12-03期间没有发布非转发的德语推文,代码直接通过response.includes['users']取值时就触发了KeyError
解决方案
修改代码,在处理用户数据和推文数据前增加空值判断:
result = [] user_dict = {} # Loop through each response object for response in csu_tweets: # 先判断是否存在用户数据,避免KeyError if 'users' in response.includes: for user in response.includes['users']: user_dict[user.id] = {'username': user.username, 'followers': user.public_metrics['followers_count'], 'tweets': user.public_metrics['tweet_count'], 'description': user.description, 'location': user.location } # 判断是否存在推文数据,避免response.data为空时报错 if response.data is not None: for tweet in response.data: author_info = user_dict[tweet.author_id] result.append({'author_id': tweet.author_id, 'username': author_info['username'], 'author_followers': author_info['followers'], 'author_tweets': author_info['tweets'], 'author_description': author_info['description'], 'author_location': author_info['location'], 'text': tweet.text, 'created_at': tweet.created_at, 'quote_count': tweet.public_metrics['quote_count'], 'retweets': tweet.public_metrics['retweet_count'], 'replies': tweet.public_metrics['reply_count'], 'likes': tweet.public_metrics['like_count'], })
也可以在抓取阶段增加日志,定位无数据的用户:
for politician in csu: print(f"开始抓取用户 {politician} 的推文...") response_count = 0 for response in tweepy.Paginator(client.search_all_tweets, query = f'from:{politician} -is:retweet lang:de', user_fields = ['username', 'public_metrics', 'description', 'location'], tweet_fields = ['created_at', 'geo', 'public_metrics', 'text'], expansions = 'author_id', start_time = '2022-12-01T00:00:00Z', end_time = '2022-12-03T00:00:00Z'): response_count +=1 time.sleep(1) csu_tweets.append(response) if response_count ==0: print(f"警告:用户 {politician} 在指定时间范围内无符合条件的推文")
内容的提问来源于stack exchange,提问作者Maxl Gemeinderat
相关产品推荐
相关产品推荐

