使用Tweepy调用Twitter API v2提取地点字段遇'places'错误求助
问题描述
尝试用Tweepy Paginator提取推文的地点字段时,Python报出和places相关的错误,但用同样方式处理用户字段完全正常,相关代码如下:
import tweepy import time import pandas as pd client = tweepy.Client(bearer_token='____') puma_tweets = [] for response in tweepy.Paginator(client.search_all_tweets, query = 'puma -is:retweet lang:en place_country:US', user_fields = ['id', 'username', 'name'], place_fields = ['id','full_name', 'country', 'geo', 'name', 'place_type'], tweet_fields = ['id', 'created_at', 'geo', 'public_metrics', 'text'], expansions = 'author_id', max_results=500, limit = 20): time.sleep(1) puma_tweets.append(response) result = [] user_dict = {} place_dict = {} for response in puma_tweets: for user in response.includes['users']: user_dict[user.id] = {'id': user.id, 'username': user.username, 'name': user.name, 'followers': user.public_metrics['followers_count'], } for place in response.includes['places']: place_dict[place.id] = {'id': place.id, 'full_name': place.full_name } for tweet in response.data: author_info = user_dict[tweet.author_id] result.append({'author_id': tweet.author_id, 'username': author_info['username'], 'text': tweet.text, 'geo': tweet.geo }) df = pd.DataFrame(result)
问题原因与解决方法
核心问题
你没在expansions参数里添加geo.place_id,Twitter API不会主动返回地点数据到response.includes['places']中。而用户字段能正常工作,是因为你把author_id加入了expansions列表。另外,部分分页响应可能没有地点数据,直接访问response.includes['places']会触发KeyError。
修复步骤
- 更新
expansions参数,把geo.place_id加进去:
expansions = ['author_id', 'geo.place_id']
注意这里要改成列表形式,因为要指定多个扩展字段。
- 给地点遍历加存在性判断,避免无地点数据时报错:
# 替换原来的place遍历代码 if 'places' in response.includes: for place in response.includes['places']: place_dict[place.id] = {'id': place.id, 'full_name': place.full_name }
- (可选)把地点信息关联到结果里,如果需要在最终DataFrame中显示地点名称:
# 在tweet循环中添加地点信息 for tweet in response.data: author_info = user_dict[tweet.author_id] # 获取地点信息(如果有的话) place_full_name = None if tweet.geo and 'place_id' in tweet.geo: place_info = place_dict.get(tweet.geo['place_id']) if place_info: place_full_name = place_info['full_name'] result.append({'author_id': tweet.author_id, 'username': author_info['username'], 'text': tweet.text, 'geo': tweet.geo, 'place_full_name': place_full_name })
修复后关键代码片段
修改后的Paginator调用部分:
for response in tweepy.Paginator(client.search_all_tweets, query = 'puma -is:retweet lang:en place_country:US', user_fields = ['id', 'username', 'name'], place_fields = ['id','full_name', 'country', 'geo', 'name', 'place_type'], tweet_fields = ['id', 'created_at', 'geo', 'public_metrics', 'text'], expansions = ['author_id', 'geo.place_id'], # 这里修改 max_results=500, limit = 20): time.sleep(1) puma_tweets.append(response)
内容的提问来源于stack exchange,提问作者user17424407
相关产品推荐
相关产品推荐

