如何用Python+Tweepy调用Twitter API v2获取超100条带用户字段的推文
解决search_recent_tweets分页及用户字段同步问题
问题拆解与修正方案
- 分页参数确认:
search_recent_tweets确实使用pagination_token参数实现分页,你代码里的参数名是对的,问题出在后续数据处理环节。 - 数据累加错误:Pandas的
append()方法已被弃用,且不会修改原DataFrame,必须用pd.concat()生成新的合并数据。 - 用户字段匹配优化:直接拼接
includes中的用户数据可能导致重复或匹配混乱,建议将用户数据转为以author_id为键的字典,再与推文数据关联,效率更高且更准确。
修正后的完整代码
import pandas as pd import tweepy BEARER_TOKEN = '你的Bearer Token' api = tweepy.Client(BEARER_TOKEN) # 初始化存储所有数据的列表 all_tweets_list = [] # 首次请求 response = api.search_recent_tweets( query='myquery', start_time='2022-09-19T00:00:00Z', end_time='2022-09-19T23:59:59Z', expansions=['author_id'], tweet_fields=['created_at'], user_fields=['username', 'name'], max_results=100 ) # 处理首次请求的数据 if response.data: # 将推文转为DataFrame tweet_df = pd.DataFrame(response.data) # 处理用户数据:转为以author_id为键的字典,方便匹配 users_dict = {user.id: {'username': user.username, 'name': user.name} for user in response.includes['users']} # 给推文DataFrame添加用户字段 tweet_df['username'] = tweet_df['author_id'].map(lambda x: users_dict[x]['username']) tweet_df['name'] = tweet_df['author_id'].map(lambda x: users_dict[x]['name']) # 添加到列表 all_tweets_list.append(tweet_df) # 获取分页token next_token = response.meta.get('next_token') # 分页循环请求 while next_token is not None: try: response = api.search_recent_tweets( query='myquery', start_time='2022-09-19T00:00:00Z', end_time='2022-09-19T23:59:59Z', expansions=['author_id'], tweet_fields=['created_at'], user_fields=['username', 'name'], pagination_token=next_token, max_results=100 ) if not response.data: break tweet_df = pd.DataFrame(response.data) users_dict = {user.id: {'username': user.username, 'name': user.name} for user in response.includes['users']} tweet_df['username'] = tweet_df['author_id'].map(lambda x: users_dict[x]['username']) tweet_df['name'] = tweet_df['author_id'].map(lambda x: users_dict[x]['name']) all_tweets_list.append(tweet_df) # 更新分页token next_token = response.meta.get('next_token') except Exception as e: print(f"请求出错: {e}") break # 合并所有数据 if all_tweets_list: all_tweets = pd.concat(all_tweets_list, ignore_index=True) else: all_tweets = pd.DataFrame() print(all_tweets)
关键修改说明
- 数据累加:用列表存储每一页的DataFrame,最后用
pd.concat()合并,替代弃用的append(),同时避免内存浪费。 - 用户字段匹配:通过字典映射的方式,直接将用户的username和name匹配到对应推文,比
merge()更高效,也避免了重复用户数据的冗余存储。 - 异常处理:增加try-except块,防止分页请求中出现错误导致程序直接崩溃。
- 空数据处理:判断response.data是否存在,避免无结果时抛出空值错误。
内容的提问来源于stack exchange,提问作者Moo
相关产品推荐
相关产品推荐

