You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python+Tweepy调用Twitter API v2获取超100条带用户字段的推文

解决search_recent_tweets分页及用户字段同步问题

问题拆解与修正方案

  1. 分页参数确认:search_recent_tweets确实使用pagination_token参数实现分页,你代码里的参数名是对的,问题出在后续数据处理环节。
  2. 数据累加错误:Pandas的append()方法已被弃用,且不会修改原DataFrame,必须用pd.concat()生成新的合并数据。
  3. 用户字段匹配优化:直接拼接includes中的用户数据可能导致重复或匹配混乱,建议将用户数据转为以author_id为键的字典,再与推文数据关联,效率更高且更准确。

修正后的完整代码

import pandas as pd
import tweepy

BEARER_TOKEN = '你的Bearer Token'
api = tweepy.Client(BEARER_TOKEN)

# 初始化存储所有数据的列表
all_tweets_list = []

# 首次请求
response = api.search_recent_tweets(
    query='myquery',
    start_time='2022-09-19T00:00:00Z',
    end_time='2022-09-19T23:59:59Z',
    expansions=['author_id'],
    tweet_fields=['created_at'],
    user_fields=['username', 'name'],
    max_results=100
)

# 处理首次请求的数据
if response.data:
    # 将推文转为DataFrame
    tweet_df = pd.DataFrame(response.data)
    # 处理用户数据:转为以author_id为键的字典,方便匹配
    users_dict = {user.id: {'username': user.username, 'name': user.name} for user in response.includes['users']}
    # 给推文DataFrame添加用户字段
    tweet_df['username'] = tweet_df['author_id'].map(lambda x: users_dict[x]['username'])
    tweet_df['name'] = tweet_df['author_id'].map(lambda x: users_dict[x]['name'])
    # 添加到列表
    all_tweets_list.append(tweet_df)

# 获取分页token
next_token = response.meta.get('next_token')

# 分页循环请求
while next_token is not None:
    try:
        response = api.search_recent_tweets(
            query='myquery',
            start_time='2022-09-19T00:00:00Z',
            end_time='2022-09-19T23:59:59Z',
            expansions=['author_id'],
            tweet_fields=['created_at'],
            user_fields=['username', 'name'],
            pagination_token=next_token,
            max_results=100
        )
        
        if not response.data:
            break
            
        tweet_df = pd.DataFrame(response.data)
        users_dict = {user.id: {'username': user.username, 'name': user.name} for user in response.includes['users']}
        tweet_df['username'] = tweet_df['author_id'].map(lambda x: users_dict[x]['username'])
        tweet_df['name'] = tweet_df['author_id'].map(lambda x: users_dict[x]['name'])
        all_tweets_list.append(tweet_df)
        
        # 更新分页token
        next_token = response.meta.get('next_token')
    except Exception as e:
        print(f"请求出错: {e}")
        break

# 合并所有数据
if all_tweets_list:
    all_tweets = pd.concat(all_tweets_list, ignore_index=True)
else:
    all_tweets = pd.DataFrame()

print(all_tweets)

关键修改说明

  • 数据累加:用列表存储每一页的DataFrame,最后用pd.concat()合并,替代弃用的append(),同时避免内存浪费。
  • 用户字段匹配:通过字典映射的方式,直接将用户的username和name匹配到对应推文,比merge()更高效,也避免了重复用户数据的冗余存储。
  • 异常处理:增加try-except块,防止分页请求中出现错误导致程序直接崩溃。
  • 空数据处理:判断response.data是否存在,避免无结果时抛出空值错误。

内容的提问来源于stack exchange,提问作者Moo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 12:10:49