You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Tweepy API V2与Pandas构建DataFrame时用户名重复错误求助

修复Tweepy Paginator获取推文作者信息重复的问题

你的代码核心问题是没有正确从API响应的includes中提取对应推文作者的用户数据,而是复用了某个未正确关联的user变量,导致所有行显示同一个错误用户名。

Tweepy的search_recent_tweets配合expansions=['author_id']时,用户信息不会直接附在tweet对象上,而是放在响应的includes['users']字段里,需要通过推文的author_id去匹配对应的用户。

修复后的完整代码

import tweepy
import pandas as pd

query = 'Wizkid'
tweet_info_ls = []

# 遍历Paginator返回的每个响应对象(而非直接遍历tweet)
for response in tweepy.Paginator(
    client.search_recent_tweets,
    query=query,
    tweet_fields=['context_annotations', 'created_at','author_id', 'public_metrics'],
    expansions=['author_id','referenced_tweets.id'],
    max_results=100,
    user_fields=['username', 'name','public_metrics']
):
    # 构建author_id到用户对象的映射字典
    user_dict = {user.id: user for user in response.includes['users']}
    
    # 遍历当前页的所有推文
    for tweet in response.data:
        # 通过推文的author_id获取对应的用户
        user = user_dict[tweet.author_id]
        
        tweet_info = {
            'created_at': tweet.created_at,
            'text': tweet.text,
            'name': user.name,
            'username': user.username,
            'followers': user.public_metrics['followers_count'],
            'repost': tweet.public_metrics['retweet_count']
        }
        tweet_info_ls.append(tweet_info)
        
        # 达到100条后停止
        if len(tweet_info_ls) >= 100:
            break
    if len(tweet_info_ls) >= 100:
        break

# 生成DataFrame
tweets_df = pd.DataFrame(tweet_info_ls)
tweets_df.head(20)

关键修改说明

  • 不再直接使用.flatten(),而是遍历每个API响应,保留includes中的用户数据
  • 用user_dict建立推文作者ID和用户对象的映射,确保每条推文匹配到正确的作者
  • 添加了条数限制判断,精准控制获取100条推文

内容的提问来源于stack exchange,提问作者Eugene

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 06:45:40