You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Tweepy 4.10.0获取推文回复失败求助(含Academic API权限)

问题描述

我正在用Python的Tweepy包获取推文的实际回复,操作流程如下,但遇到了筛选回复失败的问题:

1. 初始化Tweepy Client并搜索目标推文

使用学术研究权限的Bearer Token初始化Client,通过分页器搜索2019-12-30至2020-01-15期间带#COVID标签的英文推文:

client = tweepy.Client(bearer_token=bearer_token, wait_on_rate_limit=True)
covid_tweets = []

for mytweets in tweepy.Paginator(client.search_all_tweets, 
                                 query='#COVID lang:en', 
                                 user_fields=['username', 'public_metrics', 'description', 'location'], 
                                 tweet_fields=['created_at', 'geo', 'public_metrics', 'text'], 
                                 expansions='author_id', 
                                 start_time='2019-12-30T00:00:00Z', 
                                 end_time='2020-01-15T00:00:00Z', 
                                 max_results=10):
    time.sleep(2)
    covid_tweets.append(mytweets)

2. 转换为DataFrame提取关键字段

将搜索结果转换为DataFrame,过滤转推并提取推文及作者信息:

import pandas as pd

result = []
user_dict = {}

# 遍历分页响应
for response in covid_tweets:
    # 构建用户信息字典
    for user in response.includes['users']:
        user_dict[user.id] = {
            'username': user.username,
            'followers': user.public_metrics['followers_count'],
            'tweets': user.public_metrics['tweet_count'],
            'description': user.description,
            'location': user.location
        }
    # 提取推文信息
    for tweet in response.data:
        if 'RT @' not in tweet.text:
            author_info = user_dict[tweet.author_id]
            result.append({
                'author_id': tweet.author_id,
                'tweet_id': tweet.id,
                'username': author_info['username'],
                'author_followers': author_info['followers'],
                'author_tweets': author_info['tweets'],
                'author_description': author_info['description'],
                'author_location': author_info['location'],
                'text': tweet.text,
                'created_at': tweet.created_at,
                'retweets': tweet.public_metrics['retweet_count'],
                'replies': tweet.public_metrics['reply_count'],
                'likes': tweet.public_metrics['like_count'],
                'quote_count': tweet.public_metrics['quote_count']
            })

df_1 = pd.DataFrame(result)

3. 问题:无法获取实际回复

DataFrame中部分推文的reply_count显示非零,但调用自定义get_all_replies函数时,返回的回复列表始终为空。调试发现,回复的in_reply_to_status_id与目标tweet_id不匹配:

import sys
import json

def get_all_replies(tweet, api, fout, depth=10, Verbose=False):
    global rep
    if depth < 1:
        if Verbose:
            print('Max depth reached')
        return
    user = tweet.user.screen_name
    tweet_id = tweet.id
    search_query = '@' + user

    # 过滤转推
    retweet_filter = '-filter:retweets'
    query = search_query + retweet_filter
    
    try:
        myCursor = tweepy.Cursor(api.search_tweets, q=query,
                                 since_id=tweet_id,
                                 max_id=None,
                                 tweet_mode='extended').items()
        rep = [reply for reply in myCursor if reply.in_reply_to_status_id == tweet_id]
    except tweepy.TweepyException as e:
        sys.stderr.write(('Error get_all_replies: {}\n').format(e))
        time.sleep(60)

    if len(rep) != 0:
        if Verbose:
            if hasattr(tweet, 'full_text'):
                print('Saving replies to: %s' % tweet.full_text)
            elif hasattr(tweet, 'text'):
                print('Saving replies to: %s' % tweet.text)
            print("Output path: %s" % fout)

        # 保存到文件
        with open(fout, 'a+') as (f):
            for reply in rep:
                data_to_file = json.dumps(reply._json)
                f.write(data_to_file + '\n')

            # 递归获取回复的回复
            get_all_replies(reply, api, fout, depth=depth - 1, Verbose=False)
    return

我拥有Academic Research API权限,需要解决这个无法获取回复的问题。


解决方案

问题根源

  1. API版本不兼容:原函数使用的是Twitter API v1.1的search_tweets接口,该接口无论权限等级,仅能搜索最近7天的推文,而你的目标推文是2019-2020年的,根本无法获取历史回复。
  2. 搜索逻辑缺陷:通过@用户名搜索会包含非回复的提及内容,且v1.1接口的in_reply_to_status_id字段对历史推文的兼容性较差,容易出现匹配失败。
  3. 学术权限未正确利用:Academic Research权限仅对Twitter API v2的search_all_tweets接口生效,该接口支持全量历史数据搜索,是获取旧推文回复的唯一正确途径。

修正后的代码(基于API v2)

替换原get_all_replies函数,使用v2接口直接搜索指定推文的回复:

def get_all_replies_v2(tweet_id, client, fout, depth=10, verbose=False):
    if depth < 1:
        if verbose:
            print('Max depth reached')
        return
    
    # 直接通过推文ID筛选回复,过滤转推
    query = f'in_reply_to_tweet_id:{tweet_id} lang:en -filter:retweets'
    
    try:
        replies = []
        # 用分页器获取所有回复(学术权限支持全量搜索)
        for response in tweepy.Paginator(client.search_all_tweets,
                                         query=query,
                                         tweet_fields=['created_at', 'public_metrics', 'in_reply_to_tweet_id'],
                                         user_fields=['username'],
                                         expansions='author_id',
                                         max_results=100):
            time.sleep(2)
            if response.data:
                replies.extend(response.data)
        
        if replies:
            if verbose:
                print(f'找到 {len(replies)} 条回复,保存至: {fout}')
            
            # 保存回复数据到文件
            with open(fout, 'a+', encoding='utf-8') as f:
                for reply in replies:
                    reply_data = {
                        'tweet_id': reply.id,
                        'author_id': reply.author_id,
                        'text': reply.text,
                        'created_at': reply.created_at.isoformat(),
                        'in_reply_to_tweet_id': reply.in_reply_to_tweet_id,
                        'retweets': reply.public_metrics['retweet_count'],
                        'likes': reply.public_metrics['like_count']
                    }
                    f.write(json.dumps(reply_data, ensure_ascii=False) + '\n')
            
            # 递归获取回复的回复
            for reply in replies:
                get_all_replies_v2(reply.id, client, fout, depth=depth-1, verbose=verbose)
    
    except tweepy.TweepyException as e:
        print(f'Error get_all_replies_v2: {e}')
        time.sleep(60)

使用方法

从之前生成的DataFrame中取出tweet_id,调用修正后的函数:

# 遍历DataFrame中存在回复的推文
for idx, row in df_1.iterrows():
    if row['replies'] > 0:
        get_all_replies_v2(row['tweet_id'], client, 'covid_replies.json', verbose=True)

内容的提问来源于stack exchange,提问作者Michael Maverick

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 18:37:15