使用Tweepy 4.10.0获取推文回复失败求助(含Academic API权限)
问题描述
我正在用Python的Tweepy包获取推文的实际回复,操作流程如下,但遇到了筛选回复失败的问题:
1. 初始化Tweepy Client并搜索目标推文
使用学术研究权限的Bearer Token初始化Client,通过分页器搜索2019-12-30至2020-01-15期间带#COVID标签的英文推文:
client = tweepy.Client(bearer_token=bearer_token, wait_on_rate_limit=True) covid_tweets = [] for mytweets in tweepy.Paginator(client.search_all_tweets, query='#COVID lang:en', user_fields=['username', 'public_metrics', 'description', 'location'], tweet_fields=['created_at', 'geo', 'public_metrics', 'text'], expansions='author_id', start_time='2019-12-30T00:00:00Z', end_time='2020-01-15T00:00:00Z', max_results=10): time.sleep(2) covid_tweets.append(mytweets)
2. 转换为DataFrame提取关键字段
将搜索结果转换为DataFrame,过滤转推并提取推文及作者信息:
import pandas as pd result = [] user_dict = {} # 遍历分页响应 for response in covid_tweets: # 构建用户信息字典 for user in response.includes['users']: user_dict[user.id] = { 'username': user.username, 'followers': user.public_metrics['followers_count'], 'tweets': user.public_metrics['tweet_count'], 'description': user.description, 'location': user.location } # 提取推文信息 for tweet in response.data: if 'RT @' not in tweet.text: author_info = user_dict[tweet.author_id] result.append({ 'author_id': tweet.author_id, 'tweet_id': tweet.id, 'username': author_info['username'], 'author_followers': author_info['followers'], 'author_tweets': author_info['tweets'], 'author_description': author_info['description'], 'author_location': author_info['location'], 'text': tweet.text, 'created_at': tweet.created_at, 'retweets': tweet.public_metrics['retweet_count'], 'replies': tweet.public_metrics['reply_count'], 'likes': tweet.public_metrics['like_count'], 'quote_count': tweet.public_metrics['quote_count'] }) df_1 = pd.DataFrame(result)
3. 问题:无法获取实际回复
DataFrame中部分推文的reply_count显示非零,但调用自定义get_all_replies函数时,返回的回复列表始终为空。调试发现,回复的in_reply_to_status_id与目标tweet_id不匹配:
import sys import json def get_all_replies(tweet, api, fout, depth=10, Verbose=False): global rep if depth < 1: if Verbose: print('Max depth reached') return user = tweet.user.screen_name tweet_id = tweet.id search_query = '@' + user # 过滤转推 retweet_filter = '-filter:retweets' query = search_query + retweet_filter try: myCursor = tweepy.Cursor(api.search_tweets, q=query, since_id=tweet_id, max_id=None, tweet_mode='extended').items() rep = [reply for reply in myCursor if reply.in_reply_to_status_id == tweet_id] except tweepy.TweepyException as e: sys.stderr.write(('Error get_all_replies: {}\n').format(e)) time.sleep(60) if len(rep) != 0: if Verbose: if hasattr(tweet, 'full_text'): print('Saving replies to: %s' % tweet.full_text) elif hasattr(tweet, 'text'): print('Saving replies to: %s' % tweet.text) print("Output path: %s" % fout) # 保存到文件 with open(fout, 'a+') as (f): for reply in rep: data_to_file = json.dumps(reply._json) f.write(data_to_file + '\n') # 递归获取回复的回复 get_all_replies(reply, api, fout, depth=depth - 1, Verbose=False) return
我拥有Academic Research API权限,需要解决这个无法获取回复的问题。
解决方案
问题根源
- API版本不兼容:原函数使用的是Twitter API v1.1的
search_tweets接口,该接口无论权限等级,仅能搜索最近7天的推文,而你的目标推文是2019-2020年的,根本无法获取历史回复。 - 搜索逻辑缺陷:通过
@用户名搜索会包含非回复的提及内容,且v1.1接口的in_reply_to_status_id字段对历史推文的兼容性较差,容易出现匹配失败。 - 学术权限未正确利用:Academic Research权限仅对Twitter API v2的
search_all_tweets接口生效,该接口支持全量历史数据搜索,是获取旧推文回复的唯一正确途径。
修正后的代码(基于API v2)
替换原get_all_replies函数,使用v2接口直接搜索指定推文的回复:
def get_all_replies_v2(tweet_id, client, fout, depth=10, verbose=False): if depth < 1: if verbose: print('Max depth reached') return # 直接通过推文ID筛选回复,过滤转推 query = f'in_reply_to_tweet_id:{tweet_id} lang:en -filter:retweets' try: replies = [] # 用分页器获取所有回复(学术权限支持全量搜索) for response in tweepy.Paginator(client.search_all_tweets, query=query, tweet_fields=['created_at', 'public_metrics', 'in_reply_to_tweet_id'], user_fields=['username'], expansions='author_id', max_results=100): time.sleep(2) if response.data: replies.extend(response.data) if replies: if verbose: print(f'找到 {len(replies)} 条回复,保存至: {fout}') # 保存回复数据到文件 with open(fout, 'a+', encoding='utf-8') as f: for reply in replies: reply_data = { 'tweet_id': reply.id, 'author_id': reply.author_id, 'text': reply.text, 'created_at': reply.created_at.isoformat(), 'in_reply_to_tweet_id': reply.in_reply_to_tweet_id, 'retweets': reply.public_metrics['retweet_count'], 'likes': reply.public_metrics['like_count'] } f.write(json.dumps(reply_data, ensure_ascii=False) + '\n') # 递归获取回复的回复 for reply in replies: get_all_replies_v2(reply.id, client, fout, depth=depth-1, verbose=verbose) except tweepy.TweepyException as e: print(f'Error get_all_replies_v2: {e}') time.sleep(60)
使用方法
从之前生成的DataFrame中取出tweet_id,调用修正后的函数:
# 遍历DataFrame中存在回复的推文 for idx, row in df_1.iterrows(): if row['replies'] > 0: get_all_replies_v2(row['tweet_id'], client, 'covid_replies.json', verbose=True)
内容的提问来源于stack exchange,提问作者Michael Maverick
相关产品推荐
相关产品推荐

