使用Tweepy查询多tweet_id的Twitter API公开指标报错排查
目标
使用Tweepy调用Twitter API,获取多个tweet_id对应推文的public_metrics字段值,包含likes(点赞数)、retweets(转发数)、quotes(引用数)、replies(回复数)四类指标。
原有实现代码
client = tweepy.Client(bearer_token, wait_on_rate_limit=True) gathered_tweets = [] for response in tweepy.Paginator(client.get_tweets, query = '1537245238095323136, 1536973965394157569', tweet_fields = ['public_metrics'], max_results=100): time.sleep(1) gathered_tweets.append(response) result = [] user_dict = {} # Loop through each response object for response in gathered_tweets: # Take all of the users, and put them into a dictionary of dictionaries with the info we want to keep for user in response.includes['users']: user_dict[user.id] = {'username': user.username, 'created_at': user.created_at } for tweet in response.data: # For each tweet, find the author's information author_info = user_dict[tweet.author_id] # Put all of the information we want to keep in a single dictionary for each tweet result.append({'author_id': tweet.author_id, 'tweet_id': tweet.id, 'retweets': tweet.public_metrics['retweet_count'], 'replies': tweet.public_metrics['reply_count'], 'likes': tweet.public_metrics['like_count'], 'quotes': tweet.public_metrics['quote_count'] })
报错信息
TypeError Traceback (most recent call last) <ipython-input-31-a0842c496bea> in <module>() 9 query = '1537245238095323136,1536973965394157569', 10 tweet_fields = ['public_metrics'], ---> 11 max_results=100): 12 time.sleep(1) 13 gathered_tweets.append(response) /usr/local/lib/python3.7/dist-packages/tweepy/pagination.py in __next__(self) 96 self.kwargs["pagination_token"] = pagination_token 97 ---> 98 response = self.method(*self.args, **self.kwargs) 99 100 self.previous_token = response.meta.get("previous_token") TypeError: get_tweets() missing 1 required positional argument: 'ids'
问题原因与修正方法
报错核心是代码存在3个用法错误:
- 参数名错误:
get_tweets接口接收目标推文ID的必填参数为ids,你代码里传的query是推文搜索类接口的参数,get_tweets根本不识别这个参数,自然会提示缺少必填的ids参数。 - 传参格式错误:文档提到的「逗号分隔ID、不要加空格」是底层原生HTTP API的要求,Tweepy作为SDK已经做了封装,直接传入ID组成的Python列表即可,不需要手动拼接字符串,你之前拼接的字符串里逗号后带了空格,就算参数名改对也会触发格式错误。
- 接口逻辑误用:
get_tweets单次请求最多支持查100个ID,接口本身没有分页机制,不需要使用Paginator;另外你代码里读取作者信息的逻辑缺少必要参数,默认接口不会返回作者数据,必须加expansions='author_id'声明要展开作者ID关联的用户信息,同时指定user_fields明确要返回的用户字段,否则response.includes里根本不会有users数据,会触发二次报错。
修正后的可运行代码如下:
import time import tweepy client = tweepy.Client(bearer_token, wait_on_rate_limit=True) # 把要查询的推文ID放到列表里即可,支持整数、字符串两种格式的ID tweet_ids = [1537245238095323136, 1536973965394157569] # 单次请求最多传100个ID,ID总数超过100时手动切分批次循环调用即可,不需要分页器 response = client.get_tweets( ids=tweet_ids, tweet_fields=['public_metrics'], expansions='author_id', user_fields=['username', 'created_at'] ) time.sleep(1) result = [] user_dict = {} # 先解析返回的作者信息 for user in response.includes.get('users', []): user_dict[user.id] = { 'username': user.username, 'created_at': user.created_at } # 再解析推文指标数据 for tweet in response.data: author_info = user_dict.get(tweet.author_id, {}) result.append({ 'author_id': tweet.author_id, 'author_username': author_info.get('username'), 'tweet_id': tweet.id, 'retweets': tweet.public_metrics['retweet_count'], 'replies': tweet.public_metrics['reply_count'], 'likes': tweet.public_metrics['like_count'], 'quotes': tweet.public_metrics['quote_count'] })
内容的提问来源于stack exchange,提问作者dsx
相关产品推荐
相关产品推荐

