如何通过Twitter API v2的Tweepy StreamingClient获取更多流数据详情
问题描述
之前使用Tweepy的v1.1 Twitter流功能追踪特定账号,当该账号被提及时,能获取提及推文、含视频推文的完整详情(包括推文信息、视频链接等),v1.1核心代码如下:
class StdOutListener(tweepy.Stream): def on_data(self, data): # 处理流数据 struct = json.loads(data)
其中struct包含完整的JSON格式数据。
但切换到API v2后,使用StreamingClient仅能获取少量基础信息,核心代码如下:
class StdOutListener(tweepy.StreamingClient): def on_tweet(self, tweet): print(tweet) print(tweet.data) print(tweet.entities) print(f"{tweet.id} {tweet.created_at} ({tweet.author_id}): {tweet.text}")
输出示例:
INFO:tweepy.streaming:Stream connected @2hvQqjddfgfdUY96Ah5yW @7bdsfdsfds3h_bot {'edit_history_tweet_ids': ['1617291424910688874'], 'id': '1617291424910688874', 'text': '@2hvQqjddfgfdUY96Ah5yW @7bdsfdsfds3h_bot'} None 1617291424910688874 None (None): @2hvQqjddfgfdUY96Ah5yW @7bdsfdsfds3h_bot
初始化和启动流的代码:
printer = StdOutListener(bearer_token) # 添加规则 rule = StreamRule(value="@7bdsfdsfds3h_bot") printer.add_rules(rule) printer.filter()
请问如何在API v2中获取和v1.1相同的完整数据?
解决方案
Twitter API v2采用按需返回字段的设计,默认仅返回最基础的推文字段(如ID、文本)。要获取完整数据,需要在调用filter()时指定需要的扩展字段、推文字段、媒体字段等参数,同时改用on_response方法处理完整响应(on_tweet仅返回推文主体,不包含扩展数据)。
1. 修改监听类,使用on_response方法
on_response能拿到包含所有请求数据的响应对象,包括扩展的用户信息、媒体数据等:
class StdOutListener(tweepy.StreamingClient): def on_response(self, response): # 打印完整响应数据(可按需解析) print(response.data) print(response.includes) # 解析推文基础信息 tweet = response.data print(f"推文ID: {tweet.id}, 发布时间: {tweet.created_at}, 作者ID: {tweet.author_id}") print(f"推文内容: {tweet.text}") # 解析媒体数据(比如视频链接) if 'media' in response.includes: for media in response.includes['media']: if media.type == 'video': print(f"视频链接: {media.url}") print(f"视频变体信息: {media.video_info}")
2. 调用filter()时指定需要的字段
在启动流时,通过expansions、tweet_fields、media_fields等参数声明需要获取的内容:
printer = StdOutListener(bearer_token) # 添加规则 rule = StreamRule(value="@7bdsfdsfds3h_bot") printer.add_rules(rule) # 指定需要的扩展字段和详情字段 printer.filter( expansions=["attachments.media_keys", "author_id"], tweet_fields=["created_at", "entities", "public_metrics"], media_fields=["url", "type", "video_info", "duration_ms"] )
参数说明
expansions: 指定需要关联的扩展资源,比如attachments.media_keys(关联媒体数据)、author_id(关联作者信息)tweet_fields: 指定需要的推文详情字段,比如created_at(发布时间)、entities(实体信息)、public_metrics(互动数据)media_fields: 指定需要的媒体详情字段,比如url(媒体链接)、video_info(视频信息,包含播放链接)
通过以上配置,就能获取到和v1.1流功能类似的完整数据,包括视频链接、作者信息、推文元数据等。
内容的提问来源于stack exchange,提问作者AKMalkadi
相关产品推荐
相关产品推荐

