You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过Twitter API v2的Tweepy StreamingClient获取更多流数据详情

问题描述

之前使用Tweepy的v1.1 Twitter流功能追踪特定账号,当该账号被提及时,能获取提及推文、含视频推文的完整详情(包括推文信息、视频链接等),v1.1核心代码如下:

class StdOutListener(tweepy.Stream):
    def on_data(self, data):
        # 处理流数据
        struct = json.loads(data)

其中struct包含完整的JSON格式数据。

但切换到API v2后,使用StreamingClient仅能获取少量基础信息,核心代码如下:

class StdOutListener(tweepy.StreamingClient):
    def on_tweet(self, tweet):
        print(tweet)
        print(tweet.data)
        print(tweet.entities)
        print(f"{tweet.id} {tweet.created_at} ({tweet.author_id}): {tweet.text}")

输出示例:

INFO:tweepy.streaming:Stream connected
@2hvQqjddfgfdUY96Ah5yW @7bdsfdsfds3h_bot
{'edit_history_tweet_ids': ['1617291424910688874'], 'id': '1617291424910688874', 'text': '@2hvQqjddfgfdUY96Ah5yW @7bdsfdsfds3h_bot'}
None
1617291424910688874 None (None): @2hvQqjddfgfdUY96Ah5yW @7bdsfdsfds3h_bot

初始化和启动流的代码:

printer = StdOutListener(bearer_token)

# 添加规则
rule = StreamRule(value="@7bdsfdsfds3h_bot")
printer.add_rules(rule)

printer.filter()

请问如何在API v2中获取和v1.1相同的完整数据?

解决方案

Twitter API v2采用按需返回字段的设计,默认仅返回最基础的推文字段(如ID、文本)。要获取完整数据,需要在调用filter()时指定需要的扩展字段、推文字段、媒体字段等参数,同时改用on_response方法处理完整响应(on_tweet仅返回推文主体,不包含扩展数据)。

1. 修改监听类,使用on_response方法

on_response能拿到包含所有请求数据的响应对象,包括扩展的用户信息、媒体数据等:

class StdOutListener(tweepy.StreamingClient):
    def on_response(self, response):
        # 打印完整响应数据(可按需解析)
        print(response.data)
        print(response.includes)
        
        # 解析推文基础信息
        tweet = response.data
        print(f"推文ID: {tweet.id}, 发布时间: {tweet.created_at}, 作者ID: {tweet.author_id}")
        print(f"推文内容: {tweet.text}")
        
        # 解析媒体数据(比如视频链接)
        if 'media' in response.includes:
            for media in response.includes['media']:
                if media.type == 'video':
                    print(f"视频链接: {media.url}")
                    print(f"视频变体信息: {media.video_info}")

2. 调用filter()时指定需要的字段

在启动流时,通过expansions、tweet_fields、media_fields等参数声明需要获取的内容:

printer = StdOutListener(bearer_token)

# 添加规则
rule = StreamRule(value="@7bdsfdsfds3h_bot")
printer.add_rules(rule)

# 指定需要的扩展字段和详情字段
printer.filter(
    expansions=["attachments.media_keys", "author_id"],
    tweet_fields=["created_at", "entities", "public_metrics"],
    media_fields=["url", "type", "video_info", "duration_ms"]
)

参数说明

  • expansions: 指定需要关联的扩展资源,比如attachments.media_keys(关联媒体数据)、author_id(关联作者信息)
  • tweet_fields: 指定需要的推文详情字段,比如created_at(发布时间)、entities(实体信息)、public_metrics(互动数据)
  • media_fields: 指定需要的媒体详情字段,比如url(媒体链接)、video_info(视频信息,包含播放链接)

通过以上配置,就能获取到和v1.1流功能类似的完整数据,包括视频链接、作者信息、推文元数据等。

内容的提问来源于stack exchange,提问作者AKMalkadi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 02:00:34