You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Tweepy(含Paginator)获取带媒体的推文文本及图片URL?

解决Tweepy获取带媒体推文及图片URL的问题

原代码的核心问题

  1. 单个tweet对象不存在includes属性,includes属于分页/请求返回的response对象,不能直接从推文里取媒体数据。
  2. 没有将媒体的media_key与推文的attachments关联,导致无法匹配到对应推文的图片。
  3. 使用Paginator时,每一页的response都独立包含该页的includes数据,需要逐页处理媒体映射。

修正后的代码(含Paginator实现)

import tweepy

bearer_token = "你的Bearer Token"
client = tweepy.Client(bearer_token)

# 初始化Paginator,设置查询条件和参数
paginator = tweepy.Paginator(
    client.search_recent_tweets,
    query="#covid -is:retweet has:media",
    max_results=100,
    expansions="author_id,attachments.media_keys",
    tweet_fields="created_at,public_metrics,attachments",
    media_fields="url,type"  # 只保留需要的媒体字段,type用于筛选图片
)

# 遍历每一页结果
for response in paginator:
    # 构建媒体映射:media_key -> 图片URL(只筛选图片类型)
    media_map = {}
    if 'media' in response.includes:
        for media in response.includes['media']:
            if media.type == 'photo':  # 仅保留图片,排除视频等其他媒体
                media_map[media.media_key] = media.url

    # 处理每一条推文
    for tweet in response.data:
        # 获取推文基础信息
        tweet_id = tweet.id
        created_at = tweet.created_at
        text = tweet.text
        author_id = tweet.author_id
        retweet_count = tweet.public_metrics['retweet_count']
        like_count = tweet.public_metrics['like_count']

        # 获取当前推文对应的图片URL列表
        image_urls = []
        if hasattr(tweet, 'attachments') and tweet.attachments:
            media_keys = tweet.attachments.get('media_keys', [])
            for key in media_keys:
                if key in media_map:
                    image_urls.append(media_map[key])

        # 输出或保存数据
        print(f"推文ID: {tweet_id}")
        print(f"创建时间: {created_at}")
        print(f"作者ID: {author_id}")
        print(f"推文内容: {text}")
        print(f"转发量: {retweet_count}, 点赞量: {like_count}")
        print(f"图片URL: {image_urls}\n")

关键说明

  • 媒体映射构建:每一页都基于当前response的includes构建media_map,确保推文能匹配到对应的图片URL,同时通过media.type == 'photo'过滤掉非图片的媒体(比如视频、GIF)。
  • Paginator使用:Paginator会自动处理分页请求,不需要手动管理next_token,可以通过max_pages参数限制获取的总页数(比如max_pages=5表示获取500条推文)。
  • 字段安全访问:使用hasattr和get方法避免因部分推文没有附件导致的报错。

内容的提问来源于stack exchange,提问作者K3it4r0

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 23:40:35