如何用Tweepy(含Paginator)获取带媒体的推文文本及图片URL?
解决Tweepy获取带媒体推文及图片URL的问题
原代码的核心问题
- 单个
tweet对象不存在includes属性,includes属于分页/请求返回的response对象,不能直接从推文里取媒体数据。 - 没有将媒体的
media_key与推文的attachments关联,导致无法匹配到对应推文的图片。 - 使用
Paginator时,每一页的response都独立包含该页的includes数据,需要逐页处理媒体映射。
修正后的代码(含Paginator实现)
import tweepy bearer_token = "你的Bearer Token" client = tweepy.Client(bearer_token) # 初始化Paginator,设置查询条件和参数 paginator = tweepy.Paginator( client.search_recent_tweets, query="#covid -is:retweet has:media", max_results=100, expansions="author_id,attachments.media_keys", tweet_fields="created_at,public_metrics,attachments", media_fields="url,type" # 只保留需要的媒体字段,type用于筛选图片 ) # 遍历每一页结果 for response in paginator: # 构建媒体映射:media_key -> 图片URL(只筛选图片类型) media_map = {} if 'media' in response.includes: for media in response.includes['media']: if media.type == 'photo': # 仅保留图片,排除视频等其他媒体 media_map[media.media_key] = media.url # 处理每一条推文 for tweet in response.data: # 获取推文基础信息 tweet_id = tweet.id created_at = tweet.created_at text = tweet.text author_id = tweet.author_id retweet_count = tweet.public_metrics['retweet_count'] like_count = tweet.public_metrics['like_count'] # 获取当前推文对应的图片URL列表 image_urls = [] if hasattr(tweet, 'attachments') and tweet.attachments: media_keys = tweet.attachments.get('media_keys', []) for key in media_keys: if key in media_map: image_urls.append(media_map[key]) # 输出或保存数据 print(f"推文ID: {tweet_id}") print(f"创建时间: {created_at}") print(f"作者ID: {author_id}") print(f"推文内容: {text}") print(f"转发量: {retweet_count}, 点赞量: {like_count}") print(f"图片URL: {image_urls}\n")
关键说明
- 媒体映射构建:每一页都基于当前response的includes构建
media_map,确保推文能匹配到对应的图片URL,同时通过media.type == 'photo'过滤掉非图片的媒体(比如视频、GIF)。 - Paginator使用:Paginator会自动处理分页请求,不需要手动管理
next_token,可以通过max_pages参数限制获取的总页数(比如max_pages=5表示获取500条推文)。 - 字段安全访问:使用
hasattr和get方法避免因部分推文没有附件导致的报错。
内容的提问来源于stack exchange,提问作者K3it4r0
相关产品推荐
相关产品推荐

