You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过snscrape提取推特中的图片预览URL与完整URL

用snscrape爬取推特时提取图片URL的实现方案

在使用snscrape爬取推特数据时,发现tweet.media字段同时包含图片和视频的URL,需要在现有循环逻辑中过滤出图片类型的媒体,分别提取previewUrl(预览图链接)和fullUrl(高清图链接),并新增字段存储完整的高清图片URL。


原参考代码

import snscrape.modules.twitter as sntwitter
import pandas as pd

# Creating list to append tweet data to
attributes_container = []

# Using TwitterSearchScraper to scrape data and append tweets to list
for i,tweet in enumerate(sntwitter.TwitterSearchScraper('sex for grades since:2021-07-05 until:2022-07-06').get_items()):
    if i>150:
        break
    attributes_container.append([tweet.user.username, tweet.date, tweet.likeCount, tweet.sourceLabel, tweet.content, tweet.media])
    
# Creating a dataframe to load the list
tweets_df = pd.DataFrame(attributes_container, columns=["User", "Date Created", "Number of Likes", "Source of Tweet", "Tweet","media"])

示例media对象结构

media=[Photo(previewUrl='https://pbs.twimg.com/media/FePrYL7WQAQDKEB?format=jpg', fullUrl='https://pbs.twimg.com/media/FePrYL7WQAQDKEB?format=jpg&name=large')]

修改后的实现代码

import snscrape.modules.twitter as sntwitter
import pandas as pd

# 创建存储推文数据的列表
attributes_container = []

# 爬取推特数据并处理媒体字段
for i, tweet in enumerate(sntwitter.TwitterSearchScraper('sex for grades since:2021-07-05 until:2022-07-06').get_items()):
    if i > 150:
        break
    
    # 初始化图片链接存储列表
    preview_image_urls = []
    full_photo_urls = []
    
    # 过滤图片媒体并提取链接
    if tweet.media:
        for media_item in tweet.media:
            # 判断当前媒体是否为图片类型
            if isinstance(media_item, sntwitter.Photo):
                preview_image_urls.append(media_item.previewUrl)
                full_photo_urls.append(media_item.fullUrl)
    
    # 将处理后的数据添加到列表,新增图片链接字段
    attributes_container.append([
        tweet.user.username, 
        tweet.date, 
        tweet.likeCount, 
        tweet.sourceLabel, 
        tweet.content, 
        tweet.media,
        preview_image_urls,
        full_photo_urls
    ])

# 生成DataFrame,定义对应列名
tweets_df = pd.DataFrame(
    attributes_container, 
    columns=[
        "User", 
        "Date Created", 
        "Number of Likes", 
        "Source of Tweet", 
        "Tweet",
        "media",
        "Preview Image URLs",
        "Photo URLs"  # 对应需求的photourl字段,存储完整高清图片URL
    ]
)

关键逻辑说明

  • 用isinstance(media_item, sntwitter.Photo)精准判断媒体类型,自动过滤视频等非图片内容
  • 使用列表存储链接,适配单条推文包含多张图片的场景
  • 无图片的推文会自动生成空列表,避免空值报错

内容的提问来源于stack exchange,提问作者hard worker

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 08:50:26