You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬虫导出图片至CSV时获data:image而非图片URL的问题

解决Scrapy爬取图片时CSV导出data:image格式而非URL的问题

你遇到的情况是因为目标网站用base64编码的占位图填充了img标签的src属性,真实图片URL通常藏在data-src、data-original这类自定义属性里。以下是具体解决步骤:

1. 确认真实图片URL的属性

打开目标网页,按F12调出开发者工具,定位到图片元素,查看其属性。比如常见的结构是:

<img src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" data-src="https://example.com/real-image.jpg" />

这里data-src就是真实图片的URL所在属性。

2. 修改XPath选择器

调整你的爬虫代码,优先抓取真实URL对应的属性,没有的话再 fallback 到src:

# 优先抓data-src,不存在则用src
Feature_Image = response.xpath('//*[@id="article-wrapper"]/article/section[2]/div/div/div/img/@data-src').get() or response.xpath('//*[@id="article-wrapper"]/article/section[2]/div/div/div/img/@src').get()

3. 处理空值与URL拼接

避免抓取到空值时urljoin报错,增加判断逻辑:

if Feature_Image:
    Feature_Image = response.urljoin(Feature_Image)
else:
    Feature_Image = None  # 或者设置你需要的默认值

4. 修改后的完整代码

class NewsSpider(scrapy.Spider):
    name = "articles"

    def start_requests(self):
        url = input("Enter the article url: ")
        yield scrapy.Request(url, callback=self.parse_dir_contents)

    def parse_dir_contents(self, response):
        # 优先获取真实图片URL的属性,这里以data-src为例,根据实际情况调整
        Feature_Image = response.xpath('//*[@id="article-wrapper"]/article/section[2]/div/div/div/img/@data-src').get() or response.xpath('//*[@id="article-wrapper"]/article/section[2]/div/div/div/img/@src').get()
        
        # 处理URL拼接,空值则设为None
        if Feature_Image:
            Feature_Image = response.urljoin(Feature_Image)
        else:
            Feature_Image = None

        # 注意:确保Published_Date、Content、Category等变量已提前定义
        yield {
            'Publication Date': Published_Date,
            'Feature_Image': Feature_Image,
            'Article Content': Content
        }
        
        # 数据存储逻辑
        Data = [[Category, Headlines, Author, Source, Published_Date, Feature_Image, Content, url]]
        try:
            df = pd.DataFrame(Data, columns=['Category','Headlines','Author','Source','Published_Date','Feature_Image','Content','URL'])
            print(df)
            with open('C:/Users/Public/pagedata.csv', 'a') as f:
                df.to_csv(f, header=False)
        except Exception as e:
            print(f"存储异常: {e}")
            df = pd.DataFrame(Data, columns=['Category','Headlines','Author','Source','Published_Date','Feature_Image','Content','URL'])
            print(df)
            df.to_csv('C:/Users/Public/pagedata.csv', mode='a')

额外排查方向

  • 如果上述方法无效,可能是网站通过JavaScript动态加载图片,此时需要使用Scrapy的selenium或playwright中间件渲染页面后再抓取。
  • 部分网站会把多尺寸图片放在srcset属性中,需要解析该属性提取目标URL。

内容的提问来源于stack exchange,提问作者Info Rewind

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 06:11:14