Scrapy下载图片重命名:用商品路径名替代哈希名并获取首图
Scrapy图片爬取需求实现方案
需求1:自定义图片文件名
单图命名为LNX-500A.jpg(取自商品URL最后部分),多图命名为LNX_500A_01.jpg(横杠转下划线+两位序号)。
需求2:仅爬取页面第一张图片
一、修改Spider代码(实现仅爬第一张图+传递商品URL)
在商品详情页解析函数中,只提取第一个图片URL,同时把商品URL存入item,用于后续生成文件名:
import scrapy class ProductSpider(scrapy.Spider): name = 'product_spider' allowed_domains = ['your-target-domain.com'] start_urls = ['https://your-target-domain.com/product-list'] def parse(self, response): # 遍历商品列表链接 product_links = response.xpath('//a[contains(@class, "product-item")]/@href').getall() for link in product_links: yield response.follow(link, callback=self.parse_product) def parse_product(self, response): item = {} # 存储商品URL,用于生成文件名 item['product_url'] = response.url # 只提取页面第一张图片的URL main_image_url = response.xpath('//img[contains(@class, "main-product-img")]/@src').get() if main_image_url: item['image_urls'] = [main_image_url] yield item
说明:
- 用
get()替代getall()直接获取第一个图片URL,再包装成列表存入image_urls(ImagesPipeline要求该字段为列表格式)。 product_url字段用于后续Pipeline提取文件名的基础字符串。
二、自定义ImagesPipeline(实现文件名规则)
重写Scrapy默认的ImagesPipeline,修改file_path方法生成自定义文件名,同时重写get_media_requests传递图片索引:
from scrapy.pipelines.images import ImagesPipeline from urllib.parse import urlparse import re class CustomImagePipeline(ImagesPipeline): def get_media_requests(self, item, info): # 给每个图片请求传递索引和商品URL for idx, url in enumerate(item['image_urls']): yield scrapy.Request(url, meta={'image_index': idx, 'product_url': item['product_url']}) def file_path(self, request, response=None, info=None, *, item=None): # 提取商品URL的最后部分作为基础名称 raw_name = request.meta['product_url'].split('/')[-1] # 清理名称中的非法字符(避免文件系统报错) clean_name = re.sub(r'[^\w\-]', '', raw_name) # 获取图片扩展名 if response: ext = response.headers.get('Content-Type').decode().split('/')[-1] else: ext = urlparse(request.url).path.split('.')[-1] or 'jpg' # 根据图片索引生成文件名 img_idx = request.meta['image_index'] if img_idx == 0: # 单图:直接用清理后的名称+扩展名 return f"{clean_name}.{ext}" else: # 多图:横杠转下划线+两位序号 formatted_name = clean_name.replace('-', '_') return f"{formatted_name}_{img_idx+1:02d}.{ext}"
说明:
get_media_requests:给每个图片请求添加image_index(序号)和product_url元数据,方便file_path方法调用。file_path:- 从商品URL提取最后部分,并用正则清理非法字符(如
?、&等)。 - 优先从响应头获取图片扩展名(更准确), fallback到URL中的扩展名。
- 单图直接用清理后的名称;多图将横杠替换为下划线,添加两位序号(如
01、02)。
- 从商品URL提取最后部分,并用正则清理非法字符(如
三、配置settings.py启用自定义Pipeline
替换默认的ImagesPipeline,同时配置图片存储路径:
# 启用自定义图片Pipeline ITEM_PIPELINES = { 'your_project_name.pipelines.CustomImagePipeline': 1, } # 配置图片保存目录(可自行修改路径) IMAGES_STORE = './product_images'
内容的提问来源于stack exchange,提问作者Shadobladez
相关产品推荐
相关产品推荐

