如何用Python Scrapy XPath访问< script type="text/x-magento-init">标签并提取产品图片
解决Scrapy XPath访问Magento的
<script type="text/x-magento-init">标签并提取图片的方法 嘿,我来帮你搞定这个问题!Magento的这个script标签里藏着产品的配置数据(包括图片链接),但它不是直接的HTML元素内容,而是JSON格式的文本,所以直接用XPath拿不到里面的图片数据,得先把文本取出来再解析。下面是针对目标页面的具体步骤和代码:
步骤拆解
- 定位目标script标签:先用XPath找到带有
type="text/x-magento-init"属性的script标签,并且筛选出包含产品画廊(gallery)数据的那个(页面可能有多个同类型标签)。 - 解析JSON文本:把标签里的文本内容转换成Python字典,这样就能轻松遍历提取数据了。
- 提取完整图片URL:从解析后的字典里找到图片数组,取出每个图片的完整URL(如果是相对路径就自动拼接域名)。
具体Scrapy代码示例
import json import scrapy import re class LidlProductSpider(scrapy.Spider): name = 'lidl_product' start_urls = ['https://sortiment.lidl.ch/de/papier-tragetasche-fsc-0123997.html'] def parse(self, response): # 定位包含gallery数据的script标签(过滤掉其他同类型标签) script_text = response.xpath( '//script[@type="text/x-magento-init" and contains(text(), "gallery")]/text()' ).get() if not script_text: self.logger.warning("没找到包含gallery数据的script标签") return # 清理文本中可能存在的注释(避免JSON解析失败) cleaned_text = re.sub(r'//.*?\n|/\*[\s\S]*?\*/', '', script_text) try: # 把文本解析成Python字典 magento_config = json.loads(cleaned_text) # 找到gallery对应的配置项(Magento的键通常是带选择器的字符串) gallery_config_key = next( (key for key in magento_config if 'gallery-placeholder' in key), None ) if not gallery_config_key: self.logger.warning("没找到gallery配置项") return # 提取图片数组 gallery_data = magento_config[gallery_config_key]['mage/gallery/gallery']['data']['images'] # 提取完整图片URL(处理相对路径) full_image_urls = [ response.urljoin(img['img']) for img in gallery_data ] # 输出结果或者yield到Item里 self.logger.info(f"成功提取到{len(full_image_urls)}张图片:") for url in full_image_urls: self.logger.info(url) yield { 'product_url': response.url, 'full_image_urls': full_image_urls } except json.JSONDecodeError as e: self.logger.error(f"JSON解析失败: {str(e)}") except KeyError as e: self.logger.error(f"找不到目标字段: {str(e)}")
关键注意点
- 标签筛选:一定要用
contains(text(), "gallery")来过滤,不然可能拿到其他无关的magento-init脚本(比如导航、搜索的配置)。 - JSON清理:有些Magento页面的script里会有单行或多行注释,直接解析会报错,所以先用正则去掉注释。
- 相对路径处理:如果图片URL是相对路径(比如
/media/catalog/product/...),用response.urljoin()自动拼接成完整的HTTPS链接。
运行这个爬虫后,你就能拿到目标页面里的两张完整产品图片URL啦!
内容的提问来源于stack exchange,提问作者akmal Khan
相关产品推荐
相关产品推荐

