You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python Scrapy XPath访问< script type="text/x-magento-init">标签并提取产品图片

解决Scrapy XPath访问Magento的<script type="text/x-magento-init">标签并提取图片的方法

嘿,我来帮你搞定这个问题!Magento的这个script标签里藏着产品的配置数据(包括图片链接),但它不是直接的HTML元素内容,而是JSON格式的文本,所以直接用XPath拿不到里面的图片数据,得先把文本取出来再解析。下面是针对目标页面的具体步骤和代码:

步骤拆解

  1. 定位目标script标签:先用XPath找到带有type="text/x-magento-init"属性的script标签,并且筛选出包含产品画廊(gallery)数据的那个(页面可能有多个同类型标签)。
  2. 解析JSON文本:把标签里的文本内容转换成Python字典,这样就能轻松遍历提取数据了。
  3. 提取完整图片URL:从解析后的字典里找到图片数组,取出每个图片的完整URL(如果是相对路径就自动拼接域名)。

具体Scrapy代码示例

import json
import scrapy
import re

class LidlProductSpider(scrapy.Spider):
    name = 'lidl_product'
    start_urls = ['https://sortiment.lidl.ch/de/papier-tragetasche-fsc-0123997.html']

    def parse(self, response):
        # 定位包含gallery数据的script标签(过滤掉其他同类型标签)
        script_text = response.xpath(
            '//script[@type="text/x-magento-init" and contains(text(), "gallery")]/text()'
        ).get()

        if not script_text:
            self.logger.warning("没找到包含gallery数据的script标签")
            return

        # 清理文本中可能存在的注释(避免JSON解析失败)
        cleaned_text = re.sub(r'//.*?\n|/\*[\s\S]*?\*/', '', script_text)

        try:
            # 把文本解析成Python字典
            magento_config = json.loads(cleaned_text)

            # 找到gallery对应的配置项(Magento的键通常是带选择器的字符串)
            gallery_config_key = next(
                (key for key in magento_config if 'gallery-placeholder' in key),
                None
            )

            if not gallery_config_key:
                self.logger.warning("没找到gallery配置项")
                return

            # 提取图片数组
            gallery_data = magento_config[gallery_config_key]['mage/gallery/gallery']['data']['images']
            # 提取完整图片URL(处理相对路径)
            full_image_urls = [
                response.urljoin(img['img']) for img in gallery_data
            ]

            # 输出结果或者yield到Item里
            self.logger.info(f"成功提取到{len(full_image_urls)}张图片:")
            for url in full_image_urls:
                self.logger.info(url)

            yield {
                'product_url': response.url,
                'full_image_urls': full_image_urls
            }

        except json.JSONDecodeError as e:
            self.logger.error(f"JSON解析失败: {str(e)}")
        except KeyError as e:
            self.logger.error(f"找不到目标字段: {str(e)}")

关键注意点

  • 标签筛选:一定要用contains(text(), "gallery")来过滤,不然可能拿到其他无关的magento-init脚本(比如导航、搜索的配置)。
  • JSON清理:有些Magento页面的script里会有单行或多行注释,直接解析会报错,所以先用正则去掉注释。
  • 相对路径处理:如果图片URL是相对路径(比如/media/catalog/product/...),用response.urljoin()自动拼接成完整的HTTPS链接。

运行这个爬虫后,你就能拿到目标页面里的两张完整产品图片URL啦!

内容的提问来源于stack exchange,提问作者akmal Khan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 08:04:17