You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy亚马逊维生素爬虫修改问题:无法提取产品成分数据

解决亚马逊产品成分提取为空的问题

你的问题核心在于CSS选择器使用错误,同时亚马逊产品页面的成分信息需要更精准的定位逻辑。以下是具体修复方案:

1. 错误根源分析

你写的response.css('important-information::text')是在寻找名为<important-information>的HTML标签,但亚马逊页面中根本不存在这个标签——important-information是类名,正确的选择器应该以.开头(类选择器)。即使修正这个,也可能提取不到数据,因为成分信息通常嵌套在该板块的子元素里,而非直接在类容器的文本节点中。

2. 可行的提取方案

方案一:直接定位成分子元素

亚马逊OTC维生素产品的成分通常在「Important Information」板块下的列表项中,用以下选择器可精准提取:

# 提取板块下所有有效文本,过滤空内容并整理格式
ingredients_items = response.css('.important-information .a-list-item::text').getall()
ingredients_data = '\n'.join([item.strip() for item in ingredients_items if item.strip()])

方案二:从页面脚本提取JSON数据

亚马逊大量产品详情会嵌入在页面的<script>标签中,可通过正则匹配包含成分的JSON块(适配更多页面结构):

import json
import re

# 定位包含成分信息的脚本
script_text = response.css('script:contains("ingredients")::text').get()
ingredients_data = ''
if script_text:
    # 匹配产品详情JSON结构(根据页面实际情况调整正则)
    match = re.search(r'window\.productDetailPage\s*=\s*({.*?});', script_text, re.DOTALL)
    if match:
        product_json = json.loads(match.group(1))
        # 从JSON层级中提取成分(路径可能因产品略有差异)
        ingredients_data = product_json.get('productInfo', {}).get('content', {}).get('ingredients', '')

3. 修改后的完整parse_product_data方法

这里采用方案一的代码替换你原有的成分提取逻辑:

def parse_product_data(self, response):
    image_data = json.loads(re.findall(r"colorImages':.*'initial':\s*(\[.+?\])},\n", response.text)[0])
    variant_data = re.findall(r'dimensionValuesDisplayData"\s*:\s* ({.+?}),\n', response.text)
    feature_bullets = [bullet.strip() for bullet in response.css("#feature-bullets li ::text").getall()]
    price = response.css('.a-price span[aria-hidden="true"] ::text').get("")

    # 修正后的成分提取代码
    ingredients_items = response.css('.important-information .a-list-item::text').getall()
    ingredients_data = '\n'.join([item.strip() for item in ingredients_items if item.strip()])

    if not price:
        price = response.css('.a-price .a-offscreen ::text').get("")
    yield {
        "name": response.css("#productTitle::text").get("").strip(),
        "price": price,
        "stars": response.css("i[data-hook=average-star-rating] ::text").get("").strip(),
        "rating_count": response.css("div[data-hook=total-review-count] ::text").get("").strip(),
        "feature_bullets": feature_bullets,
        "images": image_data,
        "variant_data": variant_data,
        "ingredients": ingredients_data
    }

4. 额外注意事项

  • 反爬适配:亚马逊会检测爬虫行为,建议在settings.py中配置真实浏览器UA并添加下载延迟:
    USER_AGENT = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
    DOWNLOAD_DELAY = 2
    
  • 页面差异处理:不同产品的页面结构可能略有不同,如果方案一提取失败,可切换方案二,或用浏览器开发者工具(F12)手动定位成分元素的选择器。

内容的提问来源于stack exchange,提问作者David Lee

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 09:13:23