Scrapy亚马逊维生素爬虫修改问题:无法提取产品成分数据
解决亚马逊产品成分提取为空的问题
你的问题核心在于CSS选择器使用错误,同时亚马逊产品页面的成分信息需要更精准的定位逻辑。以下是具体修复方案:
1. 错误根源分析
你写的response.css('important-information::text')是在寻找名为<important-information>的HTML标签,但亚马逊页面中根本不存在这个标签——important-information是类名,正确的选择器应该以.开头(类选择器)。即使修正这个,也可能提取不到数据,因为成分信息通常嵌套在该板块的子元素里,而非直接在类容器的文本节点中。
2. 可行的提取方案
方案一:直接定位成分子元素
亚马逊OTC维生素产品的成分通常在「Important Information」板块下的列表项中,用以下选择器可精准提取:
# 提取板块下所有有效文本,过滤空内容并整理格式 ingredients_items = response.css('.important-information .a-list-item::text').getall() ingredients_data = '\n'.join([item.strip() for item in ingredients_items if item.strip()])
方案二:从页面脚本提取JSON数据
亚马逊大量产品详情会嵌入在页面的<script>标签中,可通过正则匹配包含成分的JSON块(适配更多页面结构):
import json import re # 定位包含成分信息的脚本 script_text = response.css('script:contains("ingredients")::text').get() ingredients_data = '' if script_text: # 匹配产品详情JSON结构(根据页面实际情况调整正则) match = re.search(r'window\.productDetailPage\s*=\s*({.*?});', script_text, re.DOTALL) if match: product_json = json.loads(match.group(1)) # 从JSON层级中提取成分(路径可能因产品略有差异) ingredients_data = product_json.get('productInfo', {}).get('content', {}).get('ingredients', '')
3. 修改后的完整parse_product_data方法
这里采用方案一的代码替换你原有的成分提取逻辑:
def parse_product_data(self, response): image_data = json.loads(re.findall(r"colorImages':.*'initial':\s*(\[.+?\])},\n", response.text)[0]) variant_data = re.findall(r'dimensionValuesDisplayData"\s*:\s* ({.+?}),\n', response.text) feature_bullets = [bullet.strip() for bullet in response.css("#feature-bullets li ::text").getall()] price = response.css('.a-price span[aria-hidden="true"] ::text').get("") # 修正后的成分提取代码 ingredients_items = response.css('.important-information .a-list-item::text').getall() ingredients_data = '\n'.join([item.strip() for item in ingredients_items if item.strip()]) if not price: price = response.css('.a-price .a-offscreen ::text').get("") yield { "name": response.css("#productTitle::text").get("").strip(), "price": price, "stars": response.css("i[data-hook=average-star-rating] ::text").get("").strip(), "rating_count": response.css("div[data-hook=total-review-count] ::text").get("").strip(), "feature_bullets": feature_bullets, "images": image_data, "variant_data": variant_data, "ingredients": ingredients_data }
4. 额外注意事项
- 反爬适配:亚马逊会检测爬虫行为,建议在
settings.py中配置真实浏览器UA并添加下载延迟:USER_AGENT = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' DOWNLOAD_DELAY = 2 - 页面差异处理:不同产品的页面结构可能略有不同,如果方案一提取失败,可切换方案二,或用浏览器开发者工具(F12)手动定位成分元素的选择器。
内容的提问来源于stack exchange,提问作者David Lee
相关产品推荐
相关产品推荐

