无法获取HTML列表正确XPath选择器,返回空列表问题求助
解决Scrapy爬取产品属性列表的XPath问题
嘿,我来帮你搞定这个属性列表爬取的问题!你现在得到空列表的核心原因是遍历的节点选错了,咱们一步步来修正:
问题分析
看你提供的目标HTML,每个产品属性都独立包裹在div.ty-product-feature标签里,所有属性项都嵌套在div#content_features容器下。但你当前的代码是直接遍历//div[@id="content_features"]——这个容器本身只有一个,所以循环只会执行一次,而且你在容器内直接取ty-product-feature__label,只能拿到第一个属性的标签,自然没法得到完整的属性列表。
修正方案
你需要遍历的是每个单独的属性项,也就是//div[@id="content_features"]/div[@class="ty-product-feature"],然后在每个属性项内部分别提取标签和对应的值。
修改后的核心代码片段
把你parse_products方法里的properties部分替换成下面的代码:
item['properties'] = [] # 遍历每个属性项,而非整个容器 for prop in response.xpath('//div[@id="content_features"]/div[@class="ty-product-feature"]'): item['properties'].append({ 'name': prop.xpath('normalize-space(.//span[@class="ty-product-feature__label"])').get(), 'value': prop.xpath('normalize-space(.//div[@class="ty-product-feature__value"])').get() })
额外优化提示
- 用
.get()替代extract_first():Scrapy 1.5+版本推荐使用get()和getall(),语法更简洁直观。 - 优化品牌字段提取:你原来的
brand字段XPath不够可靠,可以直接从属性列表中匹配提取,避免依赖其他不稳定节点:# 在属性遍历过程中同步提取品牌 for prop in response.xpath('//div[@id="content_features"]/div[@class="ty-product-feature"]'): prop_name = prop.xpath('normalize-space(.//span[@class="ty-product-feature__label"])').get() prop_value = prop.xpath('normalize-space(.//div[@class="ty-product-feature__value"])').get() item['properties'].append({'name': prop_name, 'value': prop_value}) if prop_name == 'Бренды:': item['brand'] = prop_value
完整修改后的parse_products方法
def parse_products(self, response): # 单个产品页无需循环产品区块,直接提取即可 item = dict() item['title'] = response.xpath('//h1[@class="ty-product-block-title"]/text()').get() item['price'] = response.xpath('//meta[@itemprop="price"]/@content').get() item['available'] = response.xpath('normalize-space(//span[@id="in_stock_info_5511"])').get() item['image'] = response.xpath('//meta[@property="og:image"]/@content').get() item['department'] = response.xpath('normalize-space(//a[@class="ty-breadcrumbs__a"][2]/text())').get() # 提取完整属性列表并同步获取品牌 item['properties'] = [] item['brand'] = None for prop in response.xpath('//div[@id="content_features"]/div[@class="ty-product-feature"]'): prop_name = prop.xpath('normalize-space(.//span[@class="ty-product-feature__label"])').get() prop_value = prop.xpath('normalize-space(.//div[@class="ty-product-feature__value"])').get() item['properties'].append({'name': prop_name, 'value': prop_value}) if prop_name == 'Бренды:': item['brand'] = prop_value yield item
这样修改后,你就能正确获取到所有产品属性的完整列表啦!
内容的提问来源于stack exchange,提问作者romankk
相关产品推荐
相关产品推荐

