Scrapy CSS选择器在商品详情页不通用,如何统一匹配商品描述?
解决Scrapy爬取商品描述的通用选择器问题
针对你遇到的不同商品详情页.pd_description位置不固定、固定索引选择器失效的问题,提供以下几种可靠解决方案:
方案1:通过描述标题定位目标容器
该网站商品描述区域前会有**"Beschreibung"**(德语“描述”)的标题,可先定位此标题,再选取其后续的描述容器:
- CSS选择器(若Scrapy对
:contains()支持有限,可改用XPath):
div:contains("Beschreibung") + .pd_description
- 更稳定的XPath写法:
//*[contains(text(), 'Beschreibung')]/following-sibling::div[contains(@class, 'pd_description')][1]
方案2:选取父容器内最后一个.pd_description
观察页面结构可知,商品描述是目标父容器内的最后一个.pd_description元素(前面的同class元素为产品参数项),直接选取最后一个即可:
#inner > div > div.col-lg-12-full.col-md-12-full div.pd_description:last-of-type
此选择器不依赖固定索引,能适配不同页面的元素数量变化。
方案3:基于内容长度筛选有效描述
若前两种方案仍有问题,可通过筛选文本长度区分参数项(短文本)和商品描述(长文本):
def parse_product(self, response): # 提取所有.pd_description的文本内容并去重整理 desc_candidates = [desc.strip() for desc in response.css('.pd_description::text').extract() if desc.strip()] # 筛选长度符合要求的内容(可根据实际情况调整阈值) description = next((d for d in desc_candidates if len(d) > 80), '') yield { "brand": response.css('div.pd_inforow:nth-of-type(4) span::text').extract_first(), "item_name": response.css("h1::text").extract_first(), "description": description }
代码修复提示
原代码中extract_first缺少调用括号,应改为extract_first(),否则会返回方法对象而非实际文本内容。
内容的提问来源于stack exchange,提问作者Legion
相关产品推荐
相关产品推荐

