Scrapy实现按库存状态生成缺货标记的技术问询
问题:Scrapy爬虫中根据库存状态生成缺货状态列表
我正在编写Scrapy爬虫,目标是从指定商品页提取数据,需求是仅当商品的data-inventory-qty属性值为0(缺货)时,在SKU数据中添加out_of_stock:True字段,否则不显示该字段。目前已经掌握currency、price、size的提取逻辑,只需要完善缺货状态的提取逻辑,生成对应状态列表(比如['','True','','True']这种格式)。现有部分代码:
items=response.xpath('//div//form//div//fieldset//input[@data-inventory-qty]').getall()
需要完成这部分逻辑。
解决方案
1. 生成缺货状态列表
首先需要提取data-inventory-qty的属性值,而非直接获取元素本身,再遍历判断生成目标列表:
# 提取所有input元素的data-inventory-qty属性值 inventory_qtys = response.xpath('//div//form//div//fieldset//input[@data-inventory-qty]/@data-inventory-qty').getall() # 生成对应缺货状态列表 out_of_stock_list = [] for qty in inventory_qtys: try: qty_int = int(qty) # 库存为0时添加'True',否则添加空字符串 out_of_stock_list.append('True' if qty_int == 0 else '') except ValueError: # 处理非数字的异常情况,默认按有货处理 out_of_stock_list.append('')
2. 直接整合到SKU字典(更贴合需求场景)
如果需要直接生成包含缺货字段的SKU数据(仅缺货时显示out_of_stock:True),可以结合已有的size字段一起处理:
# 提取size列表(假设size对应input的value属性) sizes = response.xpath('//div//form//div//fieldset//input[@data-inventory-qty]/@value').getall() inventory_qtys = response.xpath('//div//form//div//fieldset//input[@data-inventory-qty]/@data-inventory-qty').getall() sku_list = [] for size, qty in zip(sizes, inventory_qtys): sku_item = { 'size': size, # 这里加入你已实现的currency、price字段 } try: if int(qty) == 0: sku_item['out_of_stock'] = True except ValueError: # 异常情况不添加缺货字段 pass sku_list.append(sku_item)
关键说明
- 使用
/@data-inventory-qty直接定位属性值,避免处理整个HTML元素的冗余内容 - 加入异常处理,防止页面属性值非数字导致爬虫崩溃
- 严格遵循需求:仅在库存为0时添加
out_of_stock字段,其余情况不显示该字段
内容的提问来源于stack exchange,提问作者Malik Zohaib Mustafa
相关产品推荐
相关产品推荐

