Scrapy爬取劳力士手表数据无法通过Pipeline存入MySQL问题求助
问题根源与修复方案
1. 爬虫无法生成Item的核心原因
你写的parse和parse_categories方法缩进错误,不属于WatchbotSpider类的成员方法:Scrapy默认调用爬虫类的parse方法处理响应,你的代码里这两个方法是全局函数,爬虫运行时根本不会执行,自然没有Item传递给Pipeline。
修正方法:把两个方法的缩进调整到WatchbotSpider类内部,和name、start_urls属性对齐。
2. Pipeline的多处错误
- 依赖导入错误:只导入
mysql无法使用连接器,需要改为import mysql.connector - 字段名大小写不匹配:Items中定义的字段是小写的
year/itemnr/reference,你在Pipeline中使用大写的item["YEAR"]会直接抛出KeyError - 多余的下标取值:你在爬虫代码中已经用
extract()[0]拿到了字符串类型的字段值,Pipeline中再加[0]会取到字符串的第一个字符,完全不符合存储需求 - SQL语法错误:插入语句的格式错误,正确的单条插入语法是
insert into 表名(字段1,字段2,字段3) values (%s,%s,%s),你写的逗号分隔多组括号的格式完全错误 - 数据库连接参数为空:
host/user/passwd/database需要填入你本地MySQL的真实配置信息
3. Items.py的冗余代码
print(itemnr)是类属性定义阶段的多余代码,不会影响运行但建议删除。
修正后的代码示例
修正后的爬虫代码
import scrapy from watches.watches.items import WatchesItem class WatchbotSpider(scrapy.Spider): name = "watchbot" start_urls = ["https://www.watch.de/english/rolex.html"] # 缩进调整到类内部 def parse(self, response, **kwargs): for link in response.css("div.product-item-link a::attr(href)"): url = link.get() yield scrapy.Request(url, callback=self.parse_categories) # 缩进调整到类内部 def parse_categories(self, response): item = WatchesItem() # 用extract_first()代替extract()[0],避免空值报错 item["itemnr"] = response.xpath('//span[@itemprop="sku"]/text()').extract_first().strip() item["reference"] = response.xpath('//span[@itemprop="mpn"]/text()').extract_first().strip() item["year"] = response.xpath( '//div[@class="product-option baujahr"]/div[@class="product-option-value"]/text()' ).extract_first().strip() yield item
修正后的Pipeline代码
import mysql.connector class WatchesPipeline(object): def __init__(self): # 填入你自己的MySQL配置 self.conn = mysql.connector.connect( host="127.0.0.1", user="你的数据库用户名", passwd="你的数据库密码", database="你的数据库名", charset='utf8mb4' ) self.curr = self.conn.cursor() def process_item(self, item, spider): self.store_db(item) return item def store_db(self, item): # 字段名要和你数据库中watches表的字段名对应 self.curr.execute( """insert into watches(year, reference, itemnr) values (%s, %s, %s)""", (item["year"], item["reference"], item["itemnr"]) ) self.conn.commit() def close_spider(self, spider): self.curr.close() self.conn.close()
修正后的Items代码
import scrapy class WatchesItem(scrapy.Item): year = scrapy.Field() itemnr = scrapy.Field() reference = scrapy.Field()
验证步骤
- 先执行
scrapy crawl watchbot -o test.csv,确认csv文件中能正常输出数据,说明Item生成和传递逻辑正常 - 提前在MySQL中创建好
watches表,确保字段类型、顺序和插入语句匹配 - 重新运行爬虫,即可正常写入数据库
内容的提问来源于stack exchange,提问作者SyrixGG
相关产品推荐
相关产品推荐

