You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬取劳力士手表数据无法通过Pipeline存入MySQL问题求助

问题根源与修复方案

1. 爬虫无法生成Item的核心原因

你写的parse和parse_categories方法缩进错误,不属于WatchbotSpider类的成员方法:Scrapy默认调用爬虫类的parse方法处理响应,你的代码里这两个方法是全局函数,爬虫运行时根本不会执行,自然没有Item传递给Pipeline。
修正方法:把两个方法的缩进调整到WatchbotSpider类内部,和name、start_urls属性对齐。

2. Pipeline的多处错误

  • 依赖导入错误:只导入mysql无法使用连接器,需要改为import mysql.connector
  • 字段名大小写不匹配:Items中定义的字段是小写的year/itemnr/reference,你在Pipeline中使用大写的item["YEAR"]会直接抛出KeyError
  • 多余的下标取值:你在爬虫代码中已经用extract()[0]拿到了字符串类型的字段值,Pipeline中再加[0]会取到字符串的第一个字符,完全不符合存储需求
  • SQL语法错误:插入语句的格式错误,正确的单条插入语法是insert into 表名(字段1,字段2,字段3) values (%s,%s,%s),你写的逗号分隔多组括号的格式完全错误
  • 数据库连接参数为空:host/user/passwd/database需要填入你本地MySQL的真实配置信息

3. Items.py的冗余代码

print(itemnr)是类属性定义阶段的多余代码,不会影响运行但建议删除。

修正后的代码示例

修正后的爬虫代码

import scrapy
from watches.watches.items import WatchesItem

class WatchbotSpider(scrapy.Spider):
    name = "watchbot"
    start_urls = ["https://www.watch.de/english/rolex.html"]

    # 缩进调整到类内部
    def parse(self, response, **kwargs):
        for link in response.css("div.product-item-link a::attr(href)"):
            url = link.get()
            yield scrapy.Request(url, callback=self.parse_categories)

    # 缩进调整到类内部
    def parse_categories(self, response):
        item = WatchesItem()
        # 用extract_first()代替extract()[0],避免空值报错
        item["itemnr"] = response.xpath('//span[@itemprop="sku"]/text()').extract_first().strip()
        item["reference"] = response.xpath('//span[@itemprop="mpn"]/text()').extract_first().strip()
        item["year"] = response.xpath(
            '//div[@class="product-option baujahr"]/div[@class="product-option-value"]/text()'
        ).extract_first().strip()
        yield item

修正后的Pipeline代码

import mysql.connector

class WatchesPipeline(object):
    def __init__(self):
        # 填入你自己的MySQL配置
        self.conn = mysql.connector.connect(
            host="127.0.0.1", 
            user="你的数据库用户名", 
            passwd="你的数据库密码", 
            database="你的数据库名",
            charset='utf8mb4'
        )
        self.curr = self.conn.cursor()

    def process_item(self, item, spider):
        self.store_db(item)
        return item

    def store_db(self, item):
        # 字段名要和你数据库中watches表的字段名对应
        self.curr.execute(
            """insert into watches(year, reference, itemnr) values (%s, %s, %s)""",
            (item["year"], item["reference"], item["itemnr"])
        )
        self.conn.commit()
    
    def close_spider(self, spider):
        self.curr.close()
        self.conn.close()

修正后的Items代码

import scrapy

class WatchesItem(scrapy.Item):
    year = scrapy.Field()
    itemnr = scrapy.Field()
    reference = scrapy.Field()

验证步骤

  1. 先执行scrapy crawl watchbot -o test.csv,确认csv文件中能正常输出数据,说明Item生成和传递逻辑正常
  2. 提前在MySQL中创建好watches表,确保字段类型、顺序和插入语句匹配
  3. 重新运行爬虫,即可正常写入数据库

内容的提问来源于stack exchange,提问作者SyrixGG

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 23:27:04