Scrapy多爬虫MySQL存储旧数据丢失问题求助
解决Scrapy爬虫运行时MySQL历史数据丢失的问题
嘿,我一眼就揪出问题所在了!你的Pipeline里的create_table方法每次启动爬虫都会先删掉旧表再重建,这可不就把之前的历史数据全清空了嘛😅。咱们一步步来修复这个问题:
1. 核心问题定位
看你Pipeline代码里的这段关键逻辑:
def create_table(self): self.curr.execute(""" DROP TABLE IF EXISTS k_master_movie_reviews """) self.curr.execute(""" create table k_master_movie_reviews( id text, title text, title_es text, url text ) """)
DROP TABLE IF EXISTS这句代码会直接删除已存在的表,然后重新创建一张空表,旧数据自然就没了——这就是问题的根源!
2. 修改Pipeline代码保留历史数据
把create_table方法改成仅在表不存在时创建,删掉删表的语句,换成CREATE TABLE IF NOT EXISTS;同时还可以加个小优化,防止重复数据插入:
import mysql.connector class ReviewBotPipeline(object): def __init__(self): self.create_connection() self.create_table() def create_connection(self): self.conn = mysql.connector.connect( host='localhost', user='root', passwd='Admin', database='review_url' ) self.curr = self.conn.cursor() def create_table(self): # 只在表不存在时创建,不再删除旧表 self.curr.execute(""" CREATE TABLE IF NOT EXISTS k_master_movie_reviews( id text, title text, title_es text, url text, # 给url加唯一键,避免重复插入相同的评论链接 UNIQUE KEY unique_review_url (url) ) """) self.conn.commit() # 提交表创建的操作,确保生效 def process_item(self, outputs, spider): self.store_db(outputs) return outputs def store_db(self, outputs): # 用INSERT IGNORE,遇到重复的url就跳过插入,不会报错中断爬虫 self.curr.execute(""" INSERT IGNORE INTO k_master_movie_reviews (id, title, title_es, url) VALUES (%s, %s, %s, %s) """, (outputs['id'], outputs['title'], outputs['title_es'], outputs['url'])) self.conn.commit() def close_spider(self, spider): self.conn.close()
3. 优化点说明
- 保留历史数据:去掉
DROP TABLE语句后,每次启动爬虫只会复用已有的表,不会清空旧数据。 - 防止重复插入:给
url字段添加唯一键,配合INSERT IGNORE,如果爬虫重复爬取了同一个评论链接,数据库会自动跳过重复数据,不会抛出错误。 - 表创建提交:新增
self.conn.commit()确保表创建的修改能正确生效。
这样修改后,每次运行新爬虫时,新数据会自动追加到已有表中,历史数据就能完整保留啦!
内容的提问来源于stack exchange,提问作者Kumar Divyanshu
相关产品推荐
相关产品推荐

