You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy项目中如何初始化数据库连接实现多模块共享访问

最优实现方案:全局复用数据库连接

核心思路

通过单例模式封装数据库连接,配合Scrapy扩展管理连接生命周期,确保整个爬虫项目仅初始化一次数据库连接,在爬虫、管道、自定义StatsCollector中共享使用。


步骤1:创建独立的数据库连接管理模块

在my_project/my_project/目录下新建db_connection.py,用单例模式封装连接逻辑,同时从Scrapy配置中读取数据库参数(避免硬编码):

import psycopg2
from psycopg2.extras import RealDictCursor
from scrapy.utils.project import get_project_settings

class DBConnection:
    _instance = None
    connection = None

    def __new__(cls):
        # 单例模式:仅首次实例化时创建连接
        if cls._instance is None:
            cls._instance = super().__new__(cls)
            settings = get_project_settings()
            cls._instance.connection = psycopg2.connect(
                host=settings.get('DB_HOST'),
                user=settings.get('DB_USER'),
                password=settings.get('DB_PASSWORD'),
                dbname=settings.get('DB_NAME')
            )
        return cls._instance

    def get_cursor(self):
        # 返回支持字典格式的游标,方便数据处理
        return self.connection.cursor(cursor_factory=RealDictCursor)

    def close(self):
        # 关闭连接,重置单例实例
        if self.connection:
            self.connection.close()
            self._instance = None

步骤2:在Settings中配置数据库参数

修改my_project/my_project/settings.py,添加数据库配置:

# 数据库配置项
DB_HOST = '你的数据库地址'
DB_USER = '用户名'
DB_PASSWORD = '密码'
DB_NAME = '数据库名'

步骤3:用Scrapy扩展管理连接生命周期

新建my_project/my_project/extensions.py,通过Scrapy信号在爬虫启动时初始化连接,关闭时自动销毁:

from scrapy import signals
from scrapy.exceptions import NotConfigured
from .db_connection import DBConnection

class DBConnectionExtension:
    def __init__(self, crawler):
        self.crawler = crawler
        # 初始化数据库连接
        self.db = DBConnection()
        # 注册爬虫关闭信号,确保连接被释放
        crawler.signals.connect(self.spider_closed, signal=signals.spider_closed)

    @classmethod
    def from_crawler(cls, crawler):
        # 校验配置完整性,避免遗漏参数
        required_settings = ['DB_HOST', 'DB_USER', 'DB_PASSWORD', 'DB_NAME']
        if not all(crawler.settings.get(key) for key in required_settings):
            raise NotConfigured("数据库配置参数不全,请检查settings.py")
        return cls(crawler)

    def spider_closed(self, spider):
        self.db.close()

然后在settings.py中启用该扩展:

EXTENSIONS = {
    'my_project.extensions.DBConnectionExtension': 500,
}

步骤4:在各组件中复用连接

1. 管道(pipelines.py)

from .db_connection import DBConnection

class SaveToPostgresPipeline:
    def __init__(self):
        self.db = DBConnection()
        self.cursor = self.db.get_cursor()

    def process_item(self, item, spider):
        # 示例:插入数据到数据库
        insert_sql = """INSERT INTO target_table (field1, field2) VALUES (%s, %s)"""
        self.cursor.execute(insert_sql, (item['field1'], item['field2']))
        self.db.connection.commit()
        return item

    def close_spider(self, spider):
        self.cursor.close()

2. 爬虫(如spider1.py)

import scrapy
from my_project.db_connection import DBConnection

class Spider1(scrapy.Spider):
    name = 'spider1'

    def __init__(self, *args, **kwargs):
        super().__init__(*args, **kwargs)
        self.db = DBConnection()
        self.cursor = self.db.get_cursor()

    def start_requests(self):
        # 示例:从数据库读取起始URL
        self.cursor.execute("SELECT url FROM start_urls WHERE spider_name = %s", (self.name,))
        urls = [row['url'] for row in self.cursor.fetchall()]
        for url in urls:
            yield scrapy.Request(url, callback=self.parse)

    def parse(self, response):
        # 数据处理逻辑...
        pass

    def close(self, reason):
        self.cursor.close()

3. 自定义StatsCollector(MyStatsCollector.py)

from scrapy.statscollectors import StatsCollector
from .db_connection import DBConnection

class MyStatsCollector(StatsCollector):
    def __init__(self, crawler):
        super().__init__(crawler)
        self.db = DBConnection()
        self.cursor = self.db.get_cursor()

    def _persist_stats(self, stats, spider):
        # 示例:将爬虫统计数据存入数据库
        insert_sql = """INSERT INTO spider_stats (spider_name, stats_data, created_at) VALUES (%s, %s, NOW())"""
        self.cursor.execute(insert_sql, (spider.name, str(stats)))
        self.db.connection.commit()
        self.cursor.close()

最后在settings.py中指定自定义StatsCollector:

STATS_CLASS = 'my_project.MyStatsCollector.MyStatsCollector'

关键优势

  1. 全局单例连接:整个爬虫生命周期仅创建一次数据库连接,避免重复初始化开销
  2. 配置集中管理:数据库参数统一放在settings.py,便于维护和修改
  3. 生命周期自动管理:通过Scrapy扩展和信号,确保爬虫关闭时自动释放连接,避免资源泄漏
  4. 代码复用性高:所有组件通过导入DBConnection即可使用同一连接,无需重复编写连接逻辑

内容的提问来源于stack exchange,提问作者user984621

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 19:52:56