You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy技术问题:如何从数据库读取链接作为爬虫起始请求地址

问题分析与解决方案

我看了你的代码,核心问题出在复用Pipeline类读取数据时的意外副作用,以及不太规范的数据库操作方式上,下面一步步给你拆解和修正:

为什么你的代码无法正常运行?

  1. 表被意外删除:你的HeurekaScraperPipeline的create_table方法里执行了DROP TABLE IF EXISTS shops_tb——当第二个爬虫实例化这个Pipeline时,会直接删掉之前第一个爬虫存入数据的表,自然查不到任何链接。
  2. Pipeline职责错位:Pipeline的设计初衷是处理Item的存储,直接在爬虫里复用它来读数据,会引入不必要的初始化逻辑(比如重复建表),也不符合代码解耦的原则。

修正步骤

1. 先修复Pipeline,避免误删数据

修改pipelines.py,把强制删表的逻辑去掉,改为仅当表不存在时创建:

import sqlite3
class HeurekaScraperPipeline:
    def __init__(self):
        self.create_connection()
        self.create_table()
    def create_connection(self):
        self.conn = sqlite3.connect('shops.db')
        self.curr = self.conn.cursor()
    def create_table(self):
        # 替换成CREATE TABLE IF NOT EXISTS,不再删除已存在的表
        self.curr.execute("""create table IF NOT EXISTS shops_tb(
            product_name text,
            shop_name text,
            price text,
            link text
        )""")
    def process_item(self, item, spider):
        self.store_db(item)
        return item
    def store_db(self, item):
        self.curr.execute("""insert into shops_tb values (?, ?, ?, ?)""",(
            item['product_name'],
            item['shop_name'],
            item['price'],
            item['link'],
        ))
        self.conn.commit()

2. 写一个独立的数据库工具类,专门用于读取链接

在你的Scrapy项目根目录下新建db_utils.py,封装数据库读取逻辑:

import sqlite3

def get_shop_links():
    # 建立数据库连接
    conn = sqlite3.connect('shops.db')
    curr = conn.cursor()
    # 只查询link字段,更高效
    curr.execute("SELECT link FROM shops_tb")
    # 提取所有链接并转为列表
    links = [row[0] for row in curr.fetchall()]
    # 记得关闭连接,避免资源泄漏
    conn.close()
    return links

3. 修改爬虫代码,使用工具类读取链接

import scrapy
# 替换成你的项目名称,确保能导入db_utils
from your_project_name.db_utils import get_shop_links

class Shops_spider(scrapy.Spider):
    name = 'shops_scraper'
    custom_settings = {'DOWNLOAD_DELAY': 1}
    
    def start_requests(self):
        # 调用工具类获取链接
        links = get_shop_links()
        for url in links:
            # 过滤空链接,避免无效请求
            if url:
                print(f"准备爬取: {url}")
                yield scrapy.Request(url=url, callback=self.parse)
    
    def parse(self, response):
        url = response.request.url
        print(f'********************************{url}************************')

额外注意事项

  • 先运行第一个爬虫,确保数据成功存入shops.db后,再启动第二个爬虫。
  • 可以用SQLite可视化工具(比如DB Browser for SQLite)打开shops.db,确认shops_tb表里有数据,排除数据库本身的问题。
  • 如果你的数据库路径不是项目根目录,要注意修改sqlite3.connect()里的路径。

内容的提问来源于stack exchange,提问作者maros papan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 21:42:30