You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬虫导入主文件后运行异常(无输出/卡顿)求助

Scrapy爬虫单独运行正常,导入main.py后无输出问题排查

问题现象

  • 爬虫脚本单独执行时可正常抓取文本,但作为模块导入main.py后,Spyder IDE持续运行却无任何抓取结果,等待一小时仍无输出,内存占用无异常
  • 控制台仅显示Scrapy弃用警告:

request.py:254: ScrapyDeprecationWarning: '2.6' is a deprecated value for the 'REQUEST_FINGERPRINTER_IMPLEMENTATION' setting.

It is also the default value. In other words, it is normal to get this warning if you have not defined a value for the 'REQUEST_FINGERPRINTER_IMPLEMENTATION' setting. This is so for backward compatibility reasons, but it will change in a future version of Scrapy.

See the documentation of the 'REQUEST_FINGERPRINTER_IMPLEMENTATION' setting for information on how to handle this deprecation.
return cls(crawler)

核心问题分析

  1. stop_after_crawl=False导致进程阻塞
    你在run_wdr_spider函数中使用了process.start(stop_after_crawl=False),该参数会让Scrapy爬虫完成抓取后保持进程运行状态,不会自动退出。当从main.py导入执行时,主线程会被Twisted reactor阻塞,看起来像一直在运行,但实际抓取可能已完成,只是缺少输出反馈导致你误以为无结果。

  2. 项目配置加载异常
    从main.py导入爬虫时,get_project_settings()可能无法正确识别Scrapy项目路径,导致自定义管道bremenspider.pipelines.MongoDBPipeline2无法加载,Item无法被正常处理,没有可视化输出。

  3. Spyder与Twisted reactor兼容性冲突
    Spyder本身依赖Twisted框架,直接在Spyder中运行Scrapy爬虫可能出现reactor资源冲突,导致爬虫无法正常启动或运行无响应。

解决方案

方案1:修改进程停止参数

将run_wdr_spider中的process.start(stop_after_crawl=False)改为默认的process.start()(即stop_after_crawl=True),让爬虫完成抓取后自动终止进程,便于查看输出或错误信息:

def run_wdr_spider():
    process = CrawlerProcess(get_project_settings())
    process.crawl(SecondSpiderSpider)
    try:
        # 抓取完成后自动停止进程
        process.start()
    except KeyboardInterrupt:
        process.stop()

方案2:调试抓取过程,排除管道干扰

先注释掉自定义管道配置,添加日志和打印语句,确认Item是否正常生成:

# 爬虫类中的custom_settings修改
custom_settings = {
    'DEPTH_LIMIT': 0,
    'DOWNLOAD_DELAY': 5,
    'LOG_LEVEL': 'INFO',  # 开启日志查看抓取流程
    # 暂时注释管道,先测试基础抓取
    # 'ITEM_PIPELINES':{'bremenspider.pipelines.MongoDBPipeline2':300},
    # 'MONGO_COLLECTION_NAME2': 'human_written_texts(wdr)'
}

# text_parse方法添加打印
def text_parse(self, response):
    item = CustomItem()
    item['text'] = response.xpath('//p[@class="text small"]/descendant-or-self::*/text()').getall()
    item['links'] = response.meta['url']
    extracted_text = ' '.join(item['text']).strip()
    if extracted_text:
        # 打印抓取结果,确认是否成功
        print(f"抓取到链接:{item['links']},文本长度:{len(extracted_text)}")
        yield item

方案3:改用命令行执行

Scrapy官方不推荐在IDE中直接运行爬虫,建议在命令行执行main.py:

python main.py

或直接用Scrapy命令启动爬虫:

scrapy crawl second_spider

方案4:解决Spyder的reactor冲突

如果必须在Spyder中运行,在启动爬虫前重置Twisted reactor:

from twisted.internet import reactor

def run_wdr_spider():
    # 重置reactor避免冲突
    if reactor.running:
        reactor.stop()
    process = CrawlerProcess(get_project_settings())
    process.crawl(SecondSpiderSpider)
    process.start()

调整后的完整代码

爬虫脚本(wdrspider.py)

import scrapy
import urllib.parse
from scrapy.item import Item, Field
from scrapy.crawler import CrawlerProcess
from scrapy.utils.project import get_project_settings
from twisted.internet import reactor

# 定义Item
class CustomItem(Item):
    links = Field()
    text = Field()

# 定义爬虫
class SecondSpiderSpider(scrapy.Spider):
    name = "second_spider"
    allowed_domains = ["www1.wdr.de"]
    start_urls = ["https://www1.wdr.de/abisz120.html"]
    
    custom_settings = {
        'DEPTH_LIMIT': 0,
        'DOWNLOAD_DELAY': 5,
        'LOG_LEVEL': 'INFO'
        # 测试通过后可恢复管道配置
        # 'ITEM_PIPELINES':{'bremenspider.pipelines.MongoDBPipeline2':300},
        # 'MONGO_COLLECTION_NAME2': 'human_written_texts(wdr)'
    }
    
    def parse(self, response):
        links1 = response.css('ul.list a::attr(href)').getall()
        base_link = 'https://www1.wdr.de/'
        for link in links1:
            absolute_url = urllib.parse.urljoin(base_link, link)
            yield scrapy.Request(url=absolute_url, callback=self.link_parse)
            
    def link_parse(self, response):
        links2 = response.css('div.teaser a::attr(href)').getall()
        base_link2 = 'https://www1.wdr.de/'
        for link in links2:
            absolute_url2 = urllib.parse.urljoin(base_link2, link)
            yield scrapy.Request(absolute_url2, callback=self.text_parse,
                                 meta={'url': absolute_url2})

    def text_parse(self, response):
        item = CustomItem()
        item['text'] = response.xpath('//p[@class="text small"]/descendant-or-self::*/text()').getall()
        item['links'] = response.meta['url']
        extracted_text = ' '.join(item['text']).strip()
        if extracted_text:
            print(f"抓取到链接:{item['links']},文本长度:{len(extracted_text)}")
            yield item
            
def run_wdr_spider():
    if reactor.running:
        reactor.stop()
    process = CrawlerProcess(get_project_settings())
    process.crawl(SecondSpiderSpider)
    try:
        process.start()
    except KeyboardInterrupt:
        process.stop()

if __name__ == "__main__":
    run_wdr_spider()

main.py代码

import sys
import os
import pandas as pd

# 导入爬虫模块
import wdrspider

# 运行爬虫
wdrspider.run_wdr_spider()

内容的提问来源于stack exchange,提问作者Fangyi Chen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 22:30:56