Scrapy爬虫导入主文件后运行异常(无输出/卡顿)求助
问题现象
- 爬虫脚本单独执行时可正常抓取文本,但作为模块导入
main.py后,Spyder IDE持续运行却无任何抓取结果,等待一小时仍无输出,内存占用无异常 - 控制台仅显示Scrapy弃用警告:
request.py:254: ScrapyDeprecationWarning: '2.6' is a deprecated value for the 'REQUEST_FINGERPRINTER_IMPLEMENTATION' setting.
It is also the default value. In other words, it is normal to get this warning if you have not defined a value for the 'REQUEST_FINGERPRINTER_IMPLEMENTATION' setting. This is so for backward compatibility reasons, but it will change in a future version of Scrapy.
See the documentation of the 'REQUEST_FINGERPRINTER_IMPLEMENTATION' setting for information on how to handle this deprecation.
return cls(crawler)
核心问题分析
stop_after_crawl=False导致进程阻塞
你在run_wdr_spider函数中使用了process.start(stop_after_crawl=False),该参数会让Scrapy爬虫完成抓取后保持进程运行状态,不会自动退出。当从main.py导入执行时,主线程会被Twisted reactor阻塞,看起来像一直在运行,但实际抓取可能已完成,只是缺少输出反馈导致你误以为无结果。项目配置加载异常
从main.py导入爬虫时,get_project_settings()可能无法正确识别Scrapy项目路径,导致自定义管道bremenspider.pipelines.MongoDBPipeline2无法加载,Item无法被正常处理,没有可视化输出。Spyder与Twisted reactor兼容性冲突
Spyder本身依赖Twisted框架,直接在Spyder中运行Scrapy爬虫可能出现reactor资源冲突,导致爬虫无法正常启动或运行无响应。
解决方案
方案1:修改进程停止参数
将run_wdr_spider中的process.start(stop_after_crawl=False)改为默认的process.start()(即stop_after_crawl=True),让爬虫完成抓取后自动终止进程,便于查看输出或错误信息:
def run_wdr_spider(): process = CrawlerProcess(get_project_settings()) process.crawl(SecondSpiderSpider) try: # 抓取完成后自动停止进程 process.start() except KeyboardInterrupt: process.stop()
方案2:调试抓取过程,排除管道干扰
先注释掉自定义管道配置,添加日志和打印语句,确认Item是否正常生成:
# 爬虫类中的custom_settings修改 custom_settings = { 'DEPTH_LIMIT': 0, 'DOWNLOAD_DELAY': 5, 'LOG_LEVEL': 'INFO', # 开启日志查看抓取流程 # 暂时注释管道,先测试基础抓取 # 'ITEM_PIPELINES':{'bremenspider.pipelines.MongoDBPipeline2':300}, # 'MONGO_COLLECTION_NAME2': 'human_written_texts(wdr)' } # text_parse方法添加打印 def text_parse(self, response): item = CustomItem() item['text'] = response.xpath('//p[@class="text small"]/descendant-or-self::*/text()').getall() item['links'] = response.meta['url'] extracted_text = ' '.join(item['text']).strip() if extracted_text: # 打印抓取结果,确认是否成功 print(f"抓取到链接:{item['links']},文本长度:{len(extracted_text)}") yield item
方案3:改用命令行执行
Scrapy官方不推荐在IDE中直接运行爬虫,建议在命令行执行main.py:
python main.py
或直接用Scrapy命令启动爬虫:
scrapy crawl second_spider
方案4:解决Spyder的reactor冲突
如果必须在Spyder中运行,在启动爬虫前重置Twisted reactor:
from twisted.internet import reactor def run_wdr_spider(): # 重置reactor避免冲突 if reactor.running: reactor.stop() process = CrawlerProcess(get_project_settings()) process.crawl(SecondSpiderSpider) process.start()
调整后的完整代码
爬虫脚本(wdrspider.py)
import scrapy import urllib.parse from scrapy.item import Item, Field from scrapy.crawler import CrawlerProcess from scrapy.utils.project import get_project_settings from twisted.internet import reactor # 定义Item class CustomItem(Item): links = Field() text = Field() # 定义爬虫 class SecondSpiderSpider(scrapy.Spider): name = "second_spider" allowed_domains = ["www1.wdr.de"] start_urls = ["https://www1.wdr.de/abisz120.html"] custom_settings = { 'DEPTH_LIMIT': 0, 'DOWNLOAD_DELAY': 5, 'LOG_LEVEL': 'INFO' # 测试通过后可恢复管道配置 # 'ITEM_PIPELINES':{'bremenspider.pipelines.MongoDBPipeline2':300}, # 'MONGO_COLLECTION_NAME2': 'human_written_texts(wdr)' } def parse(self, response): links1 = response.css('ul.list a::attr(href)').getall() base_link = 'https://www1.wdr.de/' for link in links1: absolute_url = urllib.parse.urljoin(base_link, link) yield scrapy.Request(url=absolute_url, callback=self.link_parse) def link_parse(self, response): links2 = response.css('div.teaser a::attr(href)').getall() base_link2 = 'https://www1.wdr.de/' for link in links2: absolute_url2 = urllib.parse.urljoin(base_link2, link) yield scrapy.Request(absolute_url2, callback=self.text_parse, meta={'url': absolute_url2}) def text_parse(self, response): item = CustomItem() item['text'] = response.xpath('//p[@class="text small"]/descendant-or-self::*/text()').getall() item['links'] = response.meta['url'] extracted_text = ' '.join(item['text']).strip() if extracted_text: print(f"抓取到链接:{item['links']},文本长度:{len(extracted_text)}") yield item def run_wdr_spider(): if reactor.running: reactor.stop() process = CrawlerProcess(get_project_settings()) process.crawl(SecondSpiderSpider) try: process.start() except KeyboardInterrupt: process.stop() if __name__ == "__main__": run_wdr_spider()
main.py代码
import sys import os import pandas as pd # 导入爬虫模块 import wdrspider # 运行爬虫 wdrspider.run_wdr_spider()
内容的提问来源于stack exchange,提问作者Fangyi Chen

