非命令行运行Scrapy脚本失败求助(附代码及报错)
问题:直接运行Scrapy脚本失败,加载HTTPDownloadHandler出错
尝试按官方文档方式直接运行Scrapy脚本(不用scrapy crawl命令),但脚本报错无法加载HTTPDownloadHandler,核心报错如下,完整报错未全部粘贴:
我的代码
import scrapy from scrapy.crawler import CrawlerProcess class misbeneficiosSpider(scrapy.Spider): name = 'misbeneficios' start_urls = ['https://productos.misbeneficios.com.uy/tv-y-audio', 'https://productos.misbeneficios.com.uy/tv-y-audio?p=2'] def parse(self, response): for products in response.css('div.product-item-info'): yield { 'name': products.css('a.product-item-link::text').get(), 'price': products.css('span.price::text').get().replace('U$S\xa0', '')#[:-3].upper() } next_page = response.css('a.action.next').attrib['href'] if next_page is not None: yield response.follow(next_page, callback=self.parse) process = CrawlerProcess(settings={ "FEEDS": { "items.csv": {"format": "csv"}, }, }) process.crawl(misbeneficiosSpider) process.start()
报错信息
2022-11-09 00:02:08 [scrapy.utils.log] INFO: Scrapy 2.7.1 started (bot: scrapybot) 2022-11-09 00:02:08 [scrapy.utils.log] INFO: Versions: lxml 4.9.1.0, libxml2 2.9.12, cssselect 1.2.0, parsel 1.7.0, w3lib 2.0.1, Twisted 22.10.0, Python 3.8.5 (tags/v3.8.5:580fbb0, Jul 20 2020, 15:43:08) [MSC v.1926 32 bit (Intel)], pyOpenSSL 22.1.0 (OpenSSL 3.0.7 1 Nov 2022), cryptography 38.0.3, Platform Windows-10-10.0.22000-SP0 2022-11-09 00:02:08 [scrapy.crawler] INFO: Overridden settings: {} 2022-11-09 00:02:08 [py.warnings] WARNING: C:\Users\cabre\AppData\Local\Programs\Python\Python38-32\lib\site-packages\scrapy\utils\request.py:231: ScrapyDeprecationWarning: '2.6' is a deprecated value for the 'REQUEST_FINGERPRINTER_IMPLEMENTATION' setting. It is also the default value. In other words, it is normal to get this warning if you have not defined a value for the 'REQUEST_FINGERPRINTER_IMPLEMENTATION' setting. This is so for backward compatibility reasons, but it will change in a future version of Scrapy. See the documentation of the 'REQUEST_FINGERPRINTER_IMPLEMENTATION' setting for information on how to handle this deprecation. return cls(crawler) 2022-11-09 00:02:08 [scrapy.utils.log] DEBUG: Using reactor: twisted.internet.selectreactor.SelectReactor 2022-11-09 00:02:08 [scrapy.extensions.telnet] INFO: Telnet Password: 034a39dcf0704cf3 2022-11-09 00:02:08 [scrapy.middleware] INFO: Enabled extensions: ['scrapy.extensions.corestats.CoreStats', 'scrapy.extensions.telnet.TelnetConsole', 'scrapy.extensions.feedexport.FeedExporter', 'scrapy.extensions.logstats.LogStats'] 2022-11-09 00:02:09 [scrapy.core.downloader.handlers] ERROR: Loading "scrapy.core.downloader.handlers.http.HTTPDownloadHandler" for scheme "http"
解决方案
1. 解决HTTPDownloadHandler加载失败问题
报错核心是Twisted 22.10.0与Scrapy 2.7.1在32位Python环境下存在兼容性问题,执行以下命令降级Twisted:
pip install twisted==22.8.0
若问题仍存在,建议切换到64位Python环境,32位Python对部分Scrapy依赖的支持有限。
2. 修复代码中的潜在异常
原代码直接通过attrib['href']获取下一页链接,若页面无下一页元素会抛出KeyError;同时price字段未处理None值情况,修改为安全获取方式:
- 下一页链接:用
::attr(href)结合get()方法获取 - 价格字段:先判断是否为
None再处理
3. 消除DeprecationWarning
在CrawlerProcess的settings中添加REQUEST_FINGERPRINTER_IMPLEMENTATION配置,指定当前兼容的版本。
修改后的完整代码
import scrapy from scrapy.crawler import CrawlerProcess class misbeneficiosSpider(scrapy.Spider): name = 'misbeneficios' start_urls = ['https://productos.misbeneficios.com.uy/tv-y-audio', 'https://productos.misbeneficios.com.uy/tv-y-audio?p=2'] def parse(self, response): for products in response.css('div.product-item-info'): price_text = products.css('span.price::text').get() yield { 'name': products.css('a.product-item-link::text').get(), 'price': price_text.replace('U$S\xa0', '') if price_text else None } # 安全获取下一页链接 next_page = response.css('a.action.next::attr(href)').get() if next_page is not None: yield response.follow(next_page, callback=self.parse) process = CrawlerProcess(settings={ "FEEDS": { "items.csv": {"format": "csv"}, }, "REQUEST_FINGERPRINTER_IMPLEMENTATION": "2.7", }) process.crawl(misbeneficiosSpider) process.start()
内容的提问来源于stack exchange,提问作者Bruno123
相关产品推荐
相关产品推荐

