Scrapy ItemPipeline报错:RuntimeError: await wasn't used with future
Scrapy异步Pipeline使用Motor去重的RuntimeError解决方法
问题场景
为实现Scrapy爬虫的去重逻辑,通过Motor(MongoDB异步驱动)统计数据库中相同URL的文档数量,编写了如下process_item函数:
async def process_item(self, item, spider): response = self.mangas.count_documents({"url": item.url}) print(response, 'is future?', futures.isfuture(response)) # 显式类型检查 count = await response # 此处抛出异常 if count: raise DropItem(f"Duplicate found of {item}") await self.mangas.insert_one(dict(item)) return item
报错信息
运行时触发如下错误:
is future? True 2022-08-23 22:39:10 [scrapy.core.scraper] ERROR: Error processing *SOME PARSED DATA* Traceback (most recent call last): File "C:\Users\Canald\AppData\Local\pypoetry\Cache\virtualenvs\api-abuLimWH-py3.10\lib\site-packages\twisted\internet\defer.py", line 1660, in _inlineCallbacks result = current_context.run(gen.send, result) File "C:\Users\Canald\Files\VS\HChan\hentai_scrap\pipelines.py", line 48, in process_item count = await response RuntimeError: await wasn't used with future
尝试使用scrapy.utils.defer.maybe_deferred_to_future后仍出现相同异常。
解决方案
问题根源在于Motor返回的是asyncio Future对象,而Scrapy的Pipeline基于Twisted的Deferred处理异步逻辑,直接await asyncio Future会导致事件循环不兼容。需要用scrapy.utils.defer.defer_to_future将asyncio Future转换为Twisted可处理的Deferred对象。
修改后的代码如下:
from scrapy.utils.defer import defer_to_future from scrapy.exceptions import DropItem async def process_item(self, item, spider): # 将Motor的异步调用转为Twisted Deferred count = await defer_to_future(self.mangas.count_documents({"url": item.url})) if count > 0: raise DropItem(f"Duplicate found of {item}") await defer_to_future(self.mangas.insert_one(dict(item))) return item
说明
defer_to_future负责桥接asyncio和Twisted的异步模型,确保Motor的异步操作能在Scrapy的事件循环中正确执行。- 无需手动判断是否为Future对象,
defer_to_future会自动处理兼容转换。
内容的提问来源于stack exchange,提问作者Canald
相关产品推荐
相关产品推荐

