Scrapy爬虫完成后发邮件报'NoneType'无'bio_read'属性错误怎么解决?
Scrapy爬虫完成后发送邮件报错的异步改造方案
问题背景
在Scrapy项目根目录定义了MyStatsCollector类,用于爬虫完成抓取后发送邮件,代码如下:
class MyStatsCollector(StatsCollector): def _persist_stats(self, stats, spider): ... mailer = MailSender() mailer.send()
邮件可正常发送,但会抛出如下错误:
Traceback (most recent call last): File "/opt/homebrew/lib/python3.10/site-packages/twisted/python/log.py", line 96, in callWithLogger return callWithContext({"system": lp}, func, *args, **kw) File "/opt/homebrew/lib/python3.10/site-packages/twisted/python/log.py", line 80, in callWithContext return context.call({ILogContext: newCtx}, func, *args, **kw) File "/opt/homebrew/lib/python3.10/site-packages/twisted/python/context.py", line 117, in callWithContext return self.currentContext().callWithContext(ctx, func, *args, **kw) File "/opt/homebrew/lib/python3.10/site-packages/twisted/python/context.py", line 82, in callWithContext return func(*args, **kw) --- <exception caught here> --- File "/opt/homebrew/lib/python3.10/site-packages/twisted/internet/asyncioreactor.py", line 138, in _readOrWrite why = method() File "/opt/homebrew/lib/python3.10/site-packages/twisted/internet/tcp.py", line 248, in doRead return self._dataReceived(data) File "/opt/homebrew/lib/python3.10/site-packages/twisted/internet/tcp.py", line 253, in _dataReceived rval = self.protocol.dataReceived(data) File "/opt/homebrew/lib/python3.10/site-packages/twisted/protocols/tls.py", line 329, in dataReceived self._flushReceiveBIO() File "/opt/homebrew/lib/python3.10/site-packages/twisted/protocols/tls.py", line 300, in _flushReceiveBIO self._flushSendBIO() File "/opt/homebrew/lib/python3.10/site-packages/twisted/protocols/tls.py", line 253, in _flushSendBIO bytes = self._tlsConnection.bio_read(2 ** 15) builtins.AttributeError: 'NoneType' object has no attribute 'bio_read' Traceback (most recent call last): File "/opt/homebrew/lib/python3.10/site-packages/twisted/internet/asyncioreactor.py", line 138, in _readOrWrite why = method() File "/opt/homebrew/lib/python3.10/site-packages/twisted/internet/tcp.py", line 248, in doRead return self._dataReceived(data) File "/opt/homebrew/lib/python3.10/site-packages/twisted/internet/tcp.py", line 253, in _dataReceived rval = self.protocol.dataReceived(data) File "/opt/homebrew/lib/python3.10/site-packages/twisted/protocols/tls.py", line 329, in dataReceived self._flushReceiveBIO() File "/opt/homebrew/lib/python3.10/site-packages/twisted/protocols/tls.py", line 300, in _flushReceiveBIO self._flushSendBIO() File "/opt/homebrew/lib/python3.10/site-packages/twisted/protocols/tls.py", line 253, in _flushSendBIO bytes = self._tlsConnection.bio_read(2 ** 15) AttributeError: 'NoneType' object has no attribute 'bio_read'
根据相关issue提示,需使用async/await机制改造邮件发送,具体修改方案如下:
改造方案
1. 核心修改点
- 将
_persist_stats方法改为异步方法(async def) - 替换同步邮件发送方法为Scrapy内置的异步版本
send_async
完整改造代码
from scrapy.statscollectors import StatsCollector from scrapy.mail import MailSender class MyStatsCollector(StatsCollector): async def _persist_stats(self, stats, spider): # 保留原有统计逻辑... # 从项目配置初始化邮件发送器,无需硬编码参数 mailer = MailSender.from_settings(spider.settings) # 构造邮件内容 subject = f"爬虫{spider.name}抓取完成" body = f"抓取统计数据:\n{stats}" to_addrs = ["target_email@example.com"] # 异步发送邮件 await mailer.send_async( to=to_addrs, subject=subject, body=body, # 如需添加附件,可传入attach参数:attach=("filename.txt", "text/plain", b"file content") )
改造说明
- 异步方法
_persist_stats能被Scrapy正确识别并在异步上下文执行,避免阻塞Twisted Reactor循环 MailSender.from_settings从项目settings.py读取邮件配置(如SMTP服务器、端口、账号等),符合Scrapy最佳实践send_async方法以异步方式发送邮件,不会干扰TLS连接的正常生命周期,彻底解决报错中的AttributeError问题
内容的提问来源于stack exchange,提问作者user984621
相关产品推荐
相关产品推荐

