Ruby 3 + Async 2.3.0中无法捕获MetaInspector::TimeoutError的排查求助
问题原因分析
你的代码里rescue块没生效,核心原因是Async的IO调度器和MetaInspector的超时机制不兼容:
- MetaInspector内部依赖的HTTP客户端(比如Net::HTTP)使用Ruby原生
Timeout实现超时,但Async的调度器会接管所有IO操作,超时错误会从Async的事件循环(async/scheduler.rb的run_once方法)抛出,这个异常不在你写的begin/rescue代码块的上下文里。 - 嵌套Async任务的异常如果没被任务内部完全处理,会直接冒泡到顶层Async任务,导致程序崩溃,而不会被你写的rescue捕获。
解决方案
方案1:使用Async原生的超时机制(推荐)
放弃MetaInspector的超时配置,改用Async提供的Async::Timeout来控制任务超时,这样异常能被正常捕获:
require 'async' require 'async/timeout' require 'metainspector' urls = ["https://httpbin.org/delay/15", "https://www.google.com"] Async do urls.each do |url| Async do begin # 设置总超时15秒(包含连接+读取) Async::Timeout.timeout(15) do page = MetaInspector.new(url, allow_non_html_content: true) # 可以在这里添加对page的处理逻辑 end puts "made it to the end of the nested Async block with #{url}" rescue Async::Timeout::Error puts "任务超时: #{url}" rescue => e puts "rescued everything else, type: #{e.class}, message: #{e.message}" end end end end
方案2:捕获Async任务的全局异常
给每个Async任务绑定on_error回调,确保任务抛出的所有异常都能被捕获:
require 'async' require 'metainspector' urls = ["https://httpbin.org/delay/15", "https://www.google.com"] Async do urls.each do |url| task = Async do begin page = MetaInspector.new(url, allow_non_html_content: true) puts "made it to the end of the nested Async block with #{url}" rescue MetaInspector::TimeoutError puts "rescued MetaInspector's TimeoutError" rescue => e puts "rescued everything else, type: #{e.class}, message: #{e.message}" end end # 捕获任务未处理的异常 task.on_error do |error| puts "任务未捕获异常: #{error.class} - #{error.message}" end end end
方案3:让MetaInspector使用兼容Async的HTTP客户端
MetaInspector支持自定义HTTP客户端,你可以配置它使用async-http作为适配器,这样超时逻辑会和Async调度器兼容,异常能被正常rescue捕获。需要额外安装相关依赖,配置示例:
# 先安装依赖:gem install async-http faraday-async-http require 'async' require 'metainspector' require 'faraday' require 'faraday/async_http' # 配置MetaInspector使用Async HTTP客户端 MetaInspector.configure do |config| config.faraday_adapter = :async_http end urls = ["https://httpbin.org/delay/15", "https://www.google.com"] Async do urls.each do |url| Async do begin page = MetaInspector.new(url, connection_timeout:10, read_timeout:5, allow_non_html_content: true) puts "made it to the end of the nested Async block with #{url}" rescue MetaInspector::TimeoutError puts "rescued MetaInspector's TimeoutError" rescue => e puts "rescued everything else, type: #{e.class}, message: #{e.message}" end end end end
内容的提问来源于stack exchange,提问作者TepidSwitch
相关产品推荐
相关产品推荐

