You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Ruby on Rails多线程运行爬虫报错:未初始化常量ParserController::Services

解决Rails多线程爬虫中无法访问服务类的问题

这个报错uninitialized constant ParserController::Services本质上是Rails自动加载机制在子线程中的局限性导致的:Rails的常量自动加载默认在主线程上下文里工作,子线程不会自动触发未加载常量的查找和加载流程——主线程不用线程时,第一次调用服务类会触发自动加载;但到了子线程里,这个触发逻辑不生效,就找不到对应的常量了。

下面是具体的解决思路和优化方案:

一、快速修复当前线程代码

1. 提前预加载常量

在创建线程之前,先在主线程里引用一次你需要的服务类,强制Rails提前加载这些常量,这样子线程就能直接访问到了:

# 主线程预加载服务类,确保常量已加载
Services::Scrapers::FirstScraperService
Services::Scrapers::SecondScraperService

threads = []
data.each_key do |office|
  threads << Thread.new(office, data) do |office1, data1|
    # 重点:如果爬虫涉及数据库操作,必须用连接池获取线程专属连接
    ActiveRecord::Base.connection_pool.with_connection do
      data1[office1].each_key do |link|
        if (first condition)
          Services::Scrapers::FirstScraperService.new.parse(link, data1[office1][link])
        elsif (second condition)
          Services::Scrapers::SecondScraperService.new.parse(link, data1[office1][link])
        end
      end
    end
  end
end
threads.each { |thr| thr.join }

2. 检查命名空间与文件结构匹配

确保你的服务类文件命名和模块结构完全对应:

  • app/controllers/services/scrapers/first_scraper_service.rb 文件里的类定义应该是:
    module Services
      module Scrapers
        class FirstScraperService
          def parse(link, data)
            # 你的爬虫逻辑
          end
        end
      end
    end
    

Rails的自动加载依赖"约定大于配置",文件名和模块/类名不匹配也会导致常量找不到。

二、更优的并行爬虫方案

Ruby MRI的线程受GIL限制,虽然IO密集型(比如爬虫)场景下线程能提升效率,但自己管理线程容易出现连接泄漏、状态不一致等问题。更推荐以下两种方案:

1. 使用Sidekiq异步任务队列

把每个爬虫任务拆成独立的异步任务,交给Sidekiq管理,完全避免线程上下文问题,还支持任务重试、监控、并发配置:

步骤1:创建Worker

新建 app/workers/scraper_worker.rb:

class ScraperWorker
  include Sidekiq::Worker
  # 可选:设置重试次数,避免网络波动导致任务失败
  sidekiq_options retry: 3

  def perform(office, link, data_entry)
    if # 这里替换你的第一个条件(注意参数要能JSON序列化)
      Services::Scrapers::FirstScraperService.new.parse(link, data_entry)
    elsif # 替换第二个条件
      Services::Scrapers::SecondScraperService.new.parse(link, data_entry)
    end
  end
end

步骤2:在控制器中触发任务

data.each_key do |office|
  data[office].each_key do |link|
    # 把每个爬虫任务丢给Sidekiq异步处理
    ScraperWorker.perform_async(office, link, data[office][link])
  end
end

2. 使用Concurrent Ruby库的线程池

如果你更倾向于用线程而不是异步任务,可以用concurrent-ruby库提供的线程池,它能更安全地管理线程,还能处理上下文加载:

步骤1:添加依赖到Gemfile

gem 'concurrent-ruby'

然后运行bundle install。

步骤2:用线程池重构代码

require 'concurrent'

# 预加载常量
Services::Scrapers::FirstScraperService
Services::Scrapers::SecondScraperService

# 创建线程池,设置并发数(根据你的服务器配置调整)
pool = Concurrent::ThreadPoolExecutor.new(min_threads: 2, max_threads: 10)

data.each_key do |office|
  data[office].each_key do |link|
    pool.post(office, link, data[office][link]) do |office1, link1, entry1|
      ActiveRecord::Base.connection_pool.with_connection do
        if (first condition)
          Services::Scrapers::FirstScraperService.new.parse(link1, entry1)
        elsif (second condition)
          Services::Scrapers::SecondScraperService.new.parse(link1, entry1)
        end
      end
    end
  end
end

# 等待所有任务完成,关闭线程池
pool.shutdown
pool.wait_for_termination

内容的提问来源于stack exchange,提问作者Maksim

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:15:30