You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Rails基于Sidekiq的大XML后台处理任务无法完成排查求助

Hey there, let's dig into your Sidekiq XML processing issue. Your two hunches are totally on point—both the large arrays and potential timeouts could be causing the hang, plus there's a hidden inefficiency in your current code that's making big files way slower than they need to be. Let's break down how to debug and fix this:

First, Debug to Pinpoint the Root Cause

Before jumping into fixes, let's confirm exactly what's going wrong:

1. Add Detailed Logging

Stick logs in key parts of your job to track progress, memory usage, and where it gets stuck:

# In your Sidekiq job
Rails.logger.info "Starting XML processing: file size = #{xml_file.size} bytes"

# Inside get_texts, after processing each 100 nodes
if old_texts.size % 100 == 0
  memory_used = `ps -o rss= -p #{Process.pid}`.strip.to_i / 1024 # MB
  Rails.logger.info "Processed #{old_texts.size} text chunks | Memory used: #{memory_used}MB"
end

# Inside replace_texts, during each iteration
Rails.logger.info "Replacing chunk #{index+1}/#{old_texts.size} | Current inc value: #{inc}"

Check your Sidekiq logs (or Rails logs) to see if:

  • Memory usage spikes to extreme levels (signaling an OOM issue)
  • The loop gets stuck on a specific chunk (hinting at inefficient node searching)
  • The job is killed abruptly (signaling a timeout)

2. Check Sidekiq & System Configs

  • Sidekiq Timeout: By default, Sidekiq kills jobs that run longer than 30 seconds. For large XMLs, bump this in config/sidekiq.yml:
    :concurrency: 5
    :timeout: 3600 # 1 hour, adjust as needed
    :queues:
      - default
    
  • System OOM Killer: On Linux, use dmesg | grep -i oom to check if the kernel killed your Sidekiq process due to memory overload.

3. Test with Truncated Large Files

Take your problematic large XML, cut it down to 100, 500, 1000 nodes, and test each version. This will tell you if the issue scales with the number of nodes (which points to the array/loop inefficiency) or hits a hard memory limit.

Fix the Code Inefficiencies

Your current approach of collecting all texts into old_texts/new_texts then looping through them with nested searches is O(n²) time complexity—this gets exponentially slower as your XML grows. Plus, storing thousands of text chunks in arrays eats up memory fast.

Refactor to Process Nodes Streamingly

Instead of collecting all texts first, process each <w:p> node immediately when you find it. This eliminates the huge arrays and cuts your time complexity to O(n):

def process_xml(xml_doc, xml_no_includs, ancestors_excluds)
  search_param = '//w:document//w:body//w:p'
  text_params = './/text()[not(ancestor::wp14:pctHeight or ancestor::wp14:pctWidth or ancestor::wp:posOffset or ancestor::w:instrText or ancestor::w:delText or ancestor::w:delInstrText)]'

  xml_doc.xpath(search_param).each do |p_node|
    valid_text_nodes = []
    full_text = ''

    # Collect valid text nodes and build the full text
    p_node.xpath(text_params).each do |text_node|
      # Skip nodes in excluded ancestors
      next if ancestors_excluds.any? { |ancestor| text_node.ancestors(ancestor).present? }
      
      full_text += text_node.content
      valid_text_nodes << text_node
    end

    next if full_text.blank?

    # Generate the masked text
    masked_text = full_text.gsub(/(.)./, '\1*')

    # Replace the first valid node's content, clear the rest
    unless valid_text_nodes.empty?
      valid_text_nodes.first.content = masked_text
      valid_text_nodes[1..].each { |node| node.content = '' }
    end
  end
end

# Usage in your job
xml_doc = Nokogiri::XML(xml_content)
process_xml(xml_doc, xml_no_includs, ancestors_excluds)

For Ultra-Large XMLs: Use Streaming Parsing

If your XML is hundreds of MB or larger, even loading the full DOM into memory with Nokogiri will cause issues. Use Nokogiri's XML::Reader to stream the document without loading it all at once:

def stream_process_xml(xml_file_path, xml_no_includs, ancestors_excluds)
  reader = Nokogiri::XML::Reader(File.open(xml_file_path))
  current_p_text = ''
  in_p_node = false
  in_excluded_ancestor = false

  reader.each do |node|
    # Track if we're inside an excluded ancestor
    if node.node_type == Nokogiri::XML::Reader::TYPE_ELEMENT
      if ancestors_excluds.any? { |ancestor| node.name == ancestor.split('//').last }
        in_excluded_ancestor = true
      end
      in_p_node = true if node.name == 'p' # Adjust for your namespace, e.g., 'w:p'
    elsif node.node_type == Nokogiri::XML::Reader::TYPE_END_ELEMENT
      if ancestors_excluds.any? { |ancestor| node.name == ancestor.split('//').last }
        in_excluded_ancestor = false
      end

      # When we exit a <w:p> node, process the text
      if node.name == 'p' && in_p_node && !current_p_text.blank? && !in_excluded_ancestor
        masked_text = current_p_text.gsub(/(.)./, '\1*')
        # To modify the file, you'd need to write to a new file while streaming—this is a simplified example
        puts "Replacing: #{current_p_text} -> #{masked_text}"
        current_p_text = ''
      end
      in_p_node = false if node.name == 'p'
    elsif node.node_type == Nokogiri::XML::Reader::TYPE_TEXT && in_p_node && !in_excluded_ancestor
      # Skip text nodes with excluded parent tags
      parent_name = node.parent.name
      next if xml_no_includs.include?(parent_name)

      current_p_text += node.value
    end
  end
end

Final Checks

  • After refactoring, re-test with your large XML to confirm it completes.
  • Monitor Sidekiq's Web UI (mount it in your routes) to track job duration and success rates.
  • If you're still hitting memory limits, consider splitting the XML into smaller chunks and processing them as separate Sidekiq jobs.

内容的提问来源于stack exchange,提问作者echan00

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 04:34:59