Ruby Threads问题:为何我的脚本无法正常结束?
Great question! Your current thread-based approach is a solid start, but we can enhance it to handle return codes, clean shutdowns, and timeouts effectively. Let's break down the improvements step by step:
1. Track Return Codes & Node-specific Results
Instead of just spawning threads, we need a way to capture each thread's exit status, output, or errors. You can use Thread#value to retrieve a thread's return value, and wrap your node processing logic to explicitly return status codes (or structured results). We'll also use a thread-safe structure to collect results without race conditions.
2. Implement Clean Exit Handling
We need to handle system signals (like SIGINT when you press Ctrl+C) to gracefully stop all threads instead of letting them hang. Additionally, catching exceptions inside threads prevents a single failed node from taking down the entire script.
3. Add Timeouts to Avoid Hanging Threads
A stuck node can make your script run indefinitely. Using Ruby's Timeout module lets you set a maximum runtime per node, ensuring the script progresses even if one node misbehaves.
Improved Script Example
require 'timeout' # Thread-safe storage for node results node_results = {} results_mutex = Mutex.new # Flag to signal threads to exit early on interrupt shutdown = false trap('INT') do puts "\nReceived interrupt, initiating clean shutdown..." shutdown = true end def process_node(apps_to_close, node, results_mutex, node_results, shutdown_flag) Thread.new do # Exit early if shutdown was triggered before starting work return if shutdown_flag begin # Set a 5-minute timeout (adjust to your needs) Timeout.timeout(300) do puts "Now processing node #{node}" # Replace with your actual node processing logic # Example: Run a command and capture exit code # exit_code = system("ssh #{node} 'systemctl stop #{apps_to_close.join(" ")}'") ? 0 : 1 exit_code = 0 # Placeholder for real logic # Store result thread-safely results_mutex.synchronize do node_results[node] = { status: :success, exit_code: exit_code } end end rescue Timeout::Error results_mutex.synchronize do node_results[node] = { status: :timeout, error: "Processing timed out after 5 minutes" } end puts "Node #{node} processing timed out" rescue StandardError => e results_mutex.synchronize do node_results[node] = { status: :failed, error: e.message } end puts "Node #{node} processing failed: #{e.message}" end end end # Example input data (replace with your actual values) nodes = ['node-01', 'node-02', 'node-03'] apps_to_close = ['api-service', 'worker-process'] # Spawn threads for all nodes threads = nodes.map do |node| process_node(apps_to_close, node, results_mutex, node_results, shutdown) end # Wait for all threads to complete threads.each do |thr| begin thr.join rescue => e puts "Thread for node failed unexpectedly: #{e.message}" end end # Print final summary puts "\nProcessing complete. Summary:" node_results.each do |node, result| case result[:status] when :success puts "✅ Node #{node}: SUCCESS (exit code #{result[:exit_code]})" when :timeout puts "⏱️ Node #{node}: TIMEOUT - #{result[:error]}" when :failed puts "❌ Node #{node}: FAILED - #{result[:error]}" end end # Exit with non-zero code if any node failed/timed out exit 1 if node_results.values.any? { |r| r[:status] != :success }
Key Improvements Explained
- Return Code Tracking: We store each node's exit code and status in a thread-safe hash, so you can inspect detailed results after all threads finish.
- Clean Shutdown: The
trap('INT')sets a shutdown flag, which threads check before starting work (you can add periodic checks inside long-running tasks too). - Timeouts: The
Timeout.timeoutwrapper ensures no thread runs longer than the specified limit, preventing indefinite hangs. - Error Handling: Exceptions inside threads are caught and logged, so one failed node doesn't crash the entire script.
- Script Exit Code: The final exit code reflects whether all nodes succeeded, making it easy to integrate with CI/CD pipelines or monitoring tools.
内容的提问来源于stack exchange,提问作者David West

