在Ubuntu系统中使用Ray进行并行计算时Worker崩溃问题求助
I've run into similar cross-platform parallel computing crashes before, so let's walk through targeted troubleshooting steps to get to the bottom of this SIGFPE error that's plaguing your tasks on Linux/Google Cloud but working fine locally on your Mac.
1. Pinpoint the Root of the SIGFPE Exception
SIGFPE (Signal Floating-Point Exception) isn't just limited to floating-point operations—it can also trigger from integer division by zero, integer overflow, or issues with multi-precision integer libraries. The log line mpz_manager<>::machine_div() is a key clue: this points to problems with GMP (GNU Multiple Precision Arithmetic Library) multi-precision integer division.
- Audit all code paths involving division, modulus operations, or libraries that rely on GMP (e.g., SymPy, PyCryptodome, gmpy2). Look for cases where a divisor might unexpectedly become
0in parallel execution. - Confirm you're using identical input datasets across local Mac and remote servers—data discrepancies could explain why the error only appears in one environment.
2. Validate Dependency & Environment Cross-Platform Differences
Local Mac (especially ARM-based M-series) and remote Linux (x86_64) often have subtle differences in library compilation or versions that break numerical consistency:
- Compare
pip freezeoutputs from both environments to check for mismatched versions of critical packages (numpy, gmpy2, sympy, etc.). - Verify GMP library versions:
- On Debian/Ubuntu:
apt show libgmp-dev - On CentOS/RHEL:
yum info gmp - On Mac (Homebrew):
brew info gmp
Different GMP versions may handle edge-case division operations differently.
- On Debian/Ubuntu:
3. Build a Minimal Reproducible Example
Strip down your code to the smallest snippet that triggers the crash—this eliminates noise from unrelated logic. For example, with Ray:
import ray from your_library import problematic_function ray.init() @ray.remote def crash_trigger_task(input_data): # Isolate only the calculation that fails divisor = get_critical_divisor(input_data) print(f"Divisor value: {divisor}") # Add debug output return problematic_function(input_data, divisor) # Test with the exact input that caused the crash test_input = [your_faulty_input] ray.get([crash_trigger_task.remote(i) for i in test_input])
Run this minimal example on both Mac and Linux to confirm the crash is environment-specific.
4. Add Detailed Debug Logs to Parallel Tasks
Inject logging into your task functions to capture the state of variables right before the crash:
@ray.remote def debug_task(input_val): print(f"Task ID: {ray.get_runtime_context().task_id} | Input: {input_val}") # Log all variables involved in division/multi-precision operations numerator = compute_numerator(input_val) divisor = compute_divisor(input_val) print(f"Numerator: {numerator} | Divisor: {divisor}") # Execute the problematic operation result = numerator // divisor # Or your actual operation return result
This will reveal if unexpected values (like 0 for a divisor) are appearing only in the remote environment.
5. Debug the Worker Crash in Real-Time
Get deeper visibility into the crash context using debugging tools:
- For Ray, enable GDB debugging by setting the environment variable before starting Ray:
When a worker crashes, GDB will attach automatically and preserve a full stack trace.export RAY_WORKER_GDB=1 ray start --head # Or your usual Ray init command - Register a SIGFPE handler in your Python code to capture crash context:
This will dump critical state information when the exception occurs.import signal import traceback import inspect def sigfpe_handler(signum, frame): print("\n=== SIGFPE TRIGGERED ===") print("Stack Trace:") traceback.print_stack(frame) # Print local variables at the crash site local_vars = inspect.currentframe().f_back.f_locals print("\nLocal Variables:", local_vars) exit(1) signal.signal(signal.SIGFPE, sigfpe_handler)
6. Rule Out System-Level Hardware/Configuration Issues
- CPU Architecture: If your Mac uses ARM (M1/M2) and remote servers use x86_64, C extensions for numerical libraries may have compatibility bugs. Test your code on an x86 Mac or use an x86 Docker container on your local machine to see if the crash reproduces.
- Memory Stability: Remote server memory shortages or hardware faults can corrupt data and trigger unexpected numerical errors. Check memory usage with
free -hand run a memory test (e.g.,memtest86+) to rule out hardware issues. - FPU Settings: Linux systems may have different floating-point unit (FPU) control settings. Use
numpy.geterr()to compare floating-point error handling between environments, and adjust if needed.
内容的提问来源于stack exchange,提问作者Kshitij Goyal

