You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Ubuntu系统中使用Ray进行并行计算时Worker崩溃问题求助

Troubleshooting Cross-Platform SIGFPE Crash in Parallel Tasks (Ray/Dask/Scoop)

I've run into similar cross-platform parallel computing crashes before, so let's walk through targeted troubleshooting steps to get to the bottom of this SIGFPE error that's plaguing your tasks on Linux/Google Cloud but working fine locally on your Mac.

1. Pinpoint the Root of the SIGFPE Exception

SIGFPE (Signal Floating-Point Exception) isn't just limited to floating-point operations—it can also trigger from integer division by zero, integer overflow, or issues with multi-precision integer libraries. The log line mpz_manager<>::machine_div() is a key clue: this points to problems with GMP (GNU Multiple Precision Arithmetic Library) multi-precision integer division.

  • Audit all code paths involving division, modulus operations, or libraries that rely on GMP (e.g., SymPy, PyCryptodome, gmpy2). Look for cases where a divisor might unexpectedly become 0 in parallel execution.
  • Confirm you're using identical input datasets across local Mac and remote servers—data discrepancies could explain why the error only appears in one environment.

2. Validate Dependency & Environment Cross-Platform Differences

Local Mac (especially ARM-based M-series) and remote Linux (x86_64) often have subtle differences in library compilation or versions that break numerical consistency:

  • Compare pip freeze outputs from both environments to check for mismatched versions of critical packages (numpy, gmpy2, sympy, etc.).
  • Verify GMP library versions:
    • On Debian/Ubuntu: apt show libgmp-dev
    • On CentOS/RHEL: yum info gmp
    • On Mac (Homebrew): brew info gmp
      Different GMP versions may handle edge-case division operations differently.

3. Build a Minimal Reproducible Example

Strip down your code to the smallest snippet that triggers the crash—this eliminates noise from unrelated logic. For example, with Ray:

import ray
from your_library import problematic_function

ray.init()

@ray.remote
def crash_trigger_task(input_data):
    # Isolate only the calculation that fails
    divisor = get_critical_divisor(input_data)
    print(f"Divisor value: {divisor}")  # Add debug output
    return problematic_function(input_data, divisor)

# Test with the exact input that caused the crash
test_input = [your_faulty_input]
ray.get([crash_trigger_task.remote(i) for i in test_input])

Run this minimal example on both Mac and Linux to confirm the crash is environment-specific.

4. Add Detailed Debug Logs to Parallel Tasks

Inject logging into your task functions to capture the state of variables right before the crash:

@ray.remote
def debug_task(input_val):
    print(f"Task ID: {ray.get_runtime_context().task_id} | Input: {input_val}")
    # Log all variables involved in division/multi-precision operations
    numerator = compute_numerator(input_val)
    divisor = compute_divisor(input_val)
    print(f"Numerator: {numerator} | Divisor: {divisor}")
    # Execute the problematic operation
    result = numerator // divisor  # Or your actual operation
    return result

This will reveal if unexpected values (like 0 for a divisor) are appearing only in the remote environment.

5. Debug the Worker Crash in Real-Time

Get deeper visibility into the crash context using debugging tools:

  • For Ray, enable GDB debugging by setting the environment variable before starting Ray:
    export RAY_WORKER_GDB=1
    ray start --head  # Or your usual Ray init command
    
    When a worker crashes, GDB will attach automatically and preserve a full stack trace.
  • Register a SIGFPE handler in your Python code to capture crash context:
    import signal
    import traceback
    import inspect
    
    def sigfpe_handler(signum, frame):
        print("\n=== SIGFPE TRIGGERED ===")
        print("Stack Trace:")
        traceback.print_stack(frame)
        # Print local variables at the crash site
        local_vars = inspect.currentframe().f_back.f_locals
        print("\nLocal Variables:", local_vars)
        exit(1)
    
    signal.signal(signal.SIGFPE, sigfpe_handler)
    
    This will dump critical state information when the exception occurs.

6. Rule Out System-Level Hardware/Configuration Issues

  • CPU Architecture: If your Mac uses ARM (M1/M2) and remote servers use x86_64, C extensions for numerical libraries may have compatibility bugs. Test your code on an x86 Mac or use an x86 Docker container on your local machine to see if the crash reproduces.
  • Memory Stability: Remote server memory shortages or hardware faults can corrupt data and trigger unexpected numerical errors. Check memory usage with free -h and run a memory test (e.g., memtest86+) to rule out hardware issues.
  • FPU Settings: Linux systems may have different floating-point unit (FPU) control settings. Use numpy.geterr() to compare floating-point error handling between environments, and adjust if needed.

内容的提问来源于stack exchange,提问作者Kshitij Goyal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 07:17:50