You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

多线程系统中pthread_cond_timedwait与pthread_mutex_lock崩溃问题排查

Alright, let's break down what's happening here and fix this crash step by step.

Root Cause Analysis

From your code snippet and crash details, the main issues stem from incorrect condition variable usage, flawed parameter passing, and risky timeout handling:

1. You're passing the Data struct by value (fatal mistake)

Your Thread A function signature is void sleep(Data x)—this passes the entire Data struct as a copy. That means Thread A is operating on a local stack copy of the struct, not the same shared instance Thread B is using.

This creates a disjoint state:

  • Thread A locks a mutex that only exists in its own stack frame
  • When Thread A exits sleep, that copy (and its mutex/condition variable) is destroyed
  • Thread B ends up waiting on a completely different mutex, or worse—if it's referencing a destroyed mutex, that directly causes crashes. The stack trace showing value = 0 confirms this: Thread B is looking at a different value than the one Thread A modified.

2. Condition variables aren't paired with proper condition checks

POSIX condition variables have a hard rule: you must wrap pthread_cond_wait/pthread_cond_timedwait in a loop that checks your actual wait condition. This handles spurious wakeups (wait can return without a signal) and ensures you don't act on invalid state after a timeout.

In Thread A, you set value = 0 immediately after the wait finishes—no matter if it was woken up normally or timed out. This clobbers the value = 1 state Thread B sets, creating a race that breaks your synchronization logic.

3. Global timeout variable causes race hazards

It looks like T (the timespec for the timeout) is a global variable. If multiple threads call sleep, or even if Thread A calls it multiple times, they'll overwrite each other's timeout values. This makes the timeout behavior unpredictable, amplifying the race conditions that trigger your crash.


Fixes

Let's address these issues one by one:

1. Pass Data by pointer to share state

First, update both functions to accept a pointer to Data—this ensures both threads operate on the exact same shared struct:

void sleep(Data *x) {
    // All operations now use the shared struct via pointer
}

void wakeup(Data *y) {
    // Same here—operate on the shared instance
}

2. Wrap condition wait in a loop with proper checks

Modify Thread A to only wait while the condition isn't met (x->value != 1), and only reset value if the wait was successfully woken up (not timed out):

void sleep(Data *x) {
    pthread_mutex_lock( &(x->lock) );
    
    // Declare timeout as a local variable to avoid race conditions
    struct timespec T;
    struct timeval _T;
    gettimeofday( &_T, NULL );
    // Set your desired timeout here (example: 1 second from now)
    T.tv_sec = _T.tv_sec + 1;
    T.tv_nsec = _T.tv_usec * 1000; // Convert microseconds to nanoseconds

    // Loop to handle spurious wakeups and timeouts
    while (x->value != 1) {
        int ret = pthread_cond_timedwait( &(x->condition), &(x->lock), &T );
        if (ret == ETIMEDOUT) {
            // Handle timeout (e.g., log it) — don't force value to 0!
            break;
        }
    }

    // Only reset value if we were actually woken up by Thread B
    if (x->value == 1) {
        x->value = 0;
    }

    pthread_mutex_unlock( &(x->lock) );
}

3. Keep Thread B's logic (with pointer updates)

Thread B's core logic is correct—just ensure it uses the shared struct pointer:

void wakeup(Data *y) {
    pthread_mutex_lock( &(y->lock) );
    y->value = 1;
    pthread_cond_signal( &(y->condition) );
    pthread_mutex_unlock( &(y->lock) );
}

Why This Works
  • Shared state via pointer: Both threads now use the same mutex, condition variable, and value—no more disjoint synchronization primitives.
  • Condition loop: The loop ensures we only proceed when the expected state is met, avoiding spurious wakeups and incorrect state modifications after timeouts.
  • Local timeout variable: Eliminates race conditions where multiple threads overwrite the timeout value.
  • Targeted value reset: We only reset value to 0 when we know Thread B successfully woke us up, so Thread B's state changes aren't clobbered unexpectedly.

内容的提问来源于stack exchange,提问作者Anu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 06:40:25