You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

C语言Pthreads多线程动态分块迭代计算问题求助

Dynamic Work Assignment with Pthreads (Per-Block Task Claiming)

Got it, let's tackle your dynamic work distribution problem. The static block assignment you have now works, but to make threads claim a fixed-size block (10 iterations each) once they finish their current one, we need a shared, thread-safe way to track which blocks are still unprocessed. Here's how to implement it:

Key Changes Needed

  • Add a shared global counter to track the next available block's starting index
  • Protect this counter with the existing mutex to prevent race conditions when threads claim blocks
  • Rewrite the do_work function to use a while loop that keeps claiming new blocks until all iterations are done

Modified Full Code

#include <pthread.h>
#include <stdio.h>
#include <stdlib.h>

#define NTHREADS 4
#define ARRAYSIZE 1000000
#define BLOCK_SIZE 10  // Fixed block size per task

double sum = 0.0, a[ARRAYSIZE];
pthread_mutex_t sum_mutex;
int next_block_start = 0;  // Tracks the start of the next unprocessed block

void *do_work(void *tid) {
    int mytid = *(int *)tid;
    double mysum = 0.0;
    int start, end;

    printf("Thread %d started, ready to claim blocks\n", mytid);

    while (1) {
        // Lock to safely claim the next block
        pthread_mutex_lock(&sum_mutex);
        start = next_block_start;
        // Calculate end of current block, don't go beyond ARRAYSIZE
        end = (start + BLOCK_SIZE) > ARRAYSIZE ? ARRAYSIZE : start + BLOCK_SIZE;
        // Exit loop if no more work left
        if (start >= ARRAYSIZE) {
            pthread_mutex_unlock(&sum_mutex);
            break;
        }
        next_block_start = end;
        pthread_mutex_unlock(&sum_mutex);

        // Process the assigned block (no lock needed here—each thread works on unique indices)
        printf("Thread %d processing iterations %d to %d\n", mytid, start, end - 1);
        for (int i = start; i < end; i++) {
            a[i] = i * 1.0;
            mysum += a[i];
        }
    }

    // Lock to update global sum with thread's accumulated local sum
    pthread_mutex_lock(&sum_mutex);
    sum += mysum;
    pthread_mutex_unlock(&sum_mutex);

    printf("Thread %d finished, total local sum: %e\n", mytid, mysum);
    pthread_exit(NULL);
}

int main(int argc, char *argv[]) {
    int tids[NTHREADS];
    pthread_t threads[NTHREADS];
    pthread_attr_t attr;

    // Pthreads setup
    pthread_mutex_init(&sum_mutex, NULL);
    pthread_attr_init(&attr);
    pthread_attr_setdetachstate(&attr, PTHREAD_CREATE_JOINABLE);

    // Create threads
    for (int i = 0; i < NTHREADS; i++) {
        tids[i] = i;
        pthread_create(&threads[i], &attr, do_work, (void *)&tids[i]);
    }

    // Wait for all threads to complete
    for (int i = 0; i < NTHREADS; i++) {
        pthread_join(threads[i], NULL);
    }

    // Print and verify sum
    printf("\nDone. Parallel Sum= %e \n", sum);
    sum = 0.0;
    for (int i = 0; i < ARRAYSIZE; i++) {
        a[i] = i * 1.0;
        sum += a[i];
    }
    printf("Check Sum= %e\n", sum);

    // Cleanup
    pthread_attr_destroy(&attr);
    pthread_mutex_destroy(&sum_mutex);
    pthread_exit(NULL);
}

How This Works

  1. Shared Counter: next_block_start keeps track of where the next unprocessed block begins. Threads lock the mutex before accessing this value to avoid two threads claiming the same block.
  2. Block Claiming: Each thread checks if there's remaining work (start < ARRAYSIZE). If yes, it takes the current block, updates the counter for the next thread, then unlocks the mutex to let others claim blocks while it processes its own.
  3. Local Sum Calculation: Each thread accumulates its own mysum for the blocks it processes, which reduces mutex overhead—we only lock once per block claim and once at the end to update the global sum, instead of locking for every iteration.
  4. Termination: When next_block_start exceeds ARRAYSIZE, threads exit the loop, add their local sum to the global sum, and exit.

Why This Is Better Than Static Assignment

  • Load Balancing: If some threads finish faster (e.g., due to system scheduling), they'll immediately claim more work instead of idling, which is more efficient for uneven workloads.
  • Flexibility: You can easily adjust BLOCK_SIZE to tune performance—smaller blocks mean more frequent claiming (more overhead), larger blocks mean less overhead but potentially worse load balancing.

内容的提问来源于stack exchange,提问作者BloodGreen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 07:46:51