You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在C语言中实现CPU多核的真正并行计算?

C语言利用CPU多核实现计算密集型任务并行

先澄清一个关键误解

你提到“线程似乎运行在同一核心,计算密集型任务无法提升性能”——这是Python特有的GIL限制导致的,C语言中的原生线程(比如POSIX Threads/pthread)完全可以被操作系统调度到不同CPU核心上并行执行,只要你的任务是计算密集且无共享资源竞争瓶颈,线程就能带来实打实的多核性能提升。

两种核心实现方式

1. 多线程(pthread,适合内存共享场景)

POSIX线程是Unix/Linux/macOS下的标准多线程API,完全支持多核并行。计算密集型任务如果需要共享部分内存数据,用线程会更高效(不用进程间通信的开销)。

简单示例:

#include <stdio.h>
#include <pthread.h>
#include <stdlib.h>

#define NUM_THREADS 4
#define TASK_SIZE 1000000

// 每个线程要执行的计算任务
void* compute_task(void* arg) {
    int thread_id = *(int*)arg;
    long sum = 0;
    // 分配每个线程的计算区间
    int start = thread_id * (TASK_SIZE / NUM_THREADS);
    int end = (thread_id + 1) * (TASK_SIZE / NUM_THREADS);
    
    for (int i = start; i < end; i++) {
        sum += i * i; // 模拟计算密集型操作
    }
    printf("Thread %d: sum = %ld\n", thread_id, sum);
    pthread_exit((void*)sum);
}

int main() {
    pthread_t threads[NUM_THREADS];
    int thread_ids[NUM_THREADS];
    long total_sum = 0;
    
    // 创建线程
    for (int i = 0; i < NUM_THREADS; i++) {
        thread_ids[i] = i;
        pthread_create(&threads[i], NULL, compute_task, &thread_ids[i]);
    }
    
    // 等待线程完成并汇总结果
    for (int i = 0; i < NUM_THREADS; i++) {
        long thread_sum;
        pthread_join(threads[i], (void**)&thread_sum);
        total_sum += thread_sum;
    }
    
    printf("Total sum: %ld\n", total_sum);
    return 0;
}

编译运行(macOS下):gcc -o multi_thread multi_thread.c -lpthread && ./multi_thread

这个例子里,4个线程会被系统调度到你的四核MacBook的不同核心上,并行完成计算任务,性能会接近单线程的4倍(忽略调度开销)。

2. 多进程(fork,适合隔离性需求场景)

用fork()创建的子进程,每个进程都有独立的地址空间,操作系统同样会把它们调度到不同CPU核心上并行执行,完全能提升计算密集型任务的性能——并非只适用于IO密集型操作。

简单示例:

#include <stdio.h>
#include <unistd.h>
#include <sys/wait.h>
#include <stdlib.h>

#define NUM_PROCESSES 4
#define TASK_SIZE 1000000

// 子进程执行的计算任务
void compute_task(int process_id) {
    long sum = 0;
    int start = process_id * (TASK_SIZE / NUM_PROCESSES);
    int end = (process_id + 1) * (TASK_SIZE / NUM_PROCESSES);
    
    for (int i = start; i < end; i++) {
        sum += i * i;
    }
    printf("Process %d: sum = %ld\n", process_id, sum);
    exit(sum);
}

int main() {
    long total_sum = 0;
    
    for (int i = 0; i < NUM_PROCESSES; i++) {
        pid_t pid = fork();
        if (pid == 0) {
            // 子进程
            compute_task(i);
        } else if (pid > 0) {
            // 父进程,等待子进程完成并获取退出状态(这里用退出码传递结果,仅适合小数值)
            int status;
            waitpid(pid, &status, 0);
            if (WIFEXITED(status)) {
                total_sum += WEXITSTATUS(status);
            }
        } else {
            perror("fork failed");
            exit(1);
        }
    }
    
    printf("Total sum: %ld\n", total_sum);
    return 0;
}

编译运行:gcc -o multi_process multi_process.c && ./multi_process

注意:如果需要传递大量数据,进程间需要用管道、共享内存等IPC机制,这是多进程相比多线程的额外开销,但隔离性更好(一个进程崩溃不影响其他进程)。

选择建议

  • 如果你的计算任务需要频繁共享数据,优先用多线程,避免IPC开销;
  • 如果任务需要严格隔离(比如每个任务可能崩溃),或者你想避免线程同步的复杂度,用多进程;
  • 两者都能实现真正的CPU多核并行,完全满足你的需求。

内容的提问来源于stack exchange,提问作者Sebastian Gudiño

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 14:01:05