如何在C语言中实现CPU多核的真正并行计算?
C语言利用CPU多核实现计算密集型任务并行
先澄清一个关键误解
你提到“线程似乎运行在同一核心,计算密集型任务无法提升性能”——这是Python特有的GIL限制导致的,C语言中的原生线程(比如POSIX Threads/pthread)完全可以被操作系统调度到不同CPU核心上并行执行,只要你的任务是计算密集且无共享资源竞争瓶颈,线程就能带来实打实的多核性能提升。
两种核心实现方式
1. 多线程(pthread,适合内存共享场景)
POSIX线程是Unix/Linux/macOS下的标准多线程API,完全支持多核并行。计算密集型任务如果需要共享部分内存数据,用线程会更高效(不用进程间通信的开销)。
简单示例:
#include <stdio.h> #include <pthread.h> #include <stdlib.h> #define NUM_THREADS 4 #define TASK_SIZE 1000000 // 每个线程要执行的计算任务 void* compute_task(void* arg) { int thread_id = *(int*)arg; long sum = 0; // 分配每个线程的计算区间 int start = thread_id * (TASK_SIZE / NUM_THREADS); int end = (thread_id + 1) * (TASK_SIZE / NUM_THREADS); for (int i = start; i < end; i++) { sum += i * i; // 模拟计算密集型操作 } printf("Thread %d: sum = %ld\n", thread_id, sum); pthread_exit((void*)sum); } int main() { pthread_t threads[NUM_THREADS]; int thread_ids[NUM_THREADS]; long total_sum = 0; // 创建线程 for (int i = 0; i < NUM_THREADS; i++) { thread_ids[i] = i; pthread_create(&threads[i], NULL, compute_task, &thread_ids[i]); } // 等待线程完成并汇总结果 for (int i = 0; i < NUM_THREADS; i++) { long thread_sum; pthread_join(threads[i], (void**)&thread_sum); total_sum += thread_sum; } printf("Total sum: %ld\n", total_sum); return 0; }
编译运行(macOS下):gcc -o multi_thread multi_thread.c -lpthread && ./multi_thread
这个例子里,4个线程会被系统调度到你的四核MacBook的不同核心上,并行完成计算任务,性能会接近单线程的4倍(忽略调度开销)。
2. 多进程(fork,适合隔离性需求场景)
用fork()创建的子进程,每个进程都有独立的地址空间,操作系统同样会把它们调度到不同CPU核心上并行执行,完全能提升计算密集型任务的性能——并非只适用于IO密集型操作。
简单示例:
#include <stdio.h> #include <unistd.h> #include <sys/wait.h> #include <stdlib.h> #define NUM_PROCESSES 4 #define TASK_SIZE 1000000 // 子进程执行的计算任务 void compute_task(int process_id) { long sum = 0; int start = process_id * (TASK_SIZE / NUM_PROCESSES); int end = (process_id + 1) * (TASK_SIZE / NUM_PROCESSES); for (int i = start; i < end; i++) { sum += i * i; } printf("Process %d: sum = %ld\n", process_id, sum); exit(sum); } int main() { long total_sum = 0; for (int i = 0; i < NUM_PROCESSES; i++) { pid_t pid = fork(); if (pid == 0) { // 子进程 compute_task(i); } else if (pid > 0) { // 父进程,等待子进程完成并获取退出状态(这里用退出码传递结果,仅适合小数值) int status; waitpid(pid, &status, 0); if (WIFEXITED(status)) { total_sum += WEXITSTATUS(status); } } else { perror("fork failed"); exit(1); } } printf("Total sum: %ld\n", total_sum); return 0; }
编译运行:gcc -o multi_process multi_process.c && ./multi_process
注意:如果需要传递大量数据,进程间需要用管道、共享内存等IPC机制,这是多进程相比多线程的额外开销,但隔离性更好(一个进程崩溃不影响其他进程)。
选择建议
- 如果你的计算任务需要频繁共享数据,优先用多线程,避免IPC开销;
- 如果任务需要严格隔离(比如每个任务可能崩溃),或者你想避免线程同步的复杂度,用多进程;
- 两者都能实现真正的CPU多核并行,完全满足你的需求。
内容的提问来源于stack exchange,提问作者Sebastian Gudiño
相关产品推荐
相关产品推荐

