多进程与多线程性能测试异常:为何多线程运行更慢?
问题背景
我尝试用多进程和多线程执行同一任务(创建3个执行单元,每个从1打印到1000000),测试两者的运行时长。原本预期线程上下文切换更快,多线程方案速度会更快,但实际结果相反,想确认是否存在操作失误。
多进程实现代码
#include <stdio.h> #include <unistd.h> #include <wait.h> #include <semaphore.h> #include <fcntl.h> sem_t mutex; int main() { int i; int N = 1; close(STDOUT_FILENO); open("multiproc.txt", O_WRONLY | O_CREAT, 0666); sem_init(&mutex, 1, 1); for(i=0; i<2; i++) { pid_t child_pid = fork(); if(child_pid == 0) { N = i + 2; break; } } for(i=1; i<=1000000; i++) { sem_wait(&mutex); printf("Element: %d : %d\n", N, i); sem_post(&mutex); } if(N == 1) { while (wait(NULL) > 0); sem_destroy(&mutex); close(STDOUT_FILENO); } return 0; }
多线程实现代码
#include <stdio.h> #include <pthread.h> #include <semaphore.h> #include <unistd.h> #include <fcntl.h> sem_t mutex; void* mythread(void* arg) { int n = *(int*)arg; int i; for(i = 1; i <= 1000000; i++) { sem_wait(&mutex); printf("Element: %d : %d\n", n, i); sem_post(&mutex); } } int main(void) { close(STDOUT_FILENO); open("multithread.txt", O_WRONLY | O_CREAT, 0666); sem_init(&mutex, 0, 1); pthread_t thr[3]; for(int n = 0; n < 3; n++) pthread_create(&thr[n], NULL, mythread, &n); for(int n = 0; n < 3; n++) pthread_join(thr[n], NULL); sem_destroy(&mutex); close(STDOUT_FILENO); }
测试环境与结果
编译时使用gcc的-pthread和-lrt选项,测试时长如下:
abcd@abcd-Virtual-Machine:~/Desktop$ time ~/Desktop/multiproc real 0m0.173s user 0m0.281s sys 0m0.047s abcd@abcd-Virtual-Machine:~/Desktop$ time ~/Desktop/multithread real 0m1.186s user 0m0.664s sys 0m1.489s
问题分析与修正方案
1. 多进程代码的核心错误:信号量未实现跨进程互斥
你在多进程代码中使用了栈上的局部信号量sem_t mutex,并调用sem_init(&mutex, 1, 1)。这里第二个参数pshared=1表示要创建跨进程共享的信号量,但栈上的变量在fork后会被每个子进程复制一份,三个进程各自持有独立的信号量副本。这意味着sem_wait和sem_post只对当前进程的信号量生效,完全没有实现跨进程的互斥同步。
这种情况下,三个进程可以同时执行printf写入文件,相当于并行IO操作,所以总耗时极短,但输出的multiproc.txt内容会是混乱的(多个进程同时写同一个文件,内容会交织)。
修正方案:使用命名信号量实现跨进程同步
将栈上的信号量替换为命名信号量,确保所有进程共享同一个信号量:
#include <stdio.h> #include <unistd.h> #include <wait.h> #include <semaphore.h> #include <fcntl.h> int main() { int i; int N = 1; close(STDOUT_FILENO); open("multiproc.txt", O_WRONLY | O_CREAT, 0666); // 创建命名信号量,O_EXCL确保不存在则创建 sem_t *mutex = sem_open("/print_mutex", O_CREAT | O_EXCL, 0666, 1); for(i=0; i<2; i++) { pid_t child_pid = fork(); if(child_pid == 0) { N = i + 2; break; } } for(i=1; i<=1000000; i++) { sem_wait(mutex); printf("Element: %d : %d\n", N, i); sem_post(mutex); } if(N == 1) { while (wait(NULL) > 0); sem_close(mutex); sem_unlink("/print_mutex"); // 删除命名信号量 close(STDOUT_FILENO); } else { sem_close(mutex); } return 0; }
修正后,多进程会真正串行执行IO操作,此时耗时会和多线程接近甚至更长(因为进程上下文切换开销比线程大)。
2. 多线程代码的参数传递错误
创建线程时你传递的是&n(循环变量的地址),但主线程的循环会快速执行,线程启动后读取*(int*)arg时,n的值可能已经被主线程修改(比如三个线程可能都拿到n=3,或者拿到错误的编号),导致输出的编号不符合预期。
修正方案:传递独立的参数
可以用数组存储每个线程的编号,或者直接传递值(需保证int和void*在平台上大小兼容):
// 方法1:使用数组存储独立编号 #include <stdio.h> #include <pthread.h> #include <semaphore.h> #include <unistd.h> #include <fcntl.h> sem_t mutex; void* mythread(void* arg) { int n = *(int*)arg; int i; for(i = 1; i <= 1000000; i++) { sem_wait(&mutex); printf("Element: %d : %d\n", n, i); sem_post(&mutex); } return NULL; } int main(void) { close(STDOUT_FILENO); open("multithread.txt", O_WRONLY | O_CREAT, 0666); sem_init(&mutex, 0, 1); pthread_t thr[3]; int thread_nums[3] = {1, 2, 3}; // 每个线程的独立编号 for(int n = 0; n < 3; n++) pthread_create(&thr[n], NULL, mythread, &thread_nums[n]); for(int n = 0; n < 3; n++) pthread_join(thr[n], NULL); sem_destroy(&mutex); close(STDOUT_FILENO); return 0; }
3. 结果与预期相反的原因
你的测试场景中,核心任务是IO密集型的printf,并且通过信号量强制串行执行。多进程因为信号量未生效,实现了并行IO,所以速度快;而多线程正确实现了互斥,三个线程只能串行执行IO,加上线程调度的少量开销,总耗时自然更长。这种场景下,线程上下文切换快的优势完全无法体现——因为任务根本没有并发执行的机会。
内容的提问来源于stack exchange,提问作者HelloWorld

