You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

多线程环境下popen与fgets是否存在跨线程阻塞?并行执行Shell命令未提速的解决方案咨询

为什么用std::async并行执行Shell命令总耗时与串行一致?

问题描述

我尝试从5个Shell命令中提取数据,每个命令执行耗时约3秒,因此想用并行执行提升效率。我通过std::async以std::launch::async模式启动每个命令的生成器函数,但多线程版本和串行版本总耗时均约15秒。我怀疑在popen创建的管道上调用fgets会产生跨线程阻塞,想咨询这个问题是否存在,以及可行的解决办法。

我的并行实现代码如下:

std::vector<std::future<Table>> tables;
for (size_t i = 0; i < generators.size(); i++) {
    tables.push_back(std::async(std::launch::async, generators[i]));
}
for (size_t i = 0; i < tables.size(); i++) {
    tables[i].wait();
}
// 后续通过tables[i].get()获取结果

每个生成器最终调用的exec方法:

std::vector<std::string> exec(std::string const& cmd) {
    std::array<char, 128> buffer;
    std::vector<std::string> result;
    std::unique_ptr<FILE, decltype(&pclose)> pipe(popen(cmd.c_str(), "r"), pclose);
    if (pipe) {
        std::string line;
        while (fgets(buffer.data(), buffer.size(), pipe.get()) != nullptr) {
            line += buffer.data();
            if (line.back() == '\n') {
                line.pop_back();
                result.push_back(std::move(line));
            }
        }
        if (!line.empty()) {
            result.push_back(std::move(line));
        }
    }
    return result;
}

回答

首先明确:popen创建的管道是每个子进程独立持有的,每个线程中调用fgets读取自己的管道不会产生跨线程阻塞。每个管道的文件描述符是独立的,线程之间的IO操作互不干扰,所以这不是你总耗时未下降的原因。

你的并行代码没达到预期效果,大概率是以下几个原因之一,我们逐个排查:

1. 并行线程并未真正创建

虽然你指定了std::launch::async,但某些编译器/运行环境可能存在特殊处理(比如资源不足时回退到同步执行),或者你的generators函数指针指向的逻辑有隐含的串行限制。

验证方法:在每个生成器函数的开头添加线程ID打印:

#include <thread>
// ... 在生成器函数内
std::cout << "Running in thread: " << std::this_thread::get_id() << std::endl;

如果所有输出的线程ID都相同,说明std::async没有创建新线程,而是在主线程串行执行了所有任务。这种情况下,你可以尝试手动创建std::thread来代替std::async,确保并行执行:

std::vector<std::thread> threads;
std::vector<Table> results(generators.size());
for (size_t i = 0; i < generators.size(); i++) {
    threads.emplace_back([i, &generators, &results]() {
        results[i] = generators[i]();
    });
}
for (auto& t : threads) {
    t.join();
}

2. Shell命令本身无法并行执行

如果你的5个Shell命令都依赖同一个排他性资源(比如同时读取同一个机械硬盘的某个文件、访问同一个被限流的网络接口,或者调用了某个单进程服务),那么即使你开了多线程,这些命令也会串行等待资源,总耗时自然和串行执行一致。

验证方法:手动在终端同时运行其中两个命令,比如:

cmd1 & cmd2 && wait

如果总耗时接近3秒(而非6秒),说明命令可以并行;如果还是接近6秒,那问题出在命令本身的资源依赖上,需要优化命令逻辑(比如改用并行友好的工具、拆分资源访问等)。

3. exec函数的IO效率问题

虽然这不会导致总耗时等于串行时间,但如果exec的读取逻辑效率极低,可能会掩盖并行的收益。你的exec函数中用fgets逐行拼接字符串的方式没问题,但可以尝试优化读取效率,比如直接用read系统调用减少用户态/内核态切换:

std::vector<std::string> exec(std::string const& cmd) {
    std::vector<std::string> result;
    int pipefd[2];
    if (pipe(pipefd) == -1) {
        return result;
    }
    pid_t pid = fork();
    if (pid == -1) {
        close(pipefd[0]);
        close(pipefd[1]);
        return result;
    }
    if (pid == 0) {
        close(pipefd[0]);
        dup2(pipefd[1], STDOUT_FILENO);
        close(pipefd[1]);
        execl("/bin/sh", "sh", "-c", cmd.c_str(), nullptr);
        exit(EXIT_FAILURE);
    }
    close(pipefd[1]);
    std::array<char, 4096> buffer;
    ssize_t bytes_read;
    std::string line;
    while ((bytes_read = read(pipefd[0], buffer.data(), buffer.size())) > 0) {
        for (ssize_t i = 0; i < bytes_read; i++) {
            char c = buffer[i];
            if (c == '\n') {
                if (!line.empty()) {
                    result.push_back(std::move(line));
                }
            } else {
                line += c;
            }
        }
    }
    if (!line.empty()) {
        result.push_back(std::move(line));
    }
    close(pipefd[0]);
    waitpid(pid, nullptr, 0);
    return result;
}

这个版本用fork/exec代替popen,直接操作文件描述符,读取效率会更高,但前提是你的并行问题已经解决,否则优化IO也不会改变总耗时。


总结

先优先验证并行线程是否真的创建,再检查Shell命令本身的并行可行性,最后再考虑IO逻辑的优化。fgets跨线程阻塞的问题不存在,不用在这方面浪费精力。

内容的提问来源于stack exchange,提问作者Neträm

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 22:37:49