You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何多线程下locale会导致std::ostringstream性能下降?

问题:多线程下std::ostringstream因locale导致性能暴跌

我用std::ostringstream构建格式化字符串,单线程时性能无瓶颈,但多线程场景下,std::__1::locale::locale导致std::ostringstream速度骤降,线程数量越多性能下降越明显。我未做显式线程同步,但推测std::__1::locale::locale内部存在线程阻塞逻辑——单线程跑30秒的任务,10线程居然耗时长达10分钟。

涉事代码简短但调用极频繁:

static std::string to_string(const T d) {
    std::ostringstream stream;
    stream << d;

    return stream.str();
}

后来改成复用thread_local的std::ostringstream,多线程性能得以恢复,但单线程性能反而受损:

thread_local static std::ostringstream stream;
const std::string clear;

static std::string to_string(const T d) {
    stream.str(clear);
    stream << d;

    return stream.str();
}

我需要解决这个矛盾:既保证多线程性能,又不损失单线程效率。另外,这些字符串不需要可读性,只是用来解决std::complex缺少哈希函数的问题——有没有办法构建字符串时避开本地化操作?


测试代码

#include <map>
#include <sstream>
#include <complex>
#include <iostream>
#include <thread>
#include <chrono>

thread_local std::map<std::string, void *> cache;

int main(int argc, const char * argv[]) {
    for (size_t i = 1; i <= 10; i++) {
        const std::chrono::high_resolution_clock::time_point start = std::chrono::high_resolution_clock::now();
        std::vector<std::thread> threads(i);
        for (auto &t : threads) {
            t = std::thread([] () -> void {
                for (size_t j = 0; j < 1000000; j++) {
                    std::ostringstream stream;
                    stream << std::complex<double> (static_cast<double> (j));
                    cache[stream.str()] = reinterpret_cast<void *> (&j);
                }
            });
        }
        for (auto &t : threads) {
            t.join();
        }
        
        const std::chrono::high_resolution_clock::time_point end =
                  std::chrono::high_resolution_clock::now();
        const auto total_time = end - start;
        const std::chrono::nanoseconds total_time_ns =
                  std::chrono::duration_cast<std::chrono::nanoseconds> (total_time);

        if (total_time_ns.count() < 1000) {
            std::cout << total_time_ns.count()               << " ns"  << std::endl;
        } else if (total_time_ns.count() < 1000000) {
            std::cout << total_time_ns.count()/1000.0        << " μs"  << std::endl;
        } else if (total_time_ns.count() < 1000000000) {
            std::cout << total_time_ns.count()/1000000.0     << " ms"  << std::endl;
        } else if (total_time_ns.count() < 60000000000) {
            std::cout << total_time_ns.count()/1000000000.0  << " s"   << std::endl;
        } else if (total_time_ns.count() < 3600000000000) {
            std::cout << total_time_ns.count()/60000000000.0 << " min" << std::endl;
        } else {
            std::cout << total_time_ns.count()/3600000000000 << " h"   << std::endl;
        }
        std::cout << std::endl;
    }

    return 0;
}

运行环境与结果

运行环境:10核Apple M1(8性能核+2能效核),Xcode默认构建设置

Debug构建耗时

3.90096 s

4.15853 s

4.48616 s

4.843 s

6.15202 s

8.14986 s

10.6319 s

12.7732 s

16.7492 s

19.9288 s

Release构建耗时

844.28 ms

1.23803 s

2.05088 s

3.39994 s

7.43743 s

9.53968 s

11.2953 s

12.6878 s

20.3917 s

24.1944 s

解决方案

1. 避开locale:注入全局C locale

每次创建std::ostringstream时,直接设置其locale为全局C locale,避免初始化和使用系统默认locale时的线程竞争:

static std::string to_string(const T d) {
    std::ostringstream stream;
    stream.imbue(std::locale::classic()); // 使用无状态、线程安全的C locale
    stream << d;
    return stream.str();
}

C locale无本地化开销,不会触发多线程阻塞,同时单线程性能不受影响。

2. 跳过格式化:直接序列化二进制

既然字符串不需要可读性,完全可以跳过格式化,直接把std::complex的二进制内容转成字符串,性能提升显著:

static std::string to_string(const std::complex<double>& d) {
    return std::string(reinterpret_cast<const char*>(&d), sizeof(d));
}

该方式直接复用内存数据,无任何格式化和locale开销,单线程、多线程性能均拉满。若仅用于本地哈希,无需考虑字节序问题;若需跨平台兼容,可额外处理字节序。

3. 优化thread_local复用方案:重置流状态

之前的thread_local方案单线程性能下降,大概率是因为只清空了字符串,未重置流的状态(如错误位、格式标志等)。改进后:

thread_local static std::ostringstream stream;

static std::string to_string(const T d) {
    stream.str({}); // 清空字符串缓冲区
    stream.clear(); // 重置流的状态标志
    stream << d;
    return stream.str();
}

既保留了thread_local的复用优势,又避免了流状态残留导致的性能损耗或逻辑问题。


内容的提问来源于stack exchange,提问作者user1139069

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 05:45:07