为何多线程下locale会导致std::ostringstream性能下降?
我用std::ostringstream构建格式化字符串,单线程时性能无瓶颈,但多线程场景下,std::__1::locale::locale导致std::ostringstream速度骤降,线程数量越多性能下降越明显。我未做显式线程同步,但推测std::__1::locale::locale内部存在线程阻塞逻辑——单线程跑30秒的任务,10线程居然耗时长达10分钟。
涉事代码简短但调用极频繁:
static std::string to_string(const T d) { std::ostringstream stream; stream << d; return stream.str(); }
后来改成复用thread_local的std::ostringstream,多线程性能得以恢复,但单线程性能反而受损:
thread_local static std::ostringstream stream; const std::string clear; static std::string to_string(const T d) { stream.str(clear); stream << d; return stream.str(); }
我需要解决这个矛盾:既保证多线程性能,又不损失单线程效率。另外,这些字符串不需要可读性,只是用来解决std::complex缺少哈希函数的问题——有没有办法构建字符串时避开本地化操作?
测试代码
#include <map> #include <sstream> #include <complex> #include <iostream> #include <thread> #include <chrono> thread_local std::map<std::string, void *> cache; int main(int argc, const char * argv[]) { for (size_t i = 1; i <= 10; i++) { const std::chrono::high_resolution_clock::time_point start = std::chrono::high_resolution_clock::now(); std::vector<std::thread> threads(i); for (auto &t : threads) { t = std::thread([] () -> void { for (size_t j = 0; j < 1000000; j++) { std::ostringstream stream; stream << std::complex<double> (static_cast<double> (j)); cache[stream.str()] = reinterpret_cast<void *> (&j); } }); } for (auto &t : threads) { t.join(); } const std::chrono::high_resolution_clock::time_point end = std::chrono::high_resolution_clock::now(); const auto total_time = end - start; const std::chrono::nanoseconds total_time_ns = std::chrono::duration_cast<std::chrono::nanoseconds> (total_time); if (total_time_ns.count() < 1000) { std::cout << total_time_ns.count() << " ns" << std::endl; } else if (total_time_ns.count() < 1000000) { std::cout << total_time_ns.count()/1000.0 << " μs" << std::endl; } else if (total_time_ns.count() < 1000000000) { std::cout << total_time_ns.count()/1000000.0 << " ms" << std::endl; } else if (total_time_ns.count() < 60000000000) { std::cout << total_time_ns.count()/1000000000.0 << " s" << std::endl; } else if (total_time_ns.count() < 3600000000000) { std::cout << total_time_ns.count()/60000000000.0 << " min" << std::endl; } else { std::cout << total_time_ns.count()/3600000000000 << " h" << std::endl; } std::cout << std::endl; } return 0; }
运行环境与结果
运行环境:10核Apple M1(8性能核+2能效核),Xcode默认构建设置
Debug构建耗时
3.90096 s 4.15853 s 4.48616 s 4.843 s 6.15202 s 8.14986 s 10.6319 s 12.7732 s 16.7492 s 19.9288 s
Release构建耗时
844.28 ms 1.23803 s 2.05088 s 3.39994 s 7.43743 s 9.53968 s 11.2953 s 12.6878 s 20.3917 s 24.1944 s
解决方案
1. 避开locale:注入全局C locale
每次创建std::ostringstream时,直接设置其locale为全局C locale,避免初始化和使用系统默认locale时的线程竞争:
static std::string to_string(const T d) { std::ostringstream stream; stream.imbue(std::locale::classic()); // 使用无状态、线程安全的C locale stream << d; return stream.str(); }
C locale无本地化开销,不会触发多线程阻塞,同时单线程性能不受影响。
2. 跳过格式化:直接序列化二进制
既然字符串不需要可读性,完全可以跳过格式化,直接把std::complex的二进制内容转成字符串,性能提升显著:
static std::string to_string(const std::complex<double>& d) { return std::string(reinterpret_cast<const char*>(&d), sizeof(d)); }
该方式直接复用内存数据,无任何格式化和locale开销,单线程、多线程性能均拉满。若仅用于本地哈希,无需考虑字节序问题;若需跨平台兼容,可额外处理字节序。
3. 优化thread_local复用方案:重置流状态
之前的thread_local方案单线程性能下降,大概率是因为只清空了字符串,未重置流的状态(如错误位、格式标志等)。改进后:
thread_local static std::ostringstream stream; static std::string to_string(const T d) { stream.str({}); // 清空字符串缓冲区 stream.clear(); // 重置流的状态标志 stream << d; return stream.str(); }
既保留了thread_local的复用优势,又避免了流状态残留导致的性能损耗或逻辑问题。
内容的提问来源于stack exchange,提问作者user1139069

