为何fmt库fmt::output_file API未实现预期性能提升?技术求助
更新
我编写了一个C++小程序,分别使用系统fprintf和fmt::output_file API进行测试,但未观测到符合预期的运行时间差异。
使用fprintf的测试程序
#include <cstdio> #include <iostream> #include <chrono> #include <cstring> int main() { const char* filename = "large_file.txt"; const char* content = "Test"; auto start = std::chrono::high_resolution_clock::now(); FILE* fp = std::fopen(filename, "wb"); if (!fp) { std::cerr << "Error opening file: " << filename << std::endl; return 1; } size_t iterations = 1e7; for (size_t i = 0; i < iterations; ++i) { std::fprintf(fp, "%s", content); } std::fclose(fp); auto end = std::chrono::high_resolution_clock::now(); auto duration = std::chrono::duration_cast<std::chrono::milliseconds>(end - start).count(); std::cout << "Elapsed time: " << duration << " ms" << std::endl; return 0; }
输出结果:
[anhad fmt_tests]$ ./output_bench1 Elapsed time: 319 ms [anhad fmt_tests]$
使用fmt::output_file的测试程序
#include "fmt/core.h" #include "fmt/os.h" #include "fmt/ostream.h" #include "fmt/printf.h" #include <iostream> #include <chrono> int main() { const char* content = "Test"; const size_t bufferSize = 653539; auto start = std::chrono::high_resolution_clock::now(); fmt::ostream out = fmt::output_file("filename.txt", fmt::buffer_size= bufferSize); size_t iterations = 1e7; for (size_t i = 0; i < iterations; ++i) { out.print("{}", content); } auto end = std::chrono::high_resolution_clock::now(); auto duration = std::chrono::duration_cast<std::chrono::milliseconds>(end - start).count(); std::cout << "Elapsed time: " << duration << " ms" << std::endl; return 0; }
输出结果:
[anhad fmt_tests]$ ./output_fmt_bench Elapsed time: 1745 ms [anhad fmt_tests]$
生成的两个文件大小一致:
[anhad fmt_tests]$ du -sh large_file.txt filename.txt 39M large_file.txt 39M filename.txt [anhadp fmt_tests]$
问题描述
我近期在C++程序中开展性能优化探索,目标是提升程序运行效率。此前我们使用fprintf写入文件,sprintf进行字符转换。
为优化性能,我集成了fmt库(版本10),用fmt::output_file API替代fprintf,用fmt::format_to API替代sprintf。根据fmt库文档,fmt::output_file相比传统fprintf能带来5-9倍的运行效率提升,但实际测试中并未观测到预期的性能增益。
1. 文件打开方式
我在独立函数中使用fmt::ostream类结合fmt::output_file API初始化文件写入,代码片段如下:
fmt::ostream out = fmt::output_file(filename, fmt::buffer_size);
创建的out对象随后以引用方式传递给其他函数。
2. 缓冲区大小控制
我通过环境变量实现缓冲区大小的可配置性,以下代码展示了根据环境变量BUFFER_SIZE动态调整缓冲区大小的逻辑:
static int bufferSize = (bufferSizeEnv != nullptr) ? std::stoi(bufferSizeEnv) : 16384;
这让我能够根据应用需求微调缓冲区大小。
尽管我正确配置了fmt::output_file API,并尝试了从4k到65k的不同缓冲区大小,但惊讶地发现运行效率并未得到提升。fmt库文档承诺的显著性能提升并未兑现,预期与实际结果存在差距。
项目中使用fmt::output_file的示例代码片段:
FILE* outputFileApiFunc(string type, string outputfile ,fmt::ostream& out) { FILE *fp = NULL; fp = std::fopen(outputfile.c_str(),"w"); if(type == "one" || type == "two" || type == "three" || type == "all") out.print("\n type found = {}\n", func1.c_str()); else out.print("\n{} type not found\n\n",func2.c_str()); out.print("output is = {}\n", outputFunc.c_str()); return fp; }
而当我使用fmt库10版本的fmt::format_to API结合fmt::memory_buffer替代系统sprintf时,运行时间缩短了10分钟,达到了预期效果。
使用fmt::format_to的示例代码片段:
{ fmt::memory_buffer bufferString; File* fp = NULL; case name: fmt::format_to(std::back_inserter(bufferString),"{:<15} ", "Dave"); break; case age: fmt::format_to(std::back_inserter(bufferString),"{:<30} ", 17); break; fmt::print(fp,bufferString.data()); }
只有fmt::format_to API带来了性能提升,我原本预期fmt::output_file API也能达到类似效果。
我可能忽略了某些细节或存在性能瓶颈,求熟悉fmt库内部机制或有类似经验的人提供技术建议。
内容的提问来源于stack exchange,提问作者Anhad Parashar

