为何CRTP未带来显著性能提升?求原因与最佳实践
CRTP与运行时多态的性能对比实验及疑问
实验背景与目的
我做了一个小型实验,想直观感受CRTP(编译时多态)相比运行时多态的性能提升——按理论来说,编译时多态的调用开销应该更低。
实验代码
#include <iostream> #include <chrono> namespace NonCRTP { class Base { public: virtual void name() = 0; }; class D1 : public Base { public: void name() override { volatile int _ = 1; } }; } // namespace NonCRTP namespace CRTP { template <class Derived> class Base { public: void name() { static_cast<Derived*>(this)->name_impl(); } protected: Base() = default; }; class D1 : public Base<D1> { public: void name_impl() { volatile int _ = 1; } }; } // namespace CRTP void perf_test(int invocation_counts) { std::chrono::nanoseconds ns_duration_noncrtp{}; std::chrono::nanoseconds ns_duration_crtp{}; { NonCRTP::D1 obj_noncrtp; auto time_s{ std::chrono::high_resolution_clock::now() }; for (int i = 0; i < invocation_counts; ++i) { obj_noncrtp.name(); } auto time_e{ std::chrono::high_resolution_clock::now() }; ns_duration_noncrtp = time_e - time_s; } { CRTP::D1 obj_crtp; auto time_s{ std::chrono::high_resolution_clock::now() }; for (int i = 0; i < invocation_counts; ++i) { obj_crtp.name(); } auto time_e{ std::chrono::high_resolution_clock::now() }; ns_duration_crtp = time_e - time_s; } std::printf("Perf test: %d times of invocation:\n", invocation_counts); std::printf(" | -- Non-CRTP: %lu ns\n", static_cast<uint64_t>(ns_duration_noncrtp.count())); std::printf(" | ------ CRTP: %lu ns\n", static_cast<uint64_t>(ns_duration_crtp.count())); std::printf("\n"); } int main() { perf_test(1); perf_test(10); perf_test(100); perf_test(1000); perf_test(10000); perf_test(100000); perf_test(1000000); perf_test(10000000); perf_test(100000000); return 0; }
实验结果
未开启-O2优化的测试结果
CRTP实现并未始终优于非CRTP版本,甚至在多数场景下更慢:
-> % clang++ crtp.cc -std=c++20 -o crtp; ./crtp Perf test: 1 times of invocation: | -- Non-CRTP: 100 ns | ------ CRTP: 31 ns Perf test: 10 times of invocation: | -- Non-CRTP: 49 ns | ------ CRTP: 59 ns Perf test: 100 times of invocation: | -- Non-CRTP: 171 ns | ------ CRTP: 239 ns Perf test: 1000 times of invocation: | -- Non-CRTP: 1360 ns | ------ CRTP: 2142 ns Perf test: 10000 times of invocation: | -- Non-CRTP: 16609 ns | ------ CRTP: 60303 ns Perf test: 100000 times of invocation: | -- Non-CRTP: 194464 ns | ------ CRTP: 234327 ns Perf test: 1000000 times of invocation: | -- Non-CRTP: 1489131 ns | ------ CRTP: 2251113 ns Perf test: 10000000 times of invocation: | -- Non-CRTP: 14182943 ns | ------ CRTP: 23018457 ns Perf test: 100000000 times of invocation: | -- Non-CRTP: 139714335 ns | ------ CRTP: 228018231 ns
开启-O2优化的测试结果
情况有所改善,多数场景下CRTP表现更优,但仍存在CRTP更慢的情况:
-> % clang++ crtp.cc -O2 -std=c++20 -o crtp; ./crtp Perf test: 1 times of invocation: | -- Non-CRTP: 100 ns | ------ CRTP: 27 ns Perf test: 10 times of invocation: | -- Non-CRTP: 30 ns | ------ CRTP: 34 ns Perf test: 100 times of invocation: | -- Non-CRTP: 67 ns | ------ CRTP: 78 ns Perf test: 1000 times of invocation: | -- Non-CRTP: 288 ns | ------ CRTP: 289 ns Perf test: 10000 times of invocation: | -- Non-CRTP: 2856 ns | ------ CRTP: 2822 ns Perf test: 100000 times of invocation: | -- Non-CRTP: 27937 ns | ------ CRTP: 27965 ns Perf test: 1000000 times of invocation: | -- Non-CRTP: 315920 ns | ------ CRTP: 270106 ns Perf test: 10000000 times of invocation: | -- Non-CRTP: 2632259 ns | ------ CRTP: 2705106 ns Perf test: 100000000 times of invocation: | -- Non-CRTP: 24681903 ns | ------ CRTP: 22951820 ns
clang版本:18.1.3
疑问与求助
我猜测缓存可能是影响测试结果的主要因素,但不确定。如果真是这样,那CRTP对性能的影响似乎可以忽略不计。想请教大家:CRTP的最佳实践是什么?
内容的提问来源于stack exchange,提问作者Yuwei Zhao
相关产品推荐
相关产品推荐

