You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何CRTP未带来显著性能提升?求原因与最佳实践

CRTP与运行时多态的性能对比实验及疑问

实验背景与目的

我做了一个小型实验,想直观感受CRTP(编译时多态)相比运行时多态的性能提升——按理论来说,编译时多态的调用开销应该更低。

实验代码

#include <iostream>
#include <chrono>

namespace NonCRTP {

class Base {
public:
  virtual void name() = 0;
};

class D1 : public Base {
public:
  void name() override {
    volatile int _ = 1;
  }
};

} // namespace NonCRTP

namespace CRTP {

template <class Derived>
class Base {
public:
  void name() {
    static_cast<Derived*>(this)->name_impl();
  }

protected:
  Base() = default;
};

class D1 : public Base<D1> {
public:
  void name_impl() {
    volatile int _ = 1;
  }
};

} // namespace CRTP

void perf_test(int invocation_counts) {
  std::chrono::nanoseconds ns_duration_noncrtp{};
  std::chrono::nanoseconds ns_duration_crtp{};

  {
    NonCRTP::D1 obj_noncrtp;
    auto time_s{ std::chrono::high_resolution_clock::now() };
    for (int i = 0; i < invocation_counts; ++i) {
      obj_noncrtp.name();
    }

    auto time_e{ std::chrono::high_resolution_clock::now() };
    ns_duration_noncrtp = time_e - time_s;
  }

  {
    CRTP::D1 obj_crtp;
    auto time_s{ std::chrono::high_resolution_clock::now() };
    for (int i = 0; i < invocation_counts; ++i) {
      obj_crtp.name();
    }

    auto time_e{ std::chrono::high_resolution_clock::now() };
    ns_duration_crtp = time_e - time_s;
  }

  std::printf("Perf test: %d times of invocation:\n", invocation_counts);
  std::printf("  | -- Non-CRTP: %lu ns\n", static_cast<uint64_t>(ns_duration_noncrtp.count()));
  std::printf("  | ------ CRTP: %lu ns\n", static_cast<uint64_t>(ns_duration_crtp.count()));
  std::printf("\n");
}

int main() {
  perf_test(1);
  perf_test(10);
  perf_test(100);
  perf_test(1000);
  perf_test(10000);
  perf_test(100000);
  perf_test(1000000);
  perf_test(10000000);
  perf_test(100000000);
  return 0;
}

实验结果

未开启-O2优化的测试结果

CRTP实现并未始终优于非CRTP版本,甚至在多数场景下更慢:

-> % clang++ crtp.cc -std=c++20 -o crtp; ./crtp 
Perf test: 1 times of invocation:
  | -- Non-CRTP: 100 ns
  | ------ CRTP: 31 ns

Perf test: 10 times of invocation:
  | -- Non-CRTP: 49 ns
  | ------ CRTP: 59 ns

Perf test: 100 times of invocation:
  | -- Non-CRTP: 171 ns
  | ------ CRTP: 239 ns

Perf test: 1000 times of invocation:
  | -- Non-CRTP: 1360 ns
  | ------ CRTP: 2142 ns

Perf test: 10000 times of invocation:
  | -- Non-CRTP: 16609 ns
  | ------ CRTP: 60303 ns

Perf test: 100000 times of invocation:
  | -- Non-CRTP: 194464 ns
  | ------ CRTP: 234327 ns

Perf test: 1000000 times of invocation:
  | -- Non-CRTP: 1489131 ns
  | ------ CRTP: 2251113 ns

Perf test: 10000000 times of invocation:
  | -- Non-CRTP: 14182943 ns
  | ------ CRTP: 23018457 ns

Perf test: 100000000 times of invocation:
  | -- Non-CRTP: 139714335 ns
  | ------ CRTP: 228018231 ns

开启-O2优化的测试结果

情况有所改善,多数场景下CRTP表现更优,但仍存在CRTP更慢的情况:

-> % clang++ crtp.cc -O2 -std=c++20 -o crtp; ./crtp
Perf test: 1 times of invocation:
  | -- Non-CRTP: 100 ns
  | ------ CRTP: 27 ns

Perf test: 10 times of invocation:
  | -- Non-CRTP: 30 ns
  | ------ CRTP: 34 ns

Perf test: 100 times of invocation:
  | -- Non-CRTP: 67 ns
  | ------ CRTP: 78 ns

Perf test: 1000 times of invocation:
  | -- Non-CRTP: 288 ns
  | ------ CRTP: 289 ns

Perf test: 10000 times of invocation:
  | -- Non-CRTP: 2856 ns
  | ------ CRTP: 2822 ns

Perf test: 100000 times of invocation:
  | -- Non-CRTP: 27937 ns
  | ------ CRTP: 27965 ns

Perf test: 1000000 times of invocation:
  | -- Non-CRTP: 315920 ns
  | ------ CRTP: 270106 ns

Perf test: 10000000 times of invocation:
  | -- Non-CRTP: 2632259 ns
  | ------ CRTP: 2705106 ns

Perf test: 100000000 times of invocation:
  | -- Non-CRTP: 24681903 ns
  | ------ CRTP: 22951820 ns

clang版本:18.1.3

疑问与求助

我猜测缓存可能是影响测试结果的主要因素,但不确定。如果真是这样,那CRTP对性能的影响似乎可以忽略不计。想请教大家:CRTP的最佳实践是什么?


内容的提问来源于stack exchange,提问作者Yuwei Zhao

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 15:44:51