虚函数继承与std::function无继承的性能对比及基准测试合理性验证
摘要
我编写的基准测试是否能公平比较基于继承和基于std::function的多态实现?
完整问题
当需要多个以不同方式实现同一接口的对象,且需将它们放入容器并在运行时替换时,最常用的方案是使用继承:
struct Base { virtual void f() = 0; virtual ~Base() = default; }; struct Derived1 : Base { virtual void f(); }; struct Derived2 : Base { virtual void f(); };
另一种方案是使用单个类,用std::function替代虚函数:
struct Foo { std::function<void()> f{}; }; auto foo1 = Foo{[]{ return /* impl like Derived1 */; }}; auto foo2 = Foo{[]{ return /* impl like Derived2 */; }};
不过,暂不考虑两种方案的其他优缺点,我好奇如何通过基准测试衡量它们的性能差异。
我明白性能显然会受std::function的实现方式、编译器及编译选项、操作系统等因素影响。但在固定所有这些因素的前提下,我认为可以衡量是否存在性能差异。
我的目的是亲自验证:除非在特殊场景,否则两种方案的性能差异可忽略不计——这是我从相关资料中得到的结论;或是证明我的理解错误,两者确实存在显著性能差异。
我编写的基准测试如下,各部分说明:
- 所有
f函数以不同方式修改全局unsigned int变量:
我会在unsigned int RETURN{};main中返回该变量,确保函数体不会被编译器优化掉; - 修改
Derived1::f/Derived2::f和foo1/foo2的lambda函数体,使其修改上述全局变量:struct Base { virtual void f() = 0; virtual ~Base() = default; }; struct Derived1 : Base { virtual void f() { RETURN += 1; } }; struct Derived2 : Base { virtual void f() { RETURN += 2; } }; struct Foo { std::function<void()> f{}; }; auto const foo1 = Foo{[]{ RETURN += 1; }}; auto const foo2 = Foo{[]{ RETURN += 2; }}; - 在测量代码前,生成随机
bool值,用于随机选择Derived1/foo1或Derived2/foo2:std::random_device rd; std::mt19937 gen{rd()}; std::bernoulli_distribution randBool{0.5}; constexpr int N = 1000000; std::array<bool, N> bools; for (bool& b : bools) { b = randBool(gen); } - 使用Boost.Hana遍历包含两个编译期
true/false的元组,参数化两种测试场景;使用Range-v3累加每次调用虚函数/std::function的时间测量结果:using Time = duration<double, std::milli>; std::array<Time, 2> times; // 0: std::function-based, 1: inheritance-based hana::for_each(hana::make_basic_tuple(hana::false_c, hana::true_c), [&](auto hb) { constexpr bool B = hb; auto const elapsed = ranges::accumulate(bools, Time{}, [](auto acc, auto b){ /* time measurement */; }); times[!B] = elapsed; }); - 基于运行期
bool值b选择对象的函数,以编译期bool值B为模板参数区分两种场景:template<bool B> constexpr auto bool2Obj = []{ if constexpr (B) { return [](bool b){ return b ? foo1 : foo2; }; } else { using BasePtr = std::unique_ptr<Base>; return [](bool b){ return b ? BasePtr{std::make_unique<Derived1>()} : BasePtr{std::make_unique<Derived2>()}; }; } }(); - 选中对象后调用方法的函数,同样以
bool B为模板参数区分场景:template<bool B> constexpr auto call = []{ if constexpr (B) { return [](Foo const& p){ p.f(); }; } else { return [](std::unique_ptr<Base> const& p){ p->f(); }; } }(); /* time measurement */的具体代码:
其中对象的随机选择过程未纳入测量,仅测量auto obj = bool2Obj<B>(b); auto const start = high_resolution_clock::now(); call<B>(obj); auto const end = high_resolution_clock::now() - start; return acc + Time{end};call调用的耗时。
测试结果:
- 在CompilerExplorer上,因进程超时限制只能运行少量重复测试,结果显示两种方案性能大致相当——我输出的百分比
(i - f) / i(i为继承方案耗时,f为std::function方案耗时)符号频繁变化;Clang和GCC均是如此; - 在本地机器上,GCC结果与上述一致,而Clang(18.1.8)始终返回正值,表明
std::function方案更快:0.0648057 0.0716398 0.0636759 0.0649676 0.0673908 0.0756509 0.0780861 0.0890416 0.090532 0.094767 - 此外,QuickBench(需移除Boost和Range-v3)的测试结果始终支持
std::function方案更快,涉及GCC、Clang + LLVM、Clang + GNU等编译环境。
内容的提问来源于stack exchange,提问作者Enlico
相关产品推荐
相关产品推荐

