You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

虚函数继承与std::function无继承的性能对比及基准测试合理性验证

摘要

我编写的基准测试是否能公平比较基于继承和基于std::function的多态实现?

完整问题

当需要多个以不同方式实现同一接口的对象,且需将它们放入容器并在运行时替换时,最常用的方案是使用继承:

struct Base {
    virtual void f() = 0;
    virtual ~Base() = default;
};
struct Derived1 : Base {
    virtual void f();
};
struct Derived2 : Base {
    virtual void f();
};

另一种方案是使用单个类,用std::function替代虚函数:

struct Foo {
    std::function<void()> f{};
};
auto foo1 = Foo{[]{ return /* impl like Derived1 */; }};
auto foo2 = Foo{[]{ return /* impl like Derived2 */; }};

不过,暂不考虑两种方案的其他优缺点,我好奇如何通过基准测试衡量它们的性能差异。

我明白性能显然会受std::function的实现方式、编译器及编译选项、操作系统等因素影响。但在固定所有这些因素的前提下,我认为可以衡量是否存在性能差异。

我的目的是亲自验证:除非在特殊场景,否则两种方案的性能差异可忽略不计——这是我从相关资料中得到的结论;或是证明我的理解错误,两者确实存在显著性能差异。


我编写的基准测试如下,各部分说明:

  • 所有f函数以不同方式修改全局unsigned int变量:
    unsigned int RETURN{};
    
    我会在main中返回该变量,确保函数体不会被编译器优化掉;
  • 修改Derived1::f/Derived2::f和foo1/foo2的lambda函数体,使其修改上述全局变量:
    struct Base {
        virtual void f() = 0;
        virtual ~Base() = default;
    };
    struct Derived1 : Base {
        virtual void f() { RETURN += 1; }
    };
    struct Derived2 : Base {
        virtual void f() { RETURN += 2; }
    };
    
    struct Foo {
        std::function<void()> f{};
    };
    auto const foo1 = Foo{[]{ RETURN += 1; }};
    auto const foo2 = Foo{[]{ RETURN += 2; }};
    
  • 在测量代码前,生成随机bool值,用于随机选择Derived1/foo1或Derived2/foo2:
    std::random_device rd;
    std::mt19937 gen{rd()};
    std::bernoulli_distribution randBool{0.5};
    constexpr int N = 1000000;
    
    std::array<bool, N> bools;
    for (bool& b : bools) {
        b = randBool(gen);
    }
    
  • 使用Boost.Hana遍历包含两个编译期true/false的元组,参数化两种测试场景;使用Range-v3累加每次调用虚函数/std::function的时间测量结果:
    using Time = duration<double, std::milli>;
    
    std::array<Time, 2> times; // 0: std::function-based, 1: inheritance-based
    
    hana::for_each(hana::make_basic_tuple(hana::false_c, hana::true_c), [&](auto hb) {
        constexpr bool B = hb;
        auto const elapsed = ranges::accumulate(bools, Time{}, [](auto acc, auto b){
            /* time measurement */;
        });
        times[!B] = elapsed;
    });
    
  • 基于运行期bool值b选择对象的函数,以编译期bool值B为模板参数区分两种场景:
    template<bool B>
    constexpr auto bool2Obj = []{
        if constexpr (B) {
            return [](bool b){
                return b
                    ? foo1
                    : foo2;
            };
        } else {
            using BasePtr = std::unique_ptr<Base>;
            return [](bool b){
                return b
                    ? BasePtr{std::make_unique<Derived1>()}
                    : BasePtr{std::make_unique<Derived2>()};
            };
        }
    }();
    
  • 选中对象后调用方法的函数,同样以bool B为模板参数区分场景:
    template<bool B>
    constexpr auto call = []{
        if constexpr (B) {
            return [](Foo const& p){ p.f(); };
        } else {
            return [](std::unique_ptr<Base> const& p){ p->f(); };
        }
    }();
    
  • /* time measurement */的具体代码:
    auto obj = bool2Obj<B>(b);
    auto const start = high_resolution_clock::now();
    call<B>(obj);
    auto const end = high_resolution_clock::now() - start;
    return acc + Time{end};
    
    其中对象的随机选择过程未纳入测量,仅测量call调用的耗时。

测试结果:

  • 在CompilerExplorer上,因进程超时限制只能运行少量重复测试,结果显示两种方案性能大致相当——我输出的百分比(i - f) / i(i为继承方案耗时,f为std::function方案耗时)符号频繁变化;Clang和GCC均是如此;
  • 在本地机器上,GCC结果与上述一致,而Clang(18.1.8)始终返回正值,表明std::function方案更快:
    0.0648057
    0.0716398
    0.0636759
    0.0649676
    0.0673908
    0.0756509
    0.0780861
    0.0890416
    0.090532
    0.094767
    
  • 此外,QuickBench(需移除Boost和Range-v3)的测试结果始终支持std::function方案更快,涉及GCC、Clang + LLVM、Clang + GNU等编译环境。

内容的提问来源于stack exchange,提问作者Enlico

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 04:42:33