关于Google Benchmark测试结果与CPU时钟周期不符的疑问
一、为什么测试时间看起来短于单个时钟周期?
你的核心困惑其实来自对CPU执行模型和基准测试统计方式的误解,主要有这几个关键点:
CPU超标量执行能力
2GHz主频意味着单个时钟周期是0.5ns,但现代CPU是超标量处理器——每个时钟周期可以同时执行多条独立指令。比如你的加法操作和DoNotOptimize相关的内存屏障指令,能被CPU并行处理,所以实际单迭代的有效耗时会小于单个时钟周期的理论值。编译器循环优化
Google Benchmark的循环会被编译器做循环展开优化,比如把多次迭代的代码合并成一次执行,减少循环本身的开销(比如循环计数器增减、条件判断)。这样平均到每次函数调用的时间就被摊薄了,看起来比单周期还短。指令流水线的重叠执行
CPU的流水线可以让指令的取指、译码、执行等阶段重叠进行,进一步隐藏指令延迟。你的测试函数逻辑非常简单(仅成员变量加输入值),这些指令几乎没有数据依赖,流水线能高效连续执行。
对应你的测试结果来看:
BM_nonVirtualFunc的0.490ns几乎刚好是单个周期,因为非虚函数被编译器完全内联,加法指令在超标量执行下刚好占一个周期。BM_virtualFunc的0.858ns约等于1.7个周期,这是因为虚函数需要额外的虚表查找指令(读取虚表指针、跳转),这些指令的延迟被流水线部分隐藏,但还是比非虚函数多了一点开销。
如果你想验证这一点,可以查看编译后的汇编代码:非虚函数的调用会被完全内联到循环中,而虚函数会保留通过虚表指针的间接调用指令。
二、如何设计更合理的虚函数/非虚函数对比测试?
你之前用std::cout的方式虽然能体现差异,但会引入IO开销,而且输出杂乱。这里有几个更靠谱的测试方案,能避免编译器过度优化,同时清晰体现虚函数的开销:
方案1:增加函数计算复杂度,阻止虚函数去虚拟化
设计一个有实际计算量的函数,同时确保虚函数无法被编译器优化成直接调用:
#include <stdlib.h> #include <memory> #include <benchmark/benchmark.h> class Base { public: virtual int compute(int x) { // 增加计算量,避免被优化成常量 int res = x; for (int i = 0; i < 10; ++i) { res = res * 2 + i; } return res; } virtual ~Base() = default; }; class Derived : public Base { public: int compute(int x) override { int res = x; for (int i = 0; i < 10; ++i) { res = res * 3 - i; } return res; } }; // 非虚函数测试:静态类型明确,编译器可以内联 static void BM_NonVirtual(benchmark::State& state) { Derived d; volatile int x = rand(); for (auto _ : state) { auto res = d.compute(x); benchmark::DoNotOptimize(res); } } BENCHMARK(BM_NonVirtual); // 虚函数测试:使用基类指针,确保走虚表 static void BM_Virtual(benchmark::State& state) { std::unique_ptr<Base> ptr = std::make_unique<Derived>(); volatile int x = rand(); for (auto _ : state) { auto res = ptr->compute(x); benchmark::DoNotOptimize(res); } } BENCHMARK(BM_Virtual);
方案2:模拟真实多态场景(随机对象调用)
如果想模拟程序中随机使用不同派生类对象的场景,可以让编译器无法预测调用的是哪个类的函数:
static void BM_Virtual_Polymorphic(benchmark::State& state) { std::vector<std::unique_ptr<Base>> objs; objs.emplace_back(std::make_unique<Base>()); objs.emplace_back(std::make_unique<Derived>()); volatile int x = rand(); for (auto _ : state) { // 随机选择对象,阻止编译器去虚拟化 int idx = rand() % objs.size(); auto res = objs[idx]->compute(x); benchmark::DoNotOptimize(res); benchmark::ClobberMemory(); // 确保内存操作不被重排 } } BENCHMARK(BM_Virtual_Polymorphic);
方案3:直接对比虚函数与函数指针开销
你也可以对比虚函数调用和普通函数指针调用的开销,更直观地看到虚表的额外开销:
static void BM_FunctionPointer(benchmark::State& state) { Derived d; int (*fp)(Derived*, int) = &Derived::compute; volatile int x = rand(); for (auto _ : state) { auto res = fp(&d, x); benchmark::DoNotOptimize(res); } } BENCHMARK(BM_FunctionPointer);
这些方案的核心思路是:
- 给函数增加足够的计算量,让函数本身的执行时间远大于调用开销,这样虚函数和非虚函数的差异会更明显。
- 确保虚函数的调用无法被编译器优化成直接调用(去虚拟化),比如使用基类指针、随机选择对象等方式。
内容的提问来源于stack exchange,提问作者talekeDskobeDa

