You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

CPU微架构相关异常:函数指针调用开销测试疑问

函数指针调用开销测试中的异常现象求助

我正在开展一项函数指针调用开销的测试工作,但测试中发现了若干异常现象,特来求助。

测试环境与配置

  • 编译环境:VS2017 Release模式,默认配置
  • 测试设备(共4台Win10设备):
    • M1: CPU i7-7700,微架构Kaby Lake
    • M2: CPU i7-7700,微架构Kaby Lake
    • M3: CPU i7-4790,微架构Haswell
    • M4: CPU E5-2698 v3,微架构Haswell

测试结果图例说明

图例格式为machine parameter_order alias,各字段含义:

  • machine:对应上述设备标识
  • parameter_order:单次运行时传入程序的LOOP参数顺序
  • alias:计时模块说明:
    • no-exec:无函数调用的基准测试部分(对应代码第98-108行)
    • exec:包含函数调用的测试部分(对应代码第115-125行)
    • per-exec:单次函数调用的开销,该指标对应图表左Y轴;其余指标对应右Y轴,所有时间单位均为毫秒。

对比图1至图4的测试结果可以看出,结果与CPU微架构强相关——M1、M2的结果高度相似,M3、M4的结果也呈现一致的趋势。

技术疑问

我目前有以下几个核心疑问需要解答:

  1. 为何所有设备的测试结果都呈现出明显的两个阶段(LOOP < 25和LOOP > 100)?
  2. 为何所有设备的no-exec计时在32 <= LOOP <= 41区间都会出现异常峰值?
  3. 为何Kaby Lake架构的设备(M1、M2)中,no-exec和exec计时在72 <= LOOP <= 94区间会出现不连续的波动?
  4. 为何服务器处理器M4的测试结果方差比同架构的桌面处理器M3更大?

测试结果图表

图1
图2
图3
图4

测试代码

#include <cstdio> 
#include <cstdlib> 
#include <ctime> 
#include <cassert> 
#include <algorithm> 
#include <windows.h> 
using namespace std; 

const int PMAX = 11000000, ITER = 60000, RULE = 10000; 
//const int LOOP = 10; 

int func1(int a, int b, int c, int d, int e) { return 0; } 
int func2(int a, int b, int c, int d, int e) { return 0; } 
int func3(int a, int b, int c, int d, int e) { return 0; } 
int func4(int a, int b, int c, int d, int e) { return 0; } 
int func5(int a, int b, int c, int d, int e) { return 0; } 
int func6(int a, int b, int c, int d, int e) { return 0; } 

int (*init[6])(int, int, int, int, int) = { func1, func2, func3, func4, func5, func6 }; 
int (*pool[PMAX])(int, int, int, int, int); 

LARGE_INTEGER freq; 

void getTime(LARGE_INTEGER *res) { 
    QueryPerformanceCounter(res); 
} 

double delta(LARGE_INTEGER begin_time, LARGE_INTEGER end_time) { 
    return (end_time.QuadPart - begin_time.QuadPart) * 1000.0 / freq.QuadPart; 
} 

int main() { 
    char path[100], tmp[100]; 
    FILE *fin, *fout; 
    int cnt = 0; 
    int i, j, t, r; 
    int ans; 
    int LOOP; 
    LARGE_INTEGER begin_time, end_time; 
    double d1, d2, res; 

    for(i = 0;i < PMAX;i += 1) 
        pool[i] = init[i % 6]; 

    QueryPerformanceFrequency(&freq); 

    printf("file path:"); 
    scanf("%s", path); 
    fin = fopen(path, "r"); 

start: 
    if (fscanf(fin, "%d", &LOOP) == EOF) 
        goto end; 

    ans = 0; 
    getTime(&begin_time); 
    for(r = 0;r < RULE;r += 1) { 
        for(t = 0;t < ITER;t += 1) { 
            //ans ^= (pool[t])(0, 0, 0, 0, 0); 
            ans ^= pool[0](0, 0, 0, 0, 0); 
            ans = 0; 
            for(j = 0;j < LOOP;j += 1) 
                ans ^= j; 
        } 
    } 
    getTime(&end_time); 
    printf("%.10f\n", d1 = delta(begin_time, end_time)); 
    printf("ans:%d\n", ans); 

    ans = 0; 
    getTime(&begin_time); 
    for(r = 0;r < RULE;r += 1) { 
        for(t = 0;t < ITER;t += 1) { 
            ans ^= (pool[t])(0, 0, 0, 0, 0); 
            ans ^= pool[0](0, 0, 0, 0, 0); 
            ans = 0; 
            for(j = 0;j < LOOP;j += 1) 
                ans ^= j; 
        } 
    } 
    getTime(&end_time); 
    printf("%.10f\n", d2 = delta(begin_time, end_time)); 
    printf("ans:%d\n", ans); 

    printf("%.10f\n", res = (d2 - d1) / (1.0 * RULE * ITER)); 

    sprintf(tmp, "%d.txt", cnt++); 
    fout = fopen(tmp, "w"); 
    fprintf(fout, "%d,%.10f,%.10f,%.10f\n", LOOP, d1, d2, res); 
    fclose(fout); 

    goto start; 
end: 
    fclose(fin); 
    system("pause"); 
    exit(0); 
}

内容的提问来源于stack exchange,提问作者Leo Jacob

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 09:56:33