You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

gcc -fno-common选项致性能下降相关技术问询

问题背景
  • 使用UnixBench对gcc-8.5.0和gcc-10.3.0进行性能测试时,发现性能下降约10%,最终定位原因是gcc-10.3.0默认启用-fno-common选项(gcc-8.5.0默认启用-fcommon选项)。
  • 查阅gcc手册后,关于-fcommon的描述如下:

    In C code, this option controls the placement of global variables defined without an initializer, known as tentative definitions in the C standard. Tentative definitions are distinct from declarations of a variable with the extern keyword, which do not allocate storage.

    The default is -fno-common, which specifies that the compiler places uninitialized global variables in the BSS section of the object file. This inhibits the merging of tentative definitions by the linker so you get a multiple-definition error if the same variable is accidentally defined in more than one compilation unit.

    The -fcommon places uninitialized global variables in a common block. This allows the linker to resolve all tentative definitions of the same variable in different compilation units to the same object, or to a non-tentative definition. This behavior is inconsistent with C++, and on many targets implies a speed and code size penalty on global variable references. It is mainly useful to enable legacy code to link without errors.

技术问询解答

1. 同一环境下,-fcommon与-fno-common为何存在如此显著的性能差异?

两者的性能差异核心在于全局变量的存储布局与访问机制:

  • -fcommon模式下,未初始化全局变量被放入公共块,链接阶段才会统一分配地址。编译器无法提前确定变量最终地址,只能生成间接访问代码(如通过偏移量或指针),增加指令执行周期;但如果测试代码存在大量跨编译单元的同名 tentative 定义,链接器会将它们合并为单一实例,减少内存占用与缓存竞争,反而提升性能。
  • -fno-common模式下,未初始化全局变量直接放入BSS段,编译阶段即可确定相对地址,编译器能生成直接访问指令,理论上更高效。但你的场景中出现性能下降,大概率是因为该模式下每个编译单元的同名变量都是独立实例,导致内存分散、缓存命中率降低,恰好与Sapphire Rapids CPU的缓存特性不匹配。

2. 如何定位并解决Sapphire Rapids CPU与gcc-10.3.0的适配问题?

可按以下步骤排查:

  • 定位性能瓶颈:用perf record+perf report分析测试程序热点,确认是内存访问延迟、缓存命中率低还是指令执行效率问题。比如L3缓存命中率过低,说明变量布局导致缓存冲突。
  • 升级gcc版本:尝试gcc-11及以上版本,新版本通常会针对Sapphire Rapids新增架构优化,修复旧版本的适配缺陷。
  • 调整编译选项:
    • 添加-march=sapphirerapids或-m=native,强制编译器生成适配当前CPU的指令集与优化代码。
    • 搭配-fdata-sections+-Wl,--gc-sections清理冗余内存段,优化变量布局;或尝试-fno-zero-initialized-in-bss调整BSS段变量处理逻辑。
  • 对比变量布局:用objdump或readelf查看两种编译模式下目标文件的全局变量地址分布,分析-fno-common是否导致变量过度分散,增加内存访问延迟。

3. 将默认选项从-fno-common改回-fcommon存在哪些具体风险?

改回-fcommon的风险主要包括:

  • 标准兼容性问题:-fcommon是历史遗留特性,与C11及以后的标准定义冲突,且和C的严格变量定义规则不兼容,混合C/C开发时易出现链接错误或未定义行为。
  • 隐藏代码bug:-fcommon会自动合并多编译单元的同名未初始化全局变量,掩盖代码中重复定义变量的错误。比如两个源文件都定义int global_var;,-fno-common会直接抛出多重定义错误,而-fcommon会静默合并,导致逻辑隐患。
  • 潜在性能退化:gcc手册明确提到-fcommon在多数平台会导致全局变量访问的速度与代码尺寸损失,当前测试场景的性能优势可能是特殊案例,后续代码调整后劣势可能显现。
  • 维护成本上升:新版gcc默认均为-fno-common,改回该选项需要在所有编译脚本中显式配置,增加团队后续版本升级与配置维护的复杂度。

内容的提问来源于stack exchange,提问作者huyubiao

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 01:14:59