You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

C++11中返回std::map/std::vector值为何比传引用参数耗时更长?

问题原因分析
  • 测试逻辑完全不等价,这是耗时差异的核心原因
    针对vector的两组测试:

    funcVec1的测试逻辑为每次循环都创建全新的vector实例,插入10个元素后销毁,10万次循环需要执行10万次vector构造、内存分配、元素插入、析构的完整流程。
    funcVec2的测试逻辑为仅在循环外创建1次vector,后续每次循环都往同一个vector追加元素,随着vector容量逐步扩容到足够大,后续push_back操作几乎不会触发内存分配,完全没有vector构造、析构的开销,两组测试没有可比性。
    针对map的两组测试:
    funcMap1的测试逻辑为每次循环创建全新的map实例,插入10个键值对后销毁,需要执行10万次红黑树节点分配、插入、销毁操作。
    funcMap2的测试逻辑为仅在循环外创建1次map,第一次循环已经把0-9的键全部插入完成,后续99999次循环执行tmpMap[i] = i本质是对已有键的值覆盖,完全不需要新分配红黑树节点,耗时自然极低。

  • 编译器优化未开启会放大差异
    C++11标准下返回局部非拷贝可移动对象时,默认会调用移动构造函数,大部分编译器还会触发NRVO(命名返回值优化)直接消除移动开销,理论上和引用传参性能一致。但如果编译时未开启优化(如GCC默认O0级别),不会触发NRVO优化,会多一次移动构造的开销,不过移动vector、map都是O(1)操作,差异不会像你测试结果那样夸张。
修正后的对等测试代码
#include <iostream>
#include <vector>
#include <map>
#include <chrono>
using namespace std;
vector<int> funcVec1() {
    vector<int> vec;
    for (int i = 0; i < 10; ++i) {
        vec.push_back(i);
    }
    return vec;
}

void funcVec2(vector<int>& vec) {
    vec.clear(); // 每次调用前清空,保证逻辑和funcVec1一致
    for (int i = 0; i < 10; ++i) {
        vec.push_back(i);
    }
    return;
}

map<int, int> funcMap1() {
    map<int, int> tmpMap;
    for (int i = 0; i < 10; ++i) {
        tmpMap[i] = i;
    }
    return tmpMap;
}

void funcMap2(map<int, int>& tmpMap) {
    tmpMap.clear(); // 每次调用前清空,保证逻辑和funcMap1一致
    for (int i = 0; i < 10; ++i) {
        tmpMap[i] = i;
    }
}

int main()
{
    using namespace std::chrono;
    system_clock::time_point t1 = system_clock::now();
    for (int i = 0; i < 100000; ++i) {
        vector<int> vec1 = funcVec1();
    }
    auto t2 = std::chrono::system_clock::now();
    cout << "return vec takes " << (t2 - t1).count() << " tick count" << endl;
    cout << duration_cast<milliseconds>(t2 - t1).count() << " milliseconds" << endl;
    cout << " --------------------------------" << endl;
    vector<int> vec2;
    for (int i = 0; i < 100000; ++i) {
        funcVec2(vec2);
    }
    auto t3 = system_clock::now();
    cout << "reference vec takes " << (t3 - t2).count() << " tick count" << endl;
    cout << duration_cast<milliseconds>(t3 - t2).count() << " milliseconds" << endl;
    cout << " --------------------------------" << endl;
    for (int i = 0; i < 100000; ++i) {
        map<int, int> tmpMap1 = funcMap1();
    }
    auto t4 = system_clock::now();
    cout << "return map takes " << (t4 - t3).count() << " tick count" << endl;
    cout << duration_cast<milliseconds>(t4 - t3).count() << " milliseconds" << endl;
    cout << " --------------------------------" << endl;
    map<int, int> tmpMap2;
    for (int i = 0; i < 100000; ++i) {
        funcMap2(tmpMap2);
    }
    auto t5 = system_clock::now();
    cout << "reference map takes " << (t5 - t4).count() << " tick count" << endl;
    cout << duration_cast<milliseconds>(t5 - t4).count() << " milliseconds" << endl;
    cout << " --------------------------------" << endl;
    return 0;
}
测试结论

使用修正后的代码开启O2优化编译运行,两种写法的耗时差距会非常小,完全可以忽略。返回值写法代码可读性更高,不需要外部提前定义变量,日常开发更推荐使用。

内容的提问来源于stack exchange,提问作者f1msch

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.03 13:09:00