为何gcc 11.4.0中std::map<std::string,int>比std::map<std::string_view,int>更快?
疑问:为何
std::map<std::string, int>比std::map<std::string_view, int>插入更快? 测试背景
我需要对生命周期贯穿整个程序的std::string子串建立索引,通过std::map将每个子串与int关联,子串以空格' '分隔。
测试字符串生成
用以下Python代码生成5000万字符的随机字符串并保存到文件:
import random all_chars = "abcdefghijklmnopqrstuvwxyzABCDEFGHIJKLMNOPQRSTUVWXYZ 123456789!\"·$%&/()=?¿¡'|@#¬" for i in range(0, 50000000): r = int(random.random() * len(all_chars)) print(all_chars[r], end = '')
分割函数实现
我实现了两个分割函数,对比std::string和std::string_view的分割效率:
[[nodiscard]] std::vector<std::string_view> split_sv(const std::string& s, char c) noexcept { std::vector<std::string_view> res; std::size_t b = 0; std::size_t p = s.find(c); while (p != std::string_view::npos) { res.push_back( std::string_view{&s[b], p - b} ); b = p + 1; p = s.find(c, b); } return res; } [[nodiscard]] std::vector<std::string> split_s(const std::string& s, char c) noexcept { std::vector<std::string> res; std::size_t b = 0; std::size_t p = s.find(c); while (p != std::string::npos) { res.push_back( s.substr(b, p - b) ); // 修正原代码笔误:将未定义的i改为p b = p + 1; p = s.find(c, b); } return res; }
分割性能测试
编译选项为-std=c++20 -O3(gcc 11.4.0),分割测试结果显示std::string_view明显更快:
std::string_view 28.6 ms std::string 107.6 ms
存入std::map的性能对比
但将子串存入std::map时,std::string_view版本反而更慢:
测试代码片段:
{ const std::vector<std::string_view> tokens = split_sv(s, ' '); std::map<std::string_view, int> m; for (const std::string_view& t : tokens) { m.insert( {t, 0} ); } } { const std::vector<std::string> tokens = split_s(s, ' '); std::map<std::string, int> m; for (const std::string& t : tokens) { m.insert( {t, 0} ); } }
仅for循环的执行时间:
std::string_view 445.9 ms std::string 411.7 ms
差异虽不大,但我原本预期std::string_view在两部分都更快。
扩展测试结果
我还测试了3种不同空格占比(NBS)的字符串种子,各场景执行时间(秒)如下:
| NBS | 令牌数量/平均子串长度(字符) | 操作 | map<string_view> | map<string> | unordered_map<string_view> | unordered_map<string> |
|---|---|---|---|---|---|---|
| 1 | 617864 / 83.97 | 分割 | 0.010 | 0.153 | 0.023 | 0.025 |
| 存储 | 0.441 | 0.408 | 0.106 | 0.167 | ||
| 10 | 4988259 / 9.35 | 分割 | 0.594 | 0.883 | 0.469 | 0.206 |
| 存储 | 4.430 | 3.806 | 1.361 | 1.470 | ||
| 20 | 8062256 / 5.20 | 分割 | 0.696 | 0.981 | 0.559 | 0.298 |
| 存储 | 6.568 | 5.590 | 1.820 | 1.892 |
核心疑问
- 为何gcc 11.4.0中
std::map<std::string,int>的插入速度比std::map<std::string_view,int>更快? std::map<std::string,int>是否普遍比std::map<std::string_view,int>更快?
解答
1. gcc 11.4.0下std::map<string>更快的原因
std::map基于红黑树实现,插入时需要频繁比较键的大小,这是性能差异的核心来源:
- 比较逻辑的缓存友好性:
std::string的内部存储是连续内存块,gcc对其比较操作做了深度优化——先对比长度,长度不同直接返回结果;长度相同时调用memcmp或SIMD指令批量比较,缓存命中率极高。而std::string_view指向原字符串的分散子串,不同子串在内存中不连续,逐字符比较时缓存miss概率高,内存访问延迟大。 - 红黑树节点的存储特性:
std::map的节点会存储键的副本,std::string的副本是连续内存块,能更好地利用缓存;std::string_view的副本仅占两个指针/整数的空间,但比较时需要频繁跳转到原字符串的不同位置,反而放大了缓存劣势。
2. 是否普遍更快?
不是,性能差异取决于编译器版本、子串内存分布、子串长度等因素:
- 编译器版本:较新的gcc(12+)或clang对
std::string_view的比较逻辑做了更多优化,比如短字符串快速匹配、缓存预取策略,此时string_view版本性能可能反超string。 - 子串内存分布:如果子串在原字符串中是连续块(比如无分隔符场景),
string_view的内存局部性提升,缓存命中率会接近甚至超过string。 - 子串长度:极短子串的比较性能差异可忽略;超长子串场景下,
string的连续存储优势更明显,但string_view避免了内存拷贝开销,整体性能可能更优(你的测试中长串场景string更快,仅因为gcc 11的string_view优化不足)。
另外从扩展测试的unordered_map结果能看到,string_view版本存储性能明显优于string——哈希计算对内存局部性敏感度低,且string_view避免了拷贝开销,这才是它的典型优势场景。
内容的提问来源于stack exchange,提问作者llualpu
相关产品推荐
相关产品推荐

