MSVC STL中std::any的缓存行共享设计疑问
class any { // ... struct _Small_storage_t { unsigned char _Data[_Any_small_space_size]; const _Any_small_RTTI* _RTTI; }; static_assert(sizeof(_Small_storage_t) == _Any_trivial_space_size); struct _Big_storage_t { // Pad so that _Ptr and _RTTI might share _TypeData's cache line unsigned char _Padding[_Any_small_space_size - sizeof(void*)]; void* _Ptr; const _Any_big_RTTI* _RTTI; }; static_assert(sizeof(_Big_storage_t) == _Any_trivial_space_size); struct _Storage_t { union { unsigned char _TrivialData[_Any_trivial_space_size]; _Small_storage_t _SmallStorage; _Big_storage_t _BigStorage; }; uintptr_t _TypeData; }; static_assert(sizeof(_Storage_t) == _Any_trivial_space_size + sizeof(void*)); static_assert(is_standard_layout_v<_Storage_t>); union { _Storage_t _Storage{}; max_align_t _Dummy; }; };
主存储部分_Storage_t会根据std::any中存储对象的类型,对应为_Small_storage_t、_Big_storage_t或unsigned char数组。
用户平台上的参数值:
sizeof(void*); // 8字节 constexpr int _Small_object_num_ptrs = 6 + 16 / sizeof(void*); // 8 constexpr size_t _Any_small_space_size = (_Small_object_num_ptrs - 2) * sizeof(void*); // 48 constexpr size_t _Any_trivial_space_size = (_Small_object_num_ptrs - 1) * sizeof(void*); // 56 sizeof(_Small_storage_t); // 56字节 sizeof(_Big_storage_t); // 56字节
技术疑问解答
1. 为何_Padding能让_Ptr和_RTTI与_TypeData共享缓存行?
先计算_Padding的大小:_Any_small_space_size - sizeof(void*)即48-8=40字节。_Big_storage_t总大小是56字节,_Ptr(8字节)+_RTTI(8字节)=16字节,加上40字节padding正好填满56字节的union空间。
_Storage_t的结构是:union部分占56字节,后面紧跟8字节的_TypeData,总长度64字节——这正好是当前主流CPU的缓存行大小。当使用_Big_storage_t时,_Ptr和_RTTI被padding推到了union的最后16字节位置,也就是整个_Storage_t的第48-63字节区间,而_TypeData在第64-71字节?不,不对,56字节union结束后,_TypeData从第57字节开始,到第64字节结束,刚好和_Ptr、_RTTI的末尾部分落在同一个64字节缓存行里。简单说,padding把_Ptr和_RTTI挤到了union的末尾,让它们和紧接在union后面的_TypeData能被CPU一次性加载进同一个缓存行。
如果没有padding,_Ptr和_RTTI会放在union的起始位置,_TypeData就会落在下一个缓存行,无法共享。
2. 共享缓存行的设计有什么性能优势?
CPU读取数据是以缓存行为单位加载的。当操作存储大对象的std::any时,必然要读取_Ptr(指向堆上对象)和_RTTI(类型信息),同时大概率会用到_TypeData(存储类型标识等元数据)。如果这些数据在同一个缓存行里,CPU一次加载就能把它们都放进缓存,不需要多次从内存读取不同的缓存行,直接降低了缓存未命中的概率。
缓存未命中会导致CPU等待内存数据加载,这是常见的性能瓶颈。共享缓存行让这些高频访问的关联数据一次性加载完成,能提升std::any在类型检查、对象访问等操作时的效率,减少操作延迟。
内容的提问来源于stack exchange,提问作者Pluto

