如何优化含条件成员的C++聚合类型的内存对齐?
给定以下聚合类型定义:
template <class scalar_t> struct vec3 { constexpr static auto size = 3; scalar_t x, y, z; // 构造函数、赋值运算符、析构函数省略 }; template <class scalar_t> struct vec2 { constexpr static auto size = 2; scalar_t x, y; // 构造函数、赋值运算符、析构函数省略 };
希望组合出包含条件成员的聚合类型:满足条件时使用上述vecN类型,否则使用empty_t,empty_t定义如下:
struct empty { constexpr static auto size = 0; }; using empty_t = empty; // 或<variant>中的std::monostate
构造的vertex类型模板如下:
template <bool has_color = false, bool has_texture = false> struct vertex { using position_type = vec3<double>; using color_type = std::conditional_t<has_color, vec3<double>, empty_t>; using texture_coords_type = std::conditional_t<has_texture, vec2<double>, empty_t>; position_type position{}; color_type color{}; texture_coords_type texture_coords{}; // 构造函数、赋值运算符、析构函数省略 };
仅包含位置数据的vertex类型声明为:
using p_vertex = vertex<false, false>;
当前问题是编译器的类型对齐导致empty_t仍占用1字节,最终整个类型被填充至32字节:
std::cout << "position vertex:" << std::endl; std::cout << sizeof(p_vertex::position_type{}) << std::endl; // 24 std::cout << sizeof(p_vertex::color_type{}) << std::endl; // 1 std::cout << sizeof(p_vertex::texture_coords_type{}) << std::endl; // 1 std::cout << sizeof(p_vertex{}) << std::endl << std::endl; // 32
问题
- 能否通过
alignas或alignof实现24字节对齐?(注:使sizeof(empty_t)==0已被证实不可行) - 若每个
vecN的scalar_t分别模板化如下:
using position_type = vec3<A>; using color_type = std::conditional_t<has_color, vec3<B>, empty_t>; using texture_coords_type = std::conditional_t<has_texture, vec2<C>, empty_t>;
其中A、B、C为<cstdint>内置整数类型或float/double,能否实现无填充的类型对齐?
问题1:能否通过alignas/alignof实现24字节对齐?
不行。alignas只能增加类型的对齐要求,无法减少。当前p_vertex的sizeof为32字节,原因是vec3<double>的对齐要求为8字节(double的对齐标准是8),整个vertex类型的对齐要求由最大成员的对齐规则决定,即8字节。而24+1+1=26,为满足8字节对齐,编译器会填充6字节至32。
即使给empty_t加上alignas(1),或给vertex强制指定alignas(24),也无法改变empty_t占用1字节的事实——C++标准规定任何非空类型的大小至少为1,且空基类优化(EBO)仅适用于基类,不适用于成员变量。只要vertex包含empty_t类型的成员,就会占用至少1字节,导致总大小超过24,进而被对齐到更大的倍数。
正确的解决思路是改用空基类优化,将条件成员作为基类继承而非成员变量,示例如下:
template <bool has_color, bool has_texture> struct vertex_base : std::conditional_t<has_color, vec3<double>, empty_t>, std::conditional_t<has_texture, vec2<double>, empty_t> { vec3<double> position{}; };
此时当has_color为false时,继承的empty_t会被EBO优化掉,不占用空间;has_texture为false时同理。p_vertex的sizeof将等于vec3<double>的大小24字节,符合预期。
问题2:分别模板化scalar_t能否实现无填充对齐?
只要存在empty_t类型的成员,就无法完全避免填充——因为empty_t至少占用1字节,几乎不可能让所有成员的总大小刚好符合最大对齐要求的整数倍。
举个例子:若A是float(大小4,对齐4),position_type大小为12;当has_color和has_texture均为false时,vertex的总大小为12+1+1=14,会被对齐到16(满足4字节对齐规则),仍存在填充。
同样,解决核心还是用空基类优化替代成员变量,将条件成员作为基类继承。示例如下:
template <typename A, typename B, typename C, bool has_color, bool has_texture> struct vertex : std::conditional_t<has_color, vec3<B>, empty_t>, std::conditional_t<has_texture, vec2<C>, empty_t> { vec3<A> position{}; };
此时当has_color和has_texture为false时,vertex的sizeof等于vec3<A>的大小,完全无填充;若has_color为true,只要vec3<B>的对齐要求不与vec3<A>冲突,编译器会通过EBO优化掉空基类的占用,实现无填充布局。
内容的提问来源于stack exchange,提问作者matreska

