自定义STL Allocator:如何避免基础类型触发逐元素construct/destroy?
问题:自定义Allocator针对基础类型避免construct/destroy调用的优化
测试代码
auto start = high_resolution_clock::now(); std::vector<char> myBuffer(20e6); std::cout << "StandardAlloc Time:" << duration_cast<milliseconds>(high_resolution_clock::now() - start).count() << std::endl; start = high_resolution_clock::now(); std::vector<char, HeapAllocator<char>> myCustomBuffer(20e6); std::cout << "CustomAlloc Time:" << duration_cast<milliseconds>(high_resolution_clock::now() - start).count() << " CC: " << HeapAllocator<char>::constructCount << std::endl;
输出结果
StandardAlloc Time:6 CustomAlloc Time:124 CC: 20000000
自定义Allocator实现
template<class T> struct HeapAllocator { typedef T value_type; HeapAllocator(){}; template<class U> constexpr HeapAllocator(const HeapAllocator<U>&) noexcept {} [[nodiscard]] T* allocate(std::size_t n) { auto p = new T[n]; return p; } void deallocate(T* p, std::size_t n) noexcept { delete p; } template <class U> void destroy(U* p) { destroyCount++; } template< class U, class... Args > void construct(U* p, Args&&... args) { constructCount++; } static int destroyCount; static int constructCount; }; template<class T> int HeapAllocator<T>::constructCount = 0;
从测试结果可见,自定义Allocator对每个char元素都调用了construct,导致执行时间是标准分配器的20倍。核心原因是标准分配器会自动跳过可平凡构造/析构类型(如char等基础类型)的construct和destroy调用,而自定义实现没有做这个区分。
解决方案
1. 编译期区分平凡类型,跳过不必要操作
利用C类型特性std::is_trivially_constructible和std::is_trivially_destructible,结合if constexpr(C17+)在编译期判断类型,只对非平凡类型执行construct/destroy及计数逻辑:
优化后的Allocator实现
#include <type_traits> template<class T> struct HeapAllocator { using value_type = T; HeapAllocator() = default; template<class U> constexpr HeapAllocator(const HeapAllocator<U>&) noexcept {} [[nodiscard]] T* allocate(std::size_t n) { // 分配未初始化内存,和标准分配器行为一致 return static_cast<T*>(operator new(n * sizeof(T))); } void deallocate(T* p, std::size_t) noexcept { operator delete(p); } template<class U, class... Args> void construct(U* p, Args&&... args) { if constexpr (!std::is_trivially_constructible_v<U, Args...>) { constructCount++; // 用定位new构造非平凡类型 ::new(static_cast<void*>(p)) U(std::forward<Args>(args)...); } // 平凡类型跳过构造和计数 } template<class U> void destroy(U* p) { if constexpr (!std::is_trivially_destructible_v<U>) { destroyCount++; p->~U(); // 显式调用非平凡类型的析构函数 } // 平凡类型跳过析构和计数 } static inline int constructCount = 0; static inline int destroyCount = 0; }; // 分配器必要的比较运算符支持 template<class T, class U> bool operator==(const HeapAllocator<T>&, const HeapAllocator<U>&) noexcept { return true; } template<class T, class U> bool operator!=(const HeapAllocator<T>&, const HeapAllocator<U>&) noexcept { return false; }
2. 继承标准分配器复用优化逻辑
如果不需要完全自定义内存分配逻辑,可以直接继承std::allocator,只扩展计数部分,自动复用标准分配器对平凡类型的优化:
#include <type_traits> #include <memory> template<class T> struct HeapAllocator : std::allocator<T> { using base = std::allocator<T>; // 继承所有构造函数 using base::base; // 容器rebind所需的嵌套类型 template<class U> struct rebind { using other = HeapAllocator<U>; }; template<class U, class... Args> void construct(U* p, Args&&... args) { if constexpr (!std::is_trivially_constructible_v<U, Args...>) { constructCount++; base::construct(p, std::forward<Args>(args)...); } } template<class U> void destroy(U* p) { if constexpr (!std::is_trivially_destructible_v<U>) { destroyCount++; base::destroy(p); } } static inline int constructCount = 0; static inline int destroyCount = 0; };
关键注意点
- 原Allocator的
allocate用new T[n]会直接初始化所有元素,这也是性能差异的原因之一。优化后的实现改用operator new分配未初始化内存,和标准vector的行为一致。 - 必须实现分配器的
rebind和比较运算符(C++17前),确保容器能正确使用分配器的rebind版本。
内容的提问来源于stack exchange,提问作者tmanthey
相关产品推荐
相关产品推荐

