You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何实现自定义C++ PMR分配器的完全去虚拟化?

问题

我在C++多态内存资源(PMR)项目中实现了继承自std::pmr::memory_resource的自定义LoggingResource,并将其叠加在std::pmr::monotonic_buffer_resource之上,相关代码如下:

#include <vector>
#include <memory_resource>
#include <array>
#include <cstdio>

struct pmr_aware_container
{
    using allocator_type = std::pmr::polymorphic_allocator<std::byte>;

    /* ctors */

    // default
    pmr_aware_container() : pmr_aware_container{allocator_type{}} {} // delegate to aa constructor

    explicit pmr_aware_container(const allocator_type alloc)
        : str_("Hello long string!!!", alloc) {
        printf("default constructor called!\n");
    }

    // copy
    // pmr_aware_container(const pmr_aware_container&) = default;

    pmr_aware_container(const pmr_aware_container& other, allocator_type alloc = {}) 
        : str_(other.str_, alloc) {
        printf("Copy constructor called!\n");
    }

    // move
    pmr_aware_container(pmr_aware_container&& other) noexcept
        : str_{std::move(other.str_), other.get_allocator() }
    {
        printf("Noexcept move constructor called!\n");
    }

    pmr_aware_container(pmr_aware_container&& other, const allocator_type& alloc)
        : str_(std::move(other.str_), alloc)
    {
        printf("Specific move constructor called!\n");
    }

    // assignement

    pmr_aware_container& operator=(const pmr_aware_container& rhs) = default;
    pmr_aware_container& operator=(pmr_aware_container&& rhs) = default;

    ~pmr_aware_container() = default;

    allocator_type get_allocator() const {
        return str_.get_allocator();
    }

    std::pmr::string str_ = "Hello long string!!!";
};


class LoggingResource : public std::pmr::memory_resource
{
public:
    LoggingResource(std::pmr::memory_resource *underlying_resource) : underlying_resource_{underlying_resource} { }

private:
    void *do_allocate(size_t bytes, size_t align) override {
        printf("Allocating %d bytes!\n", bytes);
        return underlying_resource_->allocate(bytes, align);
    }

    void do_deallocate(void*p, size_t bytes, size_t align) {
        underlying_resource_->deallocate(p, bytes, align);
    }

    bool do_is_equal(std::pmr::memory_resource const& other) const noexcept override {
        return underlying_resource_->is_equal(other);
    }

    std::pmr::memory_resource* underlying_resource_;
};

int main()
{
    std::array<std::byte, 2024> buf;
    std::pmr::monotonic_buffer_resource mbs{buf.data(), buf.size()};

    LoggingResource log_resource{&mbs};

    std::pmr::vector<pmr_aware_container> v{ { pmr_aware_container{ &log_resource}, pmr_aware_container{ &log_resource} }, &log_resource};
}

经GCC 12.1编译后,std::pmr::monotonic_buffer_resource的虚调用已被去虚拟化,但自定义LoggingResource的虚调用未被去虚拟化(GCC 9.1时两者均未实现去虚拟化)。编译后的汇编中仍存在LoggingResource的vtable条目:

vtable for LoggingResource:
        .quad   0
        .quad   typeinfo for LoggingResource
        .quad   LoggingResource::~LoggingResource() [complete object destructor]
        .quad   LoggingResource::~LoggingResource() [deleting destructor]
        .quad   LoggingResource::do_allocate(unsigned long, unsigned long)
        .quad   LoggingResource::do_deallocate(void*, unsigned long, unsigned long)
        .quad   LoggingResource::do_is_equal(std::pmr::memory_resource const&) const

我想了解这一现象的原因,以及如何让自定义PMR分配器实现完全去虚拟化,避免嵌套分配器时的性能问题。


原因分析

  • 标准库类型的专属优化:GCC对std::pmr::monotonic_buffer_resource等标准库PMR类型有内置的去虚拟化优化——编译器会特殊识别这些标准库类型,在编译期确定其具体类型,直接将虚调用替换为直接函数调用,跳过vtable查找流程。
  • 自定义类型的编译期信息限制:自定义LoggingResource属于用户代码,编译器无法像对待标准库类型那样做硬编码式优化。虽然log_resource在main中是明确的栈上对象,但当它被抽象为std::pmr::memory_resource*指针传递给分配器或容器时,编译器无法在所有调用路径中跟踪到指针的实际类型,因此无法安全地完成去虚拟化。
  • vtable存在的必然性:只要类包含虚函数,C++标准要求生成vtable,即便部分虚调用被优化。vtable仍需为运行时类型识别(RTTI)、对象销毁等场景提供支持。

优化方案

1. 用静态多态替代动态多态

放弃继承std::pmr::memory_resource,改用模板实现静态多态,让编译期直接确定调用的函数:

template <typename UnderlyingResource>
class LoggingResource
{
public:
    explicit LoggingResource(UnderlyingResource* underlying) : underlying_(underlying) {}

    void* allocate(size_t bytes, size_t align) {
        printf("Allocating %zu bytes!\n", bytes);
        return underlying_->allocate(bytes, align);
    }

    void deallocate(void* p, size_t bytes, size_t align) {
        underlying_->deallocate(p, bytes, align);
    }

    bool is_equal(const LoggingResource& other) const noexcept {
        return underlying_->is_equal(*other.underlying_);
    }

private:
    UnderlyingResource* underlying_;
};

配合自定义静态多态分配器使用,彻底规避动态虚调用的开销。

2. 强制编译期类型推导

在调用点明确传递具体类型信息,帮助编译器完成去虚拟化。例如在构造PMR容器时,暴露LoggingResource的具体类型:

// 在main中,将log_resource的具体类型传递给容器
std::pmr::vector<pmr_aware_container> v{..., static_cast<LoggingResource*>(&log_resource)};

这种方式依赖编译器优化能力,简单场景下可能生效,但复杂调用路径中仍可能失效。

3. 使用编译器属性提示优化

在自定义LoggingResource的虚函数上添加GCC扩展属性[[gnu::always_inline]],提示编译器尽可能内联虚函数实现,间接实现去虚拟化:

void *do_allocate(size_t bytes, size_t align) override [[gnu::always_inline]] {
    printf("Allocating %zu bytes!\n", bytes);
    return underlying_resource_->allocate(bytes, align);
}

注意该属性是GCC专属,不具备跨编译器可移植性。

4. 合并资源逻辑

如果场景允许,将日志逻辑直接嵌入monotonic_buffer_resource的使用代码中,避免嵌套动态资源的开销。例如在分配前手动打印日志,无需单独的LoggingResource类。


内容的提问来源于stack exchange,提问作者glades

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.23 11:27:05