You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

预转置Eigen::Tensor能否提升张量收缩运算效率?

问题:张量收缩前预置换索引是否能提升Eigen运算效率?

我有两组Eigen::Tensor数组A、B,每组内的张量秩相同。程序中需要按固定规则(保证维度匹配)将A中的张量与B中的张量执行张量收缩操作,且收缩索引固定,该操作会被多次执行。比如A中是3阶张量,B中是4阶张量,需要将A的0、2索引与B的1、2索引进行收缩。

由于张量收缩通常先重塑为矩阵再执行矩阵乘法,我疑惑预先对张量进行索引置换是否能提升运算效率。为此我编写了极简测试代码,但结果显示各测试场景的执行耗时相近。测试环境为macOS,编译参数为-O3 -march=native -std=c++17,测试代码及耗时结果如下:

#include <unsupported/Eigen/CXX11/Tensor>
#include <Eigen/Dense>
#include <chrono>
#include <iostream>
#include <iomanip>
typedef std::complex<double> cplx;

int main()
{
    std::chrono::time_point<std::chrono::high_resolution_clock> t0,t1;
    Eigen::Tensor<cplx,3> a(200,100,90);
    a.setRandom();
    Eigen::Tensor<cplx,3> b(200,50,100);
    b.setRandom();
    t0= std::chrono::high_resolution_clock::now();

    Eigen::array<Eigen::IndexPair<int>, 1> product_dims = {Eigen::IndexPair<int>(0,0)};
    for(int i=0;i<10;++i)
        Eigen::Tensor<cplx,4> c=a.contract(b,product_dims);
    t1= std::chrono::high_resolution_clock::now();
    std::cout<<std::setprecision(12)<<" in time: "<<1*1e-10*(double)std::chrono::duration_cast<std::chrono::nanoseconds>(t1-t0).count()<<"s"<<std::endl;

    Eigen::Tensor<cplx,3> dd=a.shuffle(std::vector{1,2,0});
    Eigen::Tensor<cplx,3> ee=b.shuffle(std::vector{1,2,0});

    t0= std::chrono::high_resolution_clock::now();
    product_dims = {Eigen::IndexPair<int>(2,2)};
    for(int i=0;i<10;++i)
        Eigen::Tensor<cplx,4> c=dd.contract(ee,product_dims);
    t1= std::chrono::high_resolution_clock::now();
    std::cout<<std::setprecision(12)<<" in time: "<<1*1e-10*(double)std::chrono::duration_cast<std::chrono::nanoseconds>(t1-t0).count()<<"s"<<std::endl;

    t0= std::chrono::high_resolution_clock::now();
    product_dims = {Eigen::IndexPair<int>(2,0)};
    for(int i=0;i<10;++i)
         Eigen::Tensor<cplx,4> c=dd.contract(b,product_dims);
    t1= std::chrono::high_resolution_clock::now();
    std::cout<<std::setprecision(12)<<" in time: "<<1*1e-10*(double)std::chrono::duration_cast<std::chrono::nanoseconds>(t1-t0).count()<<"s"<<std::endl;

    t0= std::chrono::high_resolution_clock::now();
    product_dims = {Eigen::IndexPair<int>(0,2)};
    for(int i=0;i<10;++i)
        Eigen::Tensor<cplx,4> c=a.contract(ee,product_dims);
    t1= std::chrono::high_resolution_clock::now();
    std::cout<<std::setprecision(12)<<" in time: "<<1*1e-10*(double)std::chrono::duration_cast<std::chrono::nanoseconds>(t1-t0).count()<<"s"<<std::endl;

}

测试输出结果:

in time: 2.4562778s
 in time: 2.4876118s
 in time: 2.4699632s
 in time: 2.4621803s

内容的提问来源于stack exchange,提问作者jry liu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 10:02:01