如何在Eigen中避免矩阵拷贝优化函数,且不用大量if-else?
优化Eigen矩阵求和函数以避免不必要的转置拷贝
原始函数实现
Eigen::MatrixXd func(const Eigen::MatrixXd &matrix1, bool condition1, const Eigen::MatrixXd &matrix2, bool condition2, const Eigen::MatrixXd &matrix3, bool condition3) { const Eigen::MatrixXd &m1 = (condition1) ? matrix1 : matrix1.transpose(); const Eigen::MatrixXd &m2 = (condition2) ? matrix2 : matrix2.transpose(); const Eigen::MatrixXd &m3 = (condition3) ? matrix3 : matrix3.transpose(); return m1 + m2 + m3; }
性能问题
处理10000x10000规模矩阵时,不同转置场景下的耗时差异明显:
0 transpose M1+M2+M3: 2.641621 seconds 1 transpose M1+M2+M3T: 4.142240 seconds 1 transpose M1+M2T+M3: 4.335527 seconds 2 transpose M1+M2T+M3T: 5.276609 seconds 1 transpose M1T+M2+M3: 3.844055 seconds 2 transpose M1T+M2+M3T: 6.080098 seconds 2 transpose M1T+M2T+M3: 5.448039 seconds 3 transpose M1T+M2T+M3T: 6.644677 seconds
问题根源
代码中const Eigen::MatrixXd &m = (condition) ? matrix : matrix.transpose();的写法存在隐患:当条件不成立时,matrix.transpose()会生成一个临时的MatrixXd对象,而非直接返回转置视图,这会触发额外的内存分配与数据拷贝,大幅增加耗时。
虽然Eigen::Transpose<Eigen::MatrixXd>可以作为无拷贝的转置视图,但它的类型与原矩阵Eigen::MatrixXd不同,无法用同一个引用变量同时存储原矩阵和转置视图。
优化方案
经Eigen开发者在Discord上的建议,使用if-else分支是该场景下的最优解。通过显式分支处理所有转置组合,直接利用Eigen的表达式模板机制,避免临时矩阵的拷贝:
Eigen::MatrixXd func(const Eigen::MatrixXd &matrix1, bool condition1, const Eigen::MatrixXd &matrix2, bool condition2, const Eigen::MatrixXd &matrix3, bool condition3) { if (condition1) { if (condition2) { if (condition3) { return matrix1 + matrix2 + matrix3; } else { return matrix1 + matrix2 + matrix3.transpose(); } } else { if (condition3) { return matrix1 + matrix2.transpose() + matrix3; } else { return matrix1 + matrix2.transpose() + matrix3.transpose(); } } } else { if (condition2) { if (condition3) { return matrix1.transpose() + matrix2 + matrix3; } else { return matrix1.transpose() + matrix2 + matrix3.transpose(); } } else { if (condition3) { return matrix1.transpose() + matrix2.transpose() + matrix3; } else { return matrix1.transpose() + matrix2.transpose() + matrix3.transpose(); } } } }
这种写法虽然分支较多,但能确保所有转置操作都以视图形式参与表达式计算,完全避免不必要的内存分配和数据拷贝,大幅提升大矩阵场景下的性能。
内容的提问来源于stack exchange,提问作者1 1
相关产品推荐
相关产品推荐

