Eigenlib小规模矩阵运算性能疑问:手写实现为何快3倍?
Hey there, let's break down why your hand-implemented quadratic metric norm is outperforming Eigen's version by such a large margin, and how to fix it.
Background Recap
You're working with 2D vectors/matrices using Eigen, with this implementation:
const int dims = 2; template <typename T> using Vec = Eigen::Matrix<T, dims, 1>; template <typename T> using Mat = Eigen::Matrix<T, dims, dims>; template <typename T> inline auto quadraticMetricNorm(const Vec<T>& x1, const Vec<T>& x2, const Mat<T>& D) { return x1.transpose() * D * x2; }
You expected it to match (or exceed) your fully unrolled hand-written version, but in VS2017 Release mode (/O2, AVX2 enabled) on an i7-5820K, the hand-written code runs ~3x faster (440ms vs 850ms). Worse, Eigen's generated assembly doesn't use AVX2 instructions and is more complex than your tight scalar implementation.
Key Likely Causes & Fixes
1. Eigen Isn't Detecting AVX2 Support in VS2017
VS2017's compiler has quirks with how Eigen auto-detects SIMD instruction sets. Even if you enabled /arch:AVX2 in project settings, Eigen might not pick it up automatically.
Fix:
Add these preprocessor macros to your project settings (under C/C++ > Preprocessor > Preprocessor Definitions):
EIGEN_USE_AVX2: Forces Eigen to enable AVX2 optimizationsEIGEN_NO_DEBUG: Disables Eigen's debug-time assertions (critical for Release mode performance, as these add hidden overhead)
2. Alignment Issues
Eigen relies heavily on aligned memory for SIMD optimizations. For fixed-size types like your 2D Vec<T> (16 bytes for double, which fits AVX2's 256-bit registers), Eigen handles stack alignment automatically, but issues can arise if:
- You're storing
Vec<T>orMat<T>as class members without alignment - Your project's stack alignment settings are misconfigured
Fix:
- If using these types as class members, add the
EIGEN_MAKE_ALIGNED_OPERATOR_NEWmacro inside your class definition to ensure proper heap alignment:class MyClass { public: Vec<double> my_vec; EIGEN_MAKE_ALIGNED_OPERATOR_NEW }; - Verify your project's stack alignment: In VS2017, go to C/C++ > Code Generation > Struct Member Alignment and set it to
16 Bytes(matches Eigen's requirements).
3. Expression Template Overhead (For Tiny Operations)
Eigen's expression templates excel at large operations, but for extremely small computations like 2x2 quadratic forms, the template machinery can sometimes introduce unnecessary indirection that the compiler can't fully optimize away—unlike your fully unrolled hand-written code.
Fix:
Force Eigen to evaluate the expression immediately using .eval() to eliminate delayed computation overhead:
template <typename T> inline auto quadraticMetricNorm(const Vec<T>& x1, const Vec<T>& x2, const Mat<T>& D) { return (x1.transpose() * D * x2).eval(); }
4. Outdated Eigen Version
Older Eigen versions (pre-3.3) have limited support for VS2017's AVX2 implementation and compiler quirks. If you're using an older release, this could be the root cause.
Fix:
Update to the latest stable Eigen release (3.4.x or newer)—newer versions have improved VS compiler compatibility and better SIMD auto-detection.
5. Verify Eigen's Optimization Setup
To confirm Eigen is using AVX2, add the EIGEN_PRINT_SUMMARY preprocessor macro. This will make Eigen output a summary of detected instruction sets and optimizations during compilation (look for "AVX2" in the build log).
Final Notes
Once you apply these fixes, Eigen's generated assembly should match (or even surpass) your hand-written version by leveraging AVX2's 256-bit registers. Your hand-written code uses scalar SSE2 instructions (addsd, mulsd), but AVX2 can handle parallel operations—even for 2D cases, it will eliminate the extra overhead you're seeing now.
内容的提问来源于stack exchange,提问作者thogra

