如何对分块图像数组执行Numpy特定轴点积的向量化运算?
Problem
I have a 512x512 image array, and I need to perform dot products on 8x8 blocks. My current loop-based code works but is slow:
import numpy as np output = np.zeros((512, 512)) for i in range(0, 512, 8): for j in range(0, 512, 8): a = input[i:i+8, j:j+8] b = some_other_array[i:i+8, j:j+8] output[i:i+8, j:j+8] = np.dot(a, b)
I've reshaped the inputs to enable vectorization:
input_reshaped = input.reshape(64, 8, 64, 8) some_other_reshaped = some_other_array.reshape(64, 8, 64, 8)
I need to compute the dot product along the block inner axes (axis 1 and 3 of the reshaped arrays) to get an output of shape (64, 8, 64, 8). np.tensordot gave the right shape but wrong values, and I couldn't find a matching np.einsum example.
Solution
The key is to use np.einsum to explicitly define which axes to sum over (the inner axes of each 8x8 block) while preserving the block structure.
Explanation
In the loop, each np.dot(a, b) computes sum_k a[r, k] * b[k, p] for the 8x8 blocks, where:
r= row index within the block (0-7)k= column index ofa/ row index ofb(the summed axis, 0-7)p= column index within the resulting block (0-7)
For the reshaped arrays:
input_reshapedhas dimensions(block_row, row_in_block, block_col, col_in_block)→(64, 8, 64, 8)some_other_reshapedhas the same dimension structure
We need to sum over the inner axis that connects the two blocks (the col_in_block of input_reshaped and row_in_block of some_other_reshaped).
Code
import numpy as np # Assuming input and some_other_array are already reshaped to (64,8,64,8) output_reshaped = np.einsum('brck, bkcp -> brcp', input_reshaped, some_other_reshaped) # If you need to convert back to 512x512: output = output_reshaped.reshape(512, 512)
Breakdown of the Einsum Expression
brck: Representsinput_reshapeddimensions:b(block row),r(row in block),c(block column),k(column in block)bkcp: Representssome_other_reshapeddimensions:b(block row),k(row in block, matches the summed axis from input),c(block column),p(column in block)brcp: The resulting dimensions:b(block row),r(row in result block),c(block column),p(column in result block)
This exactly replicates the loop's behavior but in a fully vectorized way, which will be much faster.
Why Tensordot Didn't Work
np.tensordot collapses the summed axes, which would result in a shape like (64, 8, 64, 64, 8) if you specify axes ([3], [1]) — not the (64,8,64,8) shape you need. Einsum lets you retain all the necessary axes in their original structure.
内容的提问来源于stack exchange,提问作者Anjum Sayed

