You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Numpy中更高效地计算特定行对的平均值并与指定行拼接

Great question! You're right that np.concatenate (especially paired with np.tile) can introduce unnecessary overhead, especially with larger arrays—those intermediate copies add up quickly. Let's look at two more efficient approaches that cut down on memory usage and redundant operations:

1. Preallocate and Directly Assign (Best Performance)

The fastest way to do this is to preallocate your final result array upfront, then fill it in two parts. This avoids any intermediate array creation (like the repeated a[0] from tile) and skips the concatenation step entirely:

import numpy as np

# Your original arrays
a = np.array(
 [ [1, 2, 3],
 [4, 5, 6],
 [7, 8, 9],
 [10, 11, 12],
 [13, 14, 15],
 ]
)
b = np.array([[1,3], [1,2], [2,3]])

# Step 1: Calculate the row means first (same as your original approach)
row_means = a[b].mean(axis=1)

# Step 2: Preallocate the result array (shape: [number of row pairs, 2*columns in a])
# We use float64 to accommodate the mean's decimal values
result = np.empty((b.shape[0], a.shape[1] * 2), dtype=np.float64)

# Step 3: Fill the first half with a[0] (broadcasting handles the row repetition automatically)
result[:, :a.shape[1]] = a[0]

# Step 4: Fill the second half with our precomputed means
result[:, a.shape[1]:] = row_means

Why this works better:

  • Preallocating memory is far more efficient than dynamically combining arrays—Numpy doesn't have to reallocate and copy data multiple times.
  • The assignment operations are raw, low-overhead array writes with no extra processing.
  • We avoid creating the redundant repeated a[0] array that np.tile generates, which saves memory especially with big arrays.

2. Broadcast Instead of Tile (Simpler, Still Efficient)

If you prefer a more concise approach without preallocating, you can use Numpy's broadcasting to avoid np.tile. Broadcasting creates a "virtual" repeated version of a[0] instead of a physical copy, which saves memory and speeds things up:

row_means = a[b].mean(axis=1)
# Broadcast a[0] to match the number of rows in row_means, then stack horizontally
result = np.hstack([np.broadcast_to(a[0], (len(b), len(a[0]))), row_means])

Or even more concise (Numpy will auto-broadcast the 1D a[0] to match the 2D row_means shape):

result = np.hstack([a[0][None, :].repeat(len(b), axis=0), row_means])

Performance Comparison

To see the difference, let's run a quick benchmark with larger arrays:

import timeit

def original_approach(a, b):
    return np.concatenate((np.tile(a[0], (len(b), 1)), a[b].mean(1)), axis=1)

def prealloc_approach(a, b):
    row_means = a[b].mean(axis=1)
    result = np.empty((len(b), a.shape[1]*2), dtype=np.float64)
    result[:, :a.shape[1]] = a[0]
    result[:, a.shape[1]:] = row_means
    return result

def broadcast_approach(a, b):
    row_means = a[b].mean(axis=1)
    return np.hstack([np.broadcast_to(a[0], (len(b), len(a[0]))), row_means])

# Test with large arrays
a_large = np.random.rand(1000, 1000)
b_large = np.random.randint(0, 1000, (10000, 2))

print("Original method time:", timeit.timeit(lambda: original_approach(a_large, b_large), number=10))
print("Prealloc method time:", timeit.timeit(lambda: prealloc_approach(a_large, b_large), number=10))
print("Broadcast method time:", timeit.timeit(lambda: broadcast_approach(a_large, b_large), number=10))

In most cases, you'll see the preallocation method finish first, followed by the broadcast method, with the original concatenate/tile approach trailing behind—especially as array sizes grow.

内容的提问来源于stack exchange,提问作者SantoshGupta7

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 13:02:39