如何在二维Numpy数组中计算每行最长连续1的长度?
计算二维Numpy数组每行连续1的最大长度
给定如下二维Numpy数组:
import numpy as np a = np.array([[1, 1, 1, 1, 1], [1, 0, 1, 0, 1], [1, 1, 0, 1, 0], [0, 0, 0, 0, 0], [1, 1, 1, 0, 1], [1, 0, 0, 0, 0], [0, 1, 1, 0, 0], [1, 0, 1, 1, 0], ])
需求是计算每行中连续1的最大长度,期望输出为:[5, 1, 2, 0, 3, 1, 2, 2]
已知一维数组的解决方案:
a_1d = np.array([1, 1, 1, 1, 0, 1, 0, 0, 0, 1, 1, 0, 0]) d = np.diff(np.concatenate(([0], a_1d, [0]))) max_len = np.max(np.flatnonzero(d == -1) - np.flatnonzero(d == 1)) # 输出:4
尝试将该思路扩展到二维时,编写的代码无法正常运行:
d = np.diff(np.column_stack(([0] * a.shape[0], a, [0] * a.shape[0]))) np.max(np.flatnonzero(d == -1) - np.flatnonzero(d == 1))
解决方案
问题出在直接对二维数组整体做diff会跨行处理,导致行与行之间的元素被错误关联。需要逐行处理,保留一维思路的同时,对每行单独执行逻辑:
方法1:使用np.apply_along_axis封装一维逻辑
def max_consecutive_ones(arr): # 对单行执行原一维逻辑 d = np.diff(np.concatenate(([0], arr, [0]))) start_indices = np.flatnonzero(d == 1) end_indices = np.flatnonzero(d == -1) if len(start_indices) == 0: return 0 return np.max(end_indices - start_indices) result = np.apply_along_axis(max_consecutive_ones, axis=1, arr=a) print(result) # 输出:[5 1 2 0 3 1 2 2]
方法2:向量化实现(更高效,避免循环)
如果数据集很大,apply_along_axis效率较低,可以用向量化操作:
# 给每行前后各加一个0 padded = np.hstack([np.zeros((a.shape[0],1)), a, np.zeros((a.shape[0],1))]) # 计算每行的diff,axis=1表示每行内求差 d = np.diff(padded, axis=1) # 找到每行中d==1的位置(连续1的起始)和d==-1的位置(连续1的结束) start_mask = d == 1 end_mask = d == -1 # 对每行计算最大连续长度 result = [] for i in range(a.shape[0]): starts = np.flatnonzero(start_mask[i]) ends = np.flatnonzero(end_mask[i]) if len(starts) == 0: result.append(0) else: result.append(np.max(ends - starts)) result = np.array(result) print(result) # 输出:[5 1 2 0 3 1 2 2]
内容的提问来源于stack exchange,提问作者Abhishek Jain
相关产品推荐
相关产品推荐

