如何按出现次数移除两个numpy二维数组中的共同行?
移除NumPy二维数组中共同存在的行(按出现次数最小值匹配删除)
核心思路
先统计两个数组中每行的出现次数,对同时存在于两个数组的行,取它们在两个数组中出现次数的最小值,从两个数组中各移除对应次数的该行。
实现代码
方法1:使用collections.Counter(直观易读)
import numpy as np from collections import Counter # 示例数组 A = np.array([[1, 2], [3, 4], [1, 2], [5, 6], [3, 4], [3, 4]]) B = np.array([[1, 2], [3, 4], [3, 4], [7, 8], [1, 2], [1, 2]]) # 将数组行转为可哈希的元组,用于统计次数 rows_A = [tuple(row) for row in A] rows_B = [tuple(row) for row in B] # 统计每行出现次数 count_A = Counter(rows_A) count_B = Counter(rows_B) # 找出共同行,并确定每个行需要移除的次数(取两者次数最小值) common_rows = count_A.keys() & count_B.keys() remove_counts = {row: min(count_A[row], count_B[row]) for row in common_rows} # 过滤A数组 filtered_A = [] current_A_counts = count_A.copy() for row in rows_A: if row in remove_counts: if current_A_counts[row] > remove_counts[row]: filtered_A.append(row) current_A_counts[row] -= 1 else: filtered_A.append(row) filtered_A = np.array(filtered_A) # 过滤B数组 filtered_B = [] current_B_counts = count_B.copy() for row in rows_B: if row in remove_counts: if current_B_counts[row] > remove_counts[row]: filtered_B.append(row) current_B_counts[row] -= 1 else: filtered_B.append(row) filtered_B = np.array(filtered_B) # 输出结果 print("处理后的A数组:") print(filtered_A) print("\n处理后的B数组:") print(filtered_B)
方法2:使用np.unique(更适合大型数组,性能更优)
import numpy as np # 示例数组 A = np.array([[1, 2], [3, 4], [1, 2], [5, 6], [3, 4], [3, 4]]) B = np.array([[1, 2], [3, 4], [3, 4], [7, 8], [1, 2], [1, 2]]) # 统计A的唯一行及对应次数 unique_A, counts_A = np.unique(A, axis=0, return_counts=True) count_A_dict = {tuple(row): cnt for row, cnt in zip(unique_A, counts_A)} # 统计B的唯一行及对应次数 unique_B, counts_B = np.unique(B, axis=0, return_counts=True) count_B_dict = {tuple(row): cnt for row, cnt in zip(unique_B, counts_B)} # 确定共同行的移除次数 common_rows = count_A_dict.keys() & count_B_dict.keys() remove_counts = {row: min(count_A_dict[row], count_B_dict[row]) for row in common_rows} # 过滤A数组 filtered_A = [] temp_counts = count_A_dict.copy() for row in A: row_tuple = tuple(row) if row_tuple in remove_counts: if temp_counts[row_tuple] > remove_counts[row_tuple]: filtered_A.append(row) temp_counts[row_tuple] -= 1 else: filtered_A.append(row) filtered_A = np.array(filtered_A) # 过滤B数组 filtered_B = [] temp_counts = count_B_dict.copy() for row in B: row_tuple = tuple(row) if row_tuple in remove_counts: if temp_counts[row_tuple] > remove_counts[row_tuple]: filtered_B.append(row) temp_counts[row_tuple] -= 1 else: filtered_B.append(row) filtered_B = np.array(filtered_B) # 输出结果 print("处理后的A数组:") print(filtered_A) print("\n处理后的B数组:") print(filtered_B)
结果说明
以示例数组为例:
- 原A数组中
[1,2]出现2次,[3,4]出现3次;原B数组中[1,2]出现3次,[3,4]出现2次。 - 对共同行
[1,2],移除次数为2(取最小值),因此A中该行列全部被移除,B中保留1次;对[3,4],移除次数为2,A中保留1次,B中该行列全部被移除。 - 最终处理后的数组会保留各自独有的行,以及共同行中“剩余次数”的行。
内容的提问来源于stack exchange,提问作者BGR
相关产品推荐
相关产品推荐

