如何删除组合数组中仅顺序不同的重复二元组集合
去重逻辑实现思路
核心原理是为每个组合生成不受内部元素顺序影响的唯一标识,通过集合校验标识是否重复,实现过滤:
- 首先将每个子坐标数组转成可哈希的元组(数组本身不可哈希,无法存入集合)
- 浮点坐标要先做精度截断(比如保留6位小数),避免浮点精度误差导致相同坐标判定为不同
- 将组合内的两个元组按规则排序后拼成最终的唯一键
- 用集合存储已经出现过的键,遍历过程中如果键已存在就跳过,不存在就保留并把键加入集合
修改后完整代码
import numpy as np import itertools as it filename = "example.xyz" xyz_file = np.genfromtxt(fname=filename, skip_header=2, dtype='unicode') with open(filename) as xyz: atom_quantity = int(xyz.readline()) atom_coordinates = xyz_file[:, 1:].astype(float) print("The xyz coordinate of the each atom:", atom_coordinates) array = np.array(atom_coordinates) combinations = it.product(array, array) # 去重用的存储集合 seen = set() for combination in combinations: # 处理坐标精度+转可哈希元组 t1 = tuple(np.round(combination[0], 6)) t2 = tuple(np.round(combination[1], 6)) # 排序生成不受顺序影响的唯一键 unique_key = tuple(sorted([t1, t2])) # 重复则直接跳过 if unique_key in seen: continue seen.add(unique_key) array_comb = np.array(combination) print(array_comb)
可选优化
如果你不需要(a,a)这种自身和自身的组合,也不需要顺序相反的重复对,可以直接把it.product(array, array)替换为it.combinations(array, 2),从根源上减少冗余组合生成,后续可省掉额外去重步骤。
内容的提问来源于stack exchange,提问作者riven
相关产品推荐
相关产品推荐

