如何以Numpy标签数组将网格数组转换为指定格式的字典?
将Numpy相似度网格转换为字典列表
问题描述
现有两个Numpy数组:
label数组存储对应行/列的字母标识:['a','b','c','d']grid_Arr是4×4的相似度网格,对角线元素为1(同一字母对比的相似度),其余元素代表对应字母间的相似度。
需要将这两个数组转换为字典列表,每个字典包含word1(行对应字母)、word2(列对应字母)、score(对应位置的相似度值)三个键,格式示例如下:
[{'word1': 'a', 'word2': 'a', 'score': 1}, {'word1': 'a', 'word2': 'b', 'score': 0.3}, {'word1': 'a', 'word2': 'c', 'score': 0.5}, {'word1': 'a', 'word2': 'd', 'score': 0.6}, ...]
解决方案
基础循环实现
通过嵌套循环遍历网格每个元素,结合label数组映射字母,构建目标字典列表:
import numpy as np label = np.array(['a','b','c','d']) grid_Arr = np.array([[1. , 0.3, 0.5, 0.6], [0.3, 1. , 0.4, 0.1], [0.5, 0.4, 1. , 0.2], [0.6, 0.1, 0.2, 1. ]]) result = [] for i in range(len(label)): word1 = label[i] for j in range(len(label)): word2 = label[j] score = grid_Arr[i][j] result.append({'word1': word1, 'word2': word2, 'score': score}) # 验证结果 print(result)
高效向量化实现(适合大规模数据)
利用Numpy的向量化操作替代循环,提升处理效率:
import numpy as np label = np.array(['a','b','c','d']) grid_Arr = np.array([[1. , 0.3, 0.5, 0.6], [0.3, 1. , 0.4, 0.1], [0.5, 0.4, 1. , 0.2], [0.6, 0.1, 0.2, 1. ]]) # 生成所有字母对组合 word1s = np.repeat(label, len(label)) word2s = np.tile(label, len(label)) scores = grid_Arr.flatten() # 转换为字典列表 result = [{'word1': w1, 'word2': w2, 'score': s} for w1, w2, s in zip(word1s, word2s, scores)] print(result)
代码说明
- 基础循环:外层循环取行对应字母作为
word1,内层循环取列对应字母作为word2,匹配网格中的相似度值后构建字典,逐个加入结果列表。 - 向量化实现:用
np.repeat和np.tile快速生成所有字母对,将网格数组扁平化后与字母对一一对应,最后通过列表推导式批量生成字典,性能优于嵌套循环。
内容的提问来源于stack exchange,提问作者Salih
相关产品推荐
相关产品推荐

