字符串聚类中np.apply_along_axis报错:ValueError: too many values to unpack
解决
np.apply_along_axis的ValueError: too many values to unpack问题 我来帮你拆解一下这个报错的原因,以及怎么修复它:
报错根源
你遇到的问题出在np.triu_indices的返回值和np.apply_along_axis的参数要求不匹配上:
np.triu_indices(len(words), 1)返回的是一个元组,里面包含两个一维数组,分别存储上三角矩阵的行索引和列索引。- 当你直接把这个元组传给
np.apply_along_axis时,numpy并没有按照你预期的方式将其解析为二维数组的行,而是错误地拆解了元组结构,导致传入d函数的coord不是长度为2的数组(无法拆成i和j),而是更长的序列,于是触发了“值过多无法解包”的错误。
修复方案
方案1:转换输入结构,调整axis参数
把triu_indices的结果转换成N×2的二维数组,然后沿着行方向(axis=1)应用函数:
from jellyfish import jaro_distance import numpy as np words = 'CHEESE CHORES GEESE GLOVES'.split() def d(coord): i, j = coord return (1 - jaro_distance(words[i], words[j])) # 将索引元组转换为每行一组(i,j)的二维数组 index_pairs = np.column_stack(np.triu_indices(len(words), 1)) # 沿着行方向处理每组索引 result = np.apply_along_axis(d, axis=1, arr=index_pairs) print(result)
方案2:更高效的向量化写法(推荐)
np.apply_along_axis本质是循环,效率并不高,你可以直接利用Python的迭代或者numpy的向量化工具来实现:
from jellyfish import jaro_distance import numpy as np words = 'CHEESE CHORES GEESE GLOVES'.split() # 直接拆分索引元组 i_indices, j_indices = np.triu_indices(len(words), 1) # 迭代计算每对索引的距离 distances = np.array([1 - jaro_distance(words[i], words[j]) for i, j in zip(i_indices, j_indices)]) print(distances)
如果你的jellyfish版本支持向量化调用,还可以用np.vectorize进一步简化:
distances = 1 - np.vectorize(jaro_distance)(words[i_indices], words[j_indices])
为什么原示例看起来能运行?
大概率是原场景使用的numpy版本较旧,当时的函数对输入类型的兼容性更强,会自动将元组转换为符合要求的数组;而现在的numpy版本对参数类型的检查更严格,必须显式转换结构才能正常工作。
内容的提问来源于stack exchange,提问作者user2906657
相关产品推荐
相关产品推荐

