如何在Pandas中计算含字符串与浮点型列的成对距离并规避类型错误?
混合类型数据框的成对距离计算解决方案
一、先解决余弦距离的报错问题
余弦距离仅支持数值型数据,所以必须先处理字符串列:
- 给字符串做数值编码:用标签编码把字符串转成整数,或独热编码转成二进制向量,转换后就能和浮点列一起计算余弦距离。示例代码:
from sklearn.preprocessing import LabelEncoder import pandas as pd from sklearn.metrics.pairwise import cosine_similarity df = pd.DataFrame({ 'float_col': [1.2, 3.4, 5.6], 'str_col': ['abc-def', 'ghi-jkl', 'abc-def'] }) # 将字符串列转为数值 le = LabelEncoder() df['str_col_encoded'] = le.fit_transform(df['str_col']) # 仅使用数值列计算余弦相似度 numeric_df = df[['float_col', 'str_col_encoded']] cos_sim = cosine_similarity(numeric_df) - 移除无关字符串列:如果字符串列对相似度判断没有作用,直接删除后再计算,这是最快捷的方法,但仅适用于字符串列无意义的场景。
二、支持混合类型的现成距离指标
有专门针对混合类型数据的距离算法,无需自行从零实现:
- Gower距离:专为混合类型数据设计,对数值型采用标准化后的曼哈顿距离,对字符串这类分类型采用匹配/不匹配的0-1距离,最后加权平均各列距离。使用
gower库即可直接计算:import gower # 直接传入混合类型的DataFrame gower_dist_matrix = gower.gower_matrix(df) - 加权混合距离:分别计算数值列和字符串列的距离后按权重合并。比如数值列用欧氏距离,字符串列用Hamming距离,按比例加权:
from scipy.spatial.distance import pdist, squareform import numpy as np # 计算数值列的欧氏距离 numeric_dist = pdist(df[['float_col']], metric='euclidean') # 计算字符串列的Hamming距离 str_dist = pdist(df[['str_col']].values.astype(str), metric='hamming') # 按权重合并(示例:数值占60%,字符串占40%) combined_dist = 0.6 * numeric_dist + 0.4 * str_dist # 转换为方阵格式 combined_dist_matrix = squareform(combined_dist)
三、自行实现混合类型距离逻辑
如果需要完全自定义规则,自行实现也很简单,核心思路是按列类型分别计算距离后加权合并:
- 遍历每一对数据行
- 数值列采用归一化后的距离,字符串列采用匹配度或编辑距离
- 按预设权重将各列距离累加得到总距离
示例代码:
import pandas as pd import numpy as np def calculate_mixed_distance(row1, row2, df, weights): total_dist = 0.0 for col in row1.index: val1, val2 = row1[col], row2[col] if pd.api.types.is_numeric_dtype(df[col]): # 数值列使用归一化的绝对距离 col_range = df[col].max() - df[col].min() norm_dist = abs(val1 - val2) / col_range if col_range != 0 else 0 total_dist += norm_dist * weights[col] else: # 字符串列:相同为0,不同为1 str_dist = 0 if val1 == val2 else 1 total_dist += str_dist * weights[col] return total_dist # 测试数据 df = pd.DataFrame({ 'float_col': [1.2, 3.4, 5.6], 'str_col': ['abc-def', 'ghi-jkl', 'abc-def'] }) # 设定各列权重 weights = {'float_col': 0.7, 'str_col': 0.3} # 生成距离矩阵 n_rows = len(df) dist_matrix = np.zeros((n_rows, n_rows)) for i in range(n_rows): for j in range(n_rows): dist_matrix[i][j] = calculate_mixed_distance(df.iloc[i], df.iloc[j], df, weights)
内容的提问来源于stack exchange,提问作者TRK
相关产品推荐
相关产品推荐

