You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas中计算含字符串与浮点型列的成对距离并规避类型错误?

混合类型数据框的成对距离计算解决方案

一、先解决余弦距离的报错问题

余弦距离仅支持数值型数据,所以必须先处理字符串列:

  • 给字符串做数值编码:用标签编码把字符串转成整数,或独热编码转成二进制向量,转换后就能和浮点列一起计算余弦距离。示例代码:
    from sklearn.preprocessing import LabelEncoder
    import pandas as pd
    from sklearn.metrics.pairwise import cosine_similarity
    
    df = pd.DataFrame({
        'float_col': [1.2, 3.4, 5.6],
        'str_col': ['abc-def', 'ghi-jkl', 'abc-def']
    })
    
    # 将字符串列转为数值
    le = LabelEncoder()
    df['str_col_encoded'] = le.fit_transform(df['str_col'])
    
    # 仅使用数值列计算余弦相似度
    numeric_df = df[['float_col', 'str_col_encoded']]
    cos_sim = cosine_similarity(numeric_df)
    
  • 移除无关字符串列:如果字符串列对相似度判断没有作用,直接删除后再计算,这是最快捷的方法,但仅适用于字符串列无意义的场景。

二、支持混合类型的现成距离指标

有专门针对混合类型数据的距离算法,无需自行从零实现:

  • Gower距离:专为混合类型数据设计,对数值型采用标准化后的曼哈顿距离,对字符串这类分类型采用匹配/不匹配的0-1距离,最后加权平均各列距离。使用gower库即可直接计算:
    import gower
    
    # 直接传入混合类型的DataFrame
    gower_dist_matrix = gower.gower_matrix(df)
    
  • 加权混合距离:分别计算数值列和字符串列的距离后按权重合并。比如数值列用欧氏距离,字符串列用Hamming距离,按比例加权:
    from scipy.spatial.distance import pdist, squareform
    import numpy as np
    
    # 计算数值列的欧氏距离
    numeric_dist = pdist(df[['float_col']], metric='euclidean')
    # 计算字符串列的Hamming距离
    str_dist = pdist(df[['str_col']].values.astype(str), metric='hamming')
    # 按权重合并(示例:数值占60%,字符串占40%)
    combined_dist = 0.6 * numeric_dist + 0.4 * str_dist
    # 转换为方阵格式
    combined_dist_matrix = squareform(combined_dist)
    

三、自行实现混合类型距离逻辑

如果需要完全自定义规则,自行实现也很简单,核心思路是按列类型分别计算距离后加权合并:

  1. 遍历每一对数据行
  2. 数值列采用归一化后的距离,字符串列采用匹配度或编辑距离
  3. 按预设权重将各列距离累加得到总距离

示例代码:

import pandas as pd
import numpy as np

def calculate_mixed_distance(row1, row2, df, weights):
    total_dist = 0.0
    for col in row1.index:
        val1, val2 = row1[col], row2[col]
        if pd.api.types.is_numeric_dtype(df[col]):
            # 数值列使用归一化的绝对距离
            col_range = df[col].max() - df[col].min()
            norm_dist = abs(val1 - val2) / col_range if col_range != 0 else 0
            total_dist += norm_dist * weights[col]
        else:
            # 字符串列:相同为0,不同为1
            str_dist = 0 if val1 == val2 else 1
            total_dist += str_dist * weights[col]
    return total_dist

# 测试数据
df = pd.DataFrame({
    'float_col': [1.2, 3.4, 5.6],
    'str_col': ['abc-def', 'ghi-jkl', 'abc-def']
})

# 设定各列权重
weights = {'float_col': 0.7, 'str_col': 0.3}

# 生成距离矩阵
n_rows = len(df)
dist_matrix = np.zeros((n_rows, n_rows))
for i in range(n_rows):
    for j in range(n_rows):
        dist_matrix[i][j] = calculate_mixed_distance(df.iloc[i], df.iloc[j], df, weights)

内容的提问来源于stack exchange,提问作者TRK

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 13:50:45