You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于ID列匹配DataFrame列并计算准确率百分比的方法

针对DataFrame共同ID的列匹配准确率计算与可视化方案

原代码的问题

你的代码存在两个核心问题:

  • 使用reset_index()后比较的是整个DataFrame(含新增的索引列),而非目标列的数值,导致准确率计算完全错误
  • 仅通过排序硬对齐数据,未以ID为键确保严格匹配,若存在重复ID或排序异常会导致对齐偏差

最优实现步骤

1. 对齐共同ID数据

先提取两个DataFrame的交集ID,再分别筛选并以ID为索引对齐,确保每行ID严格对应:

import pandas as pd

# 假设df1和df2的ID列名为'id'
common_ids = df1['id'].intersection(df2['id'])

# 筛选共同ID数据,以ID为索引并排序确保对齐
df1_aligned = df1[df1['id'].isin(common_ids)].set_index('id').sort_index()
df2_aligned = df2[df2['id'].isin(common_ids)].set_index('id').sort_index()

2. 逐列计算准确率

根据你给出的示例(ID i1的a列,df1值[1,2]、df2值[1],准确率50%),定义匹配率为df2与df1的交集元素数 / df1的元素总数,编写通用计算函数并逐列统计平均准确率:

def compute_match_rate(val1, val2):
    # 处理空值:空值匹配空值算100%,否则0
    if pd.isna(val1) or pd.isna(val2):
        return 1.0 if pd.isna(val1) and pd.isna(val2) else 0.0
    # 统一转为集合处理单值/多值情况
    set1 = set(val1) if isinstance(val1, (list, tuple)) else {val1}
    set2 = set(val2) if isinstance(val2, (list, tuple)) else {val2}
    # 避免除以0
    total = len(set1)
    return len(set1 & set2) / total if total != 0 else 0.0

# 计算每列的平均准确率
accuracy = {}
for col in ['a', 'b', 'c', 'd']:
    col_match_rates = df1_aligned[col].combine(df2_aligned[col], compute_match_rate)
    accuracy[col] = col_match_rates.mean()

# 输出结果
print(pd.Series(accuracy))

3. 可视化匹配情况

热力图:展示每个ID、每列的匹配率

import seaborn as sns
import matplotlib.pyplot as plt

# 构建匹配率矩阵
match_matrix = pd.DataFrame()
for col in ['a', 'b', 'c', 'd']:
    match_matrix[col] = df1_aligned[col].combine(df2_aligned[col], compute_match_rate)

# 绘制热力图
plt.figure(figsize=(10, 6))
sns.heatmap(match_matrix, annot=True, cmap='RdYlGn', vmin=0, vmax=1)
plt.title('Match Rate by Common ID & Column')
plt.xlabel('Columns')
plt.ylabel('IDs')
plt.show()

条形图:展示每列的整体准确率

plt.figure(figsize=(8, 4))
pd.Series(accuracy).plot(kind='bar', color='deepskyblue')
plt.xticks(rotation=0)
plt.title('Column-wise Average Accuracy (Common IDs)')
plt.ylabel('Accuracy')
plt.xlabel('Columns')
# 添加百分比标签
for i, acc in enumerate(accuracy.values()):
    plt.text(i, acc + 0.02, f'{acc:.1%}', ha='center')
plt.show()

关键优势

  • 以ID交集为基准,严格过滤非共同ID,确保计算范围准确
  • 通用函数兼容单值/多值、空值场景,覆盖所有边缘情况
  • 可视化直接展示个体匹配细节与整体统计结果

内容的提问来源于stack exchange,提问作者Mejdi Dallel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 09:15:16