You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python Record Linkage模糊匹配CSV时遇KeyError: label not found问题求助

解决recordlinkage库的KeyError: 'label is not found in the dataframe'问题

问题场景

使用Python的recordlinkage库,通过公司名称和州对两个CSV文件进行模糊匹配合并时,触发KeyError错误,提示'label is not found in the dataframe'。相关代码及错误信息如下:

原代码

import pandas as pd 
import recordlinkage

reference_usa = pd.read_csv('all_reference_usa.csv', index_col='id')
oc_sample = pd.read_csv('oc_sample.csv', index_col='company_number', low_memory=False)

indexer = recordlinkage.Index()
indexer.sortedneighbourhood(left_on='state', right_on='state')
candidates = indexer.index(reference_usa, oc_sample)
print(len(candidates))

compare = recordlinkage.Compare()
compare.string('companyname',
            'name',
            threshold=0.95)
features = compare.compute(candidates, reference_usa,
                        oc_sample)

错误信息

File "/Users/Desktop/python/example.py", line 16, in <module>
    features = compare.compute(candidates, reference_usa,
  File "/Users//anaconda3/lib/python3.10/site-packages/recordlinkage/base.py", line 862, in compute
    results = self._compute(pairs, x, x_link)
  File "/Users//anaconda3/lib/python3.10/site-packages/recordlinkage/base.py", line 686, in _compute
    sublabels_left = self._get_labels_left(validate=x)
  File "/Users/anaconda3/lib/python3.10/site-packages/recordlinkage/base.py", line 652, in _get_labels_left
    raise KeyError(error_msg)
KeyError: 'label is not found in the dataframe'

错误原因

recordlinkage.Compare()默认会尝试在数据中查找名为label的列(用于监督学习的匹配标签),但当前场景是无监督模糊匹配,数据中并不存在该列,因此触发KeyError。

修复方案

初始化Compare对象时设置labels=False,明确告知库不需要标签列;同时建议给匹配规则指定自定义列名,方便后续识别结果。

修改后的代码

import pandas as pd 
import recordlinkage

reference_usa = pd.read_csv('all_reference_usa.csv', index_col='id')
oc_sample = pd.read_csv('oc_sample.csv', index_col='company_number', low_memory=False)

indexer = recordlinkage.Index()
indexer.sortedneighbourhood(left_on='state', right_on='state')
candidates = indexer.index(reference_usa, oc_sample)
print(len(candidates))

# 禁用标签列检测
compare = recordlinkage.Compare(labels=False)
# 自定义匹配结果的列名,增强可读性
compare.string('companyname',
            'name',
            threshold=0.95,
            label='company_name_match')
features = compare.compute(candidates, reference_usa, oc_sample)

# 查看匹配结果示例
print(features.head())

修改说明

  1. recordlinkage.Compare(labels=False):关闭自动检测标签列的逻辑,避免查找不存在的label列
  2. 添加label='company_name_match':给公司名称的匹配结果指定列名,让输出的特征DataFrame结构更清晰

内容的提问来源于stack exchange,提问作者newt_coding

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 08:04:55