使用Python Record Linkage模糊匹配CSV时遇KeyError: label not found问题求助
解决recordlinkage库的KeyError: 'label is not found in the dataframe'问题
问题场景
使用Python的recordlinkage库,通过公司名称和州对两个CSV文件进行模糊匹配合并时,触发KeyError错误,提示'label is not found in the dataframe'。相关代码及错误信息如下:
原代码
import pandas as pd import recordlinkage reference_usa = pd.read_csv('all_reference_usa.csv', index_col='id') oc_sample = pd.read_csv('oc_sample.csv', index_col='company_number', low_memory=False) indexer = recordlinkage.Index() indexer.sortedneighbourhood(left_on='state', right_on='state') candidates = indexer.index(reference_usa, oc_sample) print(len(candidates)) compare = recordlinkage.Compare() compare.string('companyname', 'name', threshold=0.95) features = compare.compute(candidates, reference_usa, oc_sample)
错误信息
File "/Users/Desktop/python/example.py", line 16, in <module> features = compare.compute(candidates, reference_usa, File "/Users//anaconda3/lib/python3.10/site-packages/recordlinkage/base.py", line 862, in compute results = self._compute(pairs, x, x_link) File "/Users//anaconda3/lib/python3.10/site-packages/recordlinkage/base.py", line 686, in _compute sublabels_left = self._get_labels_left(validate=x) File "/Users/anaconda3/lib/python3.10/site-packages/recordlinkage/base.py", line 652, in _get_labels_left raise KeyError(error_msg) KeyError: 'label is not found in the dataframe'
错误原因
recordlinkage.Compare()默认会尝试在数据中查找名为label的列(用于监督学习的匹配标签),但当前场景是无监督模糊匹配,数据中并不存在该列,因此触发KeyError。
修复方案
初始化Compare对象时设置labels=False,明确告知库不需要标签列;同时建议给匹配规则指定自定义列名,方便后续识别结果。
修改后的代码
import pandas as pd import recordlinkage reference_usa = pd.read_csv('all_reference_usa.csv', index_col='id') oc_sample = pd.read_csv('oc_sample.csv', index_col='company_number', low_memory=False) indexer = recordlinkage.Index() indexer.sortedneighbourhood(left_on='state', right_on='state') candidates = indexer.index(reference_usa, oc_sample) print(len(candidates)) # 禁用标签列检测 compare = recordlinkage.Compare(labels=False) # 自定义匹配结果的列名,增强可读性 compare.string('companyname', 'name', threshold=0.95, label='company_name_match') features = compare.compute(candidates, reference_usa, oc_sample) # 查看匹配结果示例 print(features.head())
修改说明
recordlinkage.Compare(labels=False):关闭自动检测标签列的逻辑,避免查找不存在的label列- 添加
label='company_name_match':给公司名称的匹配结果指定列名,让输出的特征DataFrame结构更清晰
内容的提问来源于stack exchange,提问作者newt_coding
相关产品推荐
相关产品推荐

