两数据集列模糊匹配代码报错:'str' object has no attribute 'shape' 求助
问题排查与解决:模糊匹配时出现
'str' object has no attribute 'shape'错误 错误原因
报错核心是**recordlinkage.preprocessing.clean函数的使用方式错误**:这个函数专门用于处理pandas Series(整列数据),会自动完成文本清洗(去空格、小写化等),但你在循环中把单个字符串传入该函数,而单个字符串没有shape属性,直接触发了函数内部的属性访问错误。
此外代码还有两处冗余问题:
- 重复调用
dropna(),一次调用即可完成空值移除 - 套了
ThreadPoolExecutor壳但未实现实际并行逻辑,属于无效代码
修复步骤
- 提前整列清洗:在循环匹配前,用
clean处理整个列,避免在循环中处理单个字符串 - 完善空值过滤:除了
dropna(),额外过滤掉astype(str)产生的空字符串(如''或'NaN') - 移除无效线程池:先保留串行逻辑修复核心错误,如需并行可后续单独实现
- 简化数据预处理:合并重复的空值与去重操作
修复后的完整代码
import pandas as pd from fuzzywuzzy import fuzz from recordlinkage.preprocessing import clean column1_name = 'Sold To Customer Name' column2_name = 'customer_name' try: # 读取数据集 dataset1 = pd.read_csv(r"C:\Users\JE\Downloads\edw_query_extract_distinct.csv", low_memory=False, on_bad_lines='skip', index_col=False, dtype='unicode') dataset2 = pd.read_csv(r"C:\Users\JE\Downloads\ISO_Code_joined.csv", low_memory=False, on_bad_lines='skip', index_col=False, dtype='unicode') # 选择目标列并转换为字符串类型 column1 = dataset1[column1_name].astype(str) column2 = dataset2[column2_name].astype(str) # 模糊匹配阈值 threshold = 80 # 预处理:去空值、去重、文本清洗 column1 = clean(column1.dropna().drop_duplicates()) column1 = column1[(column1 != '') & (column1 != 'NaN')].reset_index(drop=True) column2 = clean(column2.dropna().drop_duplicates()) column2 = column2[(column2 != '') & (column2 != 'NaN')].reset_index(drop=True) matches = [] # 转换为带原索引的列表,方便匹配后定位原始数据 column2_with_index = column2.reset_index().values.tolist() # 遍历第一个列的所有值,寻找最佳匹配 for idx1, value1 in column1.items(): best_match_idx = None best_score = threshold for idx2, value2 in column2_with_index: score = fuzz.token_set_ratio(value1, value2) if score > best_score: best_match_idx = idx2 best_score = score if best_match_idx is not None: matches.append((idx1, best_match_idx, best_score)) # 生成最终匹配结果数据集 matched_data = pd.DataFrame(matches, columns=['Index_1', 'Index_2', 'Similarity_Score']) matched_data['Value_1'] = column1.loc[matched_data['Index_1']].values matched_data['Value_2'] = column2.loc[matched_data['Index_2']].values print(matched_data) except FileNotFoundError: print("错误:未找到一个或两个数据集文件。") except KeyError: print("错误:一个或两个列名在数据集中不存在。") except Exception as e: print(f"发生错误:{str(e)}")
额外并行优化建议
如果数据集规模较大,串行匹配效率低,可以用以下方式实现真正的并行处理:
# 定义单个值的匹配函数 def find_best_match(value1): best_match_idx = None best_score = threshold for idx2, value2 in column2_with_index: score = fuzz.token_set_ratio(value1, value2) if score > best_score: best_match_idx = idx2 best_score = score return best_match_idx, best_score # 使用线程池并行处理所有匹配任务 with ThreadPoolExecutor() as executor: results = executor.map(find_best_match, column1.values) # 整理并行处理的结果 matches = [] for idx1, (best_match_idx, best_score) in enumerate(results): if best_match_idx is not None: matches.append((column1.index[idx1], best_match_idx, best_score))
内容的提问来源于stack exchange,提问作者med jalel
相关产品推荐
相关产品推荐

