You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

两数据集列模糊匹配代码报错:'str' object has no attribute 'shape' 求助

问题排查与解决:模糊匹配时出现'str' object has no attribute 'shape'错误

错误原因

报错核心是**recordlinkage.preprocessing.clean函数的使用方式错误**:这个函数专门用于处理pandas Series(整列数据),会自动完成文本清洗(去空格、小写化等),但你在循环中把单个字符串传入该函数,而单个字符串没有shape属性,直接触发了函数内部的属性访问错误。

此外代码还有两处冗余问题:

  • 重复调用dropna(),一次调用即可完成空值移除
  • 套了ThreadPoolExecutor壳但未实现实际并行逻辑,属于无效代码

修复步骤

  1. 提前整列清洗:在循环匹配前,用clean处理整个列,避免在循环中处理单个字符串
  2. 完善空值过滤:除了dropna(),额外过滤掉astype(str)产生的空字符串(如''或'NaN')
  3. 移除无效线程池:先保留串行逻辑修复核心错误,如需并行可后续单独实现
  4. 简化数据预处理:合并重复的空值与去重操作

修复后的完整代码

import pandas as pd
from fuzzywuzzy import fuzz
from recordlinkage.preprocessing import clean

column1_name = 'Sold To Customer Name'
column2_name = 'customer_name'

try:
    # 读取数据集
    dataset1 = pd.read_csv(r"C:\Users\JE\Downloads\edw_query_extract_distinct.csv", 
                          low_memory=False, on_bad_lines='skip', 
                          index_col=False, dtype='unicode')
    dataset2 = pd.read_csv(r"C:\Users\JE\Downloads\ISO_Code_joined.csv", 
                          low_memory=False, on_bad_lines='skip', 
                          index_col=False, dtype='unicode')

    # 选择目标列并转换为字符串类型
    column1 = dataset1[column1_name].astype(str)
    column2 = dataset2[column2_name].astype(str)

    # 模糊匹配阈值
    threshold = 80

    # 预处理:去空值、去重、文本清洗
    column1 = clean(column1.dropna().drop_duplicates())
    column1 = column1[(column1 != '') & (column1 != 'NaN')].reset_index(drop=True)
    
    column2 = clean(column2.dropna().drop_duplicates())
    column2 = column2[(column2 != '') & (column2 != 'NaN')].reset_index(drop=True)

    matches = []
    # 转换为带原索引的列表,方便匹配后定位原始数据
    column2_with_index = column2.reset_index().values.tolist()

    # 遍历第一个列的所有值,寻找最佳匹配
    for idx1, value1 in column1.items():
        best_match_idx = None
        best_score = threshold

        for idx2, value2 in column2_with_index:
            score = fuzz.token_set_ratio(value1, value2)
            if score > best_score:
                best_match_idx = idx2
                best_score = score

        if best_match_idx is not None:
            matches.append((idx1, best_match_idx, best_score))

    # 生成最终匹配结果数据集
    matched_data = pd.DataFrame(matches, columns=['Index_1', 'Index_2', 'Similarity_Score'])
    matched_data['Value_1'] = column1.loc[matched_data['Index_1']].values
    matched_data['Value_2'] = column2.loc[matched_data['Index_2']].values

    print(matched_data)

except FileNotFoundError:
    print("错误:未找到一个或两个数据集文件。")
except KeyError:
    print("错误:一个或两个列名在数据集中不存在。")
except Exception as e:
    print(f"发生错误:{str(e)}")

额外并行优化建议

如果数据集规模较大,串行匹配效率低,可以用以下方式实现真正的并行处理:

# 定义单个值的匹配函数
def find_best_match(value1):
    best_match_idx = None
    best_score = threshold
    for idx2, value2 in column2_with_index:
        score = fuzz.token_set_ratio(value1, value2)
        if score > best_score:
            best_match_idx = idx2
            best_score = score
    return best_match_idx, best_score

# 使用线程池并行处理所有匹配任务
with ThreadPoolExecutor() as executor:
    results = executor.map(find_best_match, column1.values)

# 整理并行处理的结果
matches = []
for idx1, (best_match_idx, best_score) in enumerate(results):
    if best_match_idx is not None:
        matches.append((column1.index[idx1], best_match_idx, best_score))

内容的提问来源于stack exchange,提问作者med jalel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 00:15:00