遍历列表匹配名称失败排查:为何值无法匹配?
问题分析与解决方案
第一个报错:Lengths must match to compare
失误点
你的clients是嵌套列表(每个元素是单元素列表,如['Example Name']),但names['name']是pandas的字符串列。执行names['name'] == client时,是将长度为3的字符串列和长度为1的列表做比较,维度不匹配导致报错。
修正代码
把嵌套的clients转成普通字符串列表,或者循环时取列表内的字符串:
import pandas as pd import re ## example data clients = ['Example Name','Example Name2'] # 改成普通列表 name_list = [['Example Name'],['Example Name1'],['Example Name2']] names = pd.DataFrame(data=name_list,columns=['name']) ## code matches = [] for client in clients: # 先判断是否存在匹配项,避免iloc[0]索引报错 match_row = names[names['name'] == client] if not match_row.empty: search = match_row['name'].iloc[0] if re.search(client, search, flags=re.IGNORECASE): matches.append(client) print(matches)
如果必须保留clients的嵌套结构,循环时取client[0]即可:
for client in clients: client_str = client[0] match_row = names[names['name'] == client_str] # 后续逻辑同上
第二个问题:in操作符匹配返回False
失误点
你构建的search_list是嵌套列表(每个元素是[plaintiff]),但client_是字符串类型。用字符串判断是否存在于嵌套列表中,自然返回False——因为列表里的元素是子列表,不是字符串。
修正方案
直接把search_list改成普通字符串列表,不要嵌套:
# 构建普通字符串列表 search_list = [plaintiff.upper() for plaintiff in search_df['plaintiffs']] for client in client_list: client_ = str(client).upper().replace(',', '') # 统一格式,去掉逗号 print(client_) print(client_ in search_list)
更优方案(用pandas原生方法)
不需要手动转列表,直接用pandas的isin或str.contains实现批量匹配,效率更高:
# 方案1:直接匹配完全一致的名称 client_list = ['SSN ABSECON, LLC'] # 统一格式:去掉逗号、转大写 client_clean = [c.upper().replace(',', '') for c in client_list] search_df['plaintiffs_clean'] = search_df['plaintiffs'].str.upper().str.replace(',', '') # 筛选匹配的行 matches = search_df[search_df['plaintiffs_clean'].isin(client_clean)] print(matches) # 方案2:模糊匹配(如果需要部分匹配) matches = search_df[search_df['plaintiffs_clean'].str.contains('|'.join(client_clean), regex=True)]
内容的提问来源于stack exchange,提问作者Jeff Gordon
相关产品推荐
相关产品推荐

