Python匹配数据集hypothesis列关键词返回空列表的原因排查
问题原因及解决方案
核心错误点
- 字符串匹配逻辑错误
你代码中if ('|'.join(searchfor)) in crop_dataset['hypothesis'][i]的判断逻辑完全错误:'|'.join(searchfor)会拼接出完整字符串she|he|his|her|him|boys|woman|Woman|girl|men|man|female|girls|person|horse,字符串in判断是检查目标文本中是否完整包含这一整串带|的字符,你的样本里显然没有这类内容,所以全部返回False。这里你混淆了字符串子串匹配和正则匹配的语法,|的或逻辑只在正则表达式中生效,原生字符串判断不会识别。 - Pandas方法调用逻辑错误
- pandas的
filter方法是用来筛选列名、行索引的,不支持逐行判断内容筛选数据行,你应该用apply做逐行判断再通过布尔索引过滤 searchfor in example['hypothesis']的逻辑错误:不能判断一个列表是否是字符串的子串,只能判断列表中的单个元素是否存在于字符串中。
- pandas的
修正后的代码实现
原循环逻辑修正
print(crop_dataset.shape) searchfor = ['she', 'he','his','her','him','boys','woman','Woman','girl','men','man','female','girls','person','horse'] indices_list = [] for i in range(5): hypo = crop_dataset['hypothesis'][i] print(hypo) # 改为判断任意关键词存在于文本中 if any(keyword in hypo for keyword in searchfor): indices_list.append(i) print(indices_list) rows_where_found = crop_dataset.select(indices_list) print(rows_where_found.shape)
Pandas筛选修正
df_temp = crop_dataset.to_pandas() # 逐行判断是否包含任意关键词,生成布尔索引 mask = df_temp['hypothesis'].apply(lambda x: any(keyword in x for keyword in searchfor)) rows_where_found = df_temp[mask]
HuggingFace Datasets直接筛选(更高效,无需转格式)
rows_where_found = crop_dataset.filter(lambda x: any(key in x['hypothesis'] for key in searchfor))
如果需要实现大小写不敏感的匹配,可以统一转小写后再判断:
any(key.lower() in x['hypothesis'].lower() for key in searchfor)
内容的提问来源于stack exchange,提问作者bit_by_bit
相关产品推荐
相关产品推荐

