You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python匹配数据集hypothesis列关键词返回空列表的原因排查

问题原因及解决方案

核心错误点

  • 字符串匹配逻辑错误
    你代码中if ('|'.join(searchfor)) in crop_dataset['hypothesis'][i]的判断逻辑完全错误:'|'.join(searchfor)会拼接出完整字符串she|he|his|her|him|boys|woman|Woman|girl|men|man|female|girls|person|horse,字符串in判断是检查目标文本中是否完整包含这一整串带|的字符,你的样本里显然没有这类内容,所以全部返回False。这里你混淆了字符串子串匹配和正则匹配的语法,|的或逻辑只在正则表达式中生效,原生字符串判断不会识别。
  • Pandas方法调用逻辑错误
    • pandas的filter方法是用来筛选列名、行索引的,不支持逐行判断内容筛选数据行,你应该用apply做逐行判断再通过布尔索引过滤
    • searchfor in example['hypothesis']的逻辑错误:不能判断一个列表是否是字符串的子串,只能判断列表中的单个元素是否存在于字符串中。

修正后的代码实现

原循环逻辑修正

print(crop_dataset.shape)
searchfor = ['she', 'he','his','her','him','boys','woman','Woman','girl','men','man','female','girls','person','horse']
indices_list = []
for i in range(5):
    hypo = crop_dataset['hypothesis'][i]
    print(hypo)
    # 改为判断任意关键词存在于文本中
    if any(keyword in hypo for keyword in searchfor):
        indices_list.append(i)
print(indices_list)    
rows_where_found = crop_dataset.select(indices_list)
print(rows_where_found.shape)

Pandas筛选修正

df_temp = crop_dataset.to_pandas()
# 逐行判断是否包含任意关键词,生成布尔索引
mask = df_temp['hypothesis'].apply(lambda x: any(keyword in x for keyword in searchfor))
rows_where_found = df_temp[mask]

HuggingFace Datasets直接筛选(更高效,无需转格式)

rows_where_found = crop_dataset.filter(lambda x: any(key in x['hypothesis'] for key in searchfor))

如果需要实现大小写不敏感的匹配,可以统一转小写后再判断:

any(key.lower() in x['hypothesis'].lower() for key in searchfor)

内容的提问来源于stack exchange,提问作者bit_by_bit

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.23 19:54:01