数据集hypothesis列字符串替换不生效及全词替换方案咨询
数据集修改不生效的原因
你使用的是Hugging Face Datasets库的Dataset对象,该对象默认不支持按索引直接修改列值:你通过test_small_testval['hypothesis'][i]取到的是原始数据的临时拷贝,修改这个拷贝不会同步更新到原数据集,所以看不到变化。
官方推荐使用map方法批量处理数据集,示例逻辑如下:
def replace_text(example): # 处理替换逻辑 text = example['hypothesis'] text = text.replace('she','them') text = text.replace('he','them') # 其余替换规则按你的需求补充 example['hypothesis'] = text return example # 执行批量处理,会返回处理后的新数据集 test_small_testval = test_small_testval.map(replace_text)
全词匹配替换实现方案
你之前的正则方案不生效有两个原因:
- 没有开启大小写匹配:比如示例中的
Woman、Girl首字母大写,你的正则只匹配全小写的词,无法命中 - 依旧使用了直接索引赋值的方式,修改不会同步到原数据集
可以使用如下方案实现全词匹配、忽略大小写、批量替换,且避免子串误替换:
import re # 定义替换规则字典,key是要替换的词,value是替换后的内容 replace_map = { 'she': 'them', 'he': 'them', 'her': 'them', 'him': 'them', 'cat': 'animal', 'dog': 'animal', 'woman': 'them', 'girl': 'them', 'guitar': 'instrument', 'field': 'outdoors' } def replace_whole_word(example): text = example['hypothesis'] for old_word, new_word in replace_map.items(): # \b 匹配单词边界,re.IGNORECASE 忽略大小写 text = re.sub(rf'\b{re.escape(old_word)}\b', new_word, text, flags=re.IGNORECASE) example['hypothesis'] = text return example # 执行处理得到新数据集 test_small_testval = test_small_testval.map(replace_whole_word)
如果需要保留原词的大小写格式(比如首字母大写的词替换后也首字母大写),可以给re.sub传入自定义的替换函数调整逻辑。
内容的提问来源于stack exchange,提问作者Adam Rainah
相关产品推荐
相关产品推荐

