对datasets库ArrowDataset字段做正则替换后结果未生效问题咨询
问题原因
你使用的datasets.Dataset是Hugging Face Datasets库基于Arrow格式实现的数据集对象,不支持直接通过dataset[列名][索引] = 新值的方式修改数据:你通过这种方式拿到的只是原数据的临时副本,对副本的赋值不会同步到底层的Arrow存储中,所以看起来修改没有生效。
解决方案
方法1:使用官方推荐的map方法(最稳妥)
map方法会遍历数据集的每一条样本,处理后返回新的数据集,是处理Hugging Face Dataset的标准方式:
import re def replace_keywords(example): # 处理hypothesis字段 hypo = example['hypothesis'] # 合并正则规则简化写法 hypo = re.sub(r'\b(she|he|her|him)\b', 'them', hypo) hypo = re.sub(r'\b(dog|cat)\b', 'animal', hypo) # 返回修改后的样本,其余字段会保留原值 example['hypothesis'] = hypo return example # 处理后赋值覆盖原数据集变量即可 test_small_testval = test_small_testval.map(replace_keywords)
方法2:提取整列修改后替换回去
如果更习惯操作Python列表,可以先把整列取出处理完,再替换回数据集:
import re # 提取hypothesis列的所有值为Python列表 hypo_list = test_small_testval['hypothesis'] processed_hypo = [] for s in hypo_list: s = re.sub(r'\b(she|he|her|him)\b', 'them', s) s = re.sub(r'\b(dog|cat)\b', 'animal', s) processed_hypo.append(s) # 删除原hypothesis列,添加处理后的新列 test_small_testval = test_small_testval.remove_columns(['hypothesis']) test_small_testval = test_small_testval.add_column('hypothesis', processed_hypo)
处理完成后可执行如下代码验证效果:
print(test_small_testval['hypothesis'][0:2]) # 输出示例:['The owner threw the toy', 'The animal walked across the fallen log.']
内容的提问来源于stack exchange,提问作者Adam Rainah
相关产品推荐
相关产品推荐

