You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

数据集hypothesis列字符串替换不生效及全词替换方案咨询

数据集修改不生效的原因

你使用的是Hugging Face Datasets库的Dataset对象,该对象默认不支持按索引直接修改列值:你通过test_small_testval['hypothesis'][i]取到的是原始数据的临时拷贝,修改这个拷贝不会同步更新到原数据集,所以看不到变化。

官方推荐使用map方法批量处理数据集,示例逻辑如下:

def replace_text(example):
    # 处理替换逻辑
    text = example['hypothesis']
    text = text.replace('she','them')
    text = text.replace('he','them')
    # 其余替换规则按你的需求补充
    example['hypothesis'] = text
    return example

# 执行批量处理,会返回处理后的新数据集
test_small_testval = test_small_testval.map(replace_text)
全词匹配替换实现方案

你之前的正则方案不生效有两个原因:

  1. 没有开启大小写匹配:比如示例中的Woman、Girl首字母大写,你的正则只匹配全小写的词,无法命中
  2. 依旧使用了直接索引赋值的方式,修改不会同步到原数据集

可以使用如下方案实现全词匹配、忽略大小写、批量替换,且避免子串误替换:

import re

# 定义替换规则字典,key是要替换的词,value是替换后的内容
replace_map = {
    'she': 'them',
    'he': 'them',
    'her': 'them',
    'him': 'them',
    'cat': 'animal',
    'dog': 'animal',
    'woman': 'them',
    'girl': 'them',
    'guitar': 'instrument',
    'field': 'outdoors'
}

def replace_whole_word(example):
    text = example['hypothesis']
    for old_word, new_word in replace_map.items():
        # \b 匹配单词边界,re.IGNORECASE 忽略大小写
        text = re.sub(rf'\b{re.escape(old_word)}\b', new_word, text, flags=re.IGNORECASE)
    example['hypothesis'] = text
    return example

# 执行处理得到新数据集
test_small_testval = test_small_testval.map(replace_whole_word)

如果需要保留原词的大小写格式(比如首字母大写的词替换后也首字母大写),可以给re.sub传入自定义的替换函数调整逻辑。

内容的提问来源于stack exchange,提问作者Adam Rainah

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.23 16:36:01