You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按正确词分组条件合并拼写错误语料文本文件的对应行

拼写错误语料合并实现方案

核心思路

用字典做聚合存储,以正确词为键,对应所有错误拼写的集合为值,天然去重,避免重复错误词:

  • 遍历每行拆分出错误片段和正确词
  • 把每行里的多个错误词按逗号拆分后去空格,加入对应正确词的集合
  • 最后按错误词1, 错误词2,...->正确词的格式输出即可

完整实现代码

from collections import defaultdict

# 读取原始文件
lines = []
with open('/content/drive/MyDrive/Colab Notebooks/misspellings_corpus.txt', 'r', encoding='utf-8') as fin:
    lines = [line.strip() for line in fin if line.strip()]  # 过滤空行和首尾空白

# 初始化聚合字典,value默认是set用来自动去重
correct_to_wrongs = defaultdict(set)

for line in lines:
    wrong_part, correct_word = line.split('->')
    # 拆分当前行的所有错误词,去除首尾空格
    wrong_words = [w.strip() for w in wrong_part.split(',') if w.strip()]
    # 加入对应正确词的集合
    correct_to_wrongs[correct_word.strip()].update(wrong_words)

# 生成结果行
result_lines = []
for correct, wrongs in correct_to_wrongs.items():
    result_line = f"{', '.join(wrongs)}->{correct}"
    result_lines.append(result_line)

# 可选:写入到新文件
with open('/content/drive/MyDrive/Colab Notebooks/merged_misspellings.txt', 'w', encoding='utf-8') as fout:
    fout.write('\n'.join(result_lines))

# 打印输出验证
for line in result_lines:
    print(line)

如果需要保留错误词首次出现的顺序,可以把defaultdict(set)换成defaultdict(list),加入的时候判断下是否已存在即可:

for w in wrong_words:
    if w not in correct_to_wrongs[correct_word.strip()]:
        correct_to_wrongs[correct_word.strip()].append(w)

效果验证

针对你给出的输入片段,运行后输出结果和预期完全一致:

I'de, I'ed, I'ld, Id->I'd
aple, appl->apple

内容的提问来源于stack exchange,提问作者StressedBoi69420

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 05:18:03