You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python实现拼写错误语料库文本文件每行内容的统一格式转换

我自建了一份corpus拼写错误词库。

misspellings_corpus.txt 内容示例:

English, enlist->Enlish
Hallowe'en, Halloween->Hallowean

当前文件格式是统一的,但不符合使用需求,需要进行转换:
现有格式:

correct, wrong1, wrong2->wrong3

期望转换后的格式:

wrong1,wrong2,wrong3->correct
  • wrong<N> 的顺序无需调整
  • 每行可包含任意数量用逗号 , 分隔的错误拼写
  • 每行仅有1个正确拼写 correct,需放在 -> 的右侧

最初的失败尝试代码:

with open('misspellings_corpus.txt') as oldfile, open('new.txt', 'w') as newfile:
    for line in oldfile:
      correct = line.split(', ')[0].strip()
      print(correct)
      W = line.split(', ')[1].strip()
      print(W)
      wrong_1 = W.split('->')[0] # 但这里可能存在多个错误拼写的情况
      wrong_2 = W.split('->')[1]
      newfile.write(wrong_1 + ', ' + wrong_2 + '->' + correct)

运行后输出的 new.txt 不符合预期:

enlist, Enlish->EnglishHalloween, Hallowean->Hallowe'en

最终解决方案(灵感来自@alexis):

import re

with open('misspellings_corpus.txt') as oldfile, open('new.txt', 'w') as newfile:
  for line in oldfile:
    # 输入行格式示例:'correct, wrong1, wrong2->wrong3'
    line = line.strip()
    terms = re.split(r", *|->", line)
    newfile.write(",".join(terms[1:]) + "->" + terms[0] + '\n')

运行后输出的正确 new.txt 内容:

enlist,Enlish->English
Halloween,Hallowean->Hallowe'en

内容的提问来源于stack exchange,提问作者user12264468

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 04:54:03