Python实现拼写错误语料库文本文件每行内容的统一格式转换
我自建了一份corpus拼写错误词库。
misspellings_corpus.txt 内容示例:
English, enlist->Enlish Hallowe'en, Halloween->Hallowean
当前文件格式是统一的,但不符合使用需求,需要进行转换:
现有格式:
correct, wrong1, wrong2->wrong3
期望转换后的格式:
wrong1,wrong2,wrong3->correct
wrong<N>的顺序无需调整- 每行可包含任意数量用逗号
,分隔的错误拼写 - 每行仅有1个正确拼写
correct,需放在->的右侧
最初的失败尝试代码:
with open('misspellings_corpus.txt') as oldfile, open('new.txt', 'w') as newfile: for line in oldfile: correct = line.split(', ')[0].strip() print(correct) W = line.split(', ')[1].strip() print(W) wrong_1 = W.split('->')[0] # 但这里可能存在多个错误拼写的情况 wrong_2 = W.split('->')[1] newfile.write(wrong_1 + ', ' + wrong_2 + '->' + correct)
运行后输出的 new.txt 不符合预期:
enlist, Enlish->EnglishHalloween, Hallowean->Hallowe'en
最终解决方案(灵感来自@alexis):
import re with open('misspellings_corpus.txt') as oldfile, open('new.txt', 'w') as newfile: for line in oldfile: # 输入行格式示例:'correct, wrong1, wrong2->wrong3' line = line.strip() terms = re.split(r", *|->", line) newfile.write(",".join(terms[1:]) + "->" + terms[0] + '\n')
运行后输出的正确 new.txt 内容:
enlist,Enlish->English Halloween,Hallowean->Hallowe'en
内容的提问来源于stack exchange,提问作者user12264468
相关产品推荐
相关产品推荐

