如何使用pyspellchecker修正pandas列中的拼写错误?
问题解决:用pyspellchecker修正DataFrame中文本的拼写错误
问题场景
现有如下DataFrame:
import pandas as pd df = pd.DataFrame({'id':[1,2,3],'text':['a foox juumped ovr the gate','teh car wsa bllue','why so srious']})
需要用pyspellchecker库生成包含修正后拼写的新列,但原代码无法修正任何拼写错误:
from spellchecker import SpellChecker spell = SpellChecker() def correct_spelling(word): corrected_word = spell.correction(word) return corrected_word if corrected_word is not None else word df['corrected_text'] = df['text'].apply(correct_spelling)
错误原因
原代码直接将整段文本字符串传入修正函数,而spell.correction()仅针对单个单词做拼写修正,无法识别完整句子里的错误单词,因此没有效果。
修正方案
修改函数逻辑,先把句子拆分为单个单词,逐个修正后再拼接成完整句子:
import pandas as pd from spellchecker import SpellChecker spell = SpellChecker() def correct_sentence(sentence): # 拆分句子为单词列表 words = sentence.split() # 逐个修正单词,保留无法识别的原词 corrected_words = [spell.correction(word) or word for word in words] # 拼接回完整句子 return ' '.join(corrected_words) df['corrected_text'] = df['text'].apply(correct_sentence)
验证结果
运行后得到的DataFrame与预期一致:
print(df) # 输出: # id text corrected_text # 0 1 a foox juumped ovr the gate a fox jumped over the gate # 1 2 teh car wsa bllue the car was blue # 2 3 why so srious why so serious
内容的提问来源于stack exchange,提问作者Tobes
相关产品推荐
相关产品推荐

