You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用pyspellchecker修正pandas列中的拼写错误?

问题解决:用pyspellchecker修正DataFrame中文本的拼写错误

问题场景

现有如下DataFrame:

import pandas as pd
df = pd.DataFrame({'id':[1,2,3],'text':['a foox juumped ovr the gate','teh car wsa bllue','why so srious']})

需要用pyspellchecker库生成包含修正后拼写的新列,但原代码无法修正任何拼写错误:

from spellchecker import SpellChecker

spell = SpellChecker()

def correct_spelling(word):
    corrected_word = spell.correction(word)
    return corrected_word if corrected_word is not None else word

df['corrected_text'] = df['text'].apply(correct_spelling)

错误原因

原代码直接将整段文本字符串传入修正函数,而spell.correction()仅针对单个单词做拼写修正,无法识别完整句子里的错误单词,因此没有效果。

修正方案

修改函数逻辑,先把句子拆分为单个单词,逐个修正后再拼接成完整句子:

import pandas as pd
from spellchecker import SpellChecker

spell = SpellChecker()

def correct_sentence(sentence):
    # 拆分句子为单词列表
    words = sentence.split()
    # 逐个修正单词,保留无法识别的原词
    corrected_words = [spell.correction(word) or word for word in words]
    # 拼接回完整句子
    return ' '.join(corrected_words)

df['corrected_text'] = df['text'].apply(correct_sentence)

验证结果

运行后得到的DataFrame与预期一致:

print(df)
# 输出:
#    id                          text                  corrected_text
# 0   1  a foox juumped ovr the gate  a fox jumped over the gate
# 1   2        teh car wsa bllue        the car was blue
# 2   3             why so srious             why so serious

内容的提问来源于stack exchange,提问作者Tobes

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 05:15:32