You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pyspellchecker识别拼写错误并生成DataFrame新列的问题

解决pyspellchecker字符级检测问题,实现单词级拼写错误识别

你的代码出现字符级检测结果的原因是:遍历df['ColumnA']时,每个元素是完整的句子字符串,直接对字符串执行map(spell.unknown, word)会把字符串拆分为单个字符逐一检查,而非按单词处理。

修正后的代码

!pip install pyspellchecker
from spellchecker import SpellChecker
import pandas as pd

spell = SpellChecker()

# 定义函数提取单句中的拼写错误单词
def extract_misspelled(sentence):
    # 将句子拆分为单词列表
    word_list = sentence.split()
    # 获取拼写错误的单词并转为列表
    return list(spell.unknown(word_list))

# 生成存储错误单词的新列
df['ColumnB'] = df['ColumnA'].apply(extract_misspelled)

# 查看结果
df['ColumnB']

进阶处理(清理标点干扰)

如果句子中存在附着在单词上的标点(如逗号、句号),可以先清理标点再拆分单词:

import re

def extract_misspelled(sentence):
    # 移除句子中的非单词/空格字符
    cleaned_sentence = re.sub(r'[^\w\s]', '', sentence)
    word_list = cleaned_sentence.split()
    return list(spell.unknown(word_list))

df['ColumnB'] = df['ColumnA'].apply(extract_misspelled)

内容的提问来源于stack exchange,提问作者Hugo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 19:23:22