You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python识别文本数据集中的拼写错误(仅检测不纠错)

用Python分离文本数据集中的拼写错误条目

要实现这个需求,核心是识别每个条目里的拼写错误单词,再筛选出包含错误单词的条目。可以用pyenchant库(轻量英文拼写检查工具)完成,步骤如下:

1. 安装依赖

执行命令安装拼写检查库:

pip install pyenchant

2. 代码实现

import enchant

# 加载英文(美式)词典
spell_checker = enchant.Dict("en_US")

# 样本数据集
dataset = [
    "Kurtas for women",
    "parti wear dresses",
    "denim jeans",
    "overcot"
]

# 判断条目是否包含拼写错误单词
def has_spelling_error(text):
    words = text.lower().split()
    for word in words:
        if not spell_checker.check(word):
            return True
    return False

# 筛选有拼写错误的条目
error_entries = [entry for entry in dataset if has_spelling_error(entry)]

# 格式化输出结果
for idx, entry in enumerate(error_entries, 1):
    print(f"{idx}. {entry}")

运行结果

1. parti wear dresses
2. overcot

说明

  • 若需英式英语拼写检查,可将词典替换为en_GB;
  • 转换为小写是为了避免大小写干扰检查(比如"Kurtas"这类正确单词不会被误判);
  • 只要条目包含至少一个拼写错误单词,就会被筛选出来,完全匹配需求。

内容的提问来源于stack exchange,提问作者riya23

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 19:44:56