You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python不使用任何库如何实现文本中非合法词汇的移除?

无第三方库实现英文合法词汇过滤

合法词汇判定规则

  • 仅可由英文字母、撇号'、连字符-三类字符组成,不可包含数字、标点、其他语言字符等非法内容
  • 单字符词汇仅a、I(不区分大小写)属于合法范畴

原有代码问题

原来的实现逻辑存在多个漏洞:

  1. 仅通过isalpha()判断会直接过滤掉带撇号、连字符的合法词汇,比如can't、full-time
  2. 仅裁剪末尾1个字符的逻辑无法处理前后有多个非法字符、非法字符在词中间的场景
  3. 没有做单字符合法性校验,会把其他单字符非法内容保留
  4. 没有过滤包含非英文字符、数字的无效词汇

调整后代码

def filter_legal_words(text):
    allowed_chars = set("abcdefghijklmnopqrstuvwxyzABCDEFGHIJKLMNOPQRSTUVWXYZ'-")
    words = []
    for candidate in text.split():
        # 先裁剪掉候选词两端的所有非法字符
        start = 0
        while start < len(candidate) and candidate[start] not in allowed_chars:
            start += 1
        end = len(candidate) - 1
        while end >= start and candidate[end] not in allowed_chars:
            end -= 1
        cleaned = candidate[start:end+1]
        if not cleaned:
            continue
        # 检查剩余所有字符是否都合法
        all_legal = True
        for c in cleaned:
            if c not in allowed_chars:
                all_legal = False
                break
        if not all_legal:
            continue
        # 单字符合法性校验
        if len(cleaned) == 1:
            if cleaned.lower() not in ('a', 'i'):
                continue
        # 符合要求转小写加入结果
        words.append(cleaned.lower())
    return words

使用说明

直接传入待处理的文本字符串即可,返回的列表就是全部合法的小写词汇,完全不依赖Python标准库之外的任何第三方库。

内容的提问来源于stack exchange,提问作者Rayan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 10:18:03