Python不使用任何库如何实现文本中非合法词汇的移除?
无第三方库实现英文合法词汇过滤
合法词汇判定规则
- 仅可由英文字母、撇号
'、连字符-三类字符组成,不可包含数字、标点、其他语言字符等非法内容 - 单字符词汇仅
a、I(不区分大小写)属于合法范畴
原有代码问题
原来的实现逻辑存在多个漏洞:
- 仅通过
isalpha()判断会直接过滤掉带撇号、连字符的合法词汇,比如can't、full-time - 仅裁剪末尾1个字符的逻辑无法处理前后有多个非法字符、非法字符在词中间的场景
- 没有做单字符合法性校验,会把其他单字符非法内容保留
- 没有过滤包含非英文字符、数字的无效词汇
调整后代码
def filter_legal_words(text): allowed_chars = set("abcdefghijklmnopqrstuvwxyzABCDEFGHIJKLMNOPQRSTUVWXYZ'-") words = [] for candidate in text.split(): # 先裁剪掉候选词两端的所有非法字符 start = 0 while start < len(candidate) and candidate[start] not in allowed_chars: start += 1 end = len(candidate) - 1 while end >= start and candidate[end] not in allowed_chars: end -= 1 cleaned = candidate[start:end+1] if not cleaned: continue # 检查剩余所有字符是否都合法 all_legal = True for c in cleaned: if c not in allowed_chars: all_legal = False break if not all_legal: continue # 单字符合法性校验 if len(cleaned) == 1: if cleaned.lower() not in ('a', 'i'): continue # 符合要求转小写加入结果 words.append(cleaned.lower()) return words
使用说明
直接传入待处理的文本字符串即可,返回的列表就是全部合法的小写词汇,完全不依赖Python标准库之外的任何第三方库。
内容的提问来源于stack exchange,提问作者Rayan
相关产品推荐
相关产品推荐

