Python列表去除首尾停用词的while循环方案优化及替代方案咨询
移除文本首尾停用词的优化实现
这里提供两种成熟的优化方案,均能满足「仅移除首尾停用词、保留中间停用词」的需求:
方案1:正则表达式实现
直接对原始文本做匹配替换,无需提前分词遍历,代码更简洁:
from nltk.corpus import stopwords import re cached_stop_words = set(stopwords.words("english")) # 拼接正则匹配模式,匹配开头/结尾的停用词+空格组合 stop_pat = '|'.join(re.escape(word) for word in cached_stop_words) pat = re.compile(fr'^(?:{stop_pat}\s+)*|\s+(?:{stop_pat})*$', re.IGNORECASE) text = "a the sky is the blue the a a" processed_text = pat.sub('', text) new_tokens = processed_text.split() if processed_text else [] print(new_tokens)
运行输出和原方案一致:['sky', 'is', 'the', 'blue']
方案2:集合优化版循环实现
原方案性能低的核心原因是stopwords.words("english")返回的是列表,in查询时间复杂度为O(n),转成集合后查询复杂度降为O(1),大批量处理时性能提升非常明显:
from nltk.corpus import stopwords # 停用词转集合,大幅提升查询效率 cached_stop_words = set(stopwords.words("english")) text = "a the sky is the blue the a a" tokens = text.split() start_idx, end_idx = 0, len(tokens) - 1 while start_idx < len(tokens) and tokens[start_idx] in cached_stop_words: start_idx += 1 while end_idx >= 0 and tokens[end_idx] in cached_stop_words: end_idx -= 1 new_tokens = tokens[start_idx: end_idx + 1] if start_idx <= end_idx else [] print(new_tokens)
两种方案的适用场景:
- 正则方案适合短文本、轻量处理场景,代码精简易维护
- 集合优化循环方案性能更稳定,适合长文本、大批量数据处理场景,无正则回溯带来的性能波动风险
内容的提问来源于stack exchange,提问作者Thang Pham
相关产品推荐
相关产品推荐

