Python停用词移除函数问题:句首停用词未被正确移除求助
修复停用词移除函数的句首匹配问题
问题
我写了一个Python函数remove_stopwords用来移除句子中的停用词,测试句子为"I am about to go to the store and get any snack",期望输出是"go store get snack",但实际输出是"i go store get snack"——句首的停用词"i"没被移除。
原代码的问题
原代码试图通过两种token匹配(" 词 "和"词 ")处理不同位置的停用词,但存在明显错误:
token1定义有语法错误:token1= + word + " "多了多余的加号,无法正确生成句首匹配的token- 嵌套if逻辑错误:只有当句中token(" 词 ")存在时才会检查句首token,句首的"i "根本触发不了外层判断,自然不会被处理
- 内层if里误用
token进行替换,就算触发也不会替换句首的目标
原代码:
def remove_stopwords(sentence): stopwords = ["a", "about", "above", "after", "again", "against", "all", "am", "an", "and", "any", "are", "as", "at", "be", "because", "been", "before", "being", "below", "between", "both", "but", "by", "could", "did", "do", "does", "doing", "down", "during", "each", "few", "for", "from", "further", "had", "has", "have", "having", "he", "he'd", "he'll", "he's", "her", "here", "here's", "hers", "herself", "him", "himself", "his", "how", "how's", "i", "i'd", "i'll", "i'm", "i've", "if", "in", "into", "is", "it", "it's", "its", "itself", "let's", "me", "more", "most", "my", "myself", "nor", "of", "on", "once", "only", "or", "other", "ought", "our", "ours", "ourselves", "out", "over", "own", "same", "she", "she'd", "she'll", "she's", "should", "so", "some", "such", "than", "that", "that's", "the", "their", "theirs", "them", "themselves", "then", "there", "there's", "these", "they", "they'd", "they'll", "they're", "they've", "this", "those", "through", "to", "too", "under", "until", "up", "very", "was", "we", "we'd", "we'll", "we're", "we've", "were", "what", "what's", "when", "when's", "where", "where's", "which", "while", "who", "who's", "whom", "why", "why's", "with", "would", "you", "you'd", "you'll", "you're", "you've", "your", "yours", "yourself", "yourselves" ] sentence = sentence.lower() for word in stopwords: token= " " + word + " " token1= + word + " " if (token in sentence): sentence = sentence.replace(token, " ") sentence = sentence.replace(" ", " ") if (token1 in sentence): sentence = sentence.replace(token, " ") sentence = sentence.replace(" ", " ") return sentence
修正后的代码(字符串替换版)
修复语法错误和逻辑问题,分别处理句首、句中、句尾的停用词:
def remove_stopwords(sentence): stopwords = ["a", "about", "above", "after", "again", "against", "all", "am", "an", "and", "any", "are", "as", "at", "be", "because", "been", "before", "being", "below", "between", "both", "but", "by", "could", "did", "do", "does", "doing", "down", "during", "each", "few", "for", "from", "further", "had", "has", "have", "having", "he", "he'd", "he'll", "he's", "her", "here", "here's", "hers", "herself", "him", "himself", "his", "how", "how's", "i", "i'd", "i'll", "i'm", "i've", "if", "in", "into", "is", "it", "it's", "its", "itself", "let's", "me", "more", "most", "my", "myself", "nor", "of", "on", "once", "only", "or", "other", "ought", "our", "ours", "ourselves", "out", "over", "own", "same", "she", "she'd", "she'll", "she's", "should", "so", "some", "such", "than", "that", "that's", "the", "their", "theirs", "them", "themselves", "then", "there", "there's", "these", "they", "they'd", "they'll", "they're", "they've", "this", "those", "through", "to", "too", "under", "until", "up", "very", "was", "we", "we'd", "we'll", "we're", "we've", "were", "what", "what's", "when", "when's", "where", "where's", "which", "while", "who", "who's", "whom", "why", "why's", "with", "would", "you", "you'd", "you'll", "you're", "you've", "your", "yours", "yourself", "yourselves" ] sentence = sentence.lower() for word in stopwords: # 匹配句中(前后有空格) token_mid = f" {word} " # 匹配句首(词后有空格) token_start = f"{word} " # 匹配句尾(词前有空格) token_end = f" {word}" if token_mid in sentence: sentence = sentence.replace(token_mid, " ") if token_start in sentence: sentence = sentence.replace(token_start, "") if token_end in sentence: sentence = sentence.replace(token_end, "") # 处理多余空格:先去首尾,再把多个空格换成单个 sentence = sentence.strip() while " " in sentence: sentence = sentence.replace(" ", " ") return sentence
更推荐的分词过滤版(逻辑更清晰,无边界问题)
字符串替换容易出现意外匹配风险,分词过滤的方法逻辑更简单直接,也不会有边界匹配问题:
def remove_stopwords(sentence): # 把停用词改成集合,查找速度更快 stopwords = {"a", "about", "above", "after", "again", "against", "all", "am", "an", "and", "any", "are", "as", "at", "be", "because", "been", "before", "being", "below", "between", "both", "but", "by", "could", "did", "do", "does", "doing", "down", "during", "each", "few", "for", "from", "further", "had", "has", "have", "having", "he", "he'd", "he'll", "he's", "her", "here", "here's", "hers", "herself", "him", "himself", "his", "how", "how's", "i", "i'd", "i'll", "i'm", "i've", "if", "in", "into", "is", "it", "it's", "its", "itself", "let's", "me", "more", "most", "my", "myself", "nor", "of", "on", "once", "only", "or", "other", "ought", "our", "ours", "ourselves", "out", "over", "own", "same", "she", "she'd", "she'll", "she's", "should", "so", "some", "such", "than", "that", "that's", "the", "their", "theirs", "them", "themselves", "then", "there", "there's", "these", "they", "they'd", "they'll", "they're", "they've", "this", "those", "through", "to", "too", "under", "until", "up", "very", "was", "we", "we'd", "we'll", "we're", "we've", "were", "what", "what's", "when", "when's", "where", "where's", "which", "while", "who", "who's", "whom", "why", "why's", "with", "would", "you", "you'd", "you'll", "you're", "you've", "your", "yours", "yourself", "yourselves" } # 分词并转小写 words = sentence.lower().split() # 过滤停用词 filtered_words = [word for word in words if word not in stopwords] # 拼接成句子 return " ".join(filtered_words)
测试这个版本,输入测试句会直接得到"go store get snack",完全符合预期。
内容的提问来源于stack exchange,提问作者Escanor6
相关产品推荐
相关产品推荐

