Python文本清洗脚本误删词汇无法通过评测求助
问题根源与修复方案
你的代码存在3处逻辑错误,会引发误删、漏统计、词频拆分等问题,逐一说明:
- 停用词匹配逻辑用错了判断对象
你现在拿还没去掉标点的原始拆分字符串转小写去匹配停用词表,而不是拿清洗完标点的词去匹配。这会导致非常不稳定的结果:比如不带标点的停用词会被过滤,带了句号/逗号的停用词(比如it./the,)因为多了标点匹配不上停用词表,就会漏进统计;反过来如果一个和停用词同形的有效词(比如做人名的Will、表名词“罐子”的can)刚好不带标点,就会被直接误删,和带标点时的处理结果不一致。 - 清洗后的词没有统一大小写
你现在保留了词的原始大小写,会把Hello/hello/HELLO当成三个不同的词拆分统计,不符合词频统计的常规要求,也会导致词云效果变差。 - 没有过滤空值
如果拆分出来的字符串全是标点、数字,清洗完会得到空字符串,当前逻辑会把空字符串也加入统计,产生无效的词频条目。
修复后代码
punctuations = '''!()-[]{};:'"\,<>./?@#$%^&*_~''' uninteresting_words = ["the", "a", "to", "if", "is", "it", "of", "and", "or", "an", "as", "i", "me", "my", \ "we", "our", "ours", "you", "your", "yours", "he", "she", "him", "his", "her", "hers", "its", "they", "them", \ "their", "what", "which", "who", "whom", "this", "that", "am", "are", "was", "were", "be", "been", "being", \ "have", "has", "had", "do", "does", "did", "but", "at", "by", "with", "from", "here", "when", "where", "how", \ "all", "any", "both", "each", "few", "more", "some", "such", "no", "nor", "too", "very", "can", "will", "just"] def count(file_contents): frequencies = {} word_list = file_contents.split() final_list = [] for word in word_list: new_word = "" for character in word: if character not in punctuations and character.isalpha(): new_word += character # 统一转小写后再做停用词判断,同时过滤空字符串 clean_word = new_word.lower() if clean_word and clean_word not in uninteresting_words: final_list.append(clean_word) for word in final_list: frequencies[word] = frequencies.get(word, 0) + 1 return frequencies
补充说明
如果修复后仍存在有效词被误删的情况,需要根据你的文本场景调整停用词表:当前停用词表是通用英文停用词,部分同形异义词会被默认过滤,比如will作名词表“意志/遗嘱”、can作名词表“罐子”、just作形容词表“公正的”时属于有效词,但会被当前停用词表命中过滤,按需增删停用词表即可。
内容的提问来源于stack exchange,提问作者Haroon Atif
相关产品推荐
相关产品推荐

