如何统计文本文件中包含单引号的高频词?
解决方法
要让代码能统计包含单引号的单词,核心是修改单词的提取逻辑和有效单词的判断条件——原来的isalnum()方法会直接排除带单引号的单词,具体修改如下:
修改后的完整代码
import re # 读取文件并提取包含字母、数字和单引号的单词 file_name = input('Enter the name of the file: ') with open(file_name) as f: content = f.read().lower() # 用正则匹配所有由字母、数字、单引号组成的单词 words = re.findall(r"[a-zA-Z0-9']+", content) number_of_words = int(input('Enter how many top words you want to see: ')) uniques = [] stop_words = ["a", "an", "and", "in", "is", "the"] for word in words: # 确保单词不是纯单引号,且只包含允许的字符 if re.match(r"^[a-zA-Z0-9']*[a-zA-Z0-9][a-zA-Z0-9']*$", word): check_special = True else: check_special = False if word not in uniques and word not in stop_words and check_special: uniques.append(word) counts = [] for unique in uniques: count = 0 for word in words: if word == unique: count += 1 counts.append((count, unique)) counts.sort(reverse=True) counts_dict = {} for count, word in counts: if count not in counts_dict: counts_dict[count] = [] counts_dict[count].append(word) count_num_word = 0 for count in sorted(counts_dict.keys(), reverse=True): if count_num_word >= number_of_words: break print('The following words appeared %d times each: %s' % (count, ', '.join(sorted(counts_dict[count])))) count_num_word += 1
关键修改点
单词提取逻辑:
原来直接用split()分割会把带标点的单词(比如don't.)当成一个整体,现在用正则表达式re.findall(r"[a-zA-Z0-9']+", content),可以精准提取包含单引号的有效单词,同时自动去掉单词前后的无关标点(比如句号、逗号)。有效单词判断:
替换原来的word.isalnum()判断,改用正则re.match(r"^[a-zA-Z0-9']*[a-zA-Z0-9][a-zA-Z0-9']*$", word),既允许单词中包含单引号,又能排除纯单引号的无效字符串(比如''')。排序逻辑优化:
把原来的counts.sort()+counts.reverse()改成counts.sort(reverse=True)更简洁;遍历counts_dict时按次数从高到低排序,确保输出顺序符合预期。
内容的提问来源于stack exchange,提问作者thenorthape
相关产品推荐
相关产品推荐

