You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何统计文本文件中包含单引号的高频词?

解决方法

要让代码能统计包含单引号的单词,核心是修改单词的提取逻辑和有效单词的判断条件——原来的isalnum()方法会直接排除带单引号的单词,具体修改如下:

修改后的完整代码

import re

# 读取文件并提取包含字母、数字和单引号的单词
file_name = input('Enter the name of the file: ')
with open(file_name) as f:
    content = f.read().lower()
    # 用正则匹配所有由字母、数字、单引号组成的单词
    words = re.findall(r"[a-zA-Z0-9']+", content)

number_of_words = int(input('Enter how many top words you want to see: '))
uniques = []
stop_words = ["a", "an", "and", "in", "is", "the"]

for word in words:
    # 确保单词不是纯单引号,且只包含允许的字符
    if re.match(r"^[a-zA-Z0-9']*[a-zA-Z0-9][a-zA-Z0-9']*$", word):
        check_special = True
    else:
        check_special = False
    
    if word not in uniques and word not in stop_words and check_special:
        uniques.append(word)

counts = []
for unique in uniques:
    count = 0
    for word in words:
        if word == unique:
            count += 1
    counts.append((count, unique))

counts.sort(reverse=True)

counts_dict = {}
for count, word in counts:
    if count not in counts_dict:
        counts_dict[count] = []
    counts_dict[count].append(word)

count_num_word = 0
for count in sorted(counts_dict.keys(), reverse=True):
    if count_num_word >= number_of_words:
        break
    print('The following words appeared %d times each: %s' % (count, ', '.join(sorted(counts_dict[count]))))
    count_num_word += 1

关键修改点

  1. 单词提取逻辑:
    原来直接用split()分割会把带标点的单词(比如don't.)当成一个整体,现在用正则表达式re.findall(r"[a-zA-Z0-9']+", content),可以精准提取包含单引号的有效单词,同时自动去掉单词前后的无关标点(比如句号、逗号)。

  2. 有效单词判断:
    替换原来的word.isalnum()判断,改用正则re.match(r"^[a-zA-Z0-9']*[a-zA-Z0-9][a-zA-Z0-9']*$", word),既允许单词中包含单引号,又能排除纯单引号的无效字符串(比如''')。

  3. 排序逻辑优化:
    把原来的counts.sort()+counts.reverse()改成counts.sort(reverse=True)更简洁;遍历counts_dict时按次数从高到低排序,确保输出顺序符合预期。

内容的提问来源于stack exchange,提问作者thenorthape

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 07:35:28