You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python新手求助:如何获取clean.txt中长度>3的高频词

解决获取长单词及其出现次数的问题

你的核心需求是筛选出长度大于3的单词并统计它们的出现次数,下面是简化且高效的实现方案:

优化后的代码

import re
from collections import Counter

def get_long_word_counts():
    # 读取文件并提取所有单词(转为小写,避免大小写干扰统计)
    with open('clean.txt', 'r', encoding='utf-8') as f:
        text = f.read().lower()
        words = re.findall(r'\w+', text)
    
    # 过滤出长度大于3的单词
    filtered_words = [word for word in words if len(word) > 3]
    
    # 统计单词出现次数,并按次数降序排列
    word_counts = Counter(filtered_words).most_common()
    
    # 输出结果
    for word, count in word_counts:
        print(f"{word}: {count}")

# 调用函数
get_long_word_counts()

对原代码问题的说明

  1. 你原有的read_data()函数只是逐个输出符合长度要求的单词,没有做次数统计,所以无法显示出现次数。
  2. 你已经用了Counter(words).most_common(100),这个方法本身就会返回按出现次数降序排列的结果,后续的count.sort(key=sort_key, reverse=True)是多余操作。
  3. 用re.findall(r'\w+', text)提取单词比split()更可靠,因为split()会把带标点的字符串当成一个单词(比如"hello,"会被完整保留),而正则能准确提取纯单词。

内容的提问来源于stack exchange,提问作者arketipi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 20:01:11