You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

技术问询:如何找出使用频率最高的前10个常用词及相关方法

找出使用频率最高的前10个常用词的方案及具体实现方法

通用解决方案步骤

  • 文本预处理:清理文本中的标点、特殊符号,统一转换为小写(避免大小写差异导致统计误差),过滤停用词(如“的”“是”“a”“the”这类无实际语义、高频出现的虚词)
  • 词频统计:将预处理后的文本拆分为单个词汇,统计每个词汇的出现次数
  • 排序提取:对统计结果按出现频率降序排序,提取排名前10的高频词

具体实现方法(Python示例)

以下是无需复杂第三方依赖的落地实现,适合快速验证与日常场景:

步骤说明

  1. 预处理:对输入文本做清洗、大小写统一、停用词过滤,确保统计对象为有效词汇
  2. 统计:利用collections.Counter高效完成词频统计
  3. 提取:直接调用Counter的most_common()方法获取前10高频词

代码实现

# 示例输入文本
text = "Python is a great programming language. Python is easy to learn, and Python is widely used in data science. Data science is a hot field now, and learning Python helps in getting jobs in data science."

def preprocess_text(raw_text):
    # 统一转为小写
    lower_text = raw_text.lower()
    # 移除常见标点符号
    punctuation = '''!()-[]{};:'"\,<>./?@#$%^&*_~'''
    cleaned_text = ''.join([char for char in lower_text if char not in punctuation])
    # 拆分单词
    words = cleaned_text.split()
    # 自定义停用词列表(可根据场景扩展)
    stop_words = ['is', 'a', 'and', 'in', 'to', 'now', 'the', 'of', 'for', 'on']
    # 过滤停用词
    filtered_words = [word for word in words if word not in stop_words]
    return filtered_words

# 统计词频
from collections import Counter
processed_words = preprocess_text(text)
word_frequency = Counter(processed_words)

# 获取前10个高频词
top_10 = word_frequency.most_common(10)

# 输出结果
print("前10个高频词及出现次数:")
for word, count in top_10:
    print(f"{word}: {count}")

输出结果示例

前10个高频词及出现次数:
python: 4
data: 3
science: 3
great: 1
programming: 1
language: 1
easy: 1
learn: 1
widely: 1
used: 1

扩展说明

  • 处理大规模文本时,可分块读取内容,避免内存溢出
  • 停用词列表可根据语言(中文/英文)或业务场景调整,比如中文可加入“了”“呢”等虚词
  • 处理中文文本需先完成分词(如使用jieba库),再执行后续统计流程

内容的提问来源于stack exchange,提问作者Sai Manish

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 02:01:11