技术问询:如何找出使用频率最高的前10个常用词及相关方法
找出使用频率最高的前10个常用词的方案及具体实现方法
通用解决方案步骤
- 文本预处理:清理文本中的标点、特殊符号,统一转换为小写(避免大小写差异导致统计误差),过滤停用词(如“的”“是”“a”“the”这类无实际语义、高频出现的虚词)
- 词频统计:将预处理后的文本拆分为单个词汇,统计每个词汇的出现次数
- 排序提取:对统计结果按出现频率降序排序,提取排名前10的高频词
具体实现方法(Python示例)
以下是无需复杂第三方依赖的落地实现,适合快速验证与日常场景:
步骤说明
- 预处理:对输入文本做清洗、大小写统一、停用词过滤,确保统计对象为有效词汇
- 统计:利用
collections.Counter高效完成词频统计 - 提取:直接调用Counter的
most_common()方法获取前10高频词
代码实现
# 示例输入文本 text = "Python is a great programming language. Python is easy to learn, and Python is widely used in data science. Data science is a hot field now, and learning Python helps in getting jobs in data science." def preprocess_text(raw_text): # 统一转为小写 lower_text = raw_text.lower() # 移除常见标点符号 punctuation = '''!()-[]{};:'"\,<>./?@#$%^&*_~''' cleaned_text = ''.join([char for char in lower_text if char not in punctuation]) # 拆分单词 words = cleaned_text.split() # 自定义停用词列表(可根据场景扩展) stop_words = ['is', 'a', 'and', 'in', 'to', 'now', 'the', 'of', 'for', 'on'] # 过滤停用词 filtered_words = [word for word in words if word not in stop_words] return filtered_words # 统计词频 from collections import Counter processed_words = preprocess_text(text) word_frequency = Counter(processed_words) # 获取前10个高频词 top_10 = word_frequency.most_common(10) # 输出结果 print("前10个高频词及出现次数:") for word, count in top_10: print(f"{word}: {count}")
输出结果示例
前10个高频词及出现次数: python: 4 data: 3 science: 3 great: 1 programming: 1 language: 1 easy: 1 learn: 1 widely: 1 used: 1
扩展说明
- 处理大规模文本时,可分块读取内容,避免内存溢出
- 停用词列表可根据语言(中文/英文)或业务场景调整,比如中文可加入“了”“呢”等虚词
- 处理中文文本需先完成分词(如使用
jieba库),再执行后续统计流程
内容的提问来源于stack exchange,提问作者Sai Manish
相关产品推荐
相关产品推荐

