You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中使用collections.Counter统计文本文件字符频次的简便方法

用Python统计文本文件中字符出现次数的简便方法

你当前的代码存在两个问题:一是用split()把文本拆成了单词列表,丢失了单个字符的信息;二是手动构造字符列表完全没必要,效率极低。下面是更简便的实现方式:

基础实现(统计所有字符,包括空格、标点)

直接读取整个文件内容,利用Counter可直接处理可迭代对象(字符串本身就是字符的可迭代序列)的特性,一步完成统计:

from collections import Counter

# 读取文件并转小写(统一大小写统计,不需要可去掉.lower())
with open("data/aragon.txt", 'r', encoding='utf-8') as data_file:
    full_text = data_file.read().lower()

# 统计所有字符出现次数
char_counter = Counter(full_text)

# 按出现次数从高到低输出结果
for char, count in sorted(char_counter.items(), key=lambda x: x[1], reverse=True):
    print(f"'{char}' : {count}")

进阶:只统计字母字符(排除空格、数字、标点)

如果不需要统计空格、数字或标点,可以先过滤文本:

from collections import Counter

with open("data/aragon.txt", 'r', encoding='utf-8') as data_file:
    full_text = data_file.read().lower()

# 过滤出仅字母字符
filtered_chars = [c for c in full_text if c.isalpha()]
char_counter = Counter(filtered_chars)

# 输出结果
for char, count in sorted(char_counter.items(), key=lambda x: x[1], reverse=True):
    print(f"'{char}' : {count}")

关键说明

  • data_file.read():一次性读取整个文件的所有内容,避免逐行处理的繁琐,同时保留所有字符信息(包括换行、空格)。
  • Counter直接处理字符串:字符串本身是可迭代的,Counter会自动遍历每个字符并计数,不需要手动拆分字符成列表。
  • .lower():将所有字母转为小写,确保'A'和'a'被统计为同一个字符,不需要区分大小写时可省略。

内容的提问来源于stack exchange,提问作者user19965401

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 05:40:37