You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python实现:从文本文件拆分单词为单个字母并统计字母频率

如何将文本文件中的单词拆分为单个字母?

我需要计算一段文本中各字母的出现频率,但不知道如何将文本中的单词拆分为单个字母,以便统计唯一元素并确定其频率。

抱歉无法提供文本文件,以下是给定的文本:

alice was beginning to get very tired of sitting by her sister on the bank, and of having nothing to do: once or twice she had peeped into the book her sister was reading, but it had no pictures or conversations in it, and what is the use of a book,' thought alice without pictures or conversation?' so she was considering in her own mind (as well as she could, for the hot day made her feel very sleepy and stupid), whether the pleasure of making a daisy- chain would be worth the trouble of getting up and picking the daisies, when suddenly a white rabbit with pink eyes ran close by her. there was nothing so very remarkable in that; nor did alice think it so very much out of the way to hear the rabbit say to itself, oh dear! oh dear! i shall be late!' (when she thought it over afterwards, it occurred to her that she ought to have wondered at this, but at the time it all seemed quite natural); but when the rabbit actually took a watch out of its waistcoat- pocket, and looked at it, and then hurried on, alice started to her feet, for it flashed across her mind that she had never before seen a rabbit with either a waistcoat-pocket, or a watch to take out of it, and burning with curiosity, she ran across the field after it, and fortunately was just in time to see it pop down a large rabbit-hole under the hedge.
in another moment down went alice after it, never once considering how in the world she was to get out again.
the rabbit-hole went straight on like a tunnel for some way, and then dipped suddenly down, so suddenly that alice had not a moment to think about stopping herself before she found herself falling down a very deep well.

我需要将字母拆分为a-z对应的26个变量,并确定它们的频率。

我目前编写了以下代码:

# Check where the current file you are working in, is saved. 
import os
os.getcwd()
#print(os.getcwd())

# 1. Change the current working directory to the place where you have saved the file.
os.chdir('C:/Users/Annik/Desktop/DTU/02633 Introduction to programming/Datafiles')
os.getcwd()
#print(os.chdir('C:/Users/Annik/Desktop/DTU/02633 Introduction to programming/Datafiles'))

# 2. Listing the content of current working directory type
os.listdir(os.getcwd())
#print(os.listdir(os.getcwd()))

#importing the file
filein = open("small_text.txt", "r") #opens the file for reading
lines = filein.readlines() #reads all lines into an array
smalltxt = "".join(lines) #Joins the lines into one big string.

import numpy as np

def letterFrequency(filename):
    #counts the frequency of letters in a text
    
    unique_elems, counts = np.unique(separate_words, return_counts=True)

    return unique_elems

我不清楚如何拆分文本中的字母,以便统计唯一元素。


解决方法

1. 拆分文本为单个字母

字符串在Python中本身是可迭代对象,直接遍历就能获取每个字符。但需要先过滤非字母字符(标点、空格等),并统一转为小写(避免大小写重复统计)。

2. 完善后的代码(基于numpy)

以下是修正后的代码,包含字母拆分、过滤和频率统计逻辑:

import os
import numpy as np

def letterFrequency(filename):
    # 读取文件内容,自动处理文件关闭
    with open(filename, 'r') as filein:
        smalltxt = filein.read()
    
    # 提取所有小写字母,过滤非字母字符
    letters = [c.lower() for c in smalltxt if c.isalpha()]
    
    # 统计唯一字母及出现次数
    unique_elems, counts = np.unique(letters, return_counts=True)
    
    # 整理为字典,补充a-z中未出现的字母(频率设为0)
    freq_dict = dict(zip(unique_elems, counts))
    for char in 'abcdefghijklmnopqrstuvwxyz':
        if char not in freq_dict:
            freq_dict[char] = 0
    
    # 按字母顺序排序结果
    sorted_freq = sorted(freq_dict.items())
    
    return sorted_freq

# 设置工作目录并调用函数
os.chdir('C:/Users/Annik/Desktop/DTU/02633 Introduction to programming/Datafiles')
frequency_result = letterFrequency("small_text.txt")

# 打印每个字母的频率
for char, count in frequency_result:
    print(f"{char}: {count}")

3. 纯Python替代方案(无需numpy)

如果不想依赖numpy,可以用collections.Counter实现统计,代码更简洁:

import os
from collections import Counter

def letterFrequency(filename):
    with open(filename, 'r') as filein:
        smalltxt = filein.read()
    
    # 提取小写字母
    letters = [c.lower() for c in smalltxt if c.isalpha()]
    # 统计频率
    freq_counter = Counter(letters)
    
    # 补充a-z未出现的字母
    for char in 'abcdefghijklmnopqrstuvwxyz':
        freq_counter[char] = freq_counter.get(char, 0)
    
    # 按字母顺序返回结果
    return sorted(freq_counter.items())

# 调用方式同上
os.chdir('C:/Users/Annik/Desktop/DTU/02633 Introduction to programming/Datafiles')
frequency_result = letterFrequency("small_text.txt")
for char, count in frequency_result:
    print(f"{char}: {count}")

关键说明

  • 字母过滤:用c.isalpha()判断字符是否为字母,c.lower()统一转为小写,避免大小写重复统计。
  • 文件读取优化:with open语句会自动关闭文件,避免资源泄漏。
  • 完整覆盖a-z:手动补充未出现的字母,确保结果包含全部26个字母的频率。

内容的提问来源于stack exchange,提问作者Annika

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 07:55:40