You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python读取文本文件时跳过空行以正确统计词频

解决词频统计中空字符串计数的问题

问题根源

你的代码出现空字符串("")计数的原因有两个:

  1. 文本中的空行经过line.strip()后会变成空字符串,此时用split(" ")分割会得到[""]这个列表,遍历的时候就会把空字符串加入统计
  2. 即使是非空行,用split(" ")按单个空格分割,遇到连续空格时也会生成空字符串元素

修正方案

方案1:跳过空行+使用默认split()分割

直接跳过处理后的空行,同时用不带参数的split()(会自动分割任意数量的空白字符,包括空格、换行、制表符,且忽略首尾空白),避免生成空字符串:

d = dict()
with open('/Users/admin/Desktop/textfile.txt', 'r') as f:  # 用with语句自动关闭文件,更安全
    for line in f:
        line = line.strip()
        if not line:  # 跳过空行
            continue
        words = line.split()  # 不带参数的split,分割任意空白字符
        for word in words:
            # 可选:去除单词首尾的标点,避免"advertising."和"advertising"被算作不同词
            cleaned_word = word.strip('.,')
            if cleaned_word in d:
                d[cleaned_word] += 1
            else:
                d[cleaned_word] = 1

for key, count in d.items():
    print(f"{key} : {count}")

方案2:过滤空字符串

如果不想跳过空行,也可以在遍历单词时直接过滤掉空字符串:

d = dict()
f = open('/Users/admin/Desktop/textfile.txt','r')
for line in f:
    line = line.strip()
    words = line.split(" ")
    for word in words:
        if not word:  # 跳过空字符串
            continue
        # 可选:去除标点
        cleaned_word = word.strip('.,')
        if cleaned_word in d:
            d[cleaned_word] = d[cleaned_word] + 1
        else:
            d[cleaned_word] = 1

for key in list(d.keys()):
    print(key, ":", d[key])
f.close()  # 记得手动关闭文件

额外优化说明

上面的代码里加入了cleaned_word = word.strip('.,'),是为了把单词首尾的逗号、句号去掉,避免同一个单词因为标点被统计成不同的条目。如果需要处理更多标点,可以扩展这个字符串,比如strip('.,!?;:')。

内容的提问来源于stack exchange,提问作者Ayesha

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 17:55:19