You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Google Colab中Python无法正确统计文本文件单词出现次数求助

问题排查与修正方案

核心问题

你用readlines()读取文件后得到的是每行文本组成的列表,调用myString.count(word)时,只会统计列表中完全等于该单词的整行数量,而不是单词在文本中的实际出现次数。比如文本里某行是"Hello example world",这行内容不等于"example",所以统计结果为0。

修正后的代码

import os
from google.colab import drive
import string

drive.mount('/content/drive/', force_remount=True)

# 确保目标目录存在
if not os.path.exists('/content/drive/My Drive/Miserables'):
  os.makedirs('/content/drive/My Drive/Miserables')

# 读取整个文本为单字符串并转小写,统一匹配规则
with open("/content/drive/My Drive/Miserables/miserable.txt", 'r') as f:
     full_text = f.read().lower()

# 移除所有标点符号,避免"example."和"example"被视为不同单词
full_text = full_text.translate(str.maketrans('', '', string.punctuation))
# 将文本拆分为单个单词的列表
words_list = full_text.split()

# 统计目标单词
searchWords = ["example"]
for word in searchWords:
    count = words_list.count(word.lower())
    print(f"Word '{word}' appeared {count} time/s.")

进阶优化(可选)

如果需要处理更复杂的文本场景(比如连字符单词、缩写等),可以用正则表达式精准提取单词:

import re
# 替换原有的拆分步骤
words_list = re.findall(r'\b\w+\b', full_text.lower())

内容的提问来源于stack exchange,提问作者NoobWithPython

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 18:20:31