Google Colab中Python无法正确统计文本文件单词出现次数求助
问题排查与修正方案
核心问题
你用readlines()读取文件后得到的是每行文本组成的列表,调用myString.count(word)时,只会统计列表中完全等于该单词的整行数量,而不是单词在文本中的实际出现次数。比如文本里某行是"Hello example world",这行内容不等于"example",所以统计结果为0。
修正后的代码
import os from google.colab import drive import string drive.mount('/content/drive/', force_remount=True) # 确保目标目录存在 if not os.path.exists('/content/drive/My Drive/Miserables'): os.makedirs('/content/drive/My Drive/Miserables') # 读取整个文本为单字符串并转小写,统一匹配规则 with open("/content/drive/My Drive/Miserables/miserable.txt", 'r') as f: full_text = f.read().lower() # 移除所有标点符号,避免"example."和"example"被视为不同单词 full_text = full_text.translate(str.maketrans('', '', string.punctuation)) # 将文本拆分为单个单词的列表 words_list = full_text.split() # 统计目标单词 searchWords = ["example"] for word in searchWords: count = words_list.count(word.lower()) print(f"Word '{word}' appeared {count} time/s.")
进阶优化(可选)
如果需要处理更复杂的文本场景(比如连字符单词、缩写等),可以用正则表达式精准提取单词:
import re # 替换原有的拆分步骤 words_list = re.findall(r'\b\w+\b', full_text.lower())
内容的提问来源于stack exchange,提问作者NoobWithPython
相关产品推荐
相关产品推荐

