Python词频统计脚本仅统计文本文件最后一行问题排查
歌词动词统计脚本问题排查与解决
问题描述
我是Python新手,正在开发一个脚本,需求是:
- 读取CSV文件生成以动词为键、值全为0的字典
- 遍历歌词文本文件,若字典中的动词出现在歌词中则对应值加1
但运行脚本时发现仅统计每个文本文件的最后一行,已确认路径正确,移除squeeze=True后无法计数,修改filename.endswith判断会报错,参考相关答案未解决问题。
原代码如下:
import pandas as pd import os df = pd.read_csv("List2dict.csv", header=None, index_col=0, squeeze = True) ##read list of verbs verb_dict = df.to_dict() path = "/Users/USERNAMEHERE/Desktop/Python/Notes/LYRICS" ## where lyric files are arr = os.listdir("/Users/USERNAMEHERE/Desktop/Python/Notes/LYRICS") ## list out actual filenames ## replacement_chars = {"(": "", ")": "", "!": "", ".": "", "?": "", "/": "", "'in":"ing", "in'":"ing", "a-":""} add this later ### for filename in arr: if filename.endswith(".txt"): with open(os.path.join(path, filename), "r") as lyrics: for line in lyrics: line = line.strip() line = line.lower() line = line.replace("(", "") line = line.replace(")", "") line = line.replace("!", "") line = line.replace(".", "") line = line.replace("?", "") line = line.replace("/", "") line = line.replace("'in", "ing") line = line.replace("in'", "ing") line = line.replace("a-", "") words = line.split(" ") for word in words: if word in verb_dict: verb_dict[word] = verb_dict[word] + 1 else: pass else: pass print(verb_dict)
问题原因
核心问题是缩进错误:统计单词的for word in words:循环被放在了遍历行的for line in lyrics:循环外面。程序会先遍历处理所有行,但只有最后一行的words变量会被保留,前面所有行的处理结果都被覆盖且未统计,最终仅对最后一行的单词进行计数。
解决方法
把统计单词的循环缩进一级,放到遍历行的循环内部,这样每处理完一行文本并拆分出words后,立刻对这一行的单词进行统计,就能覆盖所有行的内容。
修改后的代码:
import pandas as pd import os df = pd.read_csv("List2dict.csv", header=None, index_col=0, squeeze = True) ##read list of verbs verb_dict = df.to_dict() path = "/Users/USERNAMEHERE/Desktop/Python/Notes/LYRICS" ## where lyric files are arr = os.listdir("/Users/USERNAMEHERE/Desktop/Python/Notes/LYRICS") ## list out actual filenames ## replacement_chars = {"(": "", ")": "", "!": "", ".": "", "?": "", "/": "", "'in":"ing", "in'":"ing", "a-":""} add this later ### for filename in arr: if filename.endswith(".txt"): with open(os.path.join(path, filename), "r") as lyrics: for line in lyrics: line = line.strip() line = line.lower() line = line.replace("(", "") line = line.replace(")", "") line = line.replace("!", "") line = line.replace(".", "") line = line.replace("?", "") line = line.replace("/", "") line = line.replace("'in", "ing") line = line.replace("in'", "ing") line = line.replace("a-", "") words = line.split(" ") # 将统计循环缩进至此,每处理一行就统计一行的单词 for word in words: if word in verb_dict: verb_dict[word] += 1 else: pass print(verb_dict)
额外优化建议
- 用
str.maketrans结合translate批量处理单字符替换,再单独处理多字符替换,比多次调用replace更高效:replacement_chars = {"(": "", ")": "", "!": "", ".": "", "?": "", "/": "", "'in": "ing", "in'": "ing", "a-": ""} # 处理单字符替换 single_trans = str.maketrans({k:v for k,v in replacement_chars.items() if len(k)==1}) line = line.translate(single_trans) # 处理多字符替换 for old, new in replacement_chars.items(): if len(old) > 1: line = line.replace(old, new) - 拆分单词时使用
split()(不带参数),自动处理多个空格的情况,避免生成空字符串元素。
内容的提问来源于stack exchange,提问作者prettywack
相关产品推荐
相关产品推荐

