如何在Python中将文本文件每行转为独立字典并统计多文档词频
多文档词频统计改造方案
原代码仅支持读取单个文本文件并输出全局词频,现需适配5个包含「文档ID行+内容行」格式的文本文件,同时实现单文档独立词频统计与全局词频统计功能。
原代码
import re import string input_file = open('documents.txt', 'r') stopwords_file = open('stopwords_en.txt', 'r') stopwords_list = [] for line in stopwords_file.readlines(): stopwords_list.extend(line.split()) stopwords_set = set(stopwords_list) word_count = {} for line in input_file.readlines(): words = line.strip() words = words.translate(str.maketrans('','', string.punctuation)) words = re.findall('\w+', line) for word in words: if word.lower() in stopwords_set: continue word = word.lower() if not word in word_count: word_count[word] = 1 else: word_count[word] = word_count[word] + 1 word_index = sorted(word_count.keys()) for word in word_index: print (word, word_count[word])
改造后的代码
import re import string # 读取停用词集合 stopwords_set = set() with open('stopwords_en.txt', 'r', encoding='utf-8') as f: for line in f: stopwords_set.update(line.strip().split()) # 待处理的5个文件列表(根据实际文件名修改) target_files = ['doc1.txt', 'doc2.txt', 'doc3.txt', 'doc4.txt', 'doc5.txt'] # 存储各文档词频:key为文档ID,value为该文档的词频字典 doc_word_counts = {} # 存储全局词频 global_word_counts = {} def process_text(text): """文本预处理:转小写、移除标点、提取单词、过滤停用词""" # 转小写 text = text.lower() # 移除标点 text = text.translate(str.maketrans('', '', string.punctuation)) # 提取所有单词 words = re.findall(r'\w+', text) # 过滤停用词 return [word for word in words if word not in stopwords_set] # 批量处理每个文件 for file_path in target_files: with open(file_path, 'r', encoding='utf-8') as f: lines = [line.strip() for line in f if line.strip()] # 过滤空行 # 按ID+内容的成对格式解析文档 for i in range(0, len(lines), 2): doc_id = lines[i] doc_content = lines[i+1] # 预处理文本 filtered_words = process_text(doc_content) # 统计当前文档词频 doc_counts = {} for word in filtered_words: doc_counts[word] = doc_counts.get(word, 0) + 1 # 更新全局词频 global_word_counts[word] = global_word_counts.get(word, 0) + 1 doc_word_counts[doc_id] = doc_counts # 输出各文档词频 print("=== 各文档独立词频统计 ===") for doc_id, counts in sorted(doc_word_counts.items()): print(f"\n文档ID: {doc_id}") for word, cnt in sorted(counts.items()): print(f"{word}: {cnt}") # 输出全局词频 print("\n=== 全局词频统计 ===") for word, cnt in sorted(global_word_counts.items()): print(f"{word}: {cnt}")
关键改动说明
- 用
with语句管理文件操作,自动释放资源,避免文件句柄泄漏 - 新增
process_text函数统一文本预处理逻辑,避免重复代码 - 定义文件列表批量处理5个目标文件,无需逐个修改代码
- 按「ID行+内容行」的成对规则解析每个文件内的文档,确保每个文档的ID与内容对应
- 新增
doc_word_counts字典存储每个文档的独立词频,保留原有的全局词频统计 - 优化词频统计逻辑,使用
dict.get()简化计数代码
内容的提问来源于stack exchange,提问作者user14452102
相关产品推荐
相关产品推荐

