You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中将文本文件每行转为独立字典并统计多文档词频

多文档词频统计改造方案

原代码仅支持读取单个文本文件并输出全局词频,现需适配5个包含「文档ID行+内容行」格式的文本文件,同时实现单文档独立词频统计与全局词频统计功能。

原代码

import re 
import string

input_file = open('documents.txt', 'r')
stopwords_file = open('stopwords_en.txt', 'r')
stopwords_list = []

for line in stopwords_file.readlines():
  stopwords_list.extend(line.split())

stopwords_set = set(stopwords_list)

word_count = {}
for line in input_file.readlines():
    words = line.strip()
    words = words.translate(str.maketrans('','', string.punctuation))
    words = re.findall('\w+', line)
    for word in words: 
      if word.lower() in stopwords_set:
        continue
      word = word.lower()
      if not word in word_count: 
        word_count[word] = 1
      else: 
        word_count[word] = word_count[word] + 1

word_index = sorted(word_count.keys())
for word in word_index:
  print (word, word_count[word]) 

改造后的代码

import re
import string

# 读取停用词集合
stopwords_set = set()
with open('stopwords_en.txt', 'r', encoding='utf-8') as f:
    for line in f:
        stopwords_set.update(line.strip().split())

# 待处理的5个文件列表(根据实际文件名修改)
target_files = ['doc1.txt', 'doc2.txt', 'doc3.txt', 'doc4.txt', 'doc5.txt']

# 存储各文档词频:key为文档ID,value为该文档的词频字典
doc_word_counts = {}
# 存储全局词频
global_word_counts = {}

def process_text(text):
    """文本预处理:转小写、移除标点、提取单词、过滤停用词"""
    # 转小写
    text = text.lower()
    # 移除标点
    text = text.translate(str.maketrans('', '', string.punctuation))
    # 提取所有单词
    words = re.findall(r'\w+', text)
    # 过滤停用词
    return [word for word in words if word not in stopwords_set]

# 批量处理每个文件
for file_path in target_files:
    with open(file_path, 'r', encoding='utf-8') as f:
        lines = [line.strip() for line in f if line.strip()]  # 过滤空行
    
    # 按ID+内容的成对格式解析文档
    for i in range(0, len(lines), 2):
        doc_id = lines[i]
        doc_content = lines[i+1]
        
        # 预处理文本
        filtered_words = process_text(doc_content)
        
        # 统计当前文档词频
        doc_counts = {}
        for word in filtered_words:
            doc_counts[word] = doc_counts.get(word, 0) + 1
            # 更新全局词频
            global_word_counts[word] = global_word_counts.get(word, 0) + 1
        
        doc_word_counts[doc_id] = doc_counts

# 输出各文档词频
print("=== 各文档独立词频统计 ===")
for doc_id, counts in sorted(doc_word_counts.items()):
    print(f"\n文档ID: {doc_id}")
    for word, cnt in sorted(counts.items()):
        print(f"{word}: {cnt}")

# 输出全局词频
print("\n=== 全局词频统计 ===")
for word, cnt in sorted(global_word_counts.items()):
    print(f"{word}: {cnt}")

关键改动说明

  • 用with语句管理文件操作,自动释放资源,避免文件句柄泄漏
  • 新增process_text函数统一文本预处理逻辑,避免重复代码
  • 定义文件列表批量处理5个目标文件,无需逐个修改代码
  • 按「ID行+内容行」的成对规则解析每个文件内的文档,确保每个文档的ID与内容对应
  • 新增doc_word_counts字典存储每个文档的独立词频,保留原有的全局词频统计
  • 优化词频统计逻辑,使用dict.get()简化计数代码

内容的提问来源于stack exchange,提问作者user14452102

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 17:20:34