You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python统计多文档词频至字典并构建词-文档矩阵遇问题求助

解决思路与完整代码实现

Hey there! Let's work through this problem step by step since you're new to Python and stuck on the final if-else part. I'll break down the solution into clear, manageable parts without using any third-party libraries.

1. 核心任务拆解

我们的需求可以拆成4个关键环节:

  • 读取并拆分出11个文档的内容
  • 对每个文档单独统计词频,生成对应字典
  • 收集所有出现过的唯一词汇
  • 构建Term-Document矩阵

2. 文本预处理与词频统计函数

首先写一个通用的词频统计函数,这里要注意文本清洗(比如统一小写、去除标点)——这也是很多新手在if-else统计时出错的核心原因:同一个词因为大小写或标点差异被当成不同键,导致统计混乱。

import string

def count_words(text):
    # 初始化当前文档的词频字典
    word_counts = {}
    # 统一转小写,避免大小写造成的词重复
    text = text.lower()
    # 去除所有标点符号
    translator = str.maketrans('', '', string.punctuation)
    cleaned_text = text.translate(translator)
    # 按空格拆分单词(如果文档有换行,split()会自动处理)
    words = cleaned_text.split()
    
    # 替换你原来有问题的if-else逻辑
    for word in words:
        if word in word_counts:
            word_counts[word] += 1
        else:
            word_counts[word] = 1
        # 也可以用更简洁的写法替代if-else:
        # word_counts[word] = word_counts.get(word, 0) + 1
    return word_counts

3. 读取并拆分11个文档

假设你的txt文件中,11个文档通过特定分隔符区分(比如每个文档开头是=== DocX ===,X从1到11)。如果你的文档是连续无分隔符的,可根据实际格式调整拆分逻辑(比如按固定行数分割):

def split_docs(file_path, separator_pattern="=== Doc"):
    with open(file_path, 'r', encoding='utf-8') as f:
        content = f.read()
    # 按分隔符拆分文档,过滤掉空内容
    raw_docs = [doc.strip() for doc in content.split(separator_pattern) if doc.strip()]
    # 确保正好拆分出11个文档
    assert len(raw_docs) == 11, f"拆分出的文档数量不对,得到了{len(raw_docs)}个"
    # 生成对应Doc1到Doc11的词频字典集合
    doc_word_counts = {}
    for i in range(11):
        doc_key = f"Doc{i+1}"
        doc_word_counts[doc_key] = count_words(raw_docs[i])
    return doc_word_counts

4. 构建Term-Document矩阵

矩阵的行是所有唯一词汇,列是11个文档,每个单元格存储对应文档中该词的出现次数(无则为0):

def build_term_doc_matrix(doc_word_counts):
    # 收集所有出现过的唯一词汇,并排序(可选,让矩阵更整齐)
    all_terms = sorted({word for counts in doc_word_counts.values() for word in counts.keys()})
    # 按Doc1到Doc11的顺序整理文档名
    doc_names = sorted(doc_word_counts.keys())
    # 初始化矩阵:第一行是表头(词汇+文档名)
    matrix = [["Term"] + doc_names]
    # 遍历每个词汇,填充每行数据
    for term in all_terms:
        row = [term]
        for doc in doc_names:
            # 获取当前文档中该词的词频,没有则返回0
            row.append(doc_word_counts[doc].get(term, 0))
        matrix.append(row)
    return matrix

5. 完整调用示例

把这些函数整合起来,替换你的文件路径即可运行:

# 替换成你的txt文件实际路径
file_path = "your_docs.txt"
# 获取每个文档的词频字典(Doc1到Doc11)
doc_counts = split_docs(file_path)
# 构建Term-Document矩阵
term_doc_matrix = build_term_doc_matrix(doc_counts)

# 打印矩阵查看效果
for row in term_doc_matrix:
    print(row)

你原来if-else可能出错的常见原因

  1. 未做文本清洗:比如单词带标点(hello,和hello)、大小写不一致(Hello和hello),导致同一个词被统计成多个键,if-else无法正确累加。
  2. 单词拆分逻辑错误:比如误用split('\n')拆分换行而非空格,导致整行被当成一个词,统计完全失效。
  3. 字典初始化问题:可能未正确为每个文档创建独立的词频字典,导致if-else判断时引用了未定义的字典对象。

如果你的文档分隔方式和示例不同,可以告诉我具体的文件格式,我再帮你调整拆分逻辑~

内容的提问来源于stack exchange,提问作者WWH98932

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 08:47:47