Python统计多文档词频至字典并构建词-文档矩阵遇问题求助
解决思路与完整代码实现
Hey there! Let's work through this problem step by step since you're new to Python and stuck on the final if-else part. I'll break down the solution into clear, manageable parts without using any third-party libraries.
1. 核心任务拆解
我们的需求可以拆成4个关键环节:
- 读取并拆分出11个文档的内容
- 对每个文档单独统计词频,生成对应字典
- 收集所有出现过的唯一词汇
- 构建Term-Document矩阵
2. 文本预处理与词频统计函数
首先写一个通用的词频统计函数,这里要注意文本清洗(比如统一小写、去除标点)——这也是很多新手在if-else统计时出错的核心原因:同一个词因为大小写或标点差异被当成不同键,导致统计混乱。
import string def count_words(text): # 初始化当前文档的词频字典 word_counts = {} # 统一转小写,避免大小写造成的词重复 text = text.lower() # 去除所有标点符号 translator = str.maketrans('', '', string.punctuation) cleaned_text = text.translate(translator) # 按空格拆分单词(如果文档有换行,split()会自动处理) words = cleaned_text.split() # 替换你原来有问题的if-else逻辑 for word in words: if word in word_counts: word_counts[word] += 1 else: word_counts[word] = 1 # 也可以用更简洁的写法替代if-else: # word_counts[word] = word_counts.get(word, 0) + 1 return word_counts
3. 读取并拆分11个文档
假设你的txt文件中,11个文档通过特定分隔符区分(比如每个文档开头是=== DocX ===,X从1到11)。如果你的文档是连续无分隔符的,可根据实际格式调整拆分逻辑(比如按固定行数分割):
def split_docs(file_path, separator_pattern="=== Doc"): with open(file_path, 'r', encoding='utf-8') as f: content = f.read() # 按分隔符拆分文档,过滤掉空内容 raw_docs = [doc.strip() for doc in content.split(separator_pattern) if doc.strip()] # 确保正好拆分出11个文档 assert len(raw_docs) == 11, f"拆分出的文档数量不对,得到了{len(raw_docs)}个" # 生成对应Doc1到Doc11的词频字典集合 doc_word_counts = {} for i in range(11): doc_key = f"Doc{i+1}" doc_word_counts[doc_key] = count_words(raw_docs[i]) return doc_word_counts
4. 构建Term-Document矩阵
矩阵的行是所有唯一词汇,列是11个文档,每个单元格存储对应文档中该词的出现次数(无则为0):
def build_term_doc_matrix(doc_word_counts): # 收集所有出现过的唯一词汇,并排序(可选,让矩阵更整齐) all_terms = sorted({word for counts in doc_word_counts.values() for word in counts.keys()}) # 按Doc1到Doc11的顺序整理文档名 doc_names = sorted(doc_word_counts.keys()) # 初始化矩阵:第一行是表头(词汇+文档名) matrix = [["Term"] + doc_names] # 遍历每个词汇,填充每行数据 for term in all_terms: row = [term] for doc in doc_names: # 获取当前文档中该词的词频,没有则返回0 row.append(doc_word_counts[doc].get(term, 0)) matrix.append(row) return matrix
5. 完整调用示例
把这些函数整合起来,替换你的文件路径即可运行:
# 替换成你的txt文件实际路径 file_path = "your_docs.txt" # 获取每个文档的词频字典(Doc1到Doc11) doc_counts = split_docs(file_path) # 构建Term-Document矩阵 term_doc_matrix = build_term_doc_matrix(doc_counts) # 打印矩阵查看效果 for row in term_doc_matrix: print(row)
你原来if-else可能出错的常见原因
- 未做文本清洗:比如单词带标点(
hello,和hello)、大小写不一致(Hello和hello),导致同一个词被统计成多个键,if-else无法正确累加。 - 单词拆分逻辑错误:比如误用
split('\n')拆分换行而非空格,导致整行被当成一个词,统计完全失效。 - 字典初始化问题:可能未正确为每个文档创建独立的词频字典,导致if-else判断时引用了未定义的字典对象。
如果你的文档分隔方式和示例不同,可以告诉我具体的文件格式,我再帮你调整拆分逻辑~
内容的提问来源于stack exchange,提问作者WWH98932
相关产品推荐
相关产品推荐

