树节点文档集合中各术语的文档出现次数统计需求问询
解决单词文档出现次数统计问题(含树结构处理)
我来帮你搞定这个需求——核心是统计每个单词在多少份独立文档中出现,同一文档内重复出现的单词仅计数1次。下面分单个节点和树结构两种场景来实现:
1. 单个节点的文档统计
先从你示例里的节点A开始,实现单个文档集合的统计逻辑:
实现思路
- 用字典记录每个单词对应的文档索引集合(确保同一文档只被计数一次)
- 遍历每份文档,先对文档内的单词去重,再把文档索引添加到对应单词的集合中
- 最后将集合长度转换为该单词的文档出现次数
代码实现
def count_doc_frequency(documents): # 键:单词,值:该单词出现的文档索引集合 word_to_docs = {} for doc_idx, doc in enumerate(documents): # 对当前文档的单词去重,避免同一文档内重复计数 unique_words = set(doc) for word in unique_words: if word not in word_to_docs: word_to_docs[word] = set() word_to_docs[word].add(doc_idx) # 把集合长度转为文档出现次数 return {word: len(doc_indices) for word, doc_indices in word_to_docs.items()} # 测试示例中的节点A A_docs = [['a','b'],['a','a'],['c','d'],['a','c'],['d','e']] result = count_doc_frequency(A_docs) print(result) # 输出: {'a': 3, 'b': 1, 'c': 2, 'd': 2, 'e': 1}
2. 树结构的全局文档统计
如果需要统计整个树中所有节点的文档,我们可以通过递归遍历树节点,合并各节点的统计结果:
假设树节点结构
先定义一个简单的树节点类(你可以根据自己的实际结构调整):
class TreeNode: def __init__(self, documents): self.documents = documents # 当前节点的5份文档 self.children = [] # 子节点列表
递归合并统计的实现
这种方法不需要把所有文档都加载到内存,而是边遍历边合并统计结果,更适合大型树结构:
def count_tree_global_frequency(node): # 先统计当前节点的文档索引映射 current_map = {} for doc_idx, doc in enumerate(node.documents): unique_words = set(doc) for word in unique_words: if word not in current_map: current_map[word] = set() current_map[word].add(doc_idx) # 递归处理所有子节点 for child in node.children: child_map = count_tree_global_frequency(child) # 子节点的文档索引需要偏移(因为当前节点已有len(node.documents)份文档) offset = len(node.documents) # 偏移子节点的文档索引 shifted_child_map = {word: {idx + offset for idx in indices} for word, indices in child_map.items()} # 合并到当前节点的映射中 for word, indices in shifted_child_map.items(): if word in current_map: current_map[word].update(indices) else: current_map[word] = indices # 转换为最终的文档出现次数 return {word: len(indices) for word, indices in current_map.items()} # 构建测试树 node_A = TreeNode([['a','b'],['a','a'],['c','d'],['a','c'],['d','e']]) node_B = TreeNode([['a','e'],['f','g'],['a','f'],['h','i'],['b','c']]) node_C = TreeNode([['c','e'],['d','f'],['a','b'],['g','h'],['e','f']]) node_A.children = [node_B, node_C] # 统计整个树的结果 tree_result = count_tree_global_frequency(node_A) print(tree_result)
关键注意事项
- 必须对单份文档内的单词去重,否则会错误地把同一文档内的重复单词多次计数
- 树结构处理时,要确保不同节点的文档被视为独立的份,所以需要偏移子节点的文档索引,避免索引冲突
- 如果你的树节点结构和示例不同,只需要调整递归遍历子节点的逻辑即可
内容的提问来源于stack exchange,提问作者adi5257
相关产品推荐
相关产品推荐

