You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

树节点文档集合中各术语的文档出现次数统计需求问询

解决单词文档出现次数统计问题(含树结构处理)

我来帮你搞定这个需求——核心是统计每个单词在多少份独立文档中出现,同一文档内重复出现的单词仅计数1次。下面分单个节点和树结构两种场景来实现:

1. 单个节点的文档统计

先从你示例里的节点A开始,实现单个文档集合的统计逻辑:

实现思路

  • 用字典记录每个单词对应的文档索引集合(确保同一文档只被计数一次)
  • 遍历每份文档,先对文档内的单词去重,再把文档索引添加到对应单词的集合中
  • 最后将集合长度转换为该单词的文档出现次数

代码实现

def count_doc_frequency(documents):
    # 键:单词,值:该单词出现的文档索引集合
    word_to_docs = {}
    for doc_idx, doc in enumerate(documents):
        # 对当前文档的单词去重,避免同一文档内重复计数
        unique_words = set(doc)
        for word in unique_words:
            if word not in word_to_docs:
                word_to_docs[word] = set()
            word_to_docs[word].add(doc_idx)
    # 把集合长度转为文档出现次数
    return {word: len(doc_indices) for word, doc_indices in word_to_docs.items()}

# 测试示例中的节点A
A_docs = [['a','b'],['a','a'],['c','d'],['a','c'],['d','e']]
result = count_doc_frequency(A_docs)
print(result)  # 输出: {'a': 3, 'b': 1, 'c': 2, 'd': 2, 'e': 1}

2. 树结构的全局文档统计

如果需要统计整个树中所有节点的文档,我们可以通过递归遍历树节点,合并各节点的统计结果:

假设树节点结构

先定义一个简单的树节点类(你可以根据自己的实际结构调整):

class TreeNode:
    def __init__(self, documents):
        self.documents = documents  # 当前节点的5份文档
        self.children = []  # 子节点列表

递归合并统计的实现

这种方法不需要把所有文档都加载到内存,而是边遍历边合并统计结果,更适合大型树结构:

def count_tree_global_frequency(node):
    # 先统计当前节点的文档索引映射
    current_map = {}
    for doc_idx, doc in enumerate(node.documents):
        unique_words = set(doc)
        for word in unique_words:
            if word not in current_map:
                current_map[word] = set()
            current_map[word].add(doc_idx)
    
    # 递归处理所有子节点
    for child in node.children:
        child_map = count_tree_global_frequency(child)
        # 子节点的文档索引需要偏移(因为当前节点已有len(node.documents)份文档)
        offset = len(node.documents)
        # 偏移子节点的文档索引
        shifted_child_map = {word: {idx + offset for idx in indices} 
                            for word, indices in child_map.items()}
        # 合并到当前节点的映射中
        for word, indices in shifted_child_map.items():
            if word in current_map:
                current_map[word].update(indices)
            else:
                current_map[word] = indices
    
    # 转换为最终的文档出现次数
    return {word: len(indices) for word, indices in current_map.items()}

# 构建测试树
node_A = TreeNode([['a','b'],['a','a'],['c','d'],['a','c'],['d','e']])
node_B = TreeNode([['a','e'],['f','g'],['a','f'],['h','i'],['b','c']])
node_C = TreeNode([['c','e'],['d','f'],['a','b'],['g','h'],['e','f']])
node_A.children = [node_B, node_C]

# 统计整个树的结果
tree_result = count_tree_global_frequency(node_A)
print(tree_result)

关键注意事项

  • 必须对单份文档内的单词去重,否则会错误地把同一文档内的重复单词多次计数
  • 树结构处理时,要确保不同节点的文档被视为独立的份,所以需要偏移子节点的文档索引,避免索引冲突
  • 如果你的树节点结构和示例不同,只需要调整递归遍历子节点的逻辑即可

内容的提问来源于stack exchange,提问作者adi5257

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:44:57