You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于嵌套图数据的层级术语重复项去除方案问询

Solution: Keep Most Specific Terms by Hierarchy

First, let's clarify the hierarchy from your DataFrame: each row links a child term (via start ID) to its parent term (via end ID). For example, Colecalciferol → Vitamin D3 → Vitamin D → Vitamin. Our goal is to remove any higher-level (ancestor) terms from a list if a more specific (descendant) term is present.

Approach

  1. Map Terms to IDs: Create mappings between term names and their start IDs to work with the graph structure.
  2. Build Hierarchy Graph: Construct a directed graph where edges go from child terms to their parent terms.
  3. Filter Terms: For each term in the list, check if any other term in the list is a descendant (i.e., there's a path from the other term to the current term in the graph). If so, remove the current (ancestor) term—we only keep terms with no descendants present in the list.

Complete Code

import pandas as pd
import networkx as nx
from numpy import nan

# Sample data
d = {'start': {0:4,1:3,2:2,3:1,4:12,5:11,6:23,7:22,8:21}, 
     'name':{0:'Vitamin',1:'Vitamin D',2:'Vitamin D3',3:'Colecalciferol',4:'Vitamin D2',5:'Ergocalcifero',6:'Vitamin K',7:'Vitamin K2',8:'Menachinon'}, 
     'end':{0:nan,1:4.0,2:3.0,3:2.0,4:3.0,5:12.0,6:4.0,7:23.0,8:22.0}}
df = pd.DataFrame(d)

# Create mappings between term names and their start IDs
name_to_id = df.set_index('name')['start'].to_dict()
id_to_name = df.set_index('start')['name'].to_dict()

# Build directed graph: edges from child (start) to parent (end)
G = nx.DiGraph()
for idx, row in df.iterrows():
    start_id = row['start']
    end_id = row['end']
    if not pd.isna(end_id):
        G.add_edge(start_id, end_id)  # Child node points to parent node

def keep_most_specific_terms(term_list, graph, name_to_id):
    """Filter a list to keep only the most specific terms (no descendants present in the list)."""
    result = []
    for term in term_list:
        term_id = name_to_id[term]
        # Check if any other term in the list is a descendant of this term
        has_descendant_in_list = False
        for other_term in term_list:
            if other_term == term:
                continue
            other_id = name_to_id[other_term]
            # If other_term can reach term in the graph, term is an ancestor of other_term
            if nx.has_path(graph, other_id, term_id):
                has_descendant_in_list = True
                break
        if not has_descendant_in_list:
            result.append(term)
    return result

# Test with your example lists
l1 = ['Colecalciferol', 'Vitamin D']
l2 = ['Colecalciferol', 'Ergocalcifero', 'Vitamin D3']

processed_l1 = keep_most_specific_terms(l1, G, name_to_id)
processed_l2 = keep_most_specific_terms(l2, G, name_to_id)

print(f"l1 = {processed_l1}")  # Output: l1 = ['Colecalciferol']
print(f"l2 = {processed_l2}")  # Output: l2 = ['Colecalciferol', 'Ergocalcifero']

How It Works

  • Graph Construction: The graph links each child term to its parent. For example, Colecalciferol (ID 3) has an edge to Vitamin D3 (ID 2), which links to Vitamin D (ID 1), etc.
  • Descendant Check: Using nx.has_path(graph, other_id, term_id), we check if other_term is a descendant of term. If this path exists, term is an ancestor and should be removed if other_term is in the list.
  • Filtering: We only keep terms where no other term in the list is a descendant—these are the most specific terms in the list.

Your original code was exploring successors, but we needed to check ancestor-descendant relationships in the reverse direction (from child to parent) to identify which terms to remove.

内容的提问来源于stack exchange,提问作者Andreas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 10:57:46