基于嵌套图数据的层级术语重复项去除方案问询
Solution: Keep Most Specific Terms by Hierarchy
First, let's clarify the hierarchy from your DataFrame: each row links a child term (via start ID) to its parent term (via end ID). For example, Colecalciferol → Vitamin D3 → Vitamin D → Vitamin. Our goal is to remove any higher-level (ancestor) terms from a list if a more specific (descendant) term is present.
Approach
- Map Terms to IDs: Create mappings between term names and their
startIDs to work with the graph structure. - Build Hierarchy Graph: Construct a directed graph where edges go from child terms to their parent terms.
- Filter Terms: For each term in the list, check if any other term in the list is a descendant (i.e., there's a path from the other term to the current term in the graph). If so, remove the current (ancestor) term—we only keep terms with no descendants present in the list.
Complete Code
import pandas as pd import networkx as nx from numpy import nan # Sample data d = {'start': {0:4,1:3,2:2,3:1,4:12,5:11,6:23,7:22,8:21}, 'name':{0:'Vitamin',1:'Vitamin D',2:'Vitamin D3',3:'Colecalciferol',4:'Vitamin D2',5:'Ergocalcifero',6:'Vitamin K',7:'Vitamin K2',8:'Menachinon'}, 'end':{0:nan,1:4.0,2:3.0,3:2.0,4:3.0,5:12.0,6:4.0,7:23.0,8:22.0}} df = pd.DataFrame(d) # Create mappings between term names and their start IDs name_to_id = df.set_index('name')['start'].to_dict() id_to_name = df.set_index('start')['name'].to_dict() # Build directed graph: edges from child (start) to parent (end) G = nx.DiGraph() for idx, row in df.iterrows(): start_id = row['start'] end_id = row['end'] if not pd.isna(end_id): G.add_edge(start_id, end_id) # Child node points to parent node def keep_most_specific_terms(term_list, graph, name_to_id): """Filter a list to keep only the most specific terms (no descendants present in the list).""" result = [] for term in term_list: term_id = name_to_id[term] # Check if any other term in the list is a descendant of this term has_descendant_in_list = False for other_term in term_list: if other_term == term: continue other_id = name_to_id[other_term] # If other_term can reach term in the graph, term is an ancestor of other_term if nx.has_path(graph, other_id, term_id): has_descendant_in_list = True break if not has_descendant_in_list: result.append(term) return result # Test with your example lists l1 = ['Colecalciferol', 'Vitamin D'] l2 = ['Colecalciferol', 'Ergocalcifero', 'Vitamin D3'] processed_l1 = keep_most_specific_terms(l1, G, name_to_id) processed_l2 = keep_most_specific_terms(l2, G, name_to_id) print(f"l1 = {processed_l1}") # Output: l1 = ['Colecalciferol'] print(f"l2 = {processed_l2}") # Output: l2 = ['Colecalciferol', 'Ergocalcifero']
How It Works
- Graph Construction: The graph links each child term to its parent. For example,
Colecalciferol(ID 3) has an edge toVitamin D3(ID 2), which links toVitamin D(ID 1), etc. - Descendant Check: Using
nx.has_path(graph, other_id, term_id), we check ifother_termis a descendant ofterm. If this path exists,termis an ancestor and should be removed ifother_termis in the list. - Filtering: We only keep terms where no other term in the list is a descendant—these are the most specific terms in the list.
Your original code was exploring successors, but we needed to check ancestor-descendant relationships in the reverse direction (from child to parent) to identify which terms to remove.
内容的提问来源于stack exchange,提问作者Andreas
相关产品推荐
相关产品推荐

