You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于TSV列索引生成层级分类结构字典的实现需求

Alright, let's work through this problem to build that nested hierarchical dictionary from your TSV data. First, let's recap what we need: a function that takes a colon-separated string of column numbers (1-based) and a file path, then returns a nested dict where each level corresponds to the specified columns (highest to lowest).

First, here's your sample TSV data formatted for clarity:

9606 1 GO:0002576 TAS - platelet degranulation - Process
9606 1 GO:0003674 ND - molecular_function_z - Function
9606 1 GO:0003674 OOO - molecular_function_z - Function
9606 1 GO:0005576 IDA - extracellular region - Component
9606 1 GO:0005576 TAS - extracellular region - Component
9606 1 GO:0005576 OOO - extracellular region - Component
9606 1 GO:0005576 HDA - extracellular spaces - Component
9606 1 GO:0008150 ND - biological_processes - Process
9606 1 GO:0008150 OOO - biological_processes - Process
9606 1 GO:0008150 HHH - biological_processes - Process
9606 1 GO:0008150 YYY - biological_processes - Process
9606 1 GO:0031012 IDA - extracellular matrix - Component
9606 1 GO:0043312 TAS - neutrophil degranulat - Process

Step-by-Step Implementation

The core idea is to traverse each line of the TSV, extract the values for your specified columns, then build the dict hierarchy incrementally—avoiding duplicate entries by checking if each level already exists before creating it.

Here's a complete, robust implementation using Python's built-in csv module (great for proper TSV handling, even if fields have spaces or special characters):

import csv

def build_hierarchical_dict(column_spec, file_path):
    # Convert 1-based column numbers to 0-based indices (Python uses 0-indexing)
    hierarchy_indices = list(map(lambda x: int(x) - 1, column_spec.split(":")))
    
    # Validate input: we need at least two levels to form a hierarchy
    if len(hierarchy_indices) < 2:
        raise ValueError("Column spec must include at least two levels (e.g., '8:2:3')")
    
    # Initialize the top-level dictionary
    hierarchical_dict = {}
    
    with open(file_path, 'r', newline='') as file_handle:
        # Create a TSV reader (uses tab as delimiter)
        tsv_reader = csv.reader(file_handle, delimiter='\t')
        
        for fields in tsv_reader:
            # Extract the values that form our hierarchy (in order: highest to lowest)
            hierarchy_values = [fields[idx] for idx in hierarchy_indices]
            
            # Traverse the dictionary, creating levels as needed
            current_level = hierarchical_dict
            for level_idx, value in enumerate(hierarchy_values):
                # If the value doesn't exist at the current level, create it
                if value not in current_level:
                    # For the final level, use a set to automatically handle duplicates
                    if level_idx == len(hierarchy_values) - 1:
                        current_level[value] = set()
                    # For higher levels, create an empty dict to hold the next level
                    else:
                        current_level[value] = {}
                
                # Move down to the next level in the hierarchy
                current_level = current_level[value]
                
                # If we're at the final level (a set), we don't need to go deeper
                if isinstance(current_level, set):
                    break
    
    return hierarchical_dict

# Example usage with your column spec
three_userinput = "8:2:3"
result = build_hierarchical_dict(three_userinput, "your_data.tsv")

# To visualize the result, you can print it with indentation
import json
print(json.dumps(result, indent=2))

Key Details Explained

  • Column Index Conversion: The user input uses 1-based numbers (like "8" for the 8th column), so we subtract 1 to get Python's 0-based indices.
  • Robust TSV Parsing: The csv module handles edge cases (like fields with embedded spaces or tabs) that simple split() would mess up. If your file actually uses spaces instead of tabs, just change delimiter='\t' to delimiter=' ' (but note that this can cause issues if fields contain spaces).
  • Duplicate Handling: Using a set for the lowest level ensures that even if the same value appears multiple times in the TSV, it only gets stored once. If you don't need unique values and just want the hierarchy structure, replace the set with an empty dict ({}) or None.
  • Incremental Hierarchy Building: The current_level variable tracks our position in the nested dict, so we don't have to write repetitive code for each level of the hierarchy—it works for any number of levels (2, 3, 4, etc.).

Example Output

For your sample data and column spec 8:2:3 (0-based 7, 1, 2), the output would look like this (abbreviated):

{
  "Process": {
    "1": {
      "GO:0002576": {},
      "GO:0008150": {},
      "GO:0043312": {}
    }
  },
  "Function": {
    "1": {
      "GO:0003674": {}
    }
  },
  "Component": {
    "1": {
      "GO:0005576": {},
      "GO:0031012": {}
    }
  }
}

(Note: If you used a set for the final level, the values would be empty sets instead of empty dicts.)

内容的提问来源于stack exchange,提问作者Adrian Garcia Moreno

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 03:56:08