基于TSV列索引生成层级分类结构字典的实现需求
Alright, let's work through this problem to build that nested hierarchical dictionary from your TSV data. First, let's recap what we need: a function that takes a colon-separated string of column numbers (1-based) and a file path, then returns a nested dict where each level corresponds to the specified columns (highest to lowest).
First, here's your sample TSV data formatted for clarity:
9606 1 GO:0002576 TAS - platelet degranulation - Process 9606 1 GO:0003674 ND - molecular_function_z - Function 9606 1 GO:0003674 OOO - molecular_function_z - Function 9606 1 GO:0005576 IDA - extracellular region - Component 9606 1 GO:0005576 TAS - extracellular region - Component 9606 1 GO:0005576 OOO - extracellular region - Component 9606 1 GO:0005576 HDA - extracellular spaces - Component 9606 1 GO:0008150 ND - biological_processes - Process 9606 1 GO:0008150 OOO - biological_processes - Process 9606 1 GO:0008150 HHH - biological_processes - Process 9606 1 GO:0008150 YYY - biological_processes - Process 9606 1 GO:0031012 IDA - extracellular matrix - Component 9606 1 GO:0043312 TAS - neutrophil degranulat - Process
Step-by-Step Implementation
The core idea is to traverse each line of the TSV, extract the values for your specified columns, then build the dict hierarchy incrementally—avoiding duplicate entries by checking if each level already exists before creating it.
Here's a complete, robust implementation using Python's built-in csv module (great for proper TSV handling, even if fields have spaces or special characters):
import csv def build_hierarchical_dict(column_spec, file_path): # Convert 1-based column numbers to 0-based indices (Python uses 0-indexing) hierarchy_indices = list(map(lambda x: int(x) - 1, column_spec.split(":"))) # Validate input: we need at least two levels to form a hierarchy if len(hierarchy_indices) < 2: raise ValueError("Column spec must include at least two levels (e.g., '8:2:3')") # Initialize the top-level dictionary hierarchical_dict = {} with open(file_path, 'r', newline='') as file_handle: # Create a TSV reader (uses tab as delimiter) tsv_reader = csv.reader(file_handle, delimiter='\t') for fields in tsv_reader: # Extract the values that form our hierarchy (in order: highest to lowest) hierarchy_values = [fields[idx] for idx in hierarchy_indices] # Traverse the dictionary, creating levels as needed current_level = hierarchical_dict for level_idx, value in enumerate(hierarchy_values): # If the value doesn't exist at the current level, create it if value not in current_level: # For the final level, use a set to automatically handle duplicates if level_idx == len(hierarchy_values) - 1: current_level[value] = set() # For higher levels, create an empty dict to hold the next level else: current_level[value] = {} # Move down to the next level in the hierarchy current_level = current_level[value] # If we're at the final level (a set), we don't need to go deeper if isinstance(current_level, set): break return hierarchical_dict # Example usage with your column spec three_userinput = "8:2:3" result = build_hierarchical_dict(three_userinput, "your_data.tsv") # To visualize the result, you can print it with indentation import json print(json.dumps(result, indent=2))
Key Details Explained
- Column Index Conversion: The user input uses 1-based numbers (like "8" for the 8th column), so we subtract 1 to get Python's 0-based indices.
- Robust TSV Parsing: The
csvmodule handles edge cases (like fields with embedded spaces or tabs) that simplesplit()would mess up. If your file actually uses spaces instead of tabs, just changedelimiter='\t'todelimiter=' '(but note that this can cause issues if fields contain spaces). - Duplicate Handling: Using a
setfor the lowest level ensures that even if the same value appears multiple times in the TSV, it only gets stored once. If you don't need unique values and just want the hierarchy structure, replace the set with an empty dict ({}) orNone. - Incremental Hierarchy Building: The
current_levelvariable tracks our position in the nested dict, so we don't have to write repetitive code for each level of the hierarchy—it works for any number of levels (2, 3, 4, etc.).
Example Output
For your sample data and column spec 8:2:3 (0-based 7, 1, 2), the output would look like this (abbreviated):
{ "Process": { "1": { "GO:0002576": {}, "GO:0008150": {}, "GO:0043312": {} } }, "Function": { "1": { "GO:0003674": {} } }, "Component": { "1": { "GO:0005576": {}, "GO:0031012": {} } } }
(Note: If you used a set for the final level, the values would be empty sets instead of empty dicts.)
内容的提问来源于stack exchange,提问作者Adrian Garcia Moreno

