Python中如何将TXT文件转换为含嵌套结构的字典?
Hey Marco, let's work through this together! The trick with nested dictionaries is tracking the hierarchy as you read each line—since you mentioned you have specific data, I'll use a common real-world example that matches the kind of structure you're probably dealing with, then show you how to adapt the code to your case.
First, Let's Define the Scenario
Let's say your TXT file (data.txt) looks like this (using 2-space indentation for nested levels):
name: Alice age: 30 contact: email: alice@example.com phone: 123-456-7890 address: street: 123 Main St city: Anytown zip: 12345
Your Initial (Non-Nested) Code
I'm guessing your current code looks something like this, which works for flat key-value pairs but fails with nesting:
def txt_to_dict(file_path): result = {} with open(file_path, 'r') as f: for line in f: line = line.strip() if not line: continue if ':' in line: key, value = line.split(':', 1) key = key.strip() value = value.strip() result[key] = value return result
Current (Unwanted) Result
Running that code gives you a flat dictionary, where nested keys end up as top-level entries:
{ 'name': 'Alice', 'age': '30', 'contact': '', 'email': 'alice@example.com', 'phone': '123-456-7890', 'address': '', 'street': '123 Main St', 'city': 'Anytown', 'zip': '12345' }
Expected Nested Result
What you want is a properly nested structure:
{ 'name': 'Alice', 'age': 30, 'contact': { 'email': 'alice@example.com', 'phone': '123-456-7890' }, 'address': { 'street': '123 Main St', 'city': 'Anytown', 'zip': 12345 } }
Solution: Track Indentation for Nesting
Here's a revised function that uses an indentation stack to keep track of the current dictionary level. It handles nested structures by adjusting which dictionary we're writing to based on how indented each line is:
def txt_to_nested_dict(file_path): result = {} current_dict = result indent_stack = [] # Adjust this to match the indentation in YOUR TXT file (e.g., 4 spaces, 1 tab) indent_size = 2 with open(file_path, 'r') as f: for line_num, line in enumerate(f, 1): line = line.rstrip('\n') stripped_line = line.strip() if not stripped_line: continue # Calculate how indented this line is indent_level = (len(line) - len(stripped_line)) // indent_size # Move back up the nested structure if indent level decreases while len(indent_stack) >= indent_level: if indent_stack: current_dict = indent_stack.pop() else: break # Split key and value (only split on the first colon) if ':' in stripped_line: key, value = stripped_line.split(':', 1) key = key.strip() value = value.strip() # If value is empty, this is the start of a nested dictionary if not value: new_nested_dict = {} current_dict[key] = new_nested_dict indent_stack.append(current_dict) current_dict = new_nested_dict else: # Optional: Convert numeric values to int/float instead of keeping as strings try: value = int(value) except ValueError: try: value = float(value) except ValueError: pass current_dict[key] = value else: # Handle lines that don't follow key:value format (adjust as needed) print(f"Warning: Line {line_num} is malformed: {line}") return result
How This Works
- Indent Tracking: We calculate the indent level of each line to know if we're moving into a nested section or back out.
- Stack Management: The
indent_stackkeeps track of parent dictionaries. When we hit a line with less indentation, we pop from the stack to return to the parent dictionary. - Nested Dictionary Creation: When a key has an empty value (like
contact:), we create a new dictionary, attach it to the current key, and switch to writing to this new nested dictionary. - Optional Type Conversion: The code tries to convert string values to numbers if possible—you can remove this part if you want all values to stay as strings.
Adapt to Your Data
If your TXT file uses tabs instead of spaces, set indent_size = 1 (since each tab counts as 1 indent level). If your delimiter isn't a colon, adjust the split(':') part to match your file's structure.
内容的提问来源于stack exchange,提问作者Marco Aurélio Guerra

