如何保存加载Python复杂数据结构及将复杂行转为字典
Hey there! Let's break down your two Python problems step by step—first handling the save/load of complex data structures, then fixing that frustrating line-parsing issue you're running into.
1. Saving & Loading Complex Python Data Structures
When dealing with nested OrderedDicts, arrays, and other non-trivial structures, you’ve got two solid, battle-tested options:
Pickle (Python-Native, Supports Almost All Objects)
Pickle is Python’s built-in serialization tool that works seamlessly with nearly every Python object type—including OrderedDicts, custom classes, and nested collections. It’s super straightforward to use:
import pickle from collections import OrderedDict # Save your complex data complex_data = OrderedDict([ ('user', {'name': 'Alice', 'tags': ['admin', 'editor']}), ('metrics', [100, 200, {'avg': 150}]) ]) with open('complex_data.pkl', 'wb') as f: pickle.dump(complex_data, f) # Load it back later with open('complex_data.pkl', 'rb') as f: loaded_data = pickle.load(f) # Access values just like you would with any dict print(loaded_data['user']['tags'][0]) # Output: admin
⚠️ Quick note: Pickle files are binary and Python-specific. Also, never load pickle files from untrusted sources—they can execute malicious code.
JSON (Cross-Language Compatible)
If you need your data to work with other languages, JSON is the way to go. The catch? It doesn’t natively support OrderedDicts, but we can fix that with custom encoders/decoders:
import json from collections import OrderedDict # Custom encoder to convert OrderedDict to regular dict for JSON class OrderedDictJSONEncoder(json.JSONEncoder): def default(self, obj): if isinstance(obj, OrderedDict): return dict(obj) return super().default(obj) # Save the data complex_data = OrderedDict([ ('user', {'name': 'Bob', 'tags': ['viewer']}), ('metrics', [50, 75, {'avg': 62.5}]) ]) with open('complex_data.json', 'w', encoding='utf-8') as f: json.dump(complex_data, f, cls=OrderedDictJSONEncoder, indent=4) # Load it back as an OrderedDict (if you need to preserve order) with open('complex_data.json', 'r', encoding='utf-8') as f: loaded_data = json.load(f, object_pairs_hook=OrderedDict) # Access values normally print(loaded_data['metrics'][2]['avg']) # Output: 62.5
2. Parsing Lines into Accessible Dictionaries
Your split and ast.literal_eval errors make total sense—split can’t handle commas inside nested structures, and ast.literal_eval only recognizes basic Python literals (it doesn’t know about OrderedDict out of the box). Let’s fix this:
Option 1: Use eval (For Trusted Files Only)
If your file lines are valid Python expressions (like {'key': OrderedDict([('a', 1), ('b', 2)]), 'list': [1,2,3]}), eval can parse them correctly because it recognizes imported objects like OrderedDict. Just make sure the file is from a trusted source—eval executes arbitrary code!
from collections import OrderedDict with open('your_file.txt', 'r', encoding='utf-8') as f: for line in f: line = line.strip() if not line: continue # Parse the line into a dict my_dictionary = eval(line) # Access values as expected print(my_dictionary['key']['a']) # Output: 1
Option 2: Custom Regex Splitting (For Non-Standard Line Formats)
If your lines don’t have outer curly braces (e.g., key: OrderedDict([('a',1), ('b',2)]), list: [1,2,3]), regex can help split key-value pairs without breaking nested commas:
import re from collections import OrderedDict with open('your_file.txt', 'r', encoding='utf-8') as f: for line in f: line = line.strip() if not line: continue # Split on spaces that come before a key (avoids splitting nested commas) key_value_pairs = re.split(r'(?<!,)\s*(?=[a-zA-Z_]+:)', line) my_dictionary = {} for pair in key_value_pairs: # Split each pair into key and value (only split on the first colon) key, value_str = pair.split(':', 1) key = key.strip() value_str = value_str.strip() # Parse the value with eval (again, trust the file!) my_dictionary[key] = eval(value_str) # Access your value print(my_dictionary['list'][1]) # Output: 2
Option 3: Use pyparsing (Safe, For Super Complex Formats)
If regex and eval feel too risky or don’t handle your line structure, the pyparsing library lets you build a custom parser for nested structures. First install it with pip install pyparsing, then try this:
from pyparsing import Word, alphas, nums, Forward, Suppress, delimitedList, Group, quotedString, removeQuotes from collections import OrderedDict # Define parsing rules for our structure expr = Forward() key = Word(alphas + '_') number = Word(nums).setParseAction(lambda t: int(t[0])) string = quotedString.setParseAction(removeQuotes) # Handle dictionaries dict_entry = Group(key + Suppress(':') + expr) dict_expr = Suppress('{') + delimitedList(dict_entry) + Suppress('}') # Handle lists list_expr = Suppress('[') + delimitedList(expr) + Suppress(']') # Handle OrderedDict ordered_dict_expr = Suppress('OrderedDict(') + Group(delimitedList(Group(Suppress('(') + expr + Suppress(',') + expr + Suppress(')')))) + Suppress(')') # Tie all expressions together expr << (number | string | dict_expr | list_expr | ordered_dict_expr) # Parse each line with open('your_file.txt', 'r', encoding='utf-8') as f: for line in f: line = line.strip() if not line: continue # Wrap line in braces if it's a set of key-value pairs without them parsed_result = expr.parseString('{' + line + '}')[0] # Convert to OrderedDict to preserve order my_dictionary = OrderedDict(parsed_result) # Access your value print(my_dictionary['key']['b']) # Output: 2
内容的提问来源于stack exchange,提问作者Shahrooz Pooryousef

