使用ast.literal_eval实现文件转字典时内存占用过高问题
The issue isn't ast.literal_eval itself—it's the duplication of objects created when parsing each identical line. Here's the breakdown:
- String Duplication: Each line contains the same 5 keys and 5 values, but
ast.literal_evalcreates new string objects for every key/value pair in every line. For 10 million lines, that's 100 million total string objects (5 keys + 5 values per line × 10M lines). Each string in Python has significant overhead (~50 bytes per string, even for short ones), which adds up to ~5GB alone. - Dict Overhead: Each parsed dict has its own memory overhead (~200 bytes per dict). 10 million dicts account for another ~2GB.
Combined with the integer keys for the outer dict, this totals ~7-10GB of memory—far more than the compact text file size, since text doesn't store object overhead.
Since all lines are identical, you can eliminate nearly all duplication:
Option 1: Reuse the Same Dict Instance (Best if No Modifications Are Needed)
Parse one line once, then assign the same dict to all line numbers. This uses minimal memory because all entries reference a single dict and shared strings:
dict_of_dicts = {} # Parse once to get the prototype dict with open("each_line_is_a_python_dict.txt") as f: prototype = ast.literal_eval(next(f)) # Assign the same dict to every line number for lineno in range(10000000): dict_of_dicts[lineno] = prototype
Note: Modifying any entry will modify all others, as they point to the same object.
Option 2: Create Separate Dicts with Shared Strings (If Modifications Are Needed)
If you need independent dicts, reuse the string objects from the prototype to avoid duplication:
dict_of_dicts = {} with open("each_line_is_a_python_dict.txt") as f: prototype = ast.literal_eval(next(f)) # Create new dicts using the same key/value strings for lineno in range(10000000): # Using dict(prototype) or a comprehension reuses the existing strings dict_of_dicts[lineno] = dict(prototype)
This way, you only have 10 total string objects (5 keys +5 values) instead of 100 million, cutting memory usage by ~5GB immediately.
Text files are compact because they store raw characters, but in-memory Python objects have fixed overhead. When dealing with identical repeated data, reuse objects instead of creating duplicates—this is the biggest win for memory efficiency.
内容的提问来源于stack exchange,提问作者shrimpdrake

