如何将大字符串形式的列表转为Python可识别的列表?
Got it, let's break this down for you. Dealing with 65MB text blobs that are supposed to be lists can be a pain when standard parsers choke on memory—here are some practical fixes to get this sorted:
1. Stream the Content Instead of Parsing All at Once
The core issue with eval() and ast.literal_eval() is that they try to load the entire 65MB string into memory to parse the full list in one go, which triggers the MemoryError. Instead, parse the content in chunks or use streaming parsers:
Option A: Manual Chunked Parsing (For Simple Lists)
If your file is a straightforward Python list (starts with [, ends with ], elements separated by , ), you can manually split and process elements without loading everything at once:
def parse_large_list(file_path): with open(file_path, 'r') as f: # Skip the opening bracket f.read(1) buffer = '' while True: # Read 1MB chunks to balance speed and memory chunk = f.read(1024 * 1024) if not chunk: break buffer += chunk # Split on element separators elements = buffer.split(', ') # Yield all complete elements for elem in elements[:-1]: # Clean up quotes if your elements are string values yield elem.strip('\'"') # Keep the partial element for the next chunk buffer = elements[-1] # Handle the final element (strip closing bracket) if buffer: yield buffer.rstrip(']').strip('\'"') # Convert the generator to a list (or iterate directly to save memory) my_list = list(parse_large_list('your_large_file.txt'))
Note: This works best if your elements don't contain , (comma + space) inside them. If they do, use the option below.
Option B: Use a Streaming JSON Parser (For Complex Lists)
First, convert your Python-style list string to JSON-compatible format:
- Replace single quotes (
') with double quotes (") - Swap
Nonefornull,True/Falsefortrue/false
Then use the ijson library to stream elements one at a time (install it first with pip install ijson):
import ijson def stream_json_list(file_path): with open(file_path, 'r') as f: # Parse each item in the top-level list for elem in ijson.items(f, 'item'): yield elem # Build your list without loading everything into memory my_list = list(stream_json_list('your_converted_json_file.json'))
This method is safer for complex elements and keeps memory usage extremely low.
2. Switch to a More Efficient Format Long-Term
Once you get the list parsed, converting the file to a Python-friendly format will prevent this issue in the future:
JSON
As mentioned above, JSON is universal, human-readable, and works with streaming parsers. It's a great choice if you need to interact with other tools.
Pickle
If you only plan to use this file in Python, pickle is a binary format that loads much faster than text-based formats. Just be cautious: never load pickle files from untrusted sources.
import pickle # Save the parsed list to a pickle file with open('your_list.pkl', 'wb') as f: pickle.dump(my_list, f) # Load it later (no memory issues if the file is optimized) with open('your_list.pkl', 'rb') as f: my_list = pickle.load(f)
CSV
If your list elements are simple values (strings, numbers), CSV is lightweight and easy to parse line-by-line with Python's built-in csv module.
3. Why Wrapping in a Dictionary Didn't Help
When you wrapped the raw string in {mlst: list()}, you just stored the unparsed string as a dictionary value—Python didn't convert the string into an actual list. You'd still need to run a parser on that string, which brings you back to the original memory problem.
内容的提问来源于stack exchange,提问作者addyal

