You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将大字符串形式的列表转为Python可识别的列表?

Handling Large List-like Strings in Python (MemoryError Fixes)

Got it, let's break this down for you. Dealing with 65MB text blobs that are supposed to be lists can be a pain when standard parsers choke on memory—here are some practical fixes to get this sorted:

1. Stream the Content Instead of Parsing All at Once

The core issue with eval() and ast.literal_eval() is that they try to load the entire 65MB string into memory to parse the full list in one go, which triggers the MemoryError. Instead, parse the content in chunks or use streaming parsers:

Option A: Manual Chunked Parsing (For Simple Lists)

If your file is a straightforward Python list (starts with [, ends with ], elements separated by , ), you can manually split and process elements without loading everything at once:

def parse_large_list(file_path):
    with open(file_path, 'r') as f:
        # Skip the opening bracket
        f.read(1)
        buffer = ''
        while True:
            # Read 1MB chunks to balance speed and memory
            chunk = f.read(1024 * 1024)
            if not chunk:
                break
            buffer += chunk
            # Split on element separators
            elements = buffer.split(', ')
            # Yield all complete elements
            for elem in elements[:-1]:
                # Clean up quotes if your elements are string values
                yield elem.strip('\'"')
            # Keep the partial element for the next chunk
            buffer = elements[-1]
        # Handle the final element (strip closing bracket)
        if buffer:
            yield buffer.rstrip(']').strip('\'"')

# Convert the generator to a list (or iterate directly to save memory)
my_list = list(parse_large_list('your_large_file.txt'))

Note: This works best if your elements don't contain , (comma + space) inside them. If they do, use the option below.

Option B: Use a Streaming JSON Parser (For Complex Lists)

First, convert your Python-style list string to JSON-compatible format:

  • Replace single quotes (') with double quotes (")
  • Swap None for null, True/False for true/false

Then use the ijson library to stream elements one at a time (install it first with pip install ijson):

import ijson

def stream_json_list(file_path):
    with open(file_path, 'r') as f:
        # Parse each item in the top-level list
        for elem in ijson.items(f, 'item'):
            yield elem

# Build your list without loading everything into memory
my_list = list(stream_json_list('your_converted_json_file.json'))

This method is safer for complex elements and keeps memory usage extremely low.

2. Switch to a More Efficient Format Long-Term

Once you get the list parsed, converting the file to a Python-friendly format will prevent this issue in the future:

JSON

As mentioned above, JSON is universal, human-readable, and works with streaming parsers. It's a great choice if you need to interact with other tools.

Pickle

If you only plan to use this file in Python, pickle is a binary format that loads much faster than text-based formats. Just be cautious: never load pickle files from untrusted sources.

import pickle

# Save the parsed list to a pickle file
with open('your_list.pkl', 'wb') as f:
    pickle.dump(my_list, f)

# Load it later (no memory issues if the file is optimized)
with open('your_list.pkl', 'rb') as f:
    my_list = pickle.load(f)

CSV

If your list elements are simple values (strings, numbers), CSV is lightweight and easy to parse line-by-line with Python's built-in csv module.

3. Why Wrapping in a Dictionary Didn't Help

When you wrapped the raw string in {mlst: list()}, you just stored the unparsed string as a dictionary value—Python didn't convert the string into an actual list. You'd still need to run a parser on that string, which brings you back to the original memory problem.

内容的提问来源于stack exchange,提问作者addyal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 04:18:09