You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

list.append 内存低效问题求助:大CSV处理时内存冻结

Hey there! Let's tackle this memory freeze issue you're hitting with your large CSV dataset. The core problem is clear—you're loading all 75k+ Document objects (each with nested Token and Sentence structures) into a single list, which eats up massive amounts of RAM all at once. Here are targeted optimizations to fix this:

1. Use a Generator Instead of a List (Most Impactful Quick Fix)

Instead of storing every Document in a list, yield them one at a time with a generator. This way, you only keep one document in memory at a time (unless your downstream code needs to hold all of them, but this works great for streaming processing).

Modify your loadData function like this:

def loadData(self,datasetName,target_indx=None):
    tokens =[]
    labels = []
    ids = []
    allLabels = set()
    with open(self.dirPathName+datasetName+".csv",mode="r",encoding="UTF-8") as f:
        for line in f:
            id,t, l = self.processLine(line)
            tokens.append(t)
            labels.append(l)
            ids.append(id)
            allLabels = allLabels.union(set(l))
    labelsList = list(allLabels)
    labelsList.sort()
    targets = self.getTargets(labelsList,labels)
    
    # Replace list append with yield to generate docs on-demand
    for i in range(len(tokens)):
        doc = self.make_document(ids[i],tokens[i],labels[i],targets[i],target_indx=target_indx)
        yield doc  # Yield instead of appending to a list

When calling loadData, iterate over it like this:

for doc in your_instance.loadData("your_dataset"):
    # Process each doc one at a time
    process_doc(doc)

2. Reduce Object Memory Overhead with __slots__

Python objects use a dictionary by default to store attributes, which adds significant memory overhead for thousands of instances. Use __slots__ in your Token, Sentence, and Document classes to force fixed attribute storage, cutting down memory usage drastically.

Example adjustments:

class Token:
    __slots__ = ["text"]  # List all attributes your Token uses
    def __init__(self, text):
        self.text = text

class Sentence:
    __slots__ = ["tokens"]
    def __init__(self, tokens):
        self.tokens = tokens

class Document:
    __slots__ = ["id", "target_idx", "target_str", "sentences", "target"]
    def __init__(self, id, target_idx, target_str, sentences):
        self.id = id
        self.target_idx = target_idx
        self.target_str = target_str
        self.sentences = sentences
        self.target = None

    def set_target(self, target):
        self.target = target

3. Lazy-Load Token/Sentence Objects

If you don't need the Token and Sentence structures for every document immediately, delay their creation until they're actually needed. This avoids wasting memory on unused nested objects.

Modify make_document and the Document class:

def make_document(self,docid, tokens, labels, target,target_indx=None):
    # Store raw token strings instead of pre-generating Token objects
    doc = Document(id=docid,target_idx=target_indx,target_str=str(labels), raw_tokens=tokens)
    doc.set_target(target)
    return doc

class Document:
    __slots__ = ["id", "target_idx", "target_str", "raw_tokens", "target", "_sentences"]
    def __init__(self, id, target_idx, target_str, raw_tokens):
        self.id = id
        self.target_idx = target_idx
        self.target_str = target_str
        self.raw_tokens = raw_tokens
        self.target = None
        self._sentences = None  # Cache generated sentences

    def get_sentences(self):
        # Generate Token/Sentence only when requested
        if self._sentences is None:
            tokens = [Token(t) for t in self.raw_tokens]
            self._sentences = [Sentence(tokens=tokens)]
        return self._sentences

    def set_target(self, target):
        self.target = target

Now, call doc.get_sentences() only when you need the nested structure.

4. Process Data in Chunks

Avoid loading all 130k rows of tokens/labels/ids into memory at once. Split the CSV into smaller chunks, process each chunk, then discard it to free up RAM.

def loadData(self,datasetName,target_indx=None):
    # First pass to collect all unique labels (required for getTargets)
    allLabels = set()
    with open(self.dirPathName+datasetName+".csv",mode="r",encoding="UTF-8") as f:
        for line in f:
            _, _, l = self.processLine(line)
            allLabels.update(set(l))
    labelsList = sorted(allLabels)

    # Process in chunks of 1000 rows (adjust size based on your memory)
    chunk_size = 1000
    with open(self.dirPathName+datasetName+".csv",mode="r",encoding="UTF-8") as f:
        chunk_tokens, chunk_labels, chunk_ids = [], [], []
        for line in f:
            id,t, l = self.processLine(line)
            chunk_tokens.append(t)
            chunk_labels.append(l)
            chunk_ids.append(id)

            if len(chunk_tokens) >= chunk_size:
                # Process the chunk
                chunk_targets = self.getTargets(labelsList, chunk_labels)
                for i in range(len(chunk_tokens)):
                    doc = self.make_document(chunk_ids[i], chunk_tokens[i], chunk_labels[i], chunk_targets[i], target_indx=target_indx)
                    yield doc
                # Clear chunk to free memory
                chunk_tokens.clear()
                chunk_labels.clear()
                chunk_ids.clear()
        
        # Process remaining rows in the last partial chunk
        if chunk_tokens:
            chunk_targets = self.getTargets(labelsList, chunk_labels)
            for i in range(len(chunk_tokens)):
                doc = self.make_document(chunk_ids[i], chunk_tokens[i], chunk_labels[i], chunk_targets[i], target_indx=target_indx)
                yield doc

Start with the generator approach—it's the quickest win. Combine it with __slots__ and chunking for maximum memory savings.

内容的提问来源于stack exchange,提问作者Pratik.S

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 08:37:56