list.append 内存低效问题求助:大CSV处理时内存冻结
Hey there! Let's tackle this memory freeze issue you're hitting with your large CSV dataset. The core problem is clear—you're loading all 75k+ Document objects (each with nested Token and Sentence structures) into a single list, which eats up massive amounts of RAM all at once. Here are targeted optimizations to fix this:
1. Use a Generator Instead of a List (Most Impactful Quick Fix)
Instead of storing every Document in a list, yield them one at a time with a generator. This way, you only keep one document in memory at a time (unless your downstream code needs to hold all of them, but this works great for streaming processing).
Modify your loadData function like this:
def loadData(self,datasetName,target_indx=None): tokens =[] labels = [] ids = [] allLabels = set() with open(self.dirPathName+datasetName+".csv",mode="r",encoding="UTF-8") as f: for line in f: id,t, l = self.processLine(line) tokens.append(t) labels.append(l) ids.append(id) allLabels = allLabels.union(set(l)) labelsList = list(allLabels) labelsList.sort() targets = self.getTargets(labelsList,labels) # Replace list append with yield to generate docs on-demand for i in range(len(tokens)): doc = self.make_document(ids[i],tokens[i],labels[i],targets[i],target_indx=target_indx) yield doc # Yield instead of appending to a list
When calling loadData, iterate over it like this:
for doc in your_instance.loadData("your_dataset"): # Process each doc one at a time process_doc(doc)
2. Reduce Object Memory Overhead with __slots__
Python objects use a dictionary by default to store attributes, which adds significant memory overhead for thousands of instances. Use __slots__ in your Token, Sentence, and Document classes to force fixed attribute storage, cutting down memory usage drastically.
Example adjustments:
class Token: __slots__ = ["text"] # List all attributes your Token uses def __init__(self, text): self.text = text class Sentence: __slots__ = ["tokens"] def __init__(self, tokens): self.tokens = tokens class Document: __slots__ = ["id", "target_idx", "target_str", "sentences", "target"] def __init__(self, id, target_idx, target_str, sentences): self.id = id self.target_idx = target_idx self.target_str = target_str self.sentences = sentences self.target = None def set_target(self, target): self.target = target
3. Lazy-Load Token/Sentence Objects
If you don't need the Token and Sentence structures for every document immediately, delay their creation until they're actually needed. This avoids wasting memory on unused nested objects.
Modify make_document and the Document class:
def make_document(self,docid, tokens, labels, target,target_indx=None): # Store raw token strings instead of pre-generating Token objects doc = Document(id=docid,target_idx=target_indx,target_str=str(labels), raw_tokens=tokens) doc.set_target(target) return doc class Document: __slots__ = ["id", "target_idx", "target_str", "raw_tokens", "target", "_sentences"] def __init__(self, id, target_idx, target_str, raw_tokens): self.id = id self.target_idx = target_idx self.target_str = target_str self.raw_tokens = raw_tokens self.target = None self._sentences = None # Cache generated sentences def get_sentences(self): # Generate Token/Sentence only when requested if self._sentences is None: tokens = [Token(t) for t in self.raw_tokens] self._sentences = [Sentence(tokens=tokens)] return self._sentences def set_target(self, target): self.target = target
Now, call doc.get_sentences() only when you need the nested structure.
4. Process Data in Chunks
Avoid loading all 130k rows of tokens/labels/ids into memory at once. Split the CSV into smaller chunks, process each chunk, then discard it to free up RAM.
def loadData(self,datasetName,target_indx=None): # First pass to collect all unique labels (required for getTargets) allLabels = set() with open(self.dirPathName+datasetName+".csv",mode="r",encoding="UTF-8") as f: for line in f: _, _, l = self.processLine(line) allLabels.update(set(l)) labelsList = sorted(allLabels) # Process in chunks of 1000 rows (adjust size based on your memory) chunk_size = 1000 with open(self.dirPathName+datasetName+".csv",mode="r",encoding="UTF-8") as f: chunk_tokens, chunk_labels, chunk_ids = [], [], [] for line in f: id,t, l = self.processLine(line) chunk_tokens.append(t) chunk_labels.append(l) chunk_ids.append(id) if len(chunk_tokens) >= chunk_size: # Process the chunk chunk_targets = self.getTargets(labelsList, chunk_labels) for i in range(len(chunk_tokens)): doc = self.make_document(chunk_ids[i], chunk_tokens[i], chunk_labels[i], chunk_targets[i], target_indx=target_indx) yield doc # Clear chunk to free memory chunk_tokens.clear() chunk_labels.clear() chunk_ids.clear() # Process remaining rows in the last partial chunk if chunk_tokens: chunk_targets = self.getTargets(labelsList, chunk_labels) for i in range(len(chunk_tokens)): doc = self.make_document(chunk_ids[i], chunk_tokens[i], chunk_labels[i], chunk_targets[i], target_indx=target_indx) yield doc
Start with the generator approach—it's the quickest win. Combine it with __slots__ and chunking for maximum memory savings.
内容的提问来源于stack exchange,提问作者Pratik.S

