基于WTO报告文本的Decoder Model运行np.zeros()内存错误排查
np.zeros() throw MemoryError on a 64GB RAM machine but not on 16GB? Let’s get straight to the root of this—your code has a critical logical error that’s creating massively oversized arrays, and the reason it “works” locally is likely due to either smaller input files or your system using swap memory to compensate (even if it’s slow).
The Core Problem: You’re Creating Arrays That Are Way Too Big
Looking at your code, here’s the mistake:
- You set
max_encoder_seq_length = len(input_text)—this makes the sequence length equal to the entire length of your input text (34196 characters in your error log). - Then you create
encoder_input_datawith shape(len(input_text), max_encoder_seq_length, num_encoder_tokens)—so that’s(34196, 34196, 70).
Let’s calculate the memory footprint of this array:
Each float32 takes 4 bytes, so total bytes are:34196 * 34196 * 70 * 4 ≈ 330,000,000,000 bytes → that’s ~300 GB of RAM needed. Even your 64GB remote machine can’t handle that—hence the MemoryError.
Why Does It Work Locally?
There are two likely explanations:
- Your local input files are smaller: If the text files on your 16GB machine are way shorter than the remote ones, the array size would shrink enough to fit (even with swap).
- Swap memory is covering the gap: Your local system might have a large swap partition that’s absorbing the excess memory demand, making it seem like it works (though it’s probably running extremely slow).
How to Fix This
You’re misunderstanding how sequence-to-sequence models work—you need to split your text into individual samples/sentences, not treat the entire document as one giant sequence. Here’s what to adjust:
Split your text into individual samples
Instead of reading the entire text as a single string, split it into separate sentences (e.g., by newlines):# Split into individual input/target samples input_texts = open("/Users/zachary/Downloads/DS2_input_art_III.txt").read().split('\n') target_texts = open("/Users/zachary/Downloads/DS_2_output_art_III.txt").read().split('\n') # Filter out empty lines input_texts = [txt for txt in input_texts if txt.strip()] target_texts = [txt for txt in target_texts if txt.strip()]Calculate proper sequence lengths
Use the length of the longest individual sample, not the entire document:max_encoder_seq_length = max([len(txt) for txt in input_texts]) max_decoder_seq_length = max([len(txt) for txt in target_texts]) num_samples = len(input_texts) # This is your actual number of training samplesResize your data arrays
Now create arrays with the correct shape (samples × max_seq_length × num_tokens):encoder_input_data = np.zeros( (num_samples, max_encoder_seq_length, num_encoder_tokens), dtype='float32') decoder_input_data = np.zeros( (num_samples, max_decoder_seq_length, num_decoder_tokens), dtype='float32') decoder_target_data = np.zeros( (num_samples, max_decoder_seq_length, num_decoder_tokens), dtype='float32')Fix the data populating loop
Your original loop was iterating over individual characters (treating each as a sample)—now iterate over each sample sentence:for i, (input_txt, target_txt) in enumerate(zip(input_texts, target_texts)): for t, char in enumerate(input_txt): encoder_input_data[i, t, input_token_index[char]] = 1. for t, char in enumerate(target_txt): decoder_input_data[i, t, target_token_index[char]] = 1. if t > 0: # decoder_target_data is shifted by one timestep decoder_target_data[i, t - 1, target_token_index[char]] = 1.
With these changes, your array sizes will be manageable (e.g., 1000 samples × 100 characters × 70 tokens = ~28MB per array), and you won’t hit memory errors on either machine.
内容的提问来源于stack exchange,提问作者snapper

