使用Gensim加载GoogleNews预训练词向量时遭遇内存错误求助
Hey there, sorry you're stuck with this MemoryError when loading the GoogleNews word vectors—let's break down what's happening and walk through some fixes that should help.
First, a quick context check: the GoogleNews dataset has 3 million 300-dimensional vectors, which adds up to roughly 3.6GB of raw data (using float32). But Gensim uses extra memory during loading for parsing and temporary structures, so if your system's available RAM is lower than before (or you're running other memory-heavy apps now), that's likely why you're hitting this error even though the code worked previously.
Here are your best options to resolve this:
Load only the most common words (quick fix)
You probably don't need all 3 million words for your task. Use thelimitparameter to load just the top N most frequent terms—this cuts memory usage drastically. For example, to load the first 1 million words:import gensim.models.keyedvectors as word2vec model = word2vec.KeyedVectors.load_word2vec_format( "GoogleNews-vectors-negative300.bin", binary=True, limit=1000000 )This works for most NLP tasks since rare words rarely add meaningful value.
Convert to Gensim's native format (more efficient long-term)
The binary.binfile is great for portability but inefficient to load. If you can temporarily free up enough memory to load the full model once, save it in Gensim's optimized.kvformat. This reduces memory usage on subsequent loads and speeds up the process:# First load the bin file (free up RAM first if needed) model = word2vec.KeyedVectors.load_word2vec_format("GoogleNews-vectors-negative300.bin", binary=True) # Save to native format model.save("GoogleNews-vectors-negative300.kv") # Next time, load from this optimized file model = word2vec.KeyedVectors.load("GoogleNews-vectors-negative300.kv")Free up system memory
Chances are your system has less available RAM than when the code worked before. Close any unnecessary apps—like browser tabs, other IDEs, or running ML models. On Linux/macOS, usefree -hortopto check memory usage; on Windows, use Task Manager to see what's hogging RAM.Switch to 64-bit Python (if you're on 32-bit)
32-bit Python has a hard memory limit (usually around 4GB), which is almost exactly what the GoogleNews vectors need. If you accidentally switched to a 32-bit environment, switching back to 64-bit Python will remove this restriction.Memory-map the file (advanced, for full vector access)
If you absolutely need all 3 million vectors but don't have enough RAM, use themmap='r'parameter. This maps the file directly to your system's virtual memory instead of loading everything into RAM at once:model = word2vec.KeyedVectors.load_word2vec_format( "GoogleNews-vectors-negative300.bin", binary=True, mmap='r' )Note: This will make vector access slightly slower, but it's a lifesaver when RAM is tight.
One last thing: Make sure you're using the latest version of Gensim—older versions had less efficient memory handling. Upgrade with:
pip install --upgrade gensim
内容的提问来源于stack exchange,提问作者Mahsa

