You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Gensim加载GoogleNews预训练词向量时遭遇内存错误求助

Fixing MemoryError When Loading GoogleNews Word Vectors with Gensim

Hey there, sorry you're stuck with this MemoryError when loading the GoogleNews word vectors—let's break down what's happening and walk through some fixes that should help.

First, a quick context check: the GoogleNews dataset has 3 million 300-dimensional vectors, which adds up to roughly 3.6GB of raw data (using float32). But Gensim uses extra memory during loading for parsing and temporary structures, so if your system's available RAM is lower than before (or you're running other memory-heavy apps now), that's likely why you're hitting this error even though the code worked previously.

Here are your best options to resolve this:

  • Load only the most common words (quick fix)
    You probably don't need all 3 million words for your task. Use the limit parameter to load just the top N most frequent terms—this cuts memory usage drastically. For example, to load the first 1 million words:

    import gensim.models.keyedvectors as word2vec
    model = word2vec.KeyedVectors.load_word2vec_format(
        "GoogleNews-vectors-negative300.bin", 
        binary=True, 
        limit=1000000
    )
    

    This works for most NLP tasks since rare words rarely add meaningful value.

  • Convert to Gensim's native format (more efficient long-term)
    The binary .bin file is great for portability but inefficient to load. If you can temporarily free up enough memory to load the full model once, save it in Gensim's optimized .kv format. This reduces memory usage on subsequent loads and speeds up the process:

    # First load the bin file (free up RAM first if needed)
    model = word2vec.KeyedVectors.load_word2vec_format("GoogleNews-vectors-negative300.bin", binary=True)
    # Save to native format
    model.save("GoogleNews-vectors-negative300.kv")
    # Next time, load from this optimized file
    model = word2vec.KeyedVectors.load("GoogleNews-vectors-negative300.kv")
    
  • Free up system memory
    Chances are your system has less available RAM than when the code worked before. Close any unnecessary apps—like browser tabs, other IDEs, or running ML models. On Linux/macOS, use free -h or top to check memory usage; on Windows, use Task Manager to see what's hogging RAM.

  • Switch to 64-bit Python (if you're on 32-bit)
    32-bit Python has a hard memory limit (usually around 4GB), which is almost exactly what the GoogleNews vectors need. If you accidentally switched to a 32-bit environment, switching back to 64-bit Python will remove this restriction.

  • Memory-map the file (advanced, for full vector access)
    If you absolutely need all 3 million vectors but don't have enough RAM, use the mmap='r' parameter. This maps the file directly to your system's virtual memory instead of loading everything into RAM at once:

    model = word2vec.KeyedVectors.load_word2vec_format(
        "GoogleNews-vectors-negative300.bin", 
        binary=True, 
        mmap='r'
    )
    

    Note: This will make vector access slightly slower, but it's a lifesaver when RAM is tight.

One last thing: Make sure you're using the latest version of Gensim—older versions had less efficient memory handling. Upgrade with:

pip install --upgrade gensim

内容的提问来源于stack exchange,提问作者Mahsa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 09:37:27