如何在自有语料上训练GloVe模型并加载生成的模型文件?
Hey there! Let's work through your GloVe issues step by step—first fixing the training script confusion, then getting your trained vectors loaded properly.
First: Fixing the demo.sh Configuration Issue
Your confusion about CORPUS vs VOCAB_FILE is likely part of why you didn't get valid results earlier. Here's what you need to know:
CORPUS: This should point directly to your raw input text file (corpus.txt), not the sampletext8. The originaldemo.shusestext8as an example, so you must update this variable to your own corpus.VOCAB_FILE: This is an output file generated by thevocab_counttool, not your raw corpus. Setting it tocorpus.txtwas incorrect—you should set it to something likevocab.txt(the script will create this file automatically when it processes your corpus).
To correct your demo.sh:
- Open the script and update:
CORPUS=corpus.txt VOCAB_FILE=vocab.txt - Re-run the script to ensure your
vectors.txtis actually trained on your 900MB corpus (the existingvectors.txtmight be based on the sampletext8if you didn't fix this earlier).
Loading Your Trained GloVe Model (vectors.txt)
Once you have a valid vectors.txt generated from your corpus, you can load it using two common methods:
Method 1: Using Gensim (Simplest Approach)
Gensim has built-in support for loading GloVe vectors, since their format is compatible with Word2Vec's text format. Install gensim if you haven't already (pip install gensim), then use this code:
from gensim.models import KeyedVectors # Load the GloVe vectors glove_model = KeyedVectors.load_word2vec_format( 'vectors.txt', binary=False, # GloVe outputs text format, not binary no_header=False # Set to True if your vectors.txt doesn't have the first line with vocab size + dimension ) # Example usage # Get vector for a word print(glove_model['your_target_word']) # Find most similar words print(glove_model.most_similar('your_target_word'))
Method 2: Manual Loading (No Dependencies)
If you don't want to use gensim, you can write a simple function to load the vectors directly:
def load_glove_vectors(file_path): glove_dict = {} with open(file_path, encoding='utf-8') as f: # Check if the first line contains vocab size and dimension first_line = f.readline() try: vocab_size, embed_dim = map(int, first_line.strip().split()) except ValueError: # If not, reset the file pointer to the start f.seek(0) for line in f: parts = line.strip().split() word = parts[0] vector = [float(val) for val in parts[1:]] glove_dict[word] = vector return glove_dict # Load the model glove_model = load_glove_vectors('vectors.txt') # Example usage print(glove_model.get('your_word', 'Word not found in vocabulary'))
Quick Check
Before loading, verify your vectors.txt is valid: open it and make sure the lines start with words from your corpus followed by floating-point numbers (the vector values). If it still has words from text8, re-run the corrected demo.sh to regenerate it with your corpus.
内容的提问来源于stack exchange,提问作者Codir

