使用tf.contrib.learn.VocabularyProcessor后词汇量莫名减少的问题
tf.contrib.learn.VocabularyProcessor First, let's recap your scenario to make sure I'm aligned: you trained a Word2Vec model that generated a vocab list of 415,657 unique words, but when feeding this list into tf.contrib.learn.preprocessing.VocabularyProcessor.fit(), the resulting vocabulary only contained 412,722 words. You've already ruled out duplicates and case sensitivity, so let's dig into the most likely causes and fixes.
Why This Is Happening
Here are the top probable reasons for the missing ~2,900 words:
1. Empty or Whitespace-Only Entries
The biggest culprit here is almost certainly hidden empty strings or entries made up entirely of whitespace (spaces, tabs, newlines) in your original vocab list. VocabularyProcessor automatically ignores these during fitting—they don't get added to the vocabulary, and you won't receive a warning about it.
2. Non-Printable/Control Characters
Another possibility is that some vocab entries contain non-printable control characters (like \x00, \b, or unencoded special symbols). The default tokenizer used by VocabularyProcessor may skip these tokens entirely, as it's optimized for standard readable text.
3. Unexpected Tokenization (Less Likely)
You mentioned your vocab has no duplicates, but let's rule this out: if any entries are multi-word phrases (e.g., "machine learning" stored as a single string), the default tokenizer would split them into individual words. However, this would increase your vocab size, not decrease it—so this isn't your issue here.
How to Diagnose the Missing Words
Let's write a quick script to identify exactly which words are being excluded. This will confirm the root cause immediately:
# Convert both vocabularies to sets for easy comparison processed_vocab = set(vocab_processor.vocabulary_._mapping.keys()) original_vocab = set(vocab) # Find words present in the original but missing from the processed vocab missing_words = original_vocab - processed_vocab print(f"Total missing words: {len(missing_words)}") print("Sample of missing words:", list(missing_words)[:20])
Run this, and you'll spot patterns right away—are the missing entries blank? Full of spaces? Loaded with unreadable characters? That tells you exactly what to fix.
Fixes to Try
Based on the most common issues, here are targeted solutions:
1. Clean Your Original Vocab First
Filter out empty strings and whitespace-only entries before passing the list to fit():
# Strip whitespace from each word, and keep only non-empty entries cleaned_vocab = [word.strip() for word in vocab if word.strip()] # Now fit with the cleaned list pretrain = vocab_processor.fit(cleaned_vocab)
2. Use a Custom Tokenizer to Preserve All Entries
If you want to ensure every entry in your original vocab is treated as a single token (even if it has spaces or special characters), override the default tokenizer with an identity function:
def identity_tokenizer(document): # Treat each entire document (a single word in your case) as one token return [document] # Initialize the processor with your custom tokenizer vocab_processor = tf.contrib.learn.preprocessing.VocabularyProcessor( max_document_length, tokenizer_fn=identity_tokenizer ) # Fit with your original vocab pretrain = vocab_processor.fit(vocab)
3. Filter Out Non-Printable Characters
If your missing words have weird control characters, use regex to strip them out:
import re # Keep only printable ASCII characters (adjust if you need non-ASCII support) cleaned_vocab = [] for word in vocab: cleaned_word = re.sub(r'[^\x20-\x7E]', '', word).strip() if cleaned_word: cleaned_vocab.append(cleaned_word) pretrain = vocab_processor.fit(cleaned_vocab)
Final Thoughts
9 times out of 10, the issue is empty or whitespace-only entries in your original vocab list. Running the diagnostic script will confirm this quickly, and cleaning your vocab should resolve the size discrepancy entirely.
内容的提问来源于stack exchange,提问作者Yuguang

