调用gensim.models.Word2Vec()遇TypeError:仅可拼接元组与元组
This error almost always boils down to one thing: your input corpus format doesn't match what Word2Vec expects. Let's break down why this happens and how to fix it.
Why the error occurs
Word2Vec requires your sentences parameter to be an iterable (like a list, generator, etc.) where each element is a list of string words. When you pass something that doesn't fit this structure—like raw strings, mixed tuples and strings, or sentences that aren't split into word lists—Gensim's internal code tries to process these invalid elements and hits that tuple/string concatenation error during the build_vocab step.
Step-by-step fixes
1. Verify your corpus structure
First, double-check what you're passing to Word2Vec. A valid corpus looks like this:
# Correct format: list of word lists valid_sentences = [ ["the", "cat", "sat", "on", "the", "mat"], ["dogs", "love", "chasing", "squirrels"] ]
Common invalid formats that trigger this error:
- Raw strings instead of word lists:
["the cat sat on the mat", "dogs love chasing squirrels"] - Mixed tuples and strings:
(["the", "cat"], "sat on the mat") - Single word strings instead of lists:
"the cat sat on the mat"
2. Preprocess your corpus to fit the required format
If your raw data is in a text file (one sentence per line), process it into word lists first:
from gensim.models import Word2Vec # Process text file into valid sentences sentences = [] with open("your_corpus.txt", "r", encoding="utf-8") as f: for line in f: # Strip whitespace, split into words, add any preprocessing (lowercase, remove punctuation) cleaned_line = line.strip().lower() word_list = cleaned_line.split() # Use a proper tokenizer (like NLTK) for better results sentences.append(word_list) # Now initialize Word2Vec correctly model = Word2Vec( sentences, vector_size=100, window=5, min_count=1, workers=4 )
If you're using a generator (for large corporas), make sure each yield returns a word list:
def corpus_generator(): with open("large_corpus.txt", "r", encoding="utf-8") as f: for line in f: yield line.strip().lower().split() model = Word2Vec(corpus_generator(), vector_size=100, window=5, min_count=1, workers=4)
3. Check your trim_rule (if you're using one)
If you passed a custom trim_rule parameter, ensure it returns a tuple of (word, count, min_count). If your function accidentally returns a string or another invalid type, that can also trigger this error. For example, a valid trim_rule looks like:
def custom_trim_rule(word, count, min_count): # Keep words longer than 3 characters if len(word) > 3: return gensim.utils.RULE_KEEP else: return gensim.utils.RULE_DISCARD
If you're still stuck, sharing a snippet of how you're building your corpus would help pinpoint the exact issue!
内容的提问来源于stack exchange,提问作者isaacsultan

