You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

调用gensim.models.Word2Vec()遇TypeError:仅可拼接元组与元组

Fixing TypeError: can only concatenate tuple (not "str") to tuple in gensim Word2Vec

This error almost always boils down to one thing: your input corpus format doesn't match what Word2Vec expects. Let's break down why this happens and how to fix it.

Why the error occurs

Word2Vec requires your sentences parameter to be an iterable (like a list, generator, etc.) where each element is a list of string words. When you pass something that doesn't fit this structure—like raw strings, mixed tuples and strings, or sentences that aren't split into word lists—Gensim's internal code tries to process these invalid elements and hits that tuple/string concatenation error during the build_vocab step.

Step-by-step fixes

1. Verify your corpus structure

First, double-check what you're passing to Word2Vec. A valid corpus looks like this:

# Correct format: list of word lists
valid_sentences = [
    ["the", "cat", "sat", "on", "the", "mat"],
    ["dogs", "love", "chasing", "squirrels"]
]

Common invalid formats that trigger this error:

  • Raw strings instead of word lists: ["the cat sat on the mat", "dogs love chasing squirrels"]
  • Mixed tuples and strings: (["the", "cat"], "sat on the mat")
  • Single word strings instead of lists: "the cat sat on the mat"

2. Preprocess your corpus to fit the required format

If your raw data is in a text file (one sentence per line), process it into word lists first:

from gensim.models import Word2Vec

# Process text file into valid sentences
sentences = []
with open("your_corpus.txt", "r", encoding="utf-8") as f:
    for line in f:
        # Strip whitespace, split into words, add any preprocessing (lowercase, remove punctuation)
        cleaned_line = line.strip().lower()
        word_list = cleaned_line.split()  # Use a proper tokenizer (like NLTK) for better results
        sentences.append(word_list)

# Now initialize Word2Vec correctly
model = Word2Vec(
    sentences,
    vector_size=100,
    window=5,
    min_count=1,
    workers=4
)

If you're using a generator (for large corporas), make sure each yield returns a word list:

def corpus_generator():
    with open("large_corpus.txt", "r", encoding="utf-8") as f:
        for line in f:
            yield line.strip().lower().split()

model = Word2Vec(corpus_generator(), vector_size=100, window=5, min_count=1, workers=4)

3. Check your trim_rule (if you're using one)

If you passed a custom trim_rule parameter, ensure it returns a tuple of (word, count, min_count). If your function accidentally returns a string or another invalid type, that can also trigger this error. For example, a valid trim_rule looks like:

def custom_trim_rule(word, count, min_count):
    # Keep words longer than 3 characters
    if len(word) > 3:
        return gensim.utils.RULE_KEEP
    else:
        return gensim.utils.RULE_DISCARD

If you're still stuck, sharing a snippet of how you're building your corpus would help pinpoint the exact issue!

内容的提问来源于stack exchange,提问作者isaacsultan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 08:34:28