You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

YouTube评论情感分析程序TypeError报错求助(Python3.10)

问题解决:TypeError: TextEncodeInput must be Union[TextInputSequence, Tuple[InputSequence, InputSequence]]

错误根源

你在analyze_sentiment函数中先调用tokenizer.tokenize(text)将文本拆分为token列表,再把这个列表传给tokenizer.encode(),但encode()方法的参数要求是原始文本字符串(或字符串对),而非token列表,这直接导致了类型错误。encode()方法本身会自动完成tokenization+编码的完整流程,不需要手动提前调用tokenize()。

修复后的analyze_sentiment函数

def analyze_sentiment(text):
    # 直接传入原始文本,encode会自动完成tokenization步骤
    input_ids = tokenizer.encode(
        text,
        add_special_tokens=True,
        padding='longest',
        truncation=True,
        max_length=512,
        return_tensors='tf'
    )
    logits = bert_model.predict(input_ids)[0]
    sentiment = np.argmax(logits)
    # 将数字结果转为可读性更强的文本标签(根据你的模型训练标签调整映射关系)
    sentiment_labels = {0: '负面', 1: '中性', 2: '正面'}
    return sentiment_labels.get(sentiment, '未知')

补充:实现双模型情感分析(符合你的项目目标)

你代码中已加载Naive-Bayes模型但未调用,以下是同时输出两个模型结果的修改方案:

  1. 调整get_comments函数中的循环逻辑:
comments = []
for item in comments_data['items'][:3]:
    comment_text = item['snippet']['topLevelComment']['snippet']['textDisplay']
    # BERT模型情感分析
    bert_sentiment = analyze_sentiment(comment_text)
    # Naive-Bayes模型情感分析
    processed_text = preprocess_text(comment_text)
    nb_features = extract_features(processed_text)
    nb_sentiment = naive_bayes_model.classify(nb_features)
    # 转换为友好标签(根据你的Naive-Bayes训练标签调整)
    nb_sentiment_label = '正面' if nb_sentiment == 'pos' else '负面'
    
    comments.append({
        'comment': comment_text,
        'bert_sentiment': bert_sentiment,
        'naive_bayes_sentiment': nb_sentiment_label
    })
  1. 注意:BERT和Naive-Bayes的预处理逻辑要区分开
    你的preprocess_text返回的是带POS标签的元组列表,这是给Naive-Bayes用的;BERT需要的是纯文本字符串,建议新增一个BERT专用的预处理函数:
def preprocess_text_for_bert(text):
    # 与训练BERT时的预处理逻辑保持一致
    text = text.lower()
    text = re.sub(r'http\S+|www\S+|https\S+', '', text)
    text = re.sub(r'[%s]' % re.escape(string.punctuation), '', text)
    text = re.sub(r'[^a-zA-z\s]', '', text)
    emoji_pattern = re.compile("["
        u"\U0001F600-\U0001F64F"  # emoticons
        u"\U0001F300-\U0001F5FF"  # symbols & pictographs
        u"\U0001F680-\U0001F6FF"  # transport & map symbols
        u"\U00002702-\U000027B0"
        u"\U000024C2-\U0001F251"
        "]+", flags=re.UNICODE)
    text = emoji_pattern.sub(r'', text)
    return text

然后在analyze_sentiment中先调用这个函数处理文本,再传入encode()。

内容的提问来源于stack exchange,提问作者Rinz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 06:50:36