YouTube评论情感分析程序TypeError报错求助(Python3.10)
问题解决:TypeError: TextEncodeInput must be Union[TextInputSequence, Tuple[InputSequence, InputSequence]]
错误根源
你在analyze_sentiment函数中先调用tokenizer.tokenize(text)将文本拆分为token列表,再把这个列表传给tokenizer.encode(),但encode()方法的参数要求是原始文本字符串(或字符串对),而非token列表,这直接导致了类型错误。encode()方法本身会自动完成tokenization+编码的完整流程,不需要手动提前调用tokenize()。
修复后的analyze_sentiment函数
def analyze_sentiment(text): # 直接传入原始文本,encode会自动完成tokenization步骤 input_ids = tokenizer.encode( text, add_special_tokens=True, padding='longest', truncation=True, max_length=512, return_tensors='tf' ) logits = bert_model.predict(input_ids)[0] sentiment = np.argmax(logits) # 将数字结果转为可读性更强的文本标签(根据你的模型训练标签调整映射关系) sentiment_labels = {0: '负面', 1: '中性', 2: '正面'} return sentiment_labels.get(sentiment, '未知')
补充:实现双模型情感分析(符合你的项目目标)
你代码中已加载Naive-Bayes模型但未调用,以下是同时输出两个模型结果的修改方案:
- 调整
get_comments函数中的循环逻辑:
comments = [] for item in comments_data['items'][:3]: comment_text = item['snippet']['topLevelComment']['snippet']['textDisplay'] # BERT模型情感分析 bert_sentiment = analyze_sentiment(comment_text) # Naive-Bayes模型情感分析 processed_text = preprocess_text(comment_text) nb_features = extract_features(processed_text) nb_sentiment = naive_bayes_model.classify(nb_features) # 转换为友好标签(根据你的Naive-Bayes训练标签调整) nb_sentiment_label = '正面' if nb_sentiment == 'pos' else '负面' comments.append({ 'comment': comment_text, 'bert_sentiment': bert_sentiment, 'naive_bayes_sentiment': nb_sentiment_label })
- 注意:BERT和Naive-Bayes的预处理逻辑要区分开
你的preprocess_text返回的是带POS标签的元组列表,这是给Naive-Bayes用的;BERT需要的是纯文本字符串,建议新增一个BERT专用的预处理函数:
def preprocess_text_for_bert(text): # 与训练BERT时的预处理逻辑保持一致 text = text.lower() text = re.sub(r'http\S+|www\S+|https\S+', '', text) text = re.sub(r'[%s]' % re.escape(string.punctuation), '', text) text = re.sub(r'[^a-zA-z\s]', '', text) emoji_pattern = re.compile("[" u"\U0001F600-\U0001F64F" # emoticons u"\U0001F300-\U0001F5FF" # symbols & pictographs u"\U0001F680-\U0001F6FF" # transport & map symbols u"\U00002702-\U000027B0" u"\U000024C2-\U0001F251" "]+", flags=re.UNICODE) text = emoji_pattern.sub(r'', text) return text
然后在analyze_sentiment中先调用这个函数处理文本,再传入encode()。
内容的提问来源于stack exchange,提问作者Rinz
相关产品推荐
相关产品推荐

