You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

NLP情感分析模型训练遇UnimplementedError: Graph Execution错误求助

解决TensorFlow训练时的UnimplementedError: Graph Execution错误

问题描述

我在Google Colab上开发一个NLP情感分析模型,用于根据推文内容将情感分为中性、正面、负面三类。已使用Tokenizer模块将训练集和测试集数据转换为整数序列,这部分严格遵循TensorFlow官方YouTube教程操作,应该不存在问题。但启动模型训练时,反复遇到UnimplementedError: Graph Execution错误,尝试调整模型层结构、缩小数据集规模后,错误依然出现。需要解释该错误的含义,并定位代码中的问题。

错误截图:
UnimplementedError错误截图

错误原因分析

这个错误的核心是TensorFlow无法构建有效的计算图来执行训练,根源在于代码中存在多处与任务不匹配的配置:

  • 标签类型不兼容:当前训练标签是字符串格式(如"positive"),但TensorFlow的损失函数无法直接处理字符串标签,必须转换为数值类型。
  • 任务类型与模型配置不匹配:你的任务是三分类,但模型最后一层用了sigmoid激活+binary_crossentropy损失,这是二分类任务的配置,完全不适用于三分类场景。
  • Embedding层参数不一致:Tokenizer设置了num_words=1000,但Embedding层的输入维度设为10000,两者参数不匹配,会导致输入序列的词汇索引无法正确映射。

修复步骤及修正代码

1. 标签数值化处理

将字符串标签转换为整数编码,比如:negative=0、neutral=1、positive=2。可以用sklearn.preprocessing.LabelEncoder实现。

2. 调整模型配置适配三分类

  • 修改Embedding层的输入维度与Tokenizer的num_words一致
  • 输出层改为3个神经元,激活函数用softmax
  • 损失函数改用sparse_categorical_crossentropy(因为标签是整数编码,无需独热)

完整修正后的代码

import os
import sys
import tensorflow as tf
import numpy as np
import pandas as pd
from tensorflow import keras
from tensorflow.keras.preprocessing.text import Tokenizer
from tensorflow.keras.preprocessing.sequence import pad_sequences
from sklearn.preprocessing import LabelEncoder

device_name = tf.test.gpu_device_name()
if len(device_name) > 0:
    print("Found GPU at: {}".format(device_name))
else:
    device_name = "/device:CPU:0"
    print("No GPU, using {}.".format(device_name))
# Load dataset into a dataframe
train_data_path = "/content/drive/MyDrive/ML Datasets/tweet_sentiment_analysis/train.csv"
test_data_path = "/content/drive/MyDrive/ML Datasets/tweet_sentiment_analysis/test.csv"

train_df = pd.read_csv(train_data_path, encoding='unicode_escape')
test_df = pd.read_csv(test_data_path, encoding='unicode_escape')
# Function to convert df into a list of strings, and encode labels
def convert_and_encode_labels(df, text_col):
    selected_text_list = []
    labels = []

    for index, row in df.iterrows():
        selected_text_list.append(str(row[text_col]))
        labels.append(row['sentiment'])
    
    # 标签数值化
    le = LabelEncoder()
    encoded_labels = le.fit_transform(labels)
    
    return np.array(selected_text_list), np.array(encoded_labels)


train_sentences, train_labels = convert_and_encode_labels(train_df, 'selected_text')
test_sentences, test_labels = convert_and_encode_labels(test_df, 'text')

print(train_sentences)
print(train_labels)
# Instantiate tokenizer and create word_index
num_words = 1000
tokenizer = Tokenizer(num_words=num_words, oov_token='<oov>')
tokenizer.fit_on_texts(train_sentences)
word_index = tokenizer.word_index

# Convert sentences into a sequence 
train_sequence = tokenizer.texts_to_sequences(train_sentences)
test_sequence = tokenizer.texts_to_sequences(test_sentences)

# Padding sequences 
pad_test_seq = pad_sequences(test_sequence, padding='post')
max_len = pad_test_seq[0].size
pad_train_seq = pad_sequences(train_sequence, padding='post', maxlen=max_len)
model = tf.keras.Sequential([
    tf.keras.layers.Embedding(num_words, 24, input_length=max_len),
    tf.keras.layers.GlobalAveragePooling1D(),
    tf.keras.layers.Dense(24, activation='relu'),
    tf.keras.layers.Dense(3, activation='softmax')  # 三分类输出3个神经元,softmax激活
])

with tf.device(device_name):
    model.compile(loss='sparse_categorical_crossentropy', optimizer='adam', metrics=['accuracy'])
num_epochs = 20

with tf.device(device_name):
    history = model.fit(pad_train_seq, train_labels, epochs=num_epochs, validation_data=(pad_test_seq, test_labels), verbose=2)

额外说明

  • 如果希望使用独热编码标签,可以用tf.keras.utils.to_categorical处理编码后的标签,此时损失函数要改为categorical_crossentropy。
  • 训练前建议检查train_labels和test_labels的数值范围,确保是0-2的整数,避免出现维度不匹配的问题。

内容的提问来源于stack exchange,提问作者zolotl

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 13:10:24