aclImdb数据集上预训练静态词向量LSTM实现与报错排查
任务规划
- 阶段1:在aclImdb电影评论数据集上,实现搭载预训练静态词向量的LSTM情感分类模型
- 阶段2:基于Transformer架构开发机器翻译模型
已执行操作流程
1. 安装依赖库
!pip install ktrain !pip install tensorflow_text
2. 导入所需依赖
import pathlib import random import numpy as np from typing import Tuple, List import matplotlib.pyplot as plt import matplotlib.ticker as ticker from sklearn.model_selection import train_test_split # tensorflow imports import tensorflow as tf from tensorflow import keras from tensorflow.keras.layers import ( TextVectorization, LSTM, Dense, Embedding, Dropout, Layer, Input, MultiHeadAttention, LayerNormalization, GlobalMaxPool1D) from tensorflow.keras.models import Sequential, Model from tensorflow.keras.initializers import Constant from tensorflow.keras.optimizers import Adam from tensorflow.keras import backend as K import tensorflow_text as tf_text import ktrain from ktrain import text
3. 下载并解压数据集
使用斯坦福大学公开的aclImdb大型电影评论情感数据集:
!wget https://ai.stanford.edu/~amaas/data/sentiment/aclImdb_v1.tar.gz !tar -xzf aclImdb_v1.tar.gz
4. 初始数据预处理
调用ktrain接口从文件夹加载训练、测试数据:
%reload_ext autoreload %autoreload 2 %matplotlib inline import os os.environ["CUDA_DEVICE_ORDER"]="PCI_BUS_ID"; os.environ["CUDA_VISIBLE_DEVICES"]="0"; DATADIR ='/content/aclImdb' trn, val, preproc = text.texts_from_folder( DATADIR, max_features=20000, maxlen=400, ngram_range=1, preprocess_mode='standard', train_test_names=['train', 'test'], classes=['pos', 'neg'] )
5. 初始LSTM模型代码(存在错误)
预期模型结构:Embedding层→至少1层LSTM→Dropout正则层→Dense输出层,使用分类交叉熵损失、Adam优化器,监控分类准确率指标,初始编写代码如下:
K.clear_session() def build_LSTM_model( embedding_size: int, total_words: int, lstm_hidden_size: int, dropout_rate: float) -> Sequential: model.add(Embedding(input_dim = total_words,output_dim=embedding_size,input_length=total_words)) model.add(LSTM(lstm_hidden_size,return_sequences=True,name="lstm_layer")) model.add(GlobalMaxPool1D()) model.add(Dropout(dropout_rate)) model.add(Dense(MAX_SEQUENCE_LEN, activation="relu")) model.compile(loss='CategoricalCrossentropy', optimizer=Adam(lr=0.01), metrics=['CategoricalAccuracy']) model.summary() model = Sequential()
6. 初始训练流程代码(存在错误)
计划用ktrain封装训练流程,通过学习率查找工具选定最优学习率,控制训练轮次平衡效率和效果,初始编写代码如下:
learner: ktrain.Learner model = text.text_classifier('bert', trn , preproc=preproc) learner.lr_find() learner.lr_plot() learner.fit_onecycle(1e-4, 1)
运行报错信息
执行上述训练代码时触发如下错误:
ValueError Traceback (most recent call last) <ipython-input-12-3320d887c22c> in <module>() 6 # workers=8, use_multiprocessing=False, batch_size=64) 7 ----> 8 model = text.text_classifier('bert', trn , preproc=preproc) 9 10 # learner.lr_find() 1 frames /usr/local/lib/python3.7/dist-packages/ktrain/text/models.py in _text_model(name, train_data, preproc, multilabel, classification, metrics, verbose) 109 raise ValueError( 110 "if '%s' is selected model, then preprocess_mode='%s' should be used and vice versa" --> 111 % (BERT, BERT) 112 ) 113 is_huggingface = U.is_huggingface(data=train_data) ValueError: if 'bert' is selected model, then preprocess_mode='bert' should be used and vice versa
问题原因与修复方案
问题原因
- BERT接口调用不匹配:ktrain框架要求模型和预处理模式必须一一对应,当前预处理时设置的
preprocess_mode='standard'是为普通词嵌入+RNN/CNN类模型准备的,生成的词表、输入格式完全不符合BERT模型要求,却调用了text.text_classifier('bert')接口,直接触发参数校验报错。 - LSTM模型代码逻辑错误:
- 未先初始化
Sequential()实例就直接调用model.add()添加层,会触发变量未定义错误 - 层顺序颠倒,模型实例初始化放在了所有层添加、编译操作之后
- 输出层维度配置错误,二分类任务最终Dense层应设置为2个单元搭配softmax激活,代码里错误配置为序列长度、使用relu激活
- 缺失预训练静态词向量加载逻辑,不符合最终目标要求
- 未导入
GlobalMaxPool1D、Adam等依赖组件
- 未先初始化
修复实现(预训练静态词向量+LSTM情感分类)
由于最终目标是实现预训练静态词向量搭配LSTM的分类模型,不需要调用BERT接口,按如下步骤修正即可:
- 补充导入缺失的层和优化器组件(已在前面的导入代码中补充
GlobalMaxPool1D、Adam) - 修正LSTM模型构建逻辑,加载预训练词向量(以常用的GloVe 100维静态词向量为例),修正层顺序和输出配置:
# 加载预训练GloVe词向量,构建词向量矩阵 def load_glove_embeddings(word_index, embedding_dim=100, max_words=20000): embeddings_index = {} # 提前下载并解压glove.6B.100d.txt到对应路径 with open('glove.6B.100d.txt') as f: for line in f: word, coefs = line.split(maxsplit=1) coefs = np.fromstring(coefs, 'f', sep=' ') embeddings_index[word] = coefs # 构建嵌入矩阵 embedding_matrix = np.zeros((max_words, embedding_dim)) for word, i in word_index.items(): if i < max_words: embedding_vector = embeddings_index.get(word) if embedding_vector is not None: embedding_matrix[i] = embedding_vector return embedding_matrix K.clear_session() # 配置参数 MAX_WORDS = 20000 MAX_SEQ_LEN = 400 EMBEDDING_DIM = 100 LSTM_HIDDEN = 128 DROPOUT_RATE = 0.3 # 先初始化模型实例 model = Sequential() # 加载预训练词向量权重 embedding_matrix = load_glove_embeddings(preproc.get_word_index(), EMBEDDING_DIM, MAX_WORDS) model.add(Embedding( input_dim=MAX_WORDS, output_dim=EMBEDDING_DIM, input_length=MAX_SEQ_LEN, embeddings_initializer=Constant(embedding_matrix), trainable=False # 静态词向量设置为不训练,若要微调可设为True )) model.add(LSTM(LSTM_HIDDEN, return_sequences=True, name="lstm_layer")) model.add(GlobalMaxPool1D()) model.add(Dropout(DROPOUT_RATE)) model.add(Dense(2, activation="softmax")) # 二分类输出2维,softmax激活 model.compile( loss='categorical_crossentropy', optimizer=Adam(learning_rate=0.001), metrics=['CategoricalAccuracy'] ) model.summary()
- 修正训练流程,用ktrain封装自定义LSTM模型,执行学习率查找和训练:
# 封装自定义LSTM模型,不要调用bert分类器接口 learner = ktrain.get_learner(model, train_data=trn, val_data=val, batch_size=32) # 查找最优学习率 learner.lr_find() learner.lr_plot() # 用选定的学习率训练,示例设置3轮 learner.fit_onecycle(1e-3, 3)
内容的提问来源于stack exchange,提问作者Mohammed
相关产品推荐
相关产品推荐

