You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

aclImdb数据集上预训练静态词向量LSTM实现与报错排查

任务规划
  • 阶段1:在aclImdb电影评论数据集上,实现搭载预训练静态词向量的LSTM情感分类模型
  • 阶段2:基于Transformer架构开发机器翻译模型

已执行操作流程

1. 安装依赖库

!pip install ktrain
!pip install tensorflow_text

2. 导入所需依赖

import pathlib
import random
import numpy as np
from typing import Tuple, List

import matplotlib.pyplot as plt
import matplotlib.ticker as ticker
from sklearn.model_selection import train_test_split

# tensorflow imports
import tensorflow as tf
from tensorflow import keras
from tensorflow.keras.layers import (
    TextVectorization, LSTM, Dense, Embedding, Dropout,
    Layer, Input, MultiHeadAttention, LayerNormalization, GlobalMaxPool1D)
from tensorflow.keras.models import Sequential, Model
from tensorflow.keras.initializers import Constant
from tensorflow.keras.optimizers import Adam
from tensorflow.keras import backend as K
import tensorflow_text as tf_text 
import ktrain
from ktrain import text

3. 下载并解压数据集

使用斯坦福大学公开的aclImdb大型电影评论情感数据集:

!wget https://ai.stanford.edu/~amaas/data/sentiment/aclImdb_v1.tar.gz
!tar -xzf aclImdb_v1.tar.gz

4. 初始数据预处理

调用ktrain接口从文件夹加载训练、测试数据:

%reload_ext autoreload
%autoreload 2
%matplotlib inline
import os
os.environ["CUDA_DEVICE_ORDER"]="PCI_BUS_ID";
os.environ["CUDA_VISIBLE_DEVICES"]="0";

DATADIR ='/content/aclImdb'
trn, val, preproc = text.texts_from_folder(
    DATADIR,
    max_features=20000, 
    maxlen=400, 
    ngram_range=1,                                               
    preprocess_mode='standard', 
    train_test_names=['train', 'test'],
    classes=['pos', 'neg']
)

5. 初始LSTM模型代码(存在错误)

预期模型结构:Embedding层→至少1层LSTM→Dropout正则层→Dense输出层,使用分类交叉熵损失、Adam优化器,监控分类准确率指标,初始编写代码如下:

K.clear_session()    
def build_LSTM_model(
        embedding_size: int,
        total_words: int,
        lstm_hidden_size: int,
        dropout_rate: float) -> Sequential:
    model.add(Embedding(input_dim = total_words,output_dim=embedding_size,input_length=total_words))
    model.add(LSTM(lstm_hidden_size,return_sequences=True,name="lstm_layer"))
    model.add(GlobalMaxPool1D())
    model.add(Dropout(dropout_rate))
    model.add(Dense(MAX_SEQUENCE_LEN, activation="relu"))
    model.compile(loss='CategoricalCrossentropy', optimizer=Adam(lr=0.01), metrics=['CategoricalAccuracy'])
    model.summary()
    model = Sequential()

6. 初始训练流程代码(存在错误)

计划用ktrain封装训练流程,通过学习率查找工具选定最优学习率,控制训练轮次平衡效率和效果,初始编写代码如下:

learner: ktrain.Learner
model = text.text_classifier('bert', trn , preproc=preproc)
learner.lr_find()
learner.lr_plot()
learner.fit_onecycle(1e-4, 1)

运行报错信息

执行上述训练代码时触发如下错误:

ValueError                                Traceback (most recent call last)
<ipython-input-12-3320d887c22c> in <module>()
      6 #                              workers=8, use_multiprocessing=False, batch_size=64)
      7
----> 8 model = text.text_classifier('bert', trn , preproc=preproc)
      9 
     10 # learner.lr_find()
1 frames
/usr/local/lib/python3.7/dist-packages/ktrain/text/models.py in _text_model(name, train_data, preproc, multilabel, classification, metrics, verbose)
    109         raise ValueError(
    110             "if '%s' is selected model, then preprocess_mode='%s' should be used and vice versa"
--> 111             % (BERT, BERT)
    112         )    
    113     is_huggingface = U.is_huggingface(data=train_data)
ValueError: if 'bert' is selected model, then preprocess_mode='bert' should be used and vice versa

问题原因与修复方案

问题原因

  1. BERT接口调用不匹配:ktrain框架要求模型和预处理模式必须一一对应,当前预处理时设置的preprocess_mode='standard'是为普通词嵌入+RNN/CNN类模型准备的,生成的词表、输入格式完全不符合BERT模型要求,却调用了text.text_classifier('bert')接口,直接触发参数校验报错。
  2. LSTM模型代码逻辑错误:
    • 未先初始化Sequential()实例就直接调用model.add()添加层,会触发变量未定义错误
    • 层顺序颠倒,模型实例初始化放在了所有层添加、编译操作之后
    • 输出层维度配置错误,二分类任务最终Dense层应设置为2个单元搭配softmax激活,代码里错误配置为序列长度、使用relu激活
    • 缺失预训练静态词向量加载逻辑,不符合最终目标要求
    • 未导入GlobalMaxPool1D、Adam等依赖组件

修复实现(预训练静态词向量+LSTM情感分类)

由于最终目标是实现预训练静态词向量搭配LSTM的分类模型,不需要调用BERT接口,按如下步骤修正即可:

  1. 补充导入缺失的层和优化器组件(已在前面的导入代码中补充GlobalMaxPool1D、Adam)
  2. 修正LSTM模型构建逻辑,加载预训练词向量(以常用的GloVe 100维静态词向量为例),修正层顺序和输出配置:
# 加载预训练GloVe词向量,构建词向量矩阵
def load_glove_embeddings(word_index, embedding_dim=100, max_words=20000):
    embeddings_index = {}
    # 提前下载并解压glove.6B.100d.txt到对应路径
    with open('glove.6B.100d.txt') as f:
        for line in f:
            word, coefs = line.split(maxsplit=1)
            coefs = np.fromstring(coefs, 'f', sep=' ')
            embeddings_index[word] = coefs
    # 构建嵌入矩阵
    embedding_matrix = np.zeros((max_words, embedding_dim))
    for word, i in word_index.items():
        if i < max_words:
            embedding_vector = embeddings_index.get(word)
            if embedding_vector is not None:
                embedding_matrix[i] = embedding_vector
    return embedding_matrix

K.clear_session()
# 配置参数
MAX_WORDS = 20000
MAX_SEQ_LEN = 400
EMBEDDING_DIM = 100
LSTM_HIDDEN = 128
DROPOUT_RATE = 0.3

# 先初始化模型实例
model = Sequential()
# 加载预训练词向量权重
embedding_matrix = load_glove_embeddings(preproc.get_word_index(), EMBEDDING_DIM, MAX_WORDS)
model.add(Embedding(
    input_dim=MAX_WORDS,
    output_dim=EMBEDDING_DIM,
    input_length=MAX_SEQ_LEN,
    embeddings_initializer=Constant(embedding_matrix),
    trainable=False # 静态词向量设置为不训练,若要微调可设为True
))
model.add(LSTM(LSTM_HIDDEN, return_sequences=True, name="lstm_layer"))
model.add(GlobalMaxPool1D())
model.add(Dropout(DROPOUT_RATE))
model.add(Dense(2, activation="softmax")) # 二分类输出2维,softmax激活
model.compile(
    loss='categorical_crossentropy',
    optimizer=Adam(learning_rate=0.001),
    metrics=['CategoricalAccuracy']
)
model.summary()
  1. 修正训练流程,用ktrain封装自定义LSTM模型,执行学习率查找和训练:
# 封装自定义LSTM模型,不要调用bert分类器接口
learner = ktrain.get_learner(model, train_data=trn, val_data=val, batch_size=32)
# 查找最优学习率
learner.lr_find()
learner.lr_plot()
# 用选定的学习率训练,示例设置3轮
learner.fit_onecycle(1e-3, 3)

内容的提问来源于stack exchange,提问作者Mohammed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 09:30:44