如何解决AttributeError: 'float'对象无'lower'属性错误
错误诊断与解决
错误根源
执行x.lower()时触发AttributeError,核心原因是Plot列中存在非字符串类型的值(大概率是NaN缺失值,被pandas解析为float类型),而float对象没有lower()方法。
分步修复方案
1. 统一文本列数据类型并处理缺失值
在文本预处理前,先强制将Plot列所有值转为字符串,并填充缺失值,避免float类型混入:
# 处理缺失值(用空字符串填充NaN)+ 强制转为字符串 data['Plot'] = data['Plot'].fillna('').astype(str)
将这段代码插入到# Clean the data注释之后、原预处理代码之前。
2. 补充遗漏的库导入
代码中使用了Bidirectional、regularizers、Adam但未导入,会触发新的报错,需在导入区补充:
# 新增导入语句 from tensorflow.keras.layers import Bidirectional from tensorflow.keras import regularizers from tensorflow.keras.optimizers import Adam
3. 修正训练数据格式错误
原代码中model.fit(X, X, ...)逻辑错误:categorical_crossentropy要求标签为one-hot编码格式,而原X是整数序列。如果是做文本生成任务,需要构建输入-目标对(输入前n个词,目标为下一个词):
# 拆分输入序列和目标词 X_train = X[:, :-1] # 去掉最后一个词作为模型输入 y_train = X[:, -1] # 最后一个词作为预测目标 # 对目标词做one-hot编码 y_train = to_categorical(y_train, num_classes=5000) # 同步修正Embedding层的输入长度 model.add(Embedding(5000, 256, input_length=X_train.shape[1]))
完整修正代码
# Importing the libraries import numpy as np import pandas as pd import tensorflow as tf from tensorflow.keras.preprocessing.text import Tokenizer from tensorflow.keras.preprocessing.sequence import pad_sequences from tensorflow.keras.models import Sequential from tensorflow.keras.layers import Dense, Embedding, LSTM, SpatialDropout1D, Bidirectional from tensorflow.keras import regularizers from tensorflow.keras.optimizers import Adam from sklearn.model_selection import train_test_split from tensorflow.keras.utils import to_categorical import pickle import re # Importing the dataset filename = "MoviePlots.csv" data = pd.read_csv(filename, encoding= 'unicode_escape') # Keeping only the neccessary columns data = data[['Plot']] # Clean the data # 处理缺失值并统一为字符串类型 data['Plot'] = data['Plot'].fillna('').astype(str) data['Plot'] = data['Plot'].apply(lambda x: x.lower()) data['Plot'] = data['Plot'].apply((lambda x: re.sub('[^a-zA-z0-9\s]', '', x))) # Create the tokenizer tokenizer = Tokenizer(num_words=5000, split=" ") tokenizer.fit_on_texts(data['Plot'].values) # Save the tokenizer with open('tokenizer.pickle', 'wb') as handle: pickle.dump(tokenizer, handle, protocol=pickle.HIGHEST_PROTOCOL) # Create the sequences X = tokenizer.texts_to_sequences(data['Plot'].values) X = pad_sequences(X) # 构建输入-目标对 X_train = X[:, :-1] y_train = X[:, -1] y_train = to_categorical(y_train, num_classes=5000) # Create the model model = Sequential() model.add(Embedding(5000, 256, input_length=X_train.shape[1])) model.add(Bidirectional(LSTM(256, return_sequences=True, dropout=0.1, recurrent_dropout=0.1))) model.add(LSTM(256, return_sequences=True, dropout=0.1, recurrent_dropout=0.1)) model.add(LSTM(256, dropout=0.1, recurrent_dropout=0.1)) model.add(Dense(256, activation='relu', kernel_regularizer=regularizers.l2(0.01))) model.add(Dense(5000, activation='softmax')) # Compile the model model.compile(loss='categorical_crossentropy', optimizer=Adam(lr=0.01), metrics=['accuracy']) # Train the model model.fit(X_train, y_train, epochs=100, batch_size=128, verbose=1) # Saving the model model.save('visioniser.h5')
内容的提问来源于stack exchange,提问作者SIDHANT YADAV
相关产品推荐
相关产品推荐

