You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决AttributeError: 'float'对象无'lower'属性错误

错误诊断与解决

错误根源

执行x.lower()时触发AttributeError,核心原因是Plot列中存在非字符串类型的值(大概率是NaN缺失值,被pandas解析为float类型),而float对象没有lower()方法。

分步修复方案

1. 统一文本列数据类型并处理缺失值

在文本预处理前,先强制将Plot列所有值转为字符串,并填充缺失值,避免float类型混入:

# 处理缺失值(用空字符串填充NaN)+ 强制转为字符串
data['Plot'] = data['Plot'].fillna('').astype(str)

将这段代码插入到# Clean the data注释之后、原预处理代码之前。

2. 补充遗漏的库导入

代码中使用了Bidirectional、regularizers、Adam但未导入,会触发新的报错,需在导入区补充:

# 新增导入语句
from tensorflow.keras.layers import Bidirectional
from tensorflow.keras import regularizers
from tensorflow.keras.optimizers import Adam

3. 修正训练数据格式错误

原代码中model.fit(X, X, ...)逻辑错误:categorical_crossentropy要求标签为one-hot编码格式,而原X是整数序列。如果是做文本生成任务,需要构建输入-目标对(输入前n个词,目标为下一个词):

# 拆分输入序列和目标词
X_train = X[:, :-1]  # 去掉最后一个词作为模型输入
y_train = X[:, -1]   # 最后一个词作为预测目标
# 对目标词做one-hot编码
y_train = to_categorical(y_train, num_classes=5000)

# 同步修正Embedding层的输入长度
model.add(Embedding(5000, 256, input_length=X_train.shape[1]))

完整修正代码

# Importing the libraries
import numpy as np
import pandas as pd
import tensorflow as tf
from tensorflow.keras.preprocessing.text import Tokenizer
from tensorflow.keras.preprocessing.sequence import pad_sequences
from tensorflow.keras.models import Sequential
from tensorflow.keras.layers import Dense, Embedding, LSTM, SpatialDropout1D, Bidirectional
from tensorflow.keras import regularizers
from tensorflow.keras.optimizers import Adam
from sklearn.model_selection import train_test_split
from tensorflow.keras.utils import to_categorical
import pickle
import re

# Importing the dataset
filename = "MoviePlots.csv"
data = pd.read_csv(filename, encoding= 'unicode_escape')

# Keeping only the neccessary columns
data = data[['Plot']]

# Clean the data
# 处理缺失值并统一为字符串类型
data['Plot'] = data['Plot'].fillna('').astype(str)

data['Plot'] = data['Plot'].apply(lambda x: x.lower())
data['Plot'] = data['Plot'].apply((lambda x: re.sub('[^a-zA-z0-9\s]', '', x)))

# Create the tokenizer
tokenizer = Tokenizer(num_words=5000, split=" ")
tokenizer.fit_on_texts(data['Plot'].values)

# Save the tokenizer
with open('tokenizer.pickle', 'wb') as handle:
    pickle.dump(tokenizer, handle, protocol=pickle.HIGHEST_PROTOCOL)

# Create the sequences
X = tokenizer.texts_to_sequences(data['Plot'].values)
X = pad_sequences(X)

# 构建输入-目标对
X_train = X[:, :-1]
y_train = X[:, -1]
y_train = to_categorical(y_train, num_classes=5000)

# Create the model
model = Sequential()
model.add(Embedding(5000, 256, input_length=X_train.shape[1]))
model.add(Bidirectional(LSTM(256, return_sequences=True, dropout=0.1, recurrent_dropout=0.1)))
model.add(LSTM(256, return_sequences=True, dropout=0.1, recurrent_dropout=0.1))
model.add(LSTM(256, dropout=0.1, recurrent_dropout=0.1))
model.add(Dense(256, activation='relu', kernel_regularizer=regularizers.l2(0.01)))
model.add(Dense(5000, activation='softmax'))

# Compile the model
model.compile(loss='categorical_crossentropy', optimizer=Adam(lr=0.01), metrics=['accuracy'])

# Train the model
model.fit(X_train, y_train, epochs=100, batch_size=128, verbose=1)

# Saving the model
model.save('visioniser.h5')

内容的提问来源于stack exchange,提问作者SIDHANT YADAV

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 19:01:16