You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:AttributeError - float对象无lower属性问题排查与解决

AttributeError: 'float' object has no attribute 'lower' 问题解决

错误信息

Traceback (most recent call last):
  File "C:/Users/layan/Desktop", line 39, in <module>
    tokenizer.fit_on_texts(x_train)
  File "C:\Python311\Lib\site-packages\keras\src\preprocessing\text.py", line 293, in fit_on_texts
    seq = text_to_word_sequence(
  File "C:\Python311\Lib\site-packages\keras\src\preprocessing\text.py", line 74, in text_to_word_sequence
    input_text = input_text.lower()
AttributeError: 'float' object has no attribute 'lower'

问题原因

训练集x_train中存在float类型的数据,本质是原数据集df['text']包含缺失值(NaN),pandas会将缺失值默认存储为float类型。而Keras的Tokenizer处理文本时会调用字符串的lower()方法,遇到float类型的NaN就会触发该错误。

另外代码存在一处变量名错误:处理测试集序列时,使用了未定义的testing_sequences变量,后续执行也会报错。

解决方案

  • 处理文本列缺失值:将df['text']中的NaN替换为空字符串,并强制转换为字符串类型,确保所有数据都是Tokenizer可处理的文本格式。
  • 修正变量名错误:处理测试集序列时,使用正确的变量名。

修正后的代码

import pandas as pd
import numpy as np
import seaborn as sns
import re
import nltk

nltk.download(['stopwords','punkt','wordnet','omw-1.4'])
from nltk.corpus import stopwords
from sklearn.model_selection import train_test_split
from keras.preprocessing.text import Tokenizer
from tensorflow.keras.preprocessing.sequence import pad_sequences
from keras.callbacks import ModelCheckpoint, EarlyStopping

# 读取数据集
df = pd.read_csv("C:\\Users\\layan\\Downloads\\fake-news-master\\fake-news-master\\train.csv")

# 关键:处理文本列的缺失值,转换为字符串类型
df['text'] = df['text'].fillna('').astype(str)

# 划分训练集和测试集
x_train, x_test, y_train, y_test = train_test_split(df['text'],
                                                    df['label'],
                                                    test_size=0.2,
                                                    random_state=42)

maxlen=128
truncating='post'
padding= 'post'
oov_tok='<00V>'
vocab_size=1000

tokenizer = Tokenizer(num_words = vocab_size,
                      char_level = False,
                      oov_token = oov_tok)
tokenizer.fit_on_texts(x_train)

# 处理训练集序列
training_sequences = tokenizer.texts_to_sequences(x_train)
training_padded = pad_sequences(training_sequences,
                               maxlen = maxlen,
                               padding = padding,
                               truncating = truncating)

# 处理测试集序列,修正变量名错误
testing_sequences = tokenizer.texts_to_sequences(x_test)
testing_padded = pad_sequences(testing_sequences,
                              maxlen = maxlen,
                              padding = padding,
                              truncating = truncating)

内容的提问来源于stack exchange,提问作者Layan Alahmadi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 16:43:14