You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Keras实现情感分析遇AttributeError:'str'对象无'ndim'属性

嘿,这个错误我之前踩过坑!本质是你的模型拿到了原始文本字符串,但Keras的层(比如Embedding、Dense)只认带ndim属性的张量/数组类型数据,字符串肯定没有这个属性呀~ 毕竟示例用的Keras内置数据集已经帮你做好了文本转数值的预处理,而你用本地txt文件时刚好漏掉了这个关键环节。

问题根源拆解

AttributeError: 'str' object has no attribute 'ndim' 直白点说:你给模型喂了一堆字符串句子,但模型需要的是数字组成的矩阵/序列。Keras内置数据集(比如IMDB)默认已经完成了文本→整数序列→统一长度的步骤,而你用本地pos.txt/neg.txt时,得自己手动完成这套预处理流程。

完整解决流程(适配本地文本)

我给你整理了一套能直接对接本地txt文件的代码流程,你可以对照自己的代码找问题:

1. 先加载本地文本并打标签

先把两个文件里的评论读进来,同时给正面评论打1、负面打0:

import numpy as np

# 加载正面评论(跳过空行)
with open('pos.txt', 'r', encoding='utf-8') as f:
    pos_reviews = [line.strip() for line in f if line.strip()]
# 加载负面评论
with open('neg.txt', 'r', encoding='utf-8') as f:
    neg_reviews = [line.strip() for line in f if line.strip()]

# 合并所有文本和标签
texts = pos_reviews + neg_reviews
labels = np.array([1]*len(pos_reviews) + [0]*len(neg_reviews))

2. 核心:把文本转成模型能懂的数字序列

这一步是解决错误的关键,用Keras的Tokenizer完成文本数值化:

from tensorflow.keras.preprocessing.text import Tokenizer
from tensorflow.keras.preprocessing.sequence import pad_sequences

# 初始化Tokenizer,设置最大词汇量(按需调整)
max_words = 10000
tokenizer = Tokenizer(num_words=max_words)
tokenizer.fit_on_texts(texts)  # 学习文本中的词汇表

# 把每条评论转成整数序列
sequences = tokenizer.texts_to_sequences(texts)

# 统一所有序列的长度(模型要求输入形状一致)
max_len = 200  # 可以根据你的评论平均长度调整
x_data = pad_sequences(sequences, maxlen=max_len)

3. 划分训练/测试集

from sklearn.model_selection import train_test_split

x_train, x_test, y_train, y_test = train_test_split(
    x_data, labels, test_size=0.2, random_state=42
)

4. 构建并训练模型

现在用处理好的x_train(数组类型)喂模型就不会报错了,比如一个简单的LSTM情感分析模型:

from tensorflow.keras.models import Sequential
from tensorflow.keras.layers import Embedding, LSTM, Dense

model = Sequential([
    Embedding(max_words, 128, input_length=max_len),  # 输入是统一长度的整数序列
    LSTM(64),
    Dense(1, activation='sigmoid')
])

model.compile(optimizer='adam', loss='binary_crossentropy', metrics=['accuracy'])
model.fit(x_train, y_train, epochs=5, batch_size=32, validation_split=0.1)
快速排查你的代码

你可以先检查这几个关键点:

  • 是不是直接把原始的字符串列表传给了model.fit()?这是最常见的错误
  • 有没有用Tokenizer把文本转换成整数序列?
  • 有没有用pad_sequences统一所有输入的长度?

如果还是卡壳,可以把你预处理部分的代码贴出来,我帮你再细瞅~

内容的提问来源于stack exchange,提问作者Amy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 04:19:59