多标签电影Genre分类器训练报错:TypeError: 'FramePlotMethods'不可迭代
问题分析与解决
首先,咱们来定位你遇到的 TypeError: 'FramePlotMethods' object is not iterable 错误原因:
你在代码最后训练模型和计算准确率时,使用了 train[genre] 和 test[genre],但你从来没有定义过genre这个变量!
这里的坑在于:Pandas 的 DataFrame 有一个内置的 .plot 方法(用于绘制图表),当 Python 找不到 genre 变量时,会误把 train[genre] 解析成 train.plot(这是一个 FramePlotMethods 对象),而这个对象是不可迭代的,无法作为多标签分类的目标输入,所以就抛出了这个错误。
修复步骤
你之前其实已经正确获取了所有电影类型的列名列表——categories(通过 categories = list(df_genres.columns.values) 得到),这正是你多标签分类需要的目标列集合。只需要把代码里的 genre 替换成 categories 即可:
修改后的最后一段代码如下:
from sklearn.model_selection import train_test_split from sklearn.pipeline import Pipeline from sklearn.feature_extraction.text import TfidfVectorizer from sklearn.multiclass import OneVsRestClassifier from sklearn.naive_bayes import MultinomialNB from sklearn.metrics import accuracy_score train, test = train_test_split(df, random_state=42, test_size = 0.33, shuffle=True) x_train = train.plot x_test = test.plot # Define a pipeline combining a text feature extractor with multi label classifier NB_pipeline = Pipeline([ ('tfidf', TfidfVectorizer(stop_words='english')), ('clf', OneVsRestClassifier(MultinomialNB( fit_prior=True, class_prior=None))), ]) # 替换为已定义的categories变量 NB_pipeline.fit(x_train, train[categories]) prediction = NB_pipeline.predict(x_test) # 同样替换为categories accuracy_score(test[categories], prediction)
额外提示
在多标签分类场景下,默认的 accuracy_score 计算方式可能不是最贴合需求的,你可以考虑使用更针对性的评估指标,比如汉明损失、微平均/宏平均 F1 值,来更全面地衡量模型性能:
from sklearn.metrics import hamming_loss, f1_score print("汉明损失:", hamming_loss(test[categories], prediction)) print("微平均F1值:", f1_score(test[categories], prediction, average='micro')) print("宏平均F1值:", f1_score(test[categories], prediction, average='macro'))
内容的提问来源于stack exchange,提问作者Chat Peters
相关产品推荐
相关产品推荐

