You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

按艺术家类别提取Top C-TF-IDF关键词时索引错误排查

代码错误排查与修复:C-TF-IDF提取类别关键词报错

直接报错原因:列表索引误用字符串

报错list indices must be integers or slices, not str出现在字典推导式代码行:

words_per_class = {artists[label]: [words[index] for index in ctfidf[label].argsort()[-10:]] 
                   for label in docs_per_class.artist}
  • docs_per_class.artist遍历出的是艺术家名称字符串(如"Taylor Swift"这类)
  • artists是列表类型,仅支持整数索引(如artists[0]),用字符串label索引列表必然触发错误

其他潜在问题

  1. 冗余代码:docs = df完全多余,直接使用原DataFrame df即可
  2. CTFIDFVectorizer未定义:sklearn原生库中没有该类,需确保已参考教程完成自定义实现
  3. 索引不匹配:ctfidf的每一行对应docs_per_class中的一个艺术家,应通过行索引关联,而非艺术家名称字符串

修复后的可运行代码

import numpy as np
import pandas as pd
import scipy.sparse as sp

from sklearn.preprocessing import normalize
from sklearn.feature_extraction.text import TfidfTransformer, CountVectorizer

# 自定义实现C-TF-IDF转换器(参考常见教程逻辑)
class CTFIDFVectorizer(TfidfTransformer):
    def __init__(self, *args, **kwargs):
        super().__init__(*args, **kwargs)

    def fit(self, X, n_samples):
        super().fit(X)
        self._idf_diag = sp.diags(np.log((n_samples + 1) / (self.idf_ + 1)), 
                                  offsets=0, shape=(len(self.idf_), len(self.idf_)), 
                                  format='csr')
        return self

    def transform(self, X):
        X = super().transform(X)
        return X * self._idf_diag

# 获取唯一艺术家列表
artists = df['artist'].unique().tolist()

# 按艺术家合并所有歌词
docs_per_class = df.groupby(['artist'], as_index=False).agg({'lyrics': ' '.join})

# 构建词频矩阵
count_vectorizer = CountVectorizer().fit(docs_per_class.lyrics)
count_matrix = count_vectorizer.transform(docs_per_class.lyrics)
words = count_vectorizer.get_feature_names()

# 计算C-TF-IDF矩阵
ctfidf_transformer = CTFIDFVectorizer()
ctfidf_matrix = ctfidf_transformer.fit_transform(count_matrix, n_samples=len(df)).toarray()

# 提取每个艺术家的Top10高得分词汇
words_per_class = {}
for idx, artist_name in enumerate(docs_per_class['artist']):
    # 降序排序取前10个词汇索引,反转后按得分从高到低排列
    top_indices = ctfidf_matrix[idx].argsort()[-10:][::-1]
    top_words = [words[i] for i in top_indices]
    words_per_class[artist_name] = top_words

修复说明

  1. 改用enumerate同时获取行索引与艺术家名称,用整数索引idx访问ctfidf_matrix,规避字符串索引列表的错误
  2. 补充了CTFIDFVectorizer的自定义实现,确保代码可独立运行
  3. 删除冗余代码,调整变量命名提升可读性
  4. 对Top10索引做反转处理,让结果按得分从高到低展示

内容的提问来源于stack exchange,提问作者Jonathan Grenner

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 13:10:36