You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

创建Multi-index DataFrame时无数值显示的问题求助

问题:Multi-index DataFrame创建后无数值显示

我编写了一段用于创建Multi-index DataFrame的Python代码,但运行后DataFrame中没有显示任何数值,表格全为空值。

原代码如下:

import pandas as pd
import numpy as np

# Define the data
data = {
    ('rf', 'wv_pretrained'): (0.7392722279437006, 0.7412604086615894),
    ('rf', 'wv_custom'): (0.7746309646412634, 0.7762235207436783),
    ('rf', 'glove_pretrained'): (0.7411603158256094, 0.7427841615992615),
    ('rf', 'spacy_pretrained'): (0.731719876416066, 0.7338888018745795),
    ('rf', 'sent_trf'): (0.7229660144181257, 0.7242986991569383),
    ('rf', 'bert_trf'): (0.7126673532440783, 0.7139687043942123),
    ('rf', 'gpt_trf'): (0.7351527634740816, 0.7369294342385289),
    ('rf', 'tfidf'): (0.6920700308959835, 0.6878065672519817),
    ('Logistic_Regression', 'wv_pretrained'): (0.7392722279437006, 0.7412604086615894),
    ('Logistic_Regression', 'wv_custom'): (0.7746309646412634, 0.7762235207436783),
    ('Logistic_Regression', 'glove_pretrained'): (0.7411603158256094, 0.7427841615992615),
    ('Logistic_Regression', 'spacy_pretrained'): (0.731719876416066, 0.7338888018745795),
    ('Logistic_Regression', 'sent_trf'): (0.7229660144181257, 0.7242986991569383),
    ('Logistic_Regression', 'bert_trf'): (0.7126673532440783, 0.7139687043942123),
    ('Logistic_Regression', 'gpt_trf'): (0.7351527634740816, 0.7369294342385289),
    ('Logistic_Regression', 'tfidf'): (0.6920700308959835, 0.6878065672519817)
}

# Create the multi-index
algos = ['rf', 'Logistic_Regression']
embeddings = ['wv_pretrained', 'wv_custom', 'glove_pretrained', 'spacy_pretrained', 'sent_trf', 'bert_trf', 'gpt_trf', 'tfidf']
eval_metrics = ['accuracy', 'f1_score']
idx = pd.MultiIndex.from_product([algos, embeddings], names=['algos', 'Embedding'])
columns = pd.Index(eval_metrics, name='Metrics')

# Create the dataframe
df = pd.DataFrame(data, index=idx, columns=columns)
df

错误原因

问题出在数据字典结构与DataFrame的索引/列匹配逻辑不兼容:

  • 你定义的data字典以(算法, 嵌入方式)作为键,对应的值是(accuracy, f1_score)的元组;
  • 但创建DataFrame时,你指定了index为(算法, 嵌入方式)的MultiIndex,columns为指标列表。此时pandas会尝试用data的键去匹配columns,自然无法匹配,导致所有单元格为空。

解决方案

使用pd.DataFrame.from_dict()并指定orient='index',让字典的键作为行索引,值作为对应列的数值,就能正确匹配数据:

修正后的完整代码

import pandas as pd
import numpy as np

# Define the data
data = {
    ('rf', 'wv_pretrained'): (0.7392722279437006, 0.7412604086615894),
    ('rf', 'wv_custom'): (0.7746309646412634, 0.7762235207436783),
    ('rf', 'glove_pretrained'): (0.7411603158256094, 0.7427841615992615),
    ('rf', 'spacy_pretrained'): (0.731719876416066, 0.7338888018745795),
    ('rf', 'sent_trf'): (0.7229660144181257, 0.7242986991569383),
    ('rf', 'bert_trf'): (0.7126673532440783, 0.7139687043942123),
    ('rf', 'gpt_trf'): (0.7351527634740816, 0.7369294342385289),
    ('rf', 'tfidf'): (0.6920700308959835, 0.6878065672519817),
    ('Logistic_Regression', 'wv_pretrained'): (0.7392722279437006, 0.7412604086615894),
    ('Logistic_Regression', 'wv_custom'): (0.7746309646412634, 0.7762235207436783),
    ('Logistic_Regression', 'glove_pretrained'): (0.7411603158256094, 0.7427841615992615),
    ('Logistic_Regression', 'spacy_pretrained'): (0.731719876416066, 0.7338888018745795),
    ('Logistic_Regression', 'sent_trf'): (0.7229660144181257, 0.7242986991569383),
    ('Logistic_Regression', 'bert_trf'): (0.7126673532440783, 0.7139687043942123),
    ('Logistic_Regression', 'gpt_trf'): (0.7351527634740816, 0.7369294342385289),
    ('Logistic_Regression', 'tfidf'): (0.6920700308959835, 0.6878065672519817)
}

# Create the evaluation metrics columns
eval_metrics = ['accuracy', 'f1_score']
columns = pd.Index(eval_metrics, name='Metrics')

# Create dataframe with correct orientation
df = pd.DataFrame.from_dict(data, orient='index', columns=columns)
# 设置索引名称
df.index.names = ['algos', 'Embedding']

# 如果需要严格保证索引顺序与定义的algos、embeddings一致,可添加以下代码
algos = ['rf', 'Logistic_Regression']
embeddings = ['wv_pretrained', 'wv_custom', 'glove_pretrained', 'spacy_pretrained', 'sent_trf', 'bert_trf', 'gpt_trf', 'tfidf']
idx = pd.MultiIndex.from_product([algos, embeddings], names=['algos', 'Embedding'])
df = df.reindex(idx)

df

代码说明

  1. orient='index':指定字典的键作为DataFrame的行索引,值作为行对应的列数据;
  2. reindex(idx):可选操作,确保行索引的顺序与你预先定义的algos和embeddings顺序完全一致,避免因字典键顺序(Python 3.7+字典有序)带来的潜在问题。

内容的提问来源于stack exchange,提问作者Nishant

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 19:44:55