You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何可视化含分类变量计数数据的DataFrame并保留多维度信息?

保留刺激、情绪及关联计数的数据分析与可视化方案

1. 热力图(Heatmap)

直接基于原始计数矩阵生成热力图,通过单元格颜色深浅直观呈现关联强度,同时标注具体计数数值,完全保留所有原始信息。

  • 实现逻辑:用seaborn.heatmap,开启annot=True显示单元格数值,搭配合适的颜色映射区分计数高低。
  • 示例代码:
import seaborn as sns
import matplotlib.pyplot as plt

# 假设你的DataFrame名为df,将「刺激」列设为索引
df = df.set_index('刺激')
plt.figure(figsize=(10,6))
sns.heatmap(df, annot=True, fmt='d', cmap='YlGnBu', cbar=True)
plt.title('刺激-情绪响应计数热力图')
plt.xlabel('情绪')
plt.ylabel('刺激')
plt.show()
  • 优势:无信息损耗,可快速定位高频关联对(比如刺激b-兴奋的15次、刺激a-愤怒的10次)。

2. 双向聚类热力图

在热力图基础上对刺激和情绪同时做聚类,既保留原始计数,还能挖掘刺激与情绪的分组规律——比如哪些刺激常引发相似情绪组合,哪些情绪常被同一组刺激触发。

  • 实现逻辑:用seaborn.clustermap,保留annot=True显示数值,通过聚类树展示分组关系。
  • 示例代码:
sns.clustermap(df, annot=True, fmt='d', cmap='YlGnBu', figsize=(10,8))
plt.title('刺激-情绪响应双向聚类热力图')
plt.show()
  • 优势:在保留全量数据的同时,发现潜在的模式(比如刺激a和c因都关联愤怒、沮丧被归为一组)。

3. 桑基图(Sankey Diagram)

用“流量”大小代表计数,展示刺激到情绪的关联流向,清晰呈现单个刺激的情绪分布比例,同时保留具体计数数值。

  • 实现逻辑:先将宽格式数据转换为长格式(刺激、情绪、计数),再用plotly绘制桑基图,链接宽度对应计数,同时标注数值。
  • 示例代码:
import plotly.graph_objects as go
import pandas as pd

# 转换为长格式并过滤0值
long_df = df.reset_index().melt(id_vars='刺激', var_name='情绪', value_name='计数')
long_df = long_df[long_df['计数'] > 0]

# 构建节点与链接映射
nodes = list(set(long_df['刺激'].tolist() + long_df['情绪'].tolist()))
node_indices = {name:i for i,name in enumerate(nodes)}

fig = go.Figure(data=[go.Sankey(
    node = dict(pad=15, thickness=20, label=nodes),
    link = dict(
        source = [node_indices[s] for s in long_df['刺激']],
        target = [node_indices[e] for e in long_df['情绪']],
        value = long_df['计数'].tolist(),
        label = long_df['计数'].astype(str).tolist()
    )
)])
fig.update_layout(title_text='刺激-情绪响应桑基图')
fig.show()
  • 优势:直观展示关联的流向与强度,适合对比不同刺激的情绪分布差异。

4. 气泡图(Bubble Plot)

以X轴为情绪、Y轴为刺激,气泡大小对应计数,颜色可区分刺激或情绪,所有原始信息通过位置与气泡属性完整呈现。

  • 实现逻辑:基于长格式数据,用seaborn.scatterplot设置气泡大小为计数的倍数,增强视觉区分度。
  • 示例代码:
plt.figure(figsize=(10,6))
sns.scatterplot(data=long_df, x='情绪', y='刺激', size='计数', sizes=(50, 500), 
                hue='刺激', alpha=0.7, legend='full')
plt.title('刺激-情绪响应气泡图')
plt.xlabel('情绪')
plt.ylabel('刺激')
plt.legend(bbox_to_anchor=(1.05, 1), loc='upper left')
plt.show()
  • 优势:每个气泡对应一个刺激-情绪对,可快速对比不同关联的强度差异。

5. 增强版对应分析(CA)可视化

针对你之前尝试的对应分析,通过添加原始计数标注,避免丢失关键信息:

  • 实现逻辑:先用prince库计算对应分析结果,再在散点图上为每个刺激-情绪对标注计数。
  • 示例代码:
from prince import CA

ca = CA(n_components=2)
ca.fit(df)
stim_coords = ca.row_coordinates(df)
emo_coords = ca.column_coordinates(df)

plt.figure(figsize=(10,8))
# 绘制刺激与情绪的坐标点
plt.scatter(stim_coords['0'], stim_coords['1'], s=200, label='刺激', color='blue')
for stim, row in stim_coords.iterrows():
    plt.text(row['0'], row['1'], stim, fontsize=12)

plt.scatter(emo_coords['0'], emo_coords['1'], s=200, label='情绪', color='red')
for emo, row in emo_coords.iterrows():
    plt.text(row['0'], row['1'], emo, fontsize=12)

# 添加计数标注
for stim in df.index:
    for emo in df.columns:
        count = df.loc[stim, emo]
        if count > 0:
            stim_x, stim_y = stim_coords.loc[stim, ['0','1']]
            emo_x, emo_y = emo_coords.loc[emo, ['0','1']]
            mid_x = (stim_x + emo_x)/2
            mid_y = (stim_y + emo_y)/2
            plt.text(mid_x, mid_y, str(count), fontsize=10, color='green')

plt.title('带计数标注的对应分析图')
plt.legend()
plt.show()
  • 优势:既保留对应分析的维度降维结论(刺激与情绪的关联模式),又补充原始计数信息,避免数据损耗。

内容的提问来源于stack exchange,提问作者user15590131

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 14:53:16