You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于含逗号分隔分类值的pandas数据集绘制seaborn swarmplot

实现方案

数据预处理

当前数据为宽表格式,且分类值以逗号合并存储,需先做格式转换:

  • 拆分family/city/nation/world四个维度列的逗号分隔值,每行仅保留一个分类值
  • 转换为长表结构,保留三个核心字段:受访者id、维度(对应四个场景分类)、时间预期(对应拆分后的时间分类)
  • 为时间预期指定固定排序,保证Y轴展示顺序符合预期:next week < next few years < lifetime < children's lifetime

绘图逻辑调整

原代码的字段映射错误,正确映射规则为:

  • X轴对应维度字段(四个场景分类)
  • Y轴对应排序后的时间预期字段
  • 通过swarmplot的散点自动散开效果,展示每个分类下的受访者分布

完整可运行代码

import pandas as pd
import seaborn as sns
import matplotlib.pyplot as plt

# 构造样例数据,你可以替换为自己的数据集读取逻辑
dftest = pd.DataFrame({
    'id': [1,2,3,4,5],
    'family': ["next week, next few years", "next few years, lifetime", "next week, next few years", "next week, next few years", "next few years, lifetime"],
    'city': ["next week, next few years", "children's lifetime", "lifetime", "next few years", "children's lifetime"],
    'nation': ["lifetime", "next few years", "children's lifetime", "next week, next few years, lifetime", "children's lifetime"],
    'world': ["children's lifetime", "next week, next few years, children's lifetime", "children's lifetime", "next week, next few years, children's lifetime", "lifetime"]
})

# 宽表转长表,拆分多值
df_long = dftest.melt(id_vars='id', var_name='维度', value_name='时间预期')
df_long['时间预期'] = df_long['时间预期'].str.split(',\s*')
df_long = df_long.explode('时间预期').reset_index(drop=True)
# 清理空值和异常值
df_long = df_long[df_long['时间预期'].str.strip() != '']

# 给时间分类指定排序,匹配示例Y轴顺序
time_order = ['next week', 'next few years', 'lifetime', "children's lifetime"]
df_long['时间预期'] = pd.Categorical(df_long['时间预期'], categories=time_order, ordered=True)

# 绘图
plt.figure(figsize=(10, 6))
ax = sns.swarmplot(
    x='维度',
    y='时间预期',
    data=df_long,
    size=8,
    color='#222222',
    edgecolor='white',
    linewidth=0.5
)

# 可选:中文标签替换,按需开启
# ax.set_xlabel('视角维度')
# ax.set_ylabel('时间预期')
# plt.xticks(ticks=[0,1,2,3], labels=['家庭', '城市', '国家', '世界'])
# plt.yticks(ticks=[0,1,2,3], labels=['下周', '未来几年', '有生之年', '子辈有生之年'])

# 可选:添加Y轴网格线匹配示例样式
# ax.grid(axis='y', linestyle='-', alpha=0.3)

plt.show()

额外样式调整说明

如果需要和《增长的极限》中的示例完全匹配,可按需调整:

  • 若数据量较大出现散点重叠,可适当调小size参数,或替换为sns.stripplot并设置jitter=0.2
  • 可通过palette参数为不同维度设置不同的点颜色
  • 可通过plt.rcParams调整全局字体、字号匹配原版样式

内容的提问来源于stack exchange,提问作者orestisf

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 11:30:00