You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python实现多值字段拆分与品牌计数柱状图可视化

品牌知晓度统计绘图实现

基础数据集

现有受访者调研数据集共2个字段:受访者性别(Gender)、知晓的男士品牌(KnownBrands,多品牌用分号分隔),样例数据如下:

索引GenderKnownBrands
0ManNIVEA MEN;GATSBY;
1ManGATSBY;GARNIER MEN;L’OREAL MEN EXPERT;
2WomanCLINIQUE FOR MEN;SK-II MEN;Neutrogena MEN;
3ManNIVEA MEN;GARNIER MEN;L’OREAL MEN EXPERT;GATSBY;
4WomanNIVEA MEN;GATSBY;

需求为拆分KnownBrands字段的多值内容,按品牌维度绘制计数柱状图,原有实现代码如下:

原有数据处理代码

#split the brands
brands = Men["KnownBrands"].str.split(";").explode().astype(object).reset_index()

#use pivot to provide total for each brands
output = brandnames.pivot(index="index", columns="KnownBrands", values= "KnownBrands").reset_index(drop=True).drop('',1)

原有绘图代码

brandname=output.count().plot.bar()

#Rotate the x-axis name vertically to prevent overlapping
plt.xticks(rotation='45',horizontalalignment='right')
plt.xlabel("Brands")
plt.ylabel("Frequency")
plt.title("Brands Known by Respondents")
#Chart data labels, only seaborn version 3.4.2 have this function
plt.bar_label(brandname.containers[0])
plt.show();

原有代码问题修正

原有代码存在3个核心问题会导致运行报错或效果不佳:

  1. 变量名不一致:拆分字段后赋值的变量为brands,后续透视表操作调用了未定义的brandnames,直接触发NameError
  2. 逻辑冗余:无需将拆分后的长表转换为宽表再按列计数,拆分后过滤空值直接统计频次即可,代码可读性和运行效率更高
  3. 缺少布局适配:旋转后的x轴标签容易被画布边缘截断,需加布局自适应调整

修正后可直接运行的完整代码:

import pandas as pd
import matplotlib.pyplot as plt

# 拆分多值品牌字段,过滤分号带来的空字符串
split_brands = Men["KnownBrands"].str.split(";").explode().str.strip()
split_brands = split_brands[split_brands != ""]

# 统计品牌频次并绘制柱状图
ax = split_brands.value_counts().plot.bar()
plt.xticks(rotation=45, horizontalalignment="right")
plt.xlabel("品牌名称")
plt.ylabel("知晓人次")
plt.title("受访者知晓男士品牌分布")
plt.bar_label(ax.containers[0])
# 自动适配布局避免标签截断
plt.tight_layout()
plt.show()

运行后即可得到各品牌的知晓人次计数柱状图,柱顶会标注具体数值,x轴标签旋转45度不会重叠。

内容的提问来源于stack exchange,提问作者snoopychui

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.31 00:12:40