使用pandas合并Stack Overflow调查性别列零散取值实现三类分组展示
问题背景
我目前正在使用Jupyter Notebook分析Stack Overflow开发者调查数据,数据文件包含名为Gender的字段。
问题描述
使用df['Gender'].value_counts().plot(kind="bar")绘制柱状图时,除了Man和Woman两类取值外,还存在十余个其他取值,例如Prefer not to say、non-binary、queer等。这类边缘取值合计占比仅为4%-6%,单独为每个取值生成柱子没有实际分析价值。
需求是将所有不属于Man、Woman的取值统一合并为Undefined类别,最终可视化仅展示男性、女性、Undefined三类。
初始代码:
import pandas as pd import matplotlib.pyplot as plt df = pd.read_csv('survey.csv') df['Gender'].value_counts().plot(kind="bar", figsize=(15,7), color="#61d199")
实现方案
核心思路是对Gender字段做值重映射:通过apply方法遍历字段的每个取值,若取值不在指定的['Man', 'Woman']列表内,就统一替换为Undefined,否则保留原取值。
最终实现代码如下:
import pandas as pd import matplotlib.pyplot as plt df = pd.read_csv('survey.csv') # 重映射Gender字段取值 df['Gender'] = df['Gender'].apply(lambda x: 'Undefined' if x not in ['Man', 'Woman'] else x) # 绘制饼图展示性别占比 df['Gender'].value_counts().plot(kind="pie", figsize=(10,10), autopct='%1.0f%%', labels=['男性', '未定义', '女性'], colors = ['#98C1D9', '#3D5A80', '#E0FBFC'], title=('编程领域性别占比'))
内容的提问来源于stack exchange,提问作者Jacob Day
相关产品推荐
相关产品推荐

