基于员工信息数据集的多类统计可视化图表实现需求
员工数据可视化实现方案
以下代码可直接运行,实现所有要求的可视化效果,依赖为pandas和matplotlib,未安装的话可先执行pip install pandas matplotlib安装。
import pandas as pd import matplotlib.pyplot as plt # 原始员工数据 employee_data = { 'Age': [26, 29, 25, 26, 29, 30, 32, 31, 34, 33, 24, 27, 28], 'Home City': ['Rohtak', 'Aligarh', 'Rajkot', 'Bhilai', 'Rohtak', 'Delhi', 'Faridabad', 'Howrah', 'Delhi', 'Delhi', 'Patna', 'Patna', 'Agra'], 'Salary': [22000, 28000, 18000, 19000, 27000, 25000, 30000, 31000, 34000, 32000, 18000, 24000, 20000] } df = pd.DataFrame(employee_data) # 1. 年龄vs薪资折线图 plt.figure(figsize=(10,6)) # 先按年龄排序让折线逻辑更通顺 sorted_df = df.sort_values('Age') plt.plot(sorted_df['Age'], sorted_df['Salary'], marker='o', color='#1f77b4') plt.title('年龄 vs 薪资 折线图') plt.xlabel('年龄') plt.ylabel('薪资(元)') plt.grid(alpha=0.3) plt.show() # 2. 年龄vs薪资散点图 plt.figure(figsize=(10,6)) plt.scatter(df['Age'], df['Salary'], color='#ff7f0e', s=60) plt.title('年龄 vs 薪资 散点图') plt.xlabel('年龄') plt.ylabel('薪资(元)') plt.grid(alpha=0.3) plt.show() # 3. 年龄、薪资直方图 fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(14,6)) # 年龄直方图 ax1.hist(df['Age'], bins=8, color='skyblue', edgecolor='black') ax1.set_title('年龄分布直方图') ax1.set_xlabel('年龄') ax1.set_ylabel('员工人数') # 薪资直方图 ax2.hist(df['Salary'], bins=6, color='lightgreen', edgecolor='black') ax2.set_title('薪资分布直方图') ax2.set_xlabel('薪资(元)') ax2.set_ylabel('员工人数') plt.tight_layout() plt.show() # 4. 居住城市分布饼图 city_counts = df['Home City'].value_counts() plt.figure(figsize=(10,10)) plt.pie(city_counts, labels=city_counts.index, autopct='%1.1f%%', startangle=90) plt.title('员工居住城市占比') plt.axis('equal') # 保证饼图为正圆形 plt.show() # 5. 薪资、年龄箱线图 fig, (ax1, ax2) = plt.subplots(1,2, figsize=(14,6)) # 年龄箱线图 ax1.boxplot(df['Age']) ax1.set_title('年龄分布箱线图') ax1.set_xticklabels(['年龄']) # 薪资箱线图 ax2.boxplot(df['Salary']) ax2.set_title('薪资分布箱线图') ax2.set_xticklabels(['薪资(元)']) plt.tight_layout() plt.show()
代码说明
- 折线图提前对年龄做了升序处理,避免无规律的折线波动影响可读性
- 饼图自动计算各城市员工占比,直接显示百分比数值
- 直方图和箱线图可直观查看年龄、薪资两个数值特征的分布规律、四分位区间和异常值情况
内容的提问来源于stack exchange,提问作者md ismail
相关产品推荐
相关产品推荐

