如何仅用Pandas为条形图添加重复字符串占比标签
问题
我有如下的Pandas DataFrame(可复现代码如下),已经生成了条形图,现在想**仅使用Pandas(允许预先硬编码)**为条形图添加重复字符串的百分比标签。目前已有Matplotlib的实现方案,但优先考虑纯Pandas实现,若无法实现再使用Matplotlib。
可复现代码:
import pandas as pd import io TESTDATA="""All_services All_services Rehosting applications to AWS Replacing flexible functionalities Unaltered replatforming of underlying code structure, functionalities, features Rebuilding broken applications/software segments Optimize existing use of cloud(Cost saving) Expand use of containers Move on prem servers to Sass Expanding public clouds Implemenation CI/CD to clouds Migration Evaluator AWS Migration Hub AWS Application Discovery Services AWS Landing Zone AWS Control Tower AWS Management and Governance AWS Database Migration Services AWS Server Migration Service AWS Database Migration Service AWS Application Discovery Service AWS Direct Connect DB Migrations open-source databases to AWS. Oracle to Oracle Oracle or Microsoft SQL Server to Amazon Aurora. Migrating fileservers to Amazon S3 migrating commercial RDBMS or MySQL. Optimize existing use of cloud(Cost saving) Expand use of containers Move on prem servers to Sass Expanding public clouds Implemenation CI/CD to clouds Migration Evaluator AWS Migration Hub AWS Application Discovery Services AWS Landing Zone DB Migrations Cloud Migration Planning Replatforming Applications for Cloud Cloud Application Development Services From Monolith to Microservices Cloud Infrastructure Automation Implemenation CI/CD to clouds DB Migrations Optimize existing use of cloud(Cost saving) Implemenation CI/CD to clouds Migration Evaluator AWS Migration Hub AWS Application Discovery Services AWS Direct Connect DB Migrations open-source databases to AWS. Oracle to Oracle Oracle or Microsoft SQL Server to Amazon Aurora. Migrating fileservers to Amazon S3 Optimize existing use of cloud(Cost saving) Amazon S3 Transfer Acceleration AWS Snowball AWS Direct Connect EC2 AWS Server Migration Service AWS Database Migration Service VMWare Cloud on AWS Optimize existing use of cloud(Cost saving) Cloud Application Development Services From Monolith to Microservices Cloud Infrastructure Automation Implemenation CI/CD to clouds DB Migrations Optimize existing use of cloud(Cost saving) Implemenation CI/CD to clouds Migration Evaluator Optimize existing use of cloud(Cost saving) AWS Application Discovery Services AWS Direct Connect DB Migrations Rebuilding broken applications/software segments Optimize existing use of cloud(Cost saving) Expand use of containers Move on prem servers to Sass Expanding public clouds Implemenation CI/CD to clouds AWS Management and Governance AWS Database Migration Services AWS Server Migration Service AWS Database Migration Service AWS Application Discovery Service AWS Direct Connect DB Migrations """ df = pd.read_csv(io.StringIO(TESTDATA), sep=";") df = df.replace(r"^ +| +$", r"", regex=True) df.All_services.value_counts().sort_values().plot(kind = 'barh',figsize=(25, 15),linewidth=4)
解决方案
核心说明
纯Pandas无法直接给条形图添加自定义标签——Pandas的plot()方法只是封装了Matplotlib的绘图逻辑,本身没有提供添加标签的API。所以必须结合Matplotlib完成标签绘制,但我们可以用Pandas先完成所有数据计算,尽量减少非Pandas代码。
步骤1:用Pandas计算频次和百分比
先统计每个服务的出现次数,再算出对应的百分比:
# 统计频次并排序 counts = df['All_services'].value_counts().sort_values() # 计算总数 total = counts.sum() # 计算百分比(保留两位小数) percentages = (counts / total * 100).round(2) # 组合成要显示的标签文本(次数+百分比) label_texts = [f"{count} ({pct}%)" for count, pct in zip(counts, percentages)]
步骤2:绘图并添加标签
用Pandas绘制条形图后,获取Matplotlib的坐标轴对象,逐个添加标签:
import matplotlib.pyplot as plt # 绘制条形图,获取坐标轴实例 ax = counts.plot(kind='barh', figsize=(25, 15), linewidth=4) # 遍历每个条形,添加标签 for idx, (text, count) in enumerate(zip(label_texts, counts)): # 调整文本位置,避免和条形重叠 ax.text(count + 0.1, idx, text, va='center', fontsize=12) # 自动调整布局,防止标签被截断 plt.tight_layout() plt.show()
简化版:仅显示百分比
如果只需要展示百分比,也可以直接用百分比数据绘图并标注:
ax = percentages.plot(kind='barh', figsize=(25, 15), linewidth=4) for idx, pct in enumerate(percentages): ax.text(pct + 0.1, idx, f"{pct}%", va='center', fontsize=12) plt.tight_layout() plt.show()
内容的提问来源于stack exchange,提问作者Bhargav
相关产品推荐
相关产品推荐

