You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于多层索引DataFrame绘制Pandas/Matplotlib分组条形图

Solution for Plotting Sorted Multi-Index DataFrame as Bar Chart

Let's walk through how to create the exact bar chart you need, using your pre-sorted Cat1 DataFrame while preserving the custom order you've already established.

Step 1: Filter for Top 10 Categories

First, we'll extract the top 10 sorted categories from your multi-index DataFrame (matching your requirement to prioritize the first 10):

# Grab the first 10 unique sorted categories from your multi-index
top_10_categories = Cat1.index.get_level_values('Category').unique()[:10]

# Filter your data to only include these top 10 categories, and flatten the index
filtered_df = Cat1.loc[top_10_categories].reset_index()

The reset_index() call converts the multi-index into regular columns (Category, Content, Installs, etc.), which simplifies plotting drastically.

Seaborn handles grouped categorical data smoothly, making it ideal for this use case. We'll build a chart where:

  • X-axis displays Content values, grouped under their parent Category
  • Bar height maps to total Installs
  • Colors distinguish between different Category groups
import seaborn as sns
import matplotlib.pyplot as plt

# Set a clean, readable plotting style
sns.set_style("whitegrid")

# Create a large figure to avoid label overlap
plt.figure(figsize=(18, 8))

# Plot the bar chart: hue=Category groups content bars by their category
ax = sns.barplot(
    data=filtered_df,
    x="Content",
    y="Installs",
    hue="Category",
    dodge=False  # Keeps bars from the same category aligned; remove for side-by-side grouping
)

# Rotate X-axis labels to prevent overlap
ax.set_xticklabels(ax.get_xticklabels(), rotation=45, ha="right")

# Add clear labels and a descriptive title
ax.set_title("Top 10 Categories: Content-wise Total Installs", fontsize=16)
ax.set_xlabel("Content Type", fontsize=12)
ax.set_ylabel("Total Installs", fontsize=12)

# Move the legend outside the plot to avoid blocking bars
plt.legend(bbox_to_anchor=(1.01, 1), loc="upper left", borderaxespad=0)

# Adjust layout to fit all elements neatly
plt.tight_layout()
plt.show()

Alternative: Subplots for Each Category

If you want to avoid cluttering a single plot, create separate subplots for each top category:

# Create a grid of subplots (2 columns, 5 rows for 10 categories)
g = sns.catplot(
    data=filtered_df,
    x="Content",
    y="Installs",
    col="Category",
    col_wrap=2,
    kind="bar",
    height=4,
    aspect=1.5
)

# Rotate labels in each subplot for readability
g.set_xticklabels(rotation=45, ha="right")

# Add a main title for the entire figure
g.fig.suptitle("Top 10 Categories: Content-wise Install Breakdown", y=1.03, fontsize=14)

plt.show()

Tips for Better Readability

  • If Installs values are extremely large (common in Google Play data), add a logarithmic Y-axis with ax.set_yscale("log") to make smaller bars visible.
  • If some Content labels are too long, truncate them with filtered_df["Content"] = filtered_df["Content"].str[:20] + "..." to keep the plot clean.

Why This Preserves Your Sort Order

Your Cat1 DataFrame is already sorted by the average installs per upload metric you defined. By extracting the first 10 categories directly from its index, we retain that custom sort order—no need to re-sort or unstack in a way that breaks the order you worked hard to set up.

内容的提问来源于stack exchange,提问作者Dmitrii Ponomarev

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 18:55:14