使用Pandas绘制基于分类变量的分组箱线图技术问询
Got it, let's break down exactly how to build that 2×6 grid of boxplots you need. Since you already have your cleaned DataFrame, we'll focus on two straightforward approaches—using seaborn (my go-to for grouped plots) and matplotlib (for full manual control).
First, let's assume your DataFrame df has three key columns:
tag_count: The number of tags (1-6, we'll treat this as a categorical variable)country: The two distinct countries you're comparingfriend_count: The numerical friend count values you want to visualize
Step 1: Prep Your Data (Quick Check)
Make sure tag_count is treated as a categorical variable to ensure the subplots are ordered correctly (1 to 6, not random):
df['tag_count'] = df['tag_count'].astype('category')
Approach 1: Seaborn Catplot (Simplest Method)
Seaborn's catplot is perfect for this—it handles the grid layout automatically, so you don't have to manually create subplots. We'll split the grid by country (rows, 2 total) and tag_count (columns, 6 total):
import seaborn as sns import matplotlib.pyplot as plt # Set a clean plot style sns.set_style("whitegrid") # Create the 2x6 grid of boxplots g = sns.catplot( data=df, y="friend_count", # The value we're plotting col="tag_count", # Split columns by tag count (6 columns) row="country", # Split rows by country (2 rows) kind="box", # Specify boxplot type height=3, # Height of each subplot aspect=0.8, # Width-to-height ratio (adjust for compactness) sharey=True # Share y-axis across all subplots for easy comparison ) # Clean up labels and titles g.set_axis_labels("", "Friend Count") # Remove redundant x-label g.set_titles( row_template="{row_name}", col_template="Tag Count: {col_name}" ) # Adjust spacing so plots don't overlap plt.tight_layout() plt.show()
This will give you a neat 2-row, 6-column grid where each subplot shows the friend count distribution for a specific country and tag count group.
Approach 2: Matplotlib (Full Manual Control)
If you prefer more control over every aspect of the plot, use matplotlib's subplots to build the grid manually:
import matplotlib.pyplot as plt # Create a 2x6 grid of subplots fig, axes = plt.subplots(nrows=2, ncols=6, figsize=(18, 6), sharey=True) # Get sorted lists of unique countries and tag counts countries = sorted(df['country'].unique()) tag_counts = sorted(df['tag_count'].unique()) # Loop through each country and tag count to plot boxplots for row_idx, country in enumerate(countries): for col_idx, tag_count in enumerate(tag_counts): # Get the subset of data for this country and tag count subset = df[(df['country'] == country) & (df['tag_count'] == tag_count)] # Plot the boxplot on the corresponding subplot ax = axes[row_idx, col_idx] ax.boxplot(subset['friend_count']) # Add titles and labels ax.set_title(f"{country}\nTags: {tag_count}") ax.set_xlabel("") # Remove x-label since there's only one group per subplot # Only add y-label to the first column for readability if col_idx == 0: ax.set_ylabel("Friend Count") # Adjust spacing and display plt.tight_layout() plt.show()
Key Tips for Better Visualization
- Share the y-axis: Using
sharey=Trueensures all subplots use the same y-scale, making it easy to compare friend count distributions across groups. - Adjust subplot size: Tweak
height/aspect(seaborn) orfigsize(matplotlib) to make sure all 12 plots are readable without overlapping titles/labels. - Handle missing groups: If some (country, tag_count) combinations have no data, seaborn will leave an empty subplot, while matplotlib will throw an error—you can add a check like
if not subset.empty:before plotting to avoid this.
内容的提问来源于stack exchange,提问作者420fedoras

