基于sample_ids前缀拆分绘制多幅100%堆叠分组柱状图的技术实现问询
Fix: 100% Stacked Bar Charts Grouped by Sample Prefix
Let's break down what's off with your current code first:
- It doesn't split samples into separate groups based on the first 6 characters of
sample_ids - The 100% stacked logic is incorrect (your original perc values don't sum to 1, so we need to normalize them first)
- It's trying to overlay two bar plots instead of creating separate figures for each group
Here's a complete, working solution that meets your requirements:
import pandas as pd import matplotlib.pyplot as plt import seaborn as sns # Set default plot styling and size plt.rcParams["figure.figsize"] = (12, 8) sns.set_style("whitegrid") # Your original dataset df = pd.DataFrame( { 'sample_ids': ['sample1a_1', 'sample1b_2', 'sample1c_3', 'sample2a_2', 'sample2b_4'], 'perc_A': [0.75, 0.66667, 0.70213, 0.83333, 0.5], 'perc_B': [0.18182, 0.24747, 0.28176, 0.16529, 0.26316], 'perc_C': [0.4222, 0.15823, 0.17925, 0.55556, 0.21053] } ) # Step 1: Add a group column using the first 6 characters of sample_ids df['group'] = df['sample_ids'].str[:6] # Step 2: Normalize perc values to sum to 100% per sample (critical for 100% stacking) percentage_cols = ['perc_A', 'perc_B', 'perc_C'] df[percentage_cols] = df[percentage_cols].div(df[percentage_cols].sum(axis=1), axis=0) * 100 # Step 3: Loop through each group and create a separate stacked bar chart for group_name, group_df in df.groupby('group'): # Create a new figure for each group plt.figure() # Set up x-axis positions for samples in the group sample_positions = range(len(group_df['sample_ids'])) bar_width = 0.6 # Track the bottom position for stacking bars bottom = [0] * len(sample_positions) # Define colors for each category category_colors = ['#1f77b4', '#2ca02c', '#ff7f0e'] # Plot each category as a stacked segment for idx, category in enumerate(percentage_cols): category_values = group_df[category].values plt.bar( sample_positions, category_values, bar_width, bottom=bottom, label=category, color=category_colors[idx], alpha=0.8 ) # Update bottom position for the next stacked segment bottom = [current + val for current, val in zip(bottom, category_values)] # Customize the plot plt.title(f'100% Stacked Bar Chart - Group: {group_name}', fontsize=14) plt.xlabel('Sample ID', fontsize=12) plt.ylabel('Percentage (%)', fontsize=12) plt.xticks(sample_positions, group_df['sample_ids'], rotation=15) plt.legend(title='Categories', bbox_to_anchor=(1.02, 1), loc='upper left') plt.tight_layout() # Prevent legend/label cutoff plt.show()
Key Details Explained:
- Grouping Samples: We use
str[:6]to extract the first 6 characters ofsample_ids(e.g.,sample1fromsample1a_1) and create agroupcolumn to split our data. - Normalization: To get a true 100% stacked chart, we divide each sample's perc values by the sum of its perc columns, then multiply by 100. This ensures each sample's segments add up to 100%.
- Per-Group Plots: We loop over each group with
groupby('group'), creating a new figure for each. Using matplotlib'sbarwith thebottomparameter lets us stack each category on top of the previous one.
When you run this code, you'll get two separate figures: one for all sample1 prefixed samples, and another for sample2 prefixed samples—each showing a clean 100% stacked bar chart.
内容的提问来源于stack exchange,提问作者Naomi Sun
相关产品推荐
相关产品推荐

