You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于sample_ids前缀拆分绘制多幅100%堆叠分组柱状图的技术实现问询

Fix: 100% Stacked Bar Charts Grouped by Sample Prefix

Let's break down what's off with your current code first:

  • It doesn't split samples into separate groups based on the first 6 characters of sample_ids
  • The 100% stacked logic is incorrect (your original perc values don't sum to 1, so we need to normalize them first)
  • It's trying to overlay two bar plots instead of creating separate figures for each group

Here's a complete, working solution that meets your requirements:

import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns

# Set default plot styling and size
plt.rcParams["figure.figsize"] = (12, 8)
sns.set_style("whitegrid")

# Your original dataset
df = pd.DataFrame(
    {
        'sample_ids': ['sample1a_1', 'sample1b_2', 'sample1c_3', 'sample2a_2', 'sample2b_4'],
        'perc_A': [0.75, 0.66667, 0.70213, 0.83333, 0.5],
        'perc_B': [0.18182, 0.24747, 0.28176, 0.16529, 0.26316],
        'perc_C': [0.4222, 0.15823, 0.17925, 0.55556, 0.21053]
    }
)

# Step 1: Add a group column using the first 6 characters of sample_ids
df['group'] = df['sample_ids'].str[:6]

# Step 2: Normalize perc values to sum to 100% per sample (critical for 100% stacking)
percentage_cols = ['perc_A', 'perc_B', 'perc_C']
df[percentage_cols] = df[percentage_cols].div(df[percentage_cols].sum(axis=1), axis=0) * 100

# Step 3: Loop through each group and create a separate stacked bar chart
for group_name, group_df in df.groupby('group'):
    # Create a new figure for each group
    plt.figure()
    
    # Set up x-axis positions for samples in the group
    sample_positions = range(len(group_df['sample_ids']))
    bar_width = 0.6
    
    # Track the bottom position for stacking bars
    bottom = [0] * len(sample_positions)
    
    # Define colors for each category
    category_colors = ['#1f77b4', '#2ca02c', '#ff7f0e']
    
    # Plot each category as a stacked segment
    for idx, category in enumerate(percentage_cols):
        category_values = group_df[category].values
        plt.bar(
            sample_positions,
            category_values,
            bar_width,
            bottom=bottom,
            label=category,
            color=category_colors[idx],
            alpha=0.8
        )
        # Update bottom position for the next stacked segment
        bottom = [current + val for current, val in zip(bottom, category_values)]
    
    # Customize the plot
    plt.title(f'100% Stacked Bar Chart - Group: {group_name}', fontsize=14)
    plt.xlabel('Sample ID', fontsize=12)
    plt.ylabel('Percentage (%)', fontsize=12)
    plt.xticks(sample_positions, group_df['sample_ids'], rotation=15)
    plt.legend(title='Categories', bbox_to_anchor=(1.02, 1), loc='upper left')
    plt.tight_layout()  # Prevent legend/label cutoff
    plt.show()

Key Details Explained:

  • Grouping Samples: We use str[:6] to extract the first 6 characters of sample_ids (e.g., sample1 from sample1a_1) and create a group column to split our data.
  • Normalization: To get a true 100% stacked chart, we divide each sample's perc values by the sum of its perc columns, then multiply by 100. This ensures each sample's segments add up to 100%.
  • Per-Group Plots: We loop over each group with groupby('group'), creating a new figure for each. Using matplotlib's bar with the bottom parameter lets us stack each category on top of the previous one.

When you run this code, you'll get two separate figures: one for all sample1 prefixed samples, and another for sample2 prefixed samples—each showing a clean 100% stacked bar chart.

内容的提问来源于stack exchange,提问作者Naomi Sun

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.27 17:57:36