You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas:对含重复样本的扁平化数据集子集执行分位数归一化

Hey there! Let's work through this problem step by step. I'll break down your workflow, fix the pivot/merge hurdles you're facing, and share some optimizations to keep things clean.

First, Let's Recap the Problem

We have a flattened DataFrame with:

  • Independent variables: ind_1, ind_2, label (plus other unshown columns)
  • Response variables: out_1, out_2
  • 2 technical replicates per unique (label, ind_1, ind_2) combination
  • Need to:
    1. Group by ind_1
    2. For each group, calculate median (or mean) of out_1/out_2 per ind_2 + label combination
    3. Apply quantile normalization to these median values
    4. Broadcast the normalized values back to all original rows (keeping all columns)

Full Solution Code

First, let's import dependencies and set up the sample data:

import pandas as pd
from sklearn.preprocessing import QuantileTransformer

# Your original DataFrame (with an extra dummy column to demonstrate keeping all columns)
df = pd.DataFrame({
 'label': [1, 1, 1, 1, 1, 1, 1, 1, 2, 2, 2, 2, 2, 2, 2, 2],
 'ind_1': ['a', 'a', 'a', 'a', 'b', 'b', 'b', 'b', 'a', 'a', 'a', 'a', 'b', 'b', 'b', 'b'],
 'out_1': [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16],
 'out_2': [16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, 1],
 'ind_2': ['z', 'z','y', 'y', 'z', 'z', 'y', 'y', 'z', 'z', 'y', 'y', 'z', 'z', 'y', 'y'],
 'other_metadata_col': range(16)  # Example of an extra column to preserve
})

Next, define a function to handle each ind_1 group:

def process_ind1_subgroup(group):
    # Step 1: Calculate medians for each (label, ind_2) combination (flat format, no multi-index!)
    median_summary = group.groupby(['label', 'ind_2'])[['out_1', 'out_2']].median().reset_index()
    
    # Step 2: Apply quantile normalization to the median values
    # Use n_quantiles equal to the number of rows in our summary (matches sample size exactly)
    qt = QuantileTransformer(n_quantiles=len(median_summary), random_state=42, output_distribution='uniform')
    normalized_data = qt.fit_transform(median_summary[['out_1', 'out_2']])
    
    # Convert normalized array back to DataFrame with meaningful column names
    normalized_df = pd.DataFrame(
        normalized_data,
        columns=['norm_out_1', 'norm_out_2'],
        index=median_summary.index
    )
    
    # Step 3: Merge medians + normalized values into one intermediate flat DataFrame
    intermediate = pd.concat([median_summary, normalized_df], axis=1)
    
    # Step 4: Merge back to original group data to broadcast normalized values to all rows
    # Merge on the keys that define each replicate group (label + ind_2)
    merged_group = pd.merge(
        group,
        intermediate,
        on=['label', 'ind_2'],
        suffixes=('', '_median')  # Avoid column name conflicts
    )
    
    return merged_group

Finally, apply the function to each ind_1 group and combine results:

# Process each ind_1 group and reset index to flatten the result
final_df = df.groupby('ind_1').apply(process_ind1_subgroup).reset_index(drop=True)

# Check the output (shows original columns + median + normalized values)
print(final_df.head())

Key Fixes & Optimizations

  1. Avoiding Multi-Index Headaches
    Instead of using pivot (which can create messy multi-indexes), we use groupby(...) .median() .reset_index() to directly get a flat summary DataFrame. This eliminates the need to "convert pivot results to flat format" entirely.

  2. Clean Broadcast to Original Rows
    Using pd.merge() with label and ind_2 as keys automatically maps the normalized median values to every technical replicate row. No need to reshape data back and forth—merge handles the broadcasting for you.

  3. Stable Quantile Normalization
    We use sklearn.preprocessing.QuantileTransformer for consistent, tested normalization. Adjust output_distribution (e.g., 'normal' instead of 'uniform') if you need a different distribution.

  4. Preserving All Original Columns
    The merge operation keeps every column from your original DataFrame (including any extra metadata columns you have).


Customization Tips

  • If you meant mean instead of median for technical replicates, replace .median() with .mean() in the groupby step.
  • If you have more independent variables (e.g., ind_3), add them to both the groupby list and the on parameter in pd.merge() to maintain correct grouping.
  • To adjust the normalization range, modify the output_distribution and feature_range parameters of QuantileTransformer.

内容的提问来源于stack exchange,提问作者SumNeuron

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:19:02