Pandas:对含重复样本的扁平化数据集子集执行分位数归一化
Hey there! Let's work through this problem step by step. I'll break down your workflow, fix the pivot/merge hurdles you're facing, and share some optimizations to keep things clean.
First, Let's Recap the Problem
We have a flattened DataFrame with:
- Independent variables:
ind_1,ind_2,label(plus other unshown columns) - Response variables:
out_1,out_2 - 2 technical replicates per unique
(label, ind_1, ind_2)combination - Need to:
- Group by
ind_1 - For each group, calculate median (or mean) of
out_1/out_2perind_2+labelcombination - Apply quantile normalization to these median values
- Broadcast the normalized values back to all original rows (keeping all columns)
- Group by
Full Solution Code
First, let's import dependencies and set up the sample data:
import pandas as pd from sklearn.preprocessing import QuantileTransformer # Your original DataFrame (with an extra dummy column to demonstrate keeping all columns) df = pd.DataFrame({ 'label': [1, 1, 1, 1, 1, 1, 1, 1, 2, 2, 2, 2, 2, 2, 2, 2], 'ind_1': ['a', 'a', 'a', 'a', 'b', 'b', 'b', 'b', 'a', 'a', 'a', 'a', 'b', 'b', 'b', 'b'], 'out_1': [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16], 'out_2': [16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, 1], 'ind_2': ['z', 'z','y', 'y', 'z', 'z', 'y', 'y', 'z', 'z', 'y', 'y', 'z', 'z', 'y', 'y'], 'other_metadata_col': range(16) # Example of an extra column to preserve })
Next, define a function to handle each ind_1 group:
def process_ind1_subgroup(group): # Step 1: Calculate medians for each (label, ind_2) combination (flat format, no multi-index!) median_summary = group.groupby(['label', 'ind_2'])[['out_1', 'out_2']].median().reset_index() # Step 2: Apply quantile normalization to the median values # Use n_quantiles equal to the number of rows in our summary (matches sample size exactly) qt = QuantileTransformer(n_quantiles=len(median_summary), random_state=42, output_distribution='uniform') normalized_data = qt.fit_transform(median_summary[['out_1', 'out_2']]) # Convert normalized array back to DataFrame with meaningful column names normalized_df = pd.DataFrame( normalized_data, columns=['norm_out_1', 'norm_out_2'], index=median_summary.index ) # Step 3: Merge medians + normalized values into one intermediate flat DataFrame intermediate = pd.concat([median_summary, normalized_df], axis=1) # Step 4: Merge back to original group data to broadcast normalized values to all rows # Merge on the keys that define each replicate group (label + ind_2) merged_group = pd.merge( group, intermediate, on=['label', 'ind_2'], suffixes=('', '_median') # Avoid column name conflicts ) return merged_group
Finally, apply the function to each ind_1 group and combine results:
# Process each ind_1 group and reset index to flatten the result final_df = df.groupby('ind_1').apply(process_ind1_subgroup).reset_index(drop=True) # Check the output (shows original columns + median + normalized values) print(final_df.head())
Key Fixes & Optimizations
Avoiding Multi-Index Headaches
Instead of using pivot (which can create messy multi-indexes), we usegroupby(...) .median() .reset_index()to directly get a flat summary DataFrame. This eliminates the need to "convert pivot results to flat format" entirely.Clean Broadcast to Original Rows
Usingpd.merge()withlabelandind_2as keys automatically maps the normalized median values to every technical replicate row. No need to reshape data back and forth—merge handles the broadcasting for you.Stable Quantile Normalization
We usesklearn.preprocessing.QuantileTransformerfor consistent, tested normalization. Adjustoutput_distribution(e.g.,'normal'instead of'uniform') if you need a different distribution.Preserving All Original Columns
The merge operation keeps every column from your original DataFrame (including any extra metadata columns you have).
Customization Tips
- If you meant mean instead of median for technical replicates, replace
.median()with.mean()in the groupby step. - If you have more independent variables (e.g.,
ind_3), add them to both thegroupbylist and theonparameter inpd.merge()to maintain correct grouping. - To adjust the normalization range, modify the
output_distributionandfeature_rangeparameters ofQuantileTransformer.
内容的提问来源于stack exchange,提问作者SumNeuron

