Pandas技术实现:列值符合条件时取两列最大值否则取均值
Got it, let's figure out how to add that conditional logic to your adj_mean column. You're halfway there with the max function—we just need to layer in the condition for when both values are positive. Here are two solid approaches:
Solution for Creating adj_mean Column in Pandas DataFrame
Method 1: Using numpy.where (Best for Large Datasets)
This uses vectorized operations, which are way faster than row-by-row processing for big data. First, make sure you have numpy imported:
import numpy as np
Then write the conditional logic:
# Check if either Avg or rolling_mean is zero has_zero = (df['Avg'] == 0) | (df['rolling_mean'] == 0) # Assign adj_mean: max if there's a zero, average otherwise df['adj_mean'] = np.where( has_zero, df[['Avg', 'rolling_mean']].max(axis=1), df[['Avg', 'rolling_mean']].mean(axis=1) )
Method 2: Using pandas.apply (Simpler for Small Data)
If your dataset isn't huge, a row-wise lambda function is easy to read and implement:
df['adj_mean'] = df.apply( lambda row: max(row['Avg'], row['rolling_mean']) if row['Avg'] == 0 or row['rolling_mean'] == 0 else (row['Avg'] + row['rolling_mean']) / 2, axis=1 )
Let's Test It With Your Sample Data
Both methods will give you exactly the output you want:
| ID | Avg | rolling_mean | adj_mean |
|---|---|---|---|
| 0 | 5 | 0 | 5.0 |
| 1 | 6 | 6.3 | 6.15 |
| 2 | 5 | 8 | 6.5 |
| 3 | 4 | 0 | 4.0 |
Quick Tips
- Go with Method 1 if you're working with large datasets—vectorized operations avoid slow loops through each row.
- If your data has floating-point numbers that might be almost zero (due to precision), replace
==0with something likeabs(col) < 1e-9to catch those edge cases.
内容的提问来源于stack exchange,提问作者TroyMcClure8998
相关产品推荐
相关产品推荐

