如何优雅地为Pandas DataFrame添加基于列切片均值的新列?
Efficient Way to Compute Reverse Expanding Mean in Pandas
The most Pythonic and efficient way to achieve this is by leveraging pandas' expanding window functions combined with reversing the series. Here's how you can do it in one concise line:
import pandas as pd my_df = pd.DataFrame([10, 20, 30, 40, 50], columns=['col_1']) my_df['col_2'] = my_df['col_1'][::-1].expanding().mean()[::-1] print(my_df)
This will output exactly what you need:
col_1 col_2 0 10 30.0 1 20 35.0 2 30 40.0 3 40 45.0 4 50 50.0
How It Works:
- Reverse the Series:
my_df['col_1'][::-1]flips the 'col_1' values to[50, 40, 30, 20, 10]. - Compute Expanding Mean:
.expanding().mean()calculates the mean of all values from the start of the reversed series up to each position. This gives us[50.0, 45.0, 40.0, 35.0, 30.0]. - Reverse Back:
[::-1]flips the result again to align with the original row order, producing the desired[30.0, 35.0, 40.0, 45.0, 50.0].
Why This Is Better Than the Loop:
- Vectorized Operations: Pandas handles the calculations under the hood using optimized C-based code, which is drastically faster than a Python
forloop—especially for large datasets. - Avoids Warnings: The original loop uses chain indexing (
my_df.loc[i]['col_2']), which can triggerSettingWithCopyWarningand lead to unexpected behavior. This method modifies the DataFrame directly and safely. - Conciseness: It’s a single, readable line that clearly expresses the intent without boilerplate loop code.
内容的提问来源于stack exchange,提问作者AutomaticStatic
相关产品推荐
相关产品推荐

