You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python双时间序列场景下类似Average if功能实现的技术咨询

How to Calculate Average by Rank in a Pandas DataFrame (Like Excel's AVERAGEIF)

Hey there! Your initial approach using boolean masks is totally on the right track—no need for loops here, pandas has built-in tools to make this clean and efficient. Let's break down the best ways to get the average amount for top-ranked entries.

First, Let's Set Up a Sample DataFrame

Let's start with a sample df matching your scenario (I'll assume your amount column is named amount; adjust if yours is different):

import pandas as pd

data = {
    'ranks': [1, 2, 2, 3, 4, 1, 3],
    'amount': [100, 200, 150, 50, 30, 120, 70]
}
df = pd.DataFrame(data)

Method 1: Boolean Mask (Your Original Idea, Polished)

Your top = df['ranks'] < 3 mask is correct—it selects all rows where rank is 1 or 2. To calculate the average, just apply this mask and take the mean of the amount column. Using .loc is safer to avoid potential chained assignment issues:

# Create the mask for top 2 ranks
top_mask = df['ranks'] < 3

# Calculate average amount for top ranks
top_average = df.loc[top_mask, 'amount'].mean()
print(f"Average of top 2 ranks: {top_average}")

This will give you the average of all entries where rank is 1 or 2—exactly what you're looking for!

Method 2: Using query() for Conciseness

If you prefer a more one-liner approach, pandas' query() method lets you write the condition directly as a string:

top_average = df.query('ranks < 3')['amount'].mean()

This does the exact same thing as the boolean mask, just with a more readable syntax for simple conditions.

Bonus: Grouping for Multiple Rank Groups

If you ever need to calculate averages for multiple groups (e.g., top 2 vs. bottom 2), you can use numpy.where to create a grouping column, then use groupby():

import numpy as np

# Create a group column
df['rank_group'] = np.where(df['ranks'] < 3, 'Top 2', 'Bottom 2')

# Calculate averages per group
grouped_averages = df.groupby('rank_group')['amount'].mean()
print(grouped_averages)

This will output the average for both top and bottom rank groups in one go.

Why Your Loop Isn't Necessary

Pandas is designed for vectorized operations—meaning it can handle entire columns of data at once, without needing to loop through each row. Your boolean mask approach is already vectorized, so it's faster and more readable than a loop.

内容的提问来源于stack exchange,提问作者AlgoQuant

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:15:08