You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于双条件分组,从Pandas DataFrame生成nlargest字典及目标DataFrame

Got it, let's work through this step by step. First, let's make sure we're starting with the correct DataFrame, then we'll add the top3growth column and build those two grouping dictionaries as requested.

Step 1: Recreate the sample DataFrame

First, let's code up the DataFrame you provided to have a working base:

import pandas as pd

# Your input data
data = {
    'person_code': [231, 233, 432, 431, 654, 764, 434],
    'CNAE': [32, 43, 32, 56, 89, 32, 32],
    'growth': [0.54, 0.12, 0.44, 0.32, 0.12, 0.20, 0.82],
    'size': [32, 333, 21, 23, 89, 211, 90],
    'state': ['FR', 'LK', 'FR', 'KS', 'FR', 'TI', 'TI']
}

df = pd.DataFrame(data)

Step 2: Add the top3growth column

I'll assume you want to flag whether a person's growth is in the top 3 for their state (if you need to group by CNAE instead, just swap the group key). Here's how to do it:

# Calculate the minimum growth value in the top 3 for each state
top3_min = df.groupby('state')['growth'].transform(lambda x: x.nlargest(3).min())

# Create the flag column: True if growth is >= the top3 threshold (handles ties)
df['top3growth'] = df['growth'] >= top3_min

This will mark every row as True if its growth falls within the top 3 values of its state. For example:

  • All rows in FR are marked True since there are only 3 entries there
  • Both rows in TI are marked True (top 3 includes all entries when there are fewer than 3)

Step 3: Build dictionaries with nlargest for two grouping conditions

Let's use two common grouping keys: state and CNAE (feel free to adjust if you need different groups).

Dictionary 1: Group by state (top 3 growth per state)

This dictionary maps each state code to the rows of its top 3 growing people:

state_top3_dict = {
    state: group.nlargest(3, 'growth')
    for state, group in df.groupby('state')
}

# If you only want the person_codes instead of full rows, use this:
# state_top3_dict = {
#     state: group.nlargest(3, 'growth')['person_code'].tolist()
#     for state, group in df.groupby('state')
# }

Dictionary 2: Group by CNAE (top 3 growth per industry)

Similarly, this dictionary maps each CNAE code to its top 3 growing people:

cnae_top3_dict = {
    cnae: group.nlargest(3, 'growth')
    for cnae, group in df.groupby('CNAE')
}

# Again, for just person_codes:
# cnae_top3_dict = {
#     cnae: group.nlargest(3, 'growth')['person_code'].tolist()
#     for cnae, group in df.groupby('CNAE')
# }

For example, if you run print(state_top3_dict['FR']), you'll get the 3 rows for state FR, sorted by growth descending.

内容的提问来源于stack exchange,提问作者aabujamra

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 10:38:46