You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Pandas的groupby('A').agg('min')转换为Featuretools实现?

Replicating Pandas Groupby-Min-and-Merge in Featuretools

Hey there, let's fix this—your cum_min approach isn't working because it calculates a running cumulative minimum (based on row order), not the group-wise minimum you need. Here's how to replicate your exact Pandas workflow using Featuretools without adding extra tables:

Step 1: Prep Your Data & EntitySet

First, Featuretools requires every entity to have a unique index (a column that identifies each row uniquely). We'll add an id column to your DataFrame first:

import pandas as pd
import featuretools as ft

# Original DataFrame
df = pd.DataFrame({'A': [1, 1, 2, 2], 'B': [1, 2, 3, 4], 'C': [0.3, 0.2, 1.2, -0.5]})
# Add unique index column
df['id'] = range(len(df))

# Create EntitySet with the indexed DataFrame
es = ft.EntitySet(id='my_data_set')
es = es.entity_from_dataframe(
    entity_id='df',
    dataframe=df,
    index='id',
    make_index=False  # We already created the index manually
)

Step 2: Generate Group-Wise Minimum Features

Instead of cum_min, use agg_primitives=['min'] and specify groupby_features=['A'] in ft.dfs. This tells Featuretools to calculate the minimum of B and C for each group in A, then automatically merge those values back to every row in the original table (just like your Pandas merge step):

# Generate features
feature_matrix, feature_names = ft.dfs(
    entityset=es,
    target_entity='df',
    agg_primitives=['min'],
    groupby_features=['A']  # Group by column 'A'
)

Step 3: (Optional) Rename Features to Match Your Pandas Output

If you want the feature names to exactly match your groupby_A(min_B)/groupby_A(min_C) format, just rename the columns:

feature_matrix = feature_matrix.rename(columns={
    'MIN(df.B by A)': 'groupby_A(min_B)',
    'MIN(df.C by A)': 'groupby_A(min_C)'
})

What the Output Looks Like

The resulting feature_matrix will be identical to your df_new from Pandas:

idABCgroupby_A(min_B)groupby_A(min_C)
0110.310.2
1120.210.2
2231.23-0.5
324-0.53-0.5

Why This Works

  • agg_primitives=['min'] tells Featuretools to compute minimum values.
  • groupby_features=['A'] ensures the minimum is calculated per group in column A.
  • Featuretools automatically joins these aggregated values back to every row in the target entity (df), eliminating the need for a manual merge like in Pandas.

内容的提问来源于stack exchange,提问作者michael giacomazza

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 08:43:08