You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Featuretools生成分组聚合特征的EntitySet与原语配置方法

实现方案

1 EntitySet构造逻辑

完全基于你现有字段即可完成构造,不需要额外补充数据:

  • 首先添加主表(事实表)实体:直接用你现有的全量数据集,设置instance_id为唯一索引,date为时间索引,实体名可自定义为main_table
  • 其次构造分类维度实体:提取主表中categorical_y列的所有去重值,生成一张仅含categorical_y列的独立表,设置categorical_y为该表的唯一索引,实体名可自定义为category_dim
  • 最后添加实体关系:设置父实体为category_dim、子实体为main_table,两表关联外键为categorical_y

2 需要引入的特征原语

仅需要使用featuretools内置的聚合类原语Mean即可满足需求,如果你后续需要扩展同类分组统计特征,也可以替换为Max、Min、Sum、Std等同类型聚合原语。如果你需要生成带时间窗口的分组均值特征(比如每个categorical_y分组过去30天的numerical_x均值),不需要额外引入原语,直接在dfs方法中配置时间窗口参数即可。

3 完整实现代码示例

import featuretools as ft
import pandas as pd

# 替换为你自己的数据集
main_df = pd.DataFrame({
    "instance_id": [1,2,3,4,5,6],
    "date": pd.date_range("2024-01-01", periods=6),
    "numerical_x": [10,20,15,25,30,18],
    "categorical_y": ["a", "a", "b", "b", "a", "b"]
})

# 初始化EntitySet
es = ft.EntitySet(id="feature_gen_set")

# 添加主表实体
es = es.add_dataframe(
    dataframe_name="main_table",
    dataframe=main_df,
    index="instance_id",
    time_index="date"
)

# 构造分类维度实体
category_df = main_df[["categorical_y"]].drop_duplicates().reset_index(drop=True)
es = es.add_dataframe(
    dataframe_name="category_dim",
    dataframe=category_df,
    index="categorical_y"
)

# 绑定两表关系
es = es.add_relationship(
    ft.Relationship(
        es["category_dim"].ww.index,
        es["main_table"]["categorical_y"]
    )
)

# 生成特征,自动合并回主表
feature_matrix, feature_list = ft.dfs(
    entityset=es,
    target_dataframe_name="main_table",
    agg_primitives=["mean"]
)

# 输出的feature_matrix中已经包含按categorical_y分组的numerical_x均值特征,列名默认是`category_dim.MEAN(main_table.numerical_x)`,可自行重命名
print(feature_matrix)

内容的提问来源于stack exchange,提问作者Иван Липатов

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 00:36:10