使用Featuretools遇AttributeError:截断时间DataFrame匹配异常
问题
我正在学习Featuretools的剩余使用寿命(RUL)预测进阶教程,运行以下代码时出现错误:
from featuretools.tsfresh import CidCe import featuretools as ft fm, features = ft.dfs( entityset=es, target_dataframe_name='RUL data', agg_primitives=['last', 'max'], trans_primitives=[], chunk_size=.26, cutoff_time=cutoff_time_list[0], max_depth=3, verbose=True, ) fm.to_csv('advanced_fm.csv') fm.head()
相关信息:
- 实体集
es中的RUL data表包含engine_no(作为标识列)、time_cycles等字段 - 截断时间表
cutoff_time_list[0]包含engine_no和time_cycles两列 - 使用Featuretools版本1.27.0
遇到的错误:
AttributeError: Cutoff time DataFrame must contain a column with either the same name as the target dataframe index or a column named "instance_id"
尽管数据集和截断时间表都包含engine_no列,仍出现该错误,请问原因是什么?
错误原因及解决方法
核心原因
Featuretools的dfs方法要求截断时间表(cutoff_time参数)必须满足以下任一条件:
- 包含与目标数据帧(
target_dataframe_name指定的表)索引完全同名的列 - 包含名为
instance_id的列,用于关联目标表的实例
你遇到报错的常见原因有以下几种:
- 目标表索引未设置为
engine_no:如果在实体集es中,RUL data表的索引不是engine_no(比如使用了默认的整数索引),那么即使截断表有engine_no列,也无法与目标表的索引匹配,触发错误。 - 列数据类型不匹配:
cutoff_time_list[0]中的engine_no列数据类型,与RUL data表索引的数据类型不一致(例如一个是字符串、一个是整数),导致Featuretools无法识别两者为关联字段。 - 实体集创建时未正确配置索引:在将
RUL data添加到实体集es时,没有明确指定index='engine_no',导致Featuretools未将engine_no识别为目标表的实例标识列。
验证与解决步骤
- 检查目标表索引:执行
print(es['RUL data'].index),确认索引名称是否为engine_no。 - 修复实体集索引配置:如果索引不是
engine_no,重新创建实体集时指定索引:# 示例:重新添加目标表到实体集 es = ft.EntitySet(id='rul_dataset') es = es.add_dataframe( dataframe=rul_data_df, dataframe_name='RUL data', index='engine_no', # 明确指定engine_no为索引 time_index='time_cycles' # 时间索引根据实际情况设置 ) - 统一数据类型:对比
cutoff_time_list[0]['engine_no'].dtype和es['RUL data'].index.dtype,若不一致则转换类型:# 示例:将截断表的engine_no转为整数类型 cutoff_time_list[0]['engine_no'] = cutoff_time_list[0]['engine_no'].astype(int) - 备选方案:添加instance_id列:如果不想修改索引,可在截断时间表中添加
instance_id列并赋值为engine_no:cutoff_time_list[0]['instance_id'] = cutoff_time_list[0]['engine_no']
内容的提问来源于stack exchange,提问作者J. Maria
相关产品推荐
相关产品推荐

