FeatureTools:time_index参数是否支持指定多列时间字段?
Great question! First, a key point to clarify: in Featuretools, each entity (dataframe within an EntitySet) can only have one time_index assigned at a time. So you can't directly set both start_date and end_date as the time index for the same entity. But don't worry—there are two straightforward workarounds to get the exact features you want without merging the date columns.
Option 1: Create two separate entities in the same EntitySet
You can add two copies of your dataframe to the EntitySet, each with a different time index. Then generate features for each entity and merge the results.
Here's a concrete code example (assuming your dataframe has a primary key column named id):
import featuretools as ft import pandas as pd # Sample dataframe (replace with your actual data) df = pd.DataFrame({ "id": [1, 1, 2, 2], "start_date": pd.to_datetime(["2023-01-01", "2023-02-01", "2023-01-15", "2023-03-01"]), "end_date": pd.to_datetime(["2023-01-10", "2023-02-10", "2023-01-25", "2023-03-10"]), "value": [10, 20, 15, 25] }) # Initialize EntitySet es = ft.EntitySet(id="my_time_data") # Add first entity with start_date as time_index es = es.add_dataframe( dataframe_name="data_with_start", dataframe=df, index="id", time_index="start_date" ) # Add second entity with end_date as time_index (unique name is important) es = es.add_dataframe( dataframe_name="data_with_end", dataframe=df, index="id", time_index="end_date" ) # Add self-relationship if you want to calculate features per ID (e.g., time since first start_date for the same ID) # For the start-date entity es = es.add_relationship( ft.Relationship(es["data_with_start"]["id"], es["data_with_start"]["id"]) ) # For the end-date entity es = es.add_relationship( ft.Relationship(es["data_with_end"]["id"], es["data_with_end"]["id"]) ) # Generate features based on start_date features_start, _ = ft.dfs( entityset=es, target_dataframe_name="data_with_start", agg_primitives=["time_since_first", "time_since_last", "avg_time_between", "trend"], trans_primitives=[] # Add any transformation primitives you need ) # Generate features based on end_date features_end, _ = ft.dfs( entityset=es, target_dataframe_name="data_with_end", agg_primitives=["time_since_first", "time_since_last", "avg_time_between", "trend"], trans_primitives=[] ) # Merge the two feature dataframes into one final_features = features_start.merge(features_end, on="id")
Option 2: Reassign the time index and run DFS twice
If you don't want to duplicate entities, you can create a single entity, run DFS with one time index, then update the time index and run DFS again.
import featuretools as ft import pandas as pd # Using the same sample dataframe as above df = pd.DataFrame({ "id": [1, 1, 2, 2], "start_date": pd.to_datetime(["2023-01-01", "2023-02-01", "2023-01-15", "2023-03-01"]), "end_date": pd.to_datetime(["2023-01-10", "2023-02-10", "2023-01-25", "2023-03-10"]), "value": [10, 20, 15, 25] }) # Initialize EntitySet with start_date as time_index es = ft.EntitySet(id="my_time_data") es = es.add_dataframe( dataframe_name="data", dataframe=df, index="id", time_index="start_date" ) # Add self-relationship for per-ID calculations es = es.add_relationship( ft.Relationship(es["data"]["id"], es["data"]["id"]) ) # Generate start-date features features_start, _ = ft.dfs( entityset=es, target_dataframe_name="data", agg_primitives=["time_since_first", "time_since_last", "avg_time_between", "trend"] ) # Update the time index to end_date es["data"].time_index = "end_date" # Generate end-date features features_end, _ = ft.dfs( entityset=es, target_dataframe_name="data", agg_primitives=["time_since_first", "time_since_last", "avg_time_between", "trend"] ) # Merge results final_features = features_start.merge(features_end, on="id")
Key Notes:
- Self-Relationship: Adding a self-relationship (linking the entity to itself via the primary key) is crucial if you want to calculate features like
time_since_first_start_datefor each uniqueid—it tells Featuretools to group data by the ID when computing time-based aggregates. - Feature Naming: Featuretools will automatically name features to reflect which time index they use (e.g.,
TIME_SINCE_FIRST(data_with_start.start_date)), so you can easily distinguish between start-date and end-date derived features. - Cutoff Times: If you're working with temporal validation (using
cutoff_times), make sure to adjust the cutoff times to align with whichever time index you're using for each DFS run.
内容的提问来源于stack exchange,提问作者Ganesh Bhat

