如何确定DFS及FeatureTools中Deep Feature Synthesis使用的原语列表?
Great question! Picking the right primitives for Deep Feature Synthesis (DFS) is crucial—they directly shape the quality of the features you generate, and choosing wisely can save you from sifting through noise while capturing the signals your model needs. Let’s break this down into two clear parts:
1. How to Determine Which Primitives to Use
There’s no universal list, but these guidelines will help you narrow it down based on your data and problem:
Start with data types and entity relationships
- For numeric fields: Use aggregation primitives like
sum,mean,median,max,min, or transform primitives likerolling_mean(for time-series data). - For categorical fields: Go with
count_unique,mode, orpercent_true(if binary). - For time fields: Use
time_since_last,days_in_month, oryearto extract temporal signals. - For entity relationships (e.g., one-to-many between users and transactions): Use aggregation primitives to roll up child entity data to the parent (e.g., total spend per user).
- For numeric fields: Use aggregation primitives like
Align with business logic
Think about what features would matter for your use case. For example:- In e-commerce: "Total spend in the last 30 days" is more actionable than "lifetime total spend" for predicting near-term purchases.
- In fraud detection: "Number of failed login attempts in the last hour" is a stronger signal than total failed attempts ever.
Start simple, then iterate
Don’t overload your DFS with every primitive at once. Begin with core primitives (sum,count,mean,mode) and evaluate model performance. Then add more specific primitives (likeskew,kurtosis, ortime_since_first) based on feature importance scores.Avoid redundant primitives
If you already havesumandcountfor a field,meanis justsum/count—you might not need both unless your model explicitly benefits from redundant signals. Prioritize features that add unique information.Balance performance and interpretability
Complex primitives (like custom aggregations) might boost model accuracy, but if you need explainable results, stick to simpler, easier-to-understand primitives.
2. How to Select Primitives in FeatureTools
FeatureTools gives you full control over which primitives to use. Here’s how to implement your selection:
List available primitives first
Before choosing, check what’s built-in with this command:
import featuretools as ft # List all aggregation primitives print(ft.list_primitives("aggregation")) # List all transform primitives print(ft.list_primitives("transform"))
Manually specify primitives in dfs()
Use the agg_primitives and trans_primitives parameters to pass your chosen list:
# Define your target primitives agg_primitives = ["sum", "mean", "count_unique", "max"] trans_primitives = ["year", "month", "time_since_previous"] # Run DFS with your selected primitives features, feature_defs = ft.dfs( dataframes=your_dataframes_dict, target_dataframe_name="target_entity", agg_primitives=agg_primitives, trans_primitives=trans_primitives, verbose=True )
Exclude unwanted primitives
If you want to use most built-in primitives except a few, use exclude_agg_primitives or exclude_trans_primitives:
features, feature_defs = ft.dfs( dataframes=your_dataframes_dict, target_dataframe_name="target_entity", exclude_agg_primitives=["mode"], # Skip mode if rare categories are common verbose=True )
Add custom primitives
If built-in primitives don’t cover your use case, define your own and include them in the list:
from featuretools.primitives import AggregationPrimitive from featuretools.variable_types import Numeric # Custom aggregation to calculate average transaction amount class AvgTransaction(AggregationPrimitive): name = "avg_transaction" input_types = [Numeric] return_type = Numeric def get_function(self): def avg(x): return x.mean() return avg # Use the custom primitive alongside built-ins agg_primitives = ["sum", AvgTransaction] features, feature_defs = ft.dfs( dataframes=your_dataframes_dict, target_dataframe_name="customers", agg_primitives=agg_primitives, verbose=True )
Automate feature filtering post-DFS
After generating features, use featuretools.select_features to keep only impactful ones (based on model importance or correlation), which can inform your primitive choices for future runs:
from sklearn.ensemble import RandomForestClassifier from featuretools.selection import select_features # Train a quick model to get feature importances model = RandomForestClassifier() model.fit(features, target_labels) # Select top 20 features selected_features = select_features(features, target_labels, model=model, top_n=20)
Remember, the best primitive list comes from experimentation. Test small sets, iterate based on model performance, and keep your business context front and center.
内容的提问来源于stack exchange,提问作者Harish Rajula

